Novel RNA biomarkers for the diagnosis of prostate cancer

A novel RNA biomarker panel with rigorous quality control and machine learning models addresses the limitations of current prostate cancer screening, enhancing diagnostic accuracy and reducing invasive procedures.

JP2025529263APending Publication Date: 2025-09-04FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025513311
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-05
Filing Date
2023-09-04
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Current prostate cancer screening methods, such as PSA testing and digital rectal examinations, suffer from high false-positive rates and invasive biopsies, leading to overdiagnosis and complications, while existing RNA biomarkers like PCA3 and TMPRSS2:ERG have limitations in sensitivity and specificity, particularly for low-grade tumors.

Method used

Development of a diagnostic method using a set of RNA biomarkers, including CPNE4, PCAT14, and MSMB, analyzed through rigorous quality control and normalization, combined with machine learning models to improve the detection of prostate cancer, especially for Gleason score 6 tumors.

Benefits of technology

The new biomarkers significantly reduce false positives, enabling early and reliable diagnosis of prostate cancer, including low-grade tumors, thereby minimizing unnecessary biopsies and improving patient outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025529263000019
    Figure 2025529263000019
  • Figure 2025529263000020
    Figure 2025529263000020
  • Figure 2025529263000021
    Figure 2025529263000021
Patent Text Reader

Abstract

The present invention includes a method for ex vivo diagnosis of prostate cancer, comprising the steps of: i) providing a sample from a patient suspected of having prostate cancer; and ii) analyzing the expression level of at least one newly identified prostate cancer biomarker in the sample, wherein the sample is designated as prostate cancer positive when the expression level of the biomarker is above a threshold. Furthermore, the present invention includes strict quality control standards for the sample and method to ensure a reliable and specific method for diagnosing prostate cancer.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention is in the fields of biology and chemistry. Specifically, the present invention is in the field of molecular biology. More specifically, the present invention relates to the analysis of RNA transcripts. Most specifically, the present invention is in the field of prostate cancer diagnosis. [Background technology]

[0002] Prostate cancer (PCa) is the most common malignancy in men and the third leading cause of cancer-related deaths in Western countries. In 2020, Western Europe had the second highest annual number of newly diagnosed PCa cases. Despite widespread screening for prostate cancer and significant advances in the treatment of metastatic disease, PCa remained the third leading cause of cancer death among European men in 2020.

[0003] Currently, prostate-specific antigen (PSA) serum level testing and digital rectal examination are the two main screening methods. Patients with abnormal results are usually advised to undergo prostate biopsy. However, this has serious implications. The lack of specificity of PSA screening results in a large number of false positives, which leads to unnecessary prostate biopsies being performed on millions of men worldwide each year (overdiagnosis). In addition, taking biopsies carries a substantial risk of infectious complications. Therefore, there is an urgent need for diagnostic assays with higher sensitivity and specificity for the early diagnosis of prostate cancer in order to improve prostate cancer screening and avoid unnecessary multiple prostate biopsies. The present invention addresses this issue by providing a set of biomarkers for prostate cancer screening and diagnosis.

[0004] Prostate-specific antigen Currently, the most common screening assay for PCa is measurement of serum PSA concentrations. However, there is no single PSA threshold that can reliably distinguish PCa patients from those without PCa. The lack of specificity of the PSA test is illustrated by a high false-positive rate of up to 75% (Duffy, 2020). The area under the receiver operating characteristic curve (AUC) value for PSA testing to discriminate between PCa and non-cancer was reported to be 0.678 (Thompson, 2005). Although significant research has been conducted to improve the performance of PSA itself, including measurement of free or cleaved PSA (Jansen, 2009), the problem remains unresolved. Based on this evidence, the United States Preventive Services Task Force (USPSTF) has recommended against PSA-based prostate cancer screening due to the high false-positive rate and the associated risks of biopsy and overtreatment (Moyer, 2012).

[0005] PSA-derived biomarkers include the Prostate Health Index (PHI, Jansen, 2010) and the 4-kallikrein panel (4K score, Vickers, 2008). Both combine serum levels of various PSA protein variants, the latter also measuring kallikrein 2. Related studies have primarily focused on high-grade PCa (GS ≥ 7) to predict aggressive tumors (Duffy, 2020).

[0006] Prostate cancer antigen 3 Prostate cancer antigen 3 (PCA3), a non-coding RNA expressed exclusively in the prostate, is overexpressed in 95% of PCa samples compared with normal or benign hyperplastic prostate tissue (Salagierski, 2012). The Progensa PCA3 assay (Gen Probe Inc., San Diego, CA, USA) is a commercially available diagnostic test that quantitatively detects PCA3 in urine and prostatic fluid. A urinary PCA3 score of greater than 35 has been shown to demonstrate an average sensitivity and specificity of 66% and 76%, respectively, for diagnosing PCa (Van Gils, 2007), with a reported area under the curve (AUC) of 0.693 (Aubin, 2010). Furthermore, elevated PCA3 scores have been shown to increase the likelihood of a positive repeat biopsy in men with one or two previous negative biopsies (Haese, 2008). Although the specificity and sensitivity values ​​of the PCA3 assay are slightly better than those of the PSA serum test, they are still insufficient, and therefore the PCA3 test has not yet achieved widespread acceptance in actual clinical use.

[0007] Gleason score The so-called Gleason classification is used in combination with other parameters to predict PCa prognosis and guide treatment. The Gleason score is assigned by pathologists based on the microscopic features of a prostate biopsy and reflects the differentiation state of prostate tumor cells observed in the biopsy specimen (Gleason, 1977). Prostate cancers with a low Gleason score have a favorable prognosis, and the majority may not require treatment and should be monitored with active surveillance. However, its informative value is limited, and accurately identifying tumors that are aggressive and lethal despite a low Gleason score remains a clinical challenge (Irshad, 2013).

[0008] Novel biomarkers The TMPRSS2:ERG gene fusion is the most common genetic abnormality in PCa, occurring in approximately 40%–70% of patients. Although highly sensitive TMPRSS2-ERG assays are available for clinical use, they cannot be used to target TMPRSS2:ERG-negative PCa cases. Furthermore, the correlation between TMPRSS2:ERG and prognosis remains controversial. In addition, several novel biomarkers have been proposed in recent years as candidates for use in the diagnosis or prognosis of PCa, but these have not yet been implemented in routine clinical practice. This includes RNA biomarkers in addition to genetic biomarkers. A recent external validation study of the Select MDx test demonstrated poorer performance than previously reported tests (Lendinez-Cano, 2021). Mi-Prostate Score and ExoDx use information on the expression levels of the prostate-specific noncoding RNA PCA3 and the fusion transcript TMPRSS2:ERG. On the other hand, SelectMDx is a qPCR-based assay that measures the mRNA levels of DLX1 and HOXC6 for the detection of PCa. The three reviewed tests highlight good clinical performance in predicting high-grade PCa (GS≥7). In general, the inclusion of clinical information such as age, digital rectal examination (DRE) findings, and previous biopsies further improves test performance (Cucchiara, 2017). Summary of the Invention

[0009] The strong emphasis on high-grade tumors in the above tests improves test performance and may exclude low-risk PCa (equivalent to GS6) from biopsy. However, this carries the risk of overlooking at least some high-grade tumors. Although GS6 PCa is typically described as non-threatening, it is still cancer and has the potential to progress to GS7 or higher and metastasize. In this case, a different therapeutic approach is required. Patients with GS6 PCa are typically considered for active surveillance. It has been shown that more men who undergo active surveillance experience disease progression than those who undergo definitive treatment strategies. Because tests for early detection of PCa before biopsy identify patients who would benefit from further diagnosis, excluding GS6 tumors in test development may result in the test lacking sensitivity for the high-risk subgroup of GS6 PCa patients. Whether GS6 PCa patients should be pre-excluded from further invasive diagnosis is controversial.

[0010] The objective of this study was to identify new RNA biomarkers for the early diagnosis of PCa (GS≧6) to reduce unnecessary biopsies. This study also did not exclude patients with GS=6, so the method of the present invention can also detect early stages of PCa. In particular, the method of the present invention aims to diagnostically distinguish PCa tumors from BPH (benign prostatic hyperplasia). While the previously published application WO 2015 / 082418 also presents newly identified biomarkers for PCa diagnosis, this application provides a superior diagnostic method with improved reliability using a different model and selection of biomarkers.

[0011] Brief description of the invention Differentially expressed transcripts in tumor and control tissues were identified by next-generation sequencing from 64 prostate cancer patient and control samples, and 203 and 338 samples, respectively, were validated by microarray and qRT-PCR analysis. From these samples, a selection of potential RNA biomarkers suitable for use in prostate cancer diagnosis was identified. Evaluation and optimization of the identified biomarkers was used to develop a diagnostic method that would result in improved, more reliable PCa diagnosis. Essentially, this diagnostic method involves rigorous quality assessment of patient samples, followed by normalization of the RNA biomarkers to a set of selected and optimized reference RNAs. The normalized expression data of the RNA biomarkers is then analyzed by an algorithm to determine a diagnosis.

[0012] The present invention relates to a method for diagnosing prostate cancer, comprising analyzing the expression level of at least one nucleic acid selected from the group consisting of SEQ ID NO: 2 and SEQ ID NO: 20, wherein a sample is designated as prostate cancer positive when at least one of the nucleic acids is present and / or when the expression level of at least one of the nucleic acids is above a threshold value.

[0013] In a preferred embodiment, the present invention relates to a method for diagnosing prostate cancer, comprising a step of analyzing the expression level of at least one nucleic acid based on SEQ ID NO: 2, SEQ ID NO: 19-20, and SEQ ID NO: 35-37 in a sample derived from a patient, and designating the sample as prostate cancer positive when the expression level of the nucleic acid is above a threshold value.

[0014] In an alternative and equally preferred embodiment of the present invention, the present invention relates to a method for diagnosing prostate cancer, comprising a step of analyzing the expression level of at least one nucleic acid based on SEQ ID NOs: 1 to 3, 20 to 34, and 38 to 50 in a sample derived from a patient, wherein the sample is designated as prostate cancer positive when the expression level of the nucleic acid is above a threshold value.

[0015] In one embodiment, the present invention relates to a nucleic acid that hybridizes under stringent conditions to one of the nucleic acids having SEQ ID NOs: 1 to 3 or SEQ ID NOs: 19 to 50 or any portion thereof, or that shares preferably 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with a nucleic acid that is any one of the nucleic acids set forth in SEQ ID NOs: 1 to 3 or SEQ ID NOs: 19 to 50.

[0016] The present invention also relates to the use of a nucleic acid that hybridizes under stringent conditions with one of the nucleic acids having SEQ ID NOs: 1 to 3 or SEQ ID NOs: 19 to 50 or any portion thereof, or that shares preferably 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with a nucleic acid that is any one of the nucleic acids based on SEQ ID NOs: 1 to 3 or SEQ ID NOs: 19 to 50, for the diagnosis of prostate cancer.

[0017] The present invention relates to a probe or primer specific to a sequence in the group consisting of SEQ ID NOs: 1 to 131, and preferably the specific sequence of the probe or primer is selected from the group comprising SEQ ID NOs: 136 to 174.

[0018] The present invention relates to the use of nucleic acids having a sequence of the group consisting of SEQ ID NOs: 136 to 174 for the diagnosis of prostate cancer.

[0019] The present invention relates to nucleic acids that preferably share at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with a nucleic acid having a sequence in the group consisting of SEQ ID NOs: 1 to 131 or their reverse complements, or a nucleic acid that is any one of the nucleic acids based on SEQ ID NOs: 1 to 131.

[0020] The present invention also relates to a kit for diagnosing prostate cancer, comprising a nucleic acid that hybridizes under stringent conditions with a nucleic acid based on SEQ ID NOs: 1 to 131, or a nucleic acid that preferably shares at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with a nucleic acid that is any one of the nucleic acids based on SEQ ID NOs: 1 to 131, and reagents for amplifying and / or quantifying and / or detecting the nucleic acid.

[0021] definition The following definitions are provided for certain terms used in the body of this application.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this technology belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of this technology, the preferred methods, devices, and materials are described herein.

[0023] When an indefinite or definite article such as "a" or "an" or "the" is used to refer to a singular noun, the plural of that noun is also included unless otherwise specified. Conversely, when a plural form of a noun is used, the plural also refers to the singular. For example, when biomarkers are referred to, this should also be understood as a single biomarker.

[0024] As used herein, "nucleic acid" or "nucleic acid molecule" generally refers to any ribonucleic acid or deoxyribonucleic acid, which may be unmodified or modified. "Nucleic acid" includes, but is not limited to, single-stranded and double-stranded nucleic acids. As used herein, the term "nucleic acid" also includes the above-mentioned nucleic acids containing one or more modified bases. Thus, a nucleic acid having one or several backbone modifications for stability or other reasons is a "nucleic acid." As used herein, the term "nucleic acid" encompasses such chemically, enzymatically, or metabolically modified forms of nucleic acid as well as chemical forms of nucleic acid characteristic of viruses and cells (including, for example, simple and complex cells).

[0025] The term "transcript" refers to a nucleic acid produced by making a copy of a template nucleic acid. This definition includes nucleic acids obtained by transcribing a template nucleic acid, such as a gene locus, with a polymerase to generate a new nucleic acid having a sequence complementary to the template. The term transcript also includes any nucleic acid that has been further processed after transcription. Such further processing includes, but is not limited to, splicing, polyadenylation, editing, partial digestion, ligation, and labeling. The term also encompasses nucleic acids obtained from a transcript through, for example, further transcription, reverse transcription, amplification, or extension. A transcript can correspond to any portion of the template nucleic acid, preferably at least 20 nucleotides. A transcript also has at least 90%, preferably at least 95%, and most preferably at least 99% sequence identity with the template nucleotide or a portion of the template nucleotide, or the reverse complement of the template nucleic acid.

[0026] The terms "locus(hg38)" and the abbreviation "hg38" are used to refer to a specific locus in the hg38 assembly of the human genome. References are expressed in the form, for example, "Chr2:1,550,437-1,629,191," which refers to the sequence interval on human chromosome 2 within the genome assembly.

[0027] As used herein, "sample molecule" or "individual sample template" refers to any type of nucleic acid molecule contained in the sample being analyzed, such as single- or double-stranded DNA and / or RNA.

[0028] "Biomarker" within the present invention refers to any gene, gene transcript, gene transcript variant, or non-coding genomic sequence that is predictive of the diagnosis of PCa.

[0029] The term "level" or "expression level" in the context of the present invention refers to the level that biomarker exists in the sample from patient.The expression level of biomarker is generally measured by comparing its expression level with the expression level of one or several reference genes or housekeeping genes in the sample for normalization.If the expression level of biomarker exceeds the expression level of the same biomarker in appropriate control (for example, healthy tissue) by a set threshold, the sample from patient is designated as prostate cancer positive.Expression level is further calibrated for comparison between runs using the reference gene measured in calibrator sample.

[0030] As used herein, the term " analyzing sample for the presence and / or level of nucleic acid " or " specifically estimate the level of nucleic acid " refers to the means and method that are useful for evaluating and quantifying the level of nucleic acid.One of the useful methods is, for example, quantitative reverse transcription PCR.Similarly, the level of RNA can also be analyzed by, for example, Northern blot, next-generation sequencing, or by using the spectroscopic analysis technique that comprises measuring the absorbance at 260nm and 280nm after amplification.

[0031] "Amplification" of an individual sample template refers to any type of nucleic acid amplification method that generates multiple copies of the original template.

[0032] In this specification, analysis or testing refers to determining whether amplification has occurred, determining whether the target sequence is between the primer regions and, optionally, whether its length is appropriate, and determining the amount of amplification product having the correct target sequence. Preferably, in the method of the present invention, accurate quantification of the amplification product is desired.

[0033] As used herein, the term "amplified" when applied to a nucleic acid sequence refers to a process in which one or more copies of a specific nucleic acid sequence are generated from a nucleic acid template sequence, preferably by polymerase chain reaction (PCR). Other amplification methods include, but are not limited to, ligase chain reaction (LCR), polynucleotide-specific based amplification (NSBA), loop-mediated isothermal amplification (LAMP), quantitative reverse transcription real-time PCR (qRT-PCR), quantitative (real-time) PCR (qPCR), droplet digital PCR, digital PCR, or any other amplification method known in the art.

[0034] As used herein with respect to the use of diagnostic and prognostic markers, the term "correlating" refers to comparing the presence or amount of a marker in a sample from a patient with the presence or expression level of that marker in a sample from a person known to have or be at risk for a given condition. Marker expression levels in a patient sample can be compared to levels known to be associated with a particular diagnosis.

[0035] As used herein, the term "diagnosis" or "diagnostic" refers to identifying a disease (in this case, prostate cancer) at any stage of its development, and also includes determining a subject's predisposition to developing the disease.

[0036] "Labeling" a DNA or RNA strand can be achieved by attaching or incorporating any type of marker that is detectable by conventional imaging techniques, for example, a fluorescent marker.

[0037] As used herein, the term "fluorochrome" refers to any chemical that absorbs light energy of a particular wavelength and re-emits the light at a different wavelength.

[0038] The application of fluorescent dyes to target detection involves either using intercalating dyes in PCR reactions to detect the formation of double-stranded DNA molecules, or labeling primers or probes with suitable fluorescent molecules and observing the formation of nucleic acid molecules with complementary target sequences according to the detected fluorescent signal of the fluorophore molecule. The use of different fluorophores with distinct emission and excitation spectra makes it possible to combine multiple detection systems into one PCR reaction (multiplex PCR) and detect the fluorescent signals individually. Potential labeling molecules include, but are not limited to, fluorescent dyes or chemiluminescent dyes, especially cyanine-type dyes. In the context of the present invention, fluorescence-based assays include, for example, FAM (5-carboxyfluorescein or 6-carboxyfluorescein), VIC, NED, fluorescein, fluorescein isothiocyanate (FITC), IRD-700 / 800, cyanine dyes (e.g., CY3, CY5, CY3.5, CY5.5, Cy7), xanthene, 6-carboxy-2',4',7',4,7-hexachlorofluorescein (HEX), TE, T, 6-carboxy-4',5'-dichloro-2',7'-dimethoxyfluorescein (JOE), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 5-carboxyrhodamine-6G (R6G5), 6-carboxyrhodamine-6G (RG6), rhodamine, rhodamine green, rhodamine red, rhodamine 110, BODIPY dyes (e.g., BODIPY TMR), Oregon Green, coumarins (e.g., umbelliferone), benzimides (e.g., Hoechst 33258), phenanthridines (e.g., Texas Red), Yakima Yellow, Alexa Fluor, PET, ethidium bromide, acridinium dyes, carbazole dyes, phenoxazine dyes, porphyrin dyes, polymethine dyes, and the like.In the context of the present invention, chemiluminescence-based assays involve the use of dyes based on the physical principles described for chemiluminescent materials in Kirk-Othmer, Encyclopedia of Chemical Technology, 4th ed., executive editor, J.I. Kroschwitz; editor, M. Howe-Grant, John Wiley & Sons, 1993, vol. 15, pp. 518-562 (hereby incorporated by reference, including citations on pages 551-562). Preferred chemiluminescent dyes are acridinium esters.

[0039] As used herein, "isolated" when used with respect to nucleic acids means that a naturally occurring sequence has been removed from its normal cellular (e.g., chromosomal) environment or has been synthesized in a non-native environment (e.g., artificially synthesized). Thus, an "isolated" sequence may be in a cell-free solution or placed in a different cellular environment.

[0040] As used herein, a "kit" refers to a packaged combination, optionally including instructions for use of the combination and / or other reactants and components for such use. When a kit contains a nucleic acid, the kit may also include a synthetic or non-natural variant of the nucleic acid. It should be understood that a synthetic or non-natural nucleic acid is a nucleic acid that contains any chemical, biochemical, or biological modification such that the nucleic acid does not occur in this form in nature. Such modifications include, but are not limited to, labeling with fluorescent dyes or quencher moieties, biotin tags, and modifications in the backbone of the nucleic acid, or any other modification that distinguishes the nucleic acid from its natural counterpart. The same applies to other natural compounds, such as proteins, lipids, etc.

[0041] The term "patient," as used herein, refers to a living human or non-human organism that is receiving medical care, or that should receive medical care because of a disease, or that is suspected of having a disease. This includes people who do not have a confirmed disease but are being investigated for signs of pathology. Thus, the methods and assays described herein are applicable to both human and veterinary diseases.

[0042] As used herein, the term "primer" refers to a nucleic acid, whether naturally occurring, such as in a purified restriction enzyme digest, or synthetically produced, that can act as a point of initiation of synthesis when placed in an environment of suitable temperature and pH under conditions that induce synthesis of a primer extension product complementary to the nucleic acid strand, i.e., in the presence of nucleotides and an inducing agent (e.g., DNA polymerase). A primer can be single-stranded or double-stranded and must be long enough to prime the synthesis of the desired extension product in the presence of the inducing agent. The exact length of the primer will depend on many factors, including temperature, source of primer, and the method used. Preferably, the length of a primer is about 15-100 bases, more preferably about 20-50 bases, and most preferably about 20-40 bases. The factors involved in determining the appropriate length of a primer are readily apparent to those of skill in the art. Optionally, a primer can be synthetic, meaning that it contains chemical, biochemical, or biological modifications. Such modifications include, but are not limited to, labeling with fluorescent dyes or quencher moieties, or modifications in the backbone of the nucleic acid, or any other modification that distinguishes the primer from its natural nucleic acid counterpart.

[0043] The term "probe" refers to any element that can be used to specifically detect a biological entity, such as a nucleic acid, protein, or lipid. A probe includes a portion of the probe that allows the probe to specifically bind to the biological entity, as well as at least one modification that allows the probe to be detected in an assay. Such modifications include, but are not limited to, labels such as fluorescent dyes, quenchers, specifically introduced radioactive elements, or biotin tags. A probe can also include modifications in its structure, such as locked nucleic acids.

[0044] As used herein, the term "sample" refers to a sample of bodily fluid or tissue obtained for the purpose of diagnosing, prognosing, or evaluating a subject of interest, such as a patient. In addition, one skilled in the art will understand that some test samples are more easily analyzed after a fractionation or purification procedure, for example, separation of whole blood into serum or plasma components.

[0045] Thus, in a preferred embodiment of the present invention, the sample is selected from the group consisting of prostate tissue, biopsy material, lymph node, urine, ejaculate, blood, serum, plasma, circulating tumor cells in the blood or lymph, any tissue suspected of containing metastasis, and any material that may contain prostate tumor cells or their parts, such as vesicles, exosomes, microvesicles, and other vesicles, as well as free or protein-bound RNA molecules derived from prostate tumor cells. Preferably, the sample is a blood sample, and most preferably a serum or plasma sample. Importantly, urine (especially after digital rectal examination) and ejaculate are among the most preferred samples. The tissue sample may be a biopsy material or a tissue sample obtained during surgery. The sample for the method of the present invention is provided from the patient by suitable extraction means, so the sample is an ex vivo sample, i.e., a sample removed from the patient. The subsequent analysis of the sample is performed by the in vitro diagnostic method of the present invention.

[0046] The term "area under the curve (AUC)" as used herein refers to the area under the curve of a receiver operating characteristic (ROC) curve or ROC curve. AUC relates to the specificity and sensitivity of a biomarker. A perfect marker (AUC=1.0) will produce a point in the upper left corner of the ROC space, i.e., coordinates (0,1), which represents 100% sensitivity (no false negatives) and 100% specificity (no false positives).

[0047] The term "p-value" relates to the probability of obtaining the observed sample result (or a more extreme result) if the null hypothesis is in fact true, i.e., if there is no difference in the means between the groups. The smaller the p-value, the more likely it is that the alternative hypothesis explains the observed result better than the null hypothesis.

[0048] The term "adjusted p-value" refers to a p-value adjusted for multiple comparisons according to Benjamini & Hochberg. The applied method is detailed in the Examples section. [Brief explanation of the drawings]

[0049] [Figure 1] The cohort processed for the development of ProstaCheck. [Figure 2] Box plot of RNA-seq data for PCA3 transcripts. RNA-seq results from a retrospective PCa cohort consisting of eight prostate tissue samples as controls (C), eight PCa tumor samples each from Group V (very low risk, Gleason score <7, pN0), Group L (low risk, Gleason score 7, pN0), and Group M (intermediate risk, Gleason score ≤7, pN+), and 16 paired tumor and tumor-free tissue samples from Group H (high risk, Gleason score >7) derived from benign prostatic hyperplasia. [Figure 3]Expression in urinary sediments of transcripts per million (TPM) of known diagnostic markers (PCA3, HOXC6, DLX1) compared to genes in the ProstaCheck test (RNA-seq of urinary sediments from 6 BPH and 5 PCa cases). Genes included in ProstaCheck show higher expression, on average, in urinary sediments than known markers. DETAILED DESCRIPTION OF THE INVENTION

[0050] Detailed Description of the Invention The present invention describes a method for diagnosing prostate cancer.This method includes analyzing samples taken from patients and determining the specific level of biomarkers or combinations of biomarkers in patient samples.Then, the result is correlated with a threshold, and when the result exceeds the threshold, the patient sample is designated as prostate cancer positive.The score and threshold calculated for biomarkers are both obtained according to machine learning model.

[0051] The experimental data of the present invention demonstrates the outstanding sensitivity and specificity of the newly identified biomarkers (including two separate assay models, ProstaCheck), which clearly outperform current commercially available assays (Progensa PCA3, SelectMDx; see Table 1), as seen in the increased AUC values ​​of the ProstaCheck model. Figure 3 also shows a comparison of the newly identified biomarkers compared to the commercially available SelectMDx and PCA3 assays, which also clearly demonstrates the superiority of the newly identified biomarkers over prior art assays.

[0052] Therefore, diagnostic assays based on these new biomarkers can reduce the high false positive rate of current assays, thereby helping to avoid unnecessary invasive prostate biopsies.In addition, the method of the present invention does not exclude men with GS=6, and since men with GS=6 were also included in the experimental data set, the method of the present invention can enable early diagnosis of prostate cancer.Furthermore, the method of the present invention particularly includes the criteria for rigorously assessing the quality of samples before applying diagnostic assays, as well as the necessary normalization and calibration procedures for detected biomarkers.Sample quality assessment, normalization, and calibration are important elements of the method of the present invention, which ensure reliable and reproducible results.

[0053] [Table 1]

[0054] In addition to prostate cells, other cells may be present in urine or sediment samples. To ensure that the biomarker assay of the present invention is not adversely affected by the presence of various amounts of non-prostate cells, the detected expression of promising biomarker candidate genes was correlated with the expression of cell-specific reference genes. Typically, urine or sediment samples may contain leukocytes, kidney cells, or bladder cells in addition to prostate cells. Therefore, cell-specific reference genes for these cell types, namely, CD45 (leukocytes), UMOD (kidney cells), NPHS2 (kidney cells), and UPK2 (bladder cells), were selected. Biomarker candidates were only further investigated if their expression did not correlate with the expression of the mentioned cell-specific reference genes. Thus, the biomarkers of the present invention represent expression in prostate cells.

[0055] The term "urinary sediment" refers to cells, cell debris, and other precipitated material obtained by centrifugation of urine.

[0056] The combination of biomarkers and calculation of thresholds, as well as the calculation of sample analysis for diagnosis, are realized by a multivariate statistical model obtained by a machine learning approach using training set samples of different PCa stages compared to BPH (benign prostatic hyperplasia). The newly identified biomarkers of the present invention were found to be located in the genomic regions of the following genes: CPNE4 (biomarker: PCaD2 [SEQ ID NO: 1], PCaD4 [SEQ ID NO: 2], PCaD52 [SEQ ID NOs: 21-34]), novel ncRNA TAPIR (biomarker PCaD11 [SEQ ID NO: 3]), HOXB13 (biomarker PCaD42 [SEQ ID NO: 19]), PCAT14 (biomarker PCaD50 [SEQ ID NO: 20]), MSMB (biomarker PCaD54 [SEQ ID NOs: 35-37]), and FOLH1 (biomarker PCaD58 [SEQ ID NOs: 38-50]). Some of the listed biomarkers are located within exons of known transcript variants of their respective genes, while others are located within introns according to current annotations, potentially linking either new exons of as-yet-unknown transcript variants or novel genes encoding non-coding RNAs. For reference, PCA3, a previously described and commonly applied PCa biomarker, was also measured using either custom-designed primers or the commercially available Progensa PCA3 test (biomarker PCaD30 [SEQ ID NOs: 132-135]). Table 2 provides an overview of the newly identified biomarkers for the diagnosis of prostate cancer, as well as a statistical evaluation of these biomarkers. The performance of these biomarkers is shown in Table 2 and Figure 3.

[0057] These biomarkers were identified within the "discovery cohort" (see Figure 1) using a negative binomial log-linear model fitted to the observed data for each transcript or gene, and regression coefficients were determined using a likelihood ratio test. As shown in Table 1, the identified new biomarkers were combined with diagnostic assays in two different logistic models. These models were obtained using machine learning. The scheme for biomarker identification, evaluation, validation, and machine learning is shown in Figure 1. These models are logistic regression models fitted to the observed data for each biomarker to obtain regression coefficients. The first logistic model combined the biomarkers PCaD4, PCaD42, PCaD50, and PCaD54. The regression coefficients for these biomarkers in this model are -0.03104, 0.04972, -0.13662, and -0.14546, respectively. The second logistic model combines the biomarkers PCaD2, PCaD4, PCaD11, PCaD50, PCaD52, and PCaD58. In a preferred form of the second logistic model, PCaD2 represents the combination of markers measuring the region of the CPNE4 gene (PCaD2, PCaD4, and PCaD52). PCaD2 represents the smallest c measured after normalization and calibration for either PCaD2, PCaD4, or PCaD52. q The values ​​represent the highest expression levels of these markers observed in patients. The regression coefficients of these biomarkers in this model are -0.06262, -0.01419, -0.03030, -0.05290, -0.03138, and -0.04455, respectively. Therefore, both models include PCaD4 and PCaD50, and these two biomarkers are the essential minimum components of both models. Furthermore, these biomarkers can also be combined into additional models by themselves as an alternative embodiment of the present invention. However, the inclusion of additional biomarkers in the two described models improves the predictive value of the diagnostic assay.

[0058] The essential minimum biomarkers for both models, i.e., PCaD4 and PCaD50, were also statistically analyzed in combination with each other, excluding any additional markers related to either model of the present invention. This means that the data for the analysis of the combination of PCaD4 and PCaD50 was processed as data for each model of the present invention, including the quality control of the present invention and normalization using the reference biomarkers of the present invention. Subsequently, a regression model was calculated using the training set shown in Figure 1. Finally, the combination of PCaD4 and PCaD50 was tested in a validation cohort, as in Model 1 and Model 2, to determine the test parameters for the combination of PCaD4 and PCaD50 compared with prior art assays. The parameters indicate that the combination of PCaD4 and PCaD50 already has predictive and diagnostic advantages over prior art assays.

[0059] All applied models further include the use of reference biomarkers. These reference biomarkers are used for sample quality control and normalization of biomarker measurements. Within the assay models of the present invention, the following reference markers are used: KLK2 (biomarker PCaD91 [SEQ ID NOS: 51-73]), KLK3 (biomarker PCaD32 [SEQ ID NOS: 4-18]), EMC7 (biomarker PCaD92 [SEQ ID NOS: 74-77]), RAB7A (biomarker PCaD93 [SEQ ID NOS: 78-95]), and VCP (biomarker PCaD94 [SEQ ID NOS: 96-131]).

[0060] Thus, the present invention relates to a group of nucleic acids, the sequences of which include SEQ ID NOs: 1-131. Relevant properties of these sequences, along with further information regarding statistical analysis of individual biomarkers, are listed in Table 2 below. The sequences are separately provided in the sequence listing. These sequences represent selected biomarkers and selected reference markers (reference genes) for the diagnostic method. Performing the diagnostic method includes detecting at least one of the listed biomarkers and at least one of the listed reference markers, and using the detected expression levels to determine a patient's diagnosis of PCa.

[0061] [Table 2] TIFF2025529263000003.tif246156TIFF2025529263000004.tif247139

[0062] Specific primers and probes were designed to detect the biomarkers in Table 2. As those skilled in the art will recognize, numerous feasible design options are typically available to obtain primers and probes that are specific and reliable for a sufficiently long target sequence. Therefore, any primer or probe suitable for reliable and specific detection of the biomarkers and reference genes listed in Table 2 is part of the present invention. The inventors specifically designed and validated the primers and probes listed in Table 3. As mentioned above, these primers and probes should be understood as feasible examples; other primers and probes are available, and those skilled in the art will recognize how to design and validate them. Therefore, Table 3 does not represent a comprehensive list of all primers and probes of the present invention. The use of the primers and probes in Table 3 for detecting biomarkers and reference genes is a preferred embodiment of the present invention.

[0063] [Table 3] TIFF2025529263000006.tif181167

[0064] The biomarker PCA3 is routinely used to diagnose prostate carcinoma (PCa). Therefore, as expected, PCA3 expression levels were suggestive of PCa in subjects we tested using next-generation sequencing (Figure 2). However, we found that expression levels of this biomarker were highest in very low-risk tumors (V) and decreased as tumor risk factors increased. This finding makes PCA3 an unreliable marker for intermediate- and high-risk tumors and indicates the need for better prostate cancer biomarkers.

[0065] Therefore, the present invention relates to a method for diagnosing prostate cancer, comprising a step of analyzing the expression levels of nucleic acids based on SEQ ID NOs: 1 to 3 and SEQ ID NOs: 19 to 50, and designating a sample as prostate cancer positive when at least one of the nucleic acids is present and / or when the expression level of at least one of the nucleic acids is above a threshold value.

[0066] In an alternative embodiment, analyzing the expression level of the nucleic acid means analyzing the reverse complement of the nucleic acid or cDNA. The experimental results, as shown in Figure 3, demonstrate the high specificity and sensitivity of the novel biomarkers for PCa detection.

[0067] In a preferred embodiment, the sample is selected from the group consisting of prostate tissue, biopsy material, lymph nodes, urine, ejaculate, blood, serum, plasma, circulating tumor cells in the blood or lymph, any tissue suspected of containing metastasis, and any material that may contain prostate tumor cells or parts thereof, such as vesicles, exosomes, microvesicles, and other vesicles, and free or protein-bound RNA molecules derived from prostate tumor cells or parts thereof. More preferably, the sample is urine, even more preferably, the sample is urinary sediment, even more preferably, the sample is urine obtained from a patient after a digital rectal examination, and most preferably, the sample is urinary sediment obtained from a patient after a digital rectal examination.

[0068] Measurement of the RNA biomarkers of the present invention can be performed by any method suitable for specifically estimating RNA levels, for example, PCR-based methods such as qRT-PCR. The assays can be applied to early diagnosis of PCa (screening), predicting tumor aggressiveness (prognosis), and / or aiding in therapy selection.

[0069] In one embodiment of the present invention, the expression levels of the transcripts of the biomarker nucleic acids according to SEQ ID NOS: 1-3 and 19-50 are compared to the expression levels of one or several other gene transcripts in the sample, which are used as reference genes. Thus, the expression levels of the biomarker genes can be normalized according to the expression levels of the reference genes. Examples of suitable reference genes are shown in Table 4 below, although this list is not exhaustive and those skilled in the art will be aware of additional possible candidates.

[0070] [Table 4]

[0071] In another aspect of the present invention, the method of the present invention includes strict quality control standards for the inclusion of subject samples in diagnostic tests. The inventors of the present invention surprisingly discovered that a three-tiered quality control protocol ensures reliable and accurate diagnostic tests. The first tier involves carefully examining the amplification curve for appropriate sigmoidal curve shape, sufficiently large amplification, and sufficient amplification efficiency. Only amplifications with sufficient quality and reliability can be effectively included in diagnostic analysis. This includes ensuring that the sigmoidal shape of the amplification curve, i.e., the difference between the asymptote values, is at least 1.5, and that the minimum amplification efficiency of qPCR is at least 1.5.

[0072] Further, the quality control (QC) of the test sample includes a first QC filter. The first filter checks for a sufficient minimum amount of cells and a sufficient minimum amount of prostate cells in the sample (urine sediment sample). The first QC filter involves determining the expression of a selection of generally applicable reference genes in combination with at least one prostate-specific reference gene. The inventors have tested various possible reference genes for applicability within the method of the present invention. Those skilled in the art will recognize that generally, any gene whose expression is consistently and reliably detectable and stable across at least most cellular and culture conditions is a potential candidate for a reference gene. However, the suitability of the candidate for a particular application must always be tested and properly established. Suitable approaches for selecting candidate genes and testing their suitability are within the common knowledge of those skilled in the art. Suitable candidates for reference genes can be selected from, but are not limited to, the group of genes listed in Table 4. Several genes were examined as candidate reference genes, and the following genes were examined as general reference genes: EMC7, GAPDH, GPI, HMBS, HPRT1, RAB7A, TBP, and VCP, and the following genes were examined as prostate-specific reference genes (Table 4).All of these reference genes can be used in alternative embodiments of the present invention.In a preferred embodiment of the present invention, the reference genes of the first QC filter are selected from the group comprising general reference genes: EMC7, RAB7A, and VCP, and prostate-specific reference gene: KLK2.Further relevant information about these reference markers is provided in Table 2 and Table 3.

[0073] For a sample to pass the first QC filter, all four reference genes tested (EMC7, RAB7A, VCP, KLK2) must show expression levels above a threshold level. We determined the following threshold levels: for EMC7, expression levels above c q The sample passed the threshold if the value was at least 28. (All c qThe values ​​are calculated according to the Cy0 method (Guescini, 2008), where c q Alternative approaches for calculating values ​​are known in the art and are equally valid for the present invention.) For RAB7A, the expression level is c q The threshold is at least 30. For VCP, the expression level is q The threshold is at least 28. For KLK2, the expression level is q The threshold is at least as low as 28.

[0074] To further improve the reliability of the first QC filter, expression levels are determined by triplicate amplification reactions. Each test is run with a variance of c between the three replicates. q The maximum difference in values ​​is at most 0.5, or c between at least two of three replicates q Only samples with a maximum difference of at most 0.25 are suitable for analysis. If a sample does not pass the first QC filter, it is invalid and excluded from further analysis. The first QC filter described is then combined with a second QC filter.

[0075] The second QC filter is applied after normalizing and calibrating the observed diagnostic results.This quality control aspect of the present invention ensures that only reliable and reproducible data without large variations are included in the analysis.Those skilled in the art will recognize that this aspect of the present invention depends on the method and device used to quantify the expression level of biomarkers, and it is within the skill of those skilled in the art to appropriately adjust these quality control methods to methods other than RT-qPCR.In addition, minimal variations in the number of replicates (4, 5, or 6 replicates instead of 3) or slight variations in cut-off values ​​only have minimal impact on this quality control test, and these aspects are also included in the present invention.

[0076] Similar to the reference genes, the selected biomarkers were determined by triplicate amplification reactions for each sample to ensure reliable diagnosis. q Each test is suitable for analysis when the interindividual c of the three replicates is at least less than the selected threshold. q The maximum difference in values ​​is at most 0.5 or between at least two of three replicates. q The maximum difference in values ​​is at most 0.25. If a sample does not pass this test, it is invalid and excluded from further analysis. Alternatively, at least two of the three replicates of each test have a c q Each test is considered a valid c when the value is not at least less than the selected threshold, or when no threshold is preselected. q It can still be analyzed using only the value of q The value is the c value of the sample that passed all the QC checks. q In this alternative case, however, increased variability in the results can be expected. As noted above, those skilled in the art will recognize that this quality control can be adapted to different methods with slight variations, which are also included in the present invention.

[0077] The measured c of valid replicates of each sample, i.e., replicates that have passed all previous QC steps q The values ​​are used in the diagnostic analysis, which involves calculating the arithmetic mean for each set of technical replicate measurements (Equation 1).

number

[0078] For relative quantification, the average expression level across all samples on the same plate is calculated for each gene separately (Equation 2). Reference samples for inter-run calibration are excluded from this calculation.

number

[0079] average c q value (Equation 1) and the average expression level of each gene (Cq 参照,j , Equation 2), which represents the relative quantification of the expression level of the biomarker or reference gene in each sample (Equation 3).

number

[0080] These relative quantifications are normalized according to standard practice in the art by first calculating a normalization factor, which is calculated as the arithmetic mean of the relative quantifications of the reference genes (Equation 4).

number

[0081] In a preferred embodiment, the reference gene used is the general reference gene of the first QC filter.Nevertheless, those skilled in the art will recognize how to select and verify alternative sets of reference genes that can be realized.The normalization coefficient thus obtained is then used to calculate the normalized relative quantification of selected biomarkers by calculating the difference between the relative quantification of biomarkers and normalization coefficient (Equation 5).This normalization ensures the comparability of test samples within a run (intra-run normalization).

number

[0082] Inter-run calibration is necessary for reproducible diagnostic assays. Inter-run calibration is calculated in a similar manner to normalization. First, a calibration coefficient is determined, which is the arithmetic mean of the normalized relative quantification of the calibrators between runs. For inter-run calibration, each run / test plate contained a pool of three PCa tissue samples and a pool of three BPH patient tissue samples. For each gene, the average expression level in these two calibrator samples was calculated (Equation 6).

number

[0083] At least one calibrator sample must be validly detected. The calibration factor is then used to calibrate the normalized relative quantification of the selected biomarker by calculating the difference between the normalized relative quantification of the biomarker and the calibration factor (Equation 7). This calibration ensures comparability of test samples between separate runs (inter-run calibration).

number

[0084] Those skilled in the art will recognize alternative mathematical approaches to calculating gene normalization and calibration, which are also encompassed by the present invention.

[0085] The normalized and calibrated results are then confirmed by applying a second QC filter. This second QC filter determines whether the target sample contains a minimum amount of prostate cells sufficient to detect the selected prostate cancer biomarker. Similar to the first QC filter, those skilled in the art will recognize how to select and validate suitable reference genes. For prostate-specific reference genes, the inventors have considered and tested several possible candidates. Suitable reference genes can be selected, but they are not limited to the following group: ACPP, KLK2, KLK3, and RDH11 (Table 4). In a preferred embodiment of the present invention, the reference genes of the second QC filter are selected from the group consisting of prostate-specific reference genes KLK2 and KLK3.

[0086] For a sample to pass the second QC filter, the tested reference gene must show an expression level above a threshold level. We determined the following threshold levels: for KLK2, the sample passes the threshold when the expression level is at least less than the calibrated relative quantification of 9.6676667, and for KLK3, the threshold is when the expression level is at least less than the calibrated relative quantification of 11.113. If a sample does not pass the second QC filter, it is invalid and is excluded from further analysis.

[0087] The calibrated relative quantification of the selected biomarkers is then combined with a mathematical prediction model for patient diagnosis. The inventors have found two applicable and predictive models using different sets of biomarkers. These two models have surprisingly been found to achieve better diagnostic predictive power than prior art diagnostic methods (Table 1). The identified models share two common biomarkers essential to both models. The common biomarkers are PCaD4 and PCaD50. Therefore, in one embodiment of the present invention, a method for diagnosing PCa includes detecting at least the biomarkers PCaD4 and PCaD50. The determined expression levels are then used in the mathematical model to obtain a patient diagnosis.

[0088] Alternatively, the inventors identified a first model for diagnosis. This model combined calibrated and normalized relative quantification of the following biomarkers: PCaD4, PCaD54, PCaD50, and PCaD42. This model was found to have an AUC of 0.817 and a specificity of 0.512 at 90% sensitivity. A logistic regression model was fitted to the observed data for each transcript or gene to obtain regression coefficients.

[0089] Alternatively, the inventors identified a second model for diagnosis. This model combines the calibrated and normalized relative quantification of the following biomarkers: PCaD4, PCaD11, PCaD50, PCaD52, PCaD58, and PCaD2. This model was found to have an AUC of 0.829 and a specificity of 0.605 at 90% sensitivity. A logistic regression model was fitted to the observed data of each transcript or gene to obtain regression coefficients. The threshold is the minimum score calculated from the calibrated and normalized relative expression levels of selected biomarkers between test and control samples according to the applied model, at which the corresponding sample is designated as cancer positive.

[0090] The test parameters for the essential minimum sets of both models, consisting of biomarkers PCaD4 and PcaD50, which were statistically analyzed in the same manner as the regression models described above, were also determined within the validation cohort. For this essential minimum set of PCaD4 and PcaD50, the AUC was determined to be 0.832, with a 95% CI of 0.74-0.92. The specificity corresponding to a sensitivity of 90% for the essential minimum set was 0.721. These parameters demonstrate that even the essential minimum set already achieves superior predictive value compared to the reference tests SelectMDx and Progensa PCA3 (see Table 1).

[0091] The logistic regression model allows for the prediction of the tumor status of a test sample by combining the expression values ​​measured from the sample with the regression coefficients determined for the model. When the estimated tumor probability is 50% or higher, the sample is classified as tumor-positive in each model (Model 1 or Model 2). The present invention relates to the quantification of the expression levels of RNA biomarkers. After amplification, quantification is easy and can be achieved by several methods. When primers are used and at least one primer is attached to a fluorescent dye, quantification is possible using the fluorescent signal derived from the dye. Various primer systems and dyes are available, such as SYBR Green, multiplex probes, TaqMan probes, molecular beacons, and Scorpion primers. These are suitable for performing PCR-based methods such as quantitative reverse transcription PCR (qRT-PCR). Other possible quantification methods include, for example, Northern blotting, next-generation sequencing, or measuring absorbance at 260 nm and 280 nm.

[0092] Any suitable method for quantifying nucleic acids can be used to analyze the expression level of nucleic acids. In one embodiment of the present invention, the analysis in the method is performed by a fluorescence-based assay. In a preferred embodiment, the analysis is performed by measuring the fluorescence of a labeled primer, a labeled probe, or a fluorescent detection agent (e.g., SYBR Green). More preferably, the analysis of the expression level is performed by RT-qPCR. In this method, after reverse transcription, the sample is mixed with forward and reverse primers specific to at least one nucleic acid selected from the group of SEQ ID NOS: 1-131, preferably selected from the group of primers and probes listed in Table 2 (SEQ ID NOS: 136-174), followed by amplification. The probes or primers are designed to hybridize to their respective target sequences under stringent conditions.

[0093] In one embodiment, the analysis of expression levels is performed by next generation sequencing.

[0094] The present invention also relates to a nucleic acid that hybridizes under stringent conditions to one of the nucleic acids based on SEQ ID NOs: 1 to 131 or any part thereof, wherein the nucleic acid is a primer or a probe, preferably a labeled probe.

[0095] In a preferred embodiment of the present invention, a nucleic acid that hybridizes under stringent conditions to one of the nucleic acids based on SEQ ID NOs: 1 to 131 is about 5 to 500 nt in length, more preferably 10 to 200 nt, and even more preferably 10 to 100 nt in length. In a most preferred embodiment, the nucleic acid is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nt in length.

[0096] In one embodiment of the invention, the nucleic acid that hybridizes under stringent conditions to one of the nucleic acids based on SEQ ID NOs: 1 to 131 comprises a detectable label. In an even more preferred embodiment, the nucleic acid that hybridizes under stringent conditions to one of the nucleic acids based on SEQ ID NOs: 1 to 131 further comprises a quencher moiety.

[0097] The present invention also relates to the use of nucleic acids that hybridize under stringent conditions with one of the nucleic acids based on SEQ ID NOs: 1 to 131 for the diagnosis of prostate cancer.

[0098] The present invention further relates to nucleic acids having a sequence of the group consisting of SEQ ID NOs: 1 to 131, or their reverse complements, or any part thereof, or nucleic acids sharing preferably at least 85%, 90%, 95%, or 99% sequence identity with any one of the nucleic acids based on SEQ ID NOs: 1 to 131, or their reverse complements, or any part thereof, preferably the nucleic acid is selected from the group comprising SEQ ID NOs: 136 to 174.

[0099] The present invention further relates to the use of nucleic acids having a sequence of the group consisting of SEQ ID NOs: 1 to 131 for the diagnosis of prostate cancer.

[0100] The present invention also relates to a kit for diagnosing prostate cancer, comprising at least one nucleic acid that hybridizes under stringent conditions with one of the nucleic acids based on SEQ ID NOs: 1 to 131. The kit may contain multiple nucleic acids. In a preferred embodiment, the kit further comprises reagents for amplifying and / or quantifying and / or detecting the nucleic acid. In another embodiment, the kit comprises a control sample.

[0101] In an alternative embodiment, the present invention relates to a method for treating and diagnosing prostate cancer, comprising analyzing the expression level of nucleic acids according to SEQ ID NOs: 1-3 and SEQ ID NOs: 19-50 in a sample from a patient, designating the sample as prostate cancer positive when the expression level of the nucleic acid is above a threshold, and administering to the patient one or more prostate cancer therapeutic agents.

[0102] In one embodiment, prostate cancer therapeutic agents include docetaxel (Taxotere®); cabazitaxel (Jevtana®); mitoxantrone (Novantrone®); estramustine (Emcyt®); doxorubicin (Adriamycin®); etoposide (VP-16); vinblastine (Velban®); paclitaxel (Taxol®); carboplatin (Paraplatin®); abiraterone acetate, bicalutamide, casodex, degarelix, enzalutamide, goserelin acetate, leuprolide acetate, prednisone, sipuleucel-T, radium-223 dichloride, and / or vinorelbine (Navelbine®).

[0103] One aspect of the present invention is a) providing a sample from a patient suspected of having prostate cancer; b) analyzing the expression level of at least nucleic acids according to SEQ ID NO: 2 and SEQ ID NO: 20 in the sample; wherein the sample is designated as prostate cancer positive when the expression level of the nucleic acid is above a threshold.

[0104] In another alternative aspect, the present invention provides a method for producing a medicament for the treatment of a pulmonary arthritis, comprising: a) providing a sample from a patient suspected of having prostate cancer; b) analyzing the expression level of at least nucleic acids based on SEQ ID NO: 2, SEQ ID NOs: 19-20, and SEQ ID NOs: 35-37 in the sample; wherein the sample is designated as prostate cancer positive when the expression level of the nucleic acid is above a threshold.

[0105] In another alternative aspect, the present invention provides a method for producing a medicament for the treatment of a pulmonary arthritis, comprising: a) providing a sample from a patient suspected of having prostate cancer; b) analyzing the expression levels of nucleic acids based on at least SEQ ID NOs: 1 to 3, 20 to 34, and 38 to 50 in the sample; wherein the sample is designated as prostate cancer positive when the expression level of the nucleic acid is above a threshold.

[0106] In one embodiment of the invention, the sample for the method of the invention is selected from the group comprising prostate tissue, biopsy material, lymph nodes, urine, ejaculate, blood, serum, plasma, circulating tumor cells in the blood or lymph, any tissue suspected of containing metastasis, and any material that may contain prostate tumor cells or parts thereof, such as vesicles, including exosomes, microvesicles, and other vesicles, and free or protein-bound RNA molecules derived from prostate tumor cells.

[0107] In a preferred embodiment of the present invention, the sample is a urine sample, and in a more preferred embodiment, the sample is a urine sediment. In an even more preferred embodiment, the sample is a urine sample obtained after digital rectal examination (DRE) (DRE-urine), and in a most preferred embodiment, the sample is a urine sediment obtained after DRE (DRE-urine sediment). DRE is a routinely performed diagnostic method that allows the collection of a urine sample containing a certain amount of prostate cells. The obtained DRE urine sample is centrifuged and washed twice with PBS to isolate cells. Subsequently, the isolated cells are lysed, and the contained RNA is isolated according to a well-established technique that is generally known. Those skilled in the art are well aware of various applicable approaches to obtain purified RNA from cell samples.

[0108] In another embodiment of the invention, the method comprises analyzing expression levels by measuring the fluorescence of labeled primers, labeled probes, or fluorescent detection agents.

[0109] In another embodiment of the invention, the method comprises analysis of expression levels by qRT-PCR.

[0110] In another embodiment of the invention, the method comprises normalizing the expression levels of the biomarkers to general reference genes and prostate cell-specific reference genes.

[0111] In a preferred embodiment of the present invention, the general and prostate cell-specific reference genes are selected from the group comprising EMC7, RAB7A, VCP, KLK2, and KLK3.

[0112] In another embodiment of the method of the present invention, the method includes using a general reference gene and a prostate cell-specific reference gene to assess the quality of the sample. The expression levels of the general reference gene and the prostate cell-specific reference gene are used to provide at least one quality control filter, and only samples that pass at least one quality control filter are included in the analysis.

[0113] In another aspect of the present invention, the present invention includes nucleic acids that hybridize under stringent conditions to one of the nucleic acids set forth in SEQ ID NOS: 1-3 and 19-50. This means that the present invention includes nucleic acids that share at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with one of the sequences set forth in SEQ ID NOS: 1-3 and 19-50, or their reverse complements, or any subsequences thereof. Preferably, the nucleic acids are selected from the group comprising SEQ ID NOS: 136-144 and 148-162, and these nucleic acids can be used as primers or probes for detecting the sequences set forth in SEQ ID NOS: 1-3 and 19-50, or sequences that share at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity therewith.

[0114] The above-mentioned nucleic acids that can be used as primers or probes have a length of about 10 to 100 nucleotides, more preferably about 10 to 50 nucleotides, even more preferably about 15 to 40 nucleotides, even more preferably about 15 to 35 nucleotides, and most preferably about 15 to 25 nucleotides.

[0115] The above-mentioned nucleic acid may further comprise a detectable label.In a preferred embodiment, the detectable label is a fluorescent label that is covalently bound to nucleic acid.The fluorescent label may further be combined with at least one quencher molecule that is also covalently bound to nucleic acid and can quench the fluorescent signal of the fluorescent label.

[0116] The nucleic acids described above can be used in the diagnosis of prostate cancer.

[0117] Nucleic acids having one of the sequences of SEQ ID NOs: 1 to 131, or their reverse complements, or any part thereof, or nucleic acids sharing preferably at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with one of the sequences of SEQ ID NOs: 1 to 131, or their reverse complements, or any part thereof, preferably wherein the nucleic acids are selected from the group comprising SEQ ID NOs: 136 to 174, can be used as primers or probes for detecting the sequences of SEQ ID NOs: 1 to 131, or sequences sharing at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity therewith.

[0118] The above-mentioned nucleic acid can further comprise detectable label.In a preferred embodiment, the detectable label is a fluorescent label that is covalently bound to nucleic acid.The fluorescent label can also be combined with at least one quencher molecule that is also covalently bound to nucleic acid and can quench the fluorescent signal of fluorescent label.The labeled primer or labeled probe described can be used for diagnosing prostate cancer.

[0119] In another aspect of the invention, the invention also includes a kit for diagnosing prostate cancer, the kit comprising any of the above nucleic acids according to any one of SEQ ID NOs: 1 to 3 and SEQ ID NOs: 19 to 50, or SEQ ID NOs: 1 to 131, or SEQ ID NOs: 136 to 144 and SEQ ID NOs: 148 to 162, or SEQ ID NOs: 136 to 174, and optionally also reagents for amplifying and / or quantitating and / or detecting the nucleic acid.

[0120] In a further aspect of the invention, the present invention also provides a method for producing a composition comprising: a) analyzing the expression level of a nucleic acid based on any one of SEQ ID NO: 2 and SEQ ID NO: 20, or SEQ ID NO: 2, SEQ ID NO: 19-20, and SEQ ID NO: 35-37, or SEQ ID NO: 1-3, SEQ ID NO: 20-34, and SEQ ID NO: 38-50 in a patient-derived sample; b) designating the sample as prostate cancer positive if the expression level of the nucleic acid is above a threshold; c) administering to the patient one or more prostate cancer therapeutic agents; The present invention includes methods for the treatment and diagnosis of prostate cancer, including:

[0121] The agent may be selected from known prostate cancer therapeutic agents that may be used, but is not limited to the group including docetaxel (Taxotere®); cabazitaxel (Jevtana®); mitoxantrone (Novantrone®); estramustine (Emcyt®); doxorubicin (Adriamycin®); etoposide (VP-16); vinblastine (Velban®); paclitaxel (Taxol®); carboplatin (Paraplatin®); abiraterone acetate, bicalutamide, casodex, degarelix, enzalutamide, goserelin acetate, leuprolide acetate, prednisone, sipuleucel-T, radium-223 dichloride, and / or vinorelbine (Navelbine®).

[0122] In a further aspect of the invention, the present invention provides a method for producing a composition comprising: a) analyzing the expression level of at least nucleic acids based on SEQ ID NO: 2 and SEQ ID NO: 20, or SEQ ID NO: 2, SEQ ID NO: 19-20, and SEQ ID NO: 35-37, or SEQ ID NO: 1-3, SEQ ID NO: 20-34, and SEQ ID NO: 38-50 in a patient-derived sample; b) designating the sample as prostate cancer positive when the expression level of the nucleic acid is above a threshold; The present invention also includes a method for diagnosing prostate cancer, comprising:

[0123] The following examples detail the methods applied to identify and evaluate novel biomarkers and reference genes for the diagnosis of PCa. The newly identified biomarkers and the diagnostic methods based on these biomarkers described above enable more accurate and sensitive disease diagnosis than current biomarkers. [Example]

[0124] material and method Clinical cohort Patients with prostate carcinoma (PCa) who underwent radical prostatectomy (RPE) or surgery to remove benign prostatic hyperplasia (BPH) were included in a retrospective clinical cohort aimed at identifying novel biomarkers for PCa. In accordance with legal regulations, approval from the local ethics committee and informed consent from patients were obtained. Clinical follow-up data were collected for at least 5 years for PCa patients.

[0125] For the identification of diagnostically relevant biomarkers by genome-wide RNA sequencing, prostate tissue samples from a cohort of 40 PCa and 8 BPH patients were used. Four PCa groups were defined based on Gleason staging (Tannenbaum, M. Urologic Pathology: The Prostate, Philadelphia: Lea and Febiger, pp. 171-198; The Veterans Administration Cooperative Urologic Research Group: histologic grading and clinical staging of prostatic carcinoma) and the presence of adjacent lymph node metastasis at RPE (see Table 5). The composition of the test cohort is also shown in Figure 1.

[0126] Selected biomarker candidates were further validated by custom microarray and quantitative reverse transcription real-time PCR (qRT-PCR) in cohorts consisting of 203 patients (39 control BPH samples and 164 tumor samples) and 338 patients (126 control BPH samples and 212 tumor samples), respectively.

[0127] [Table 5]

[0128] Prostate tissue samples Prostate tissue samples were obtained by surgery and stored in liquid nitrogen. Prostate tissue samples obtained from radical prostatectomy (RPE) of patients with prostate carcinoma (PCa) were divided into tumor and tumor-free samples. Prostate tissue samples from patients with benign prostatic hyperplasia (BPH) were used as controls. Patient consent was always obtained.

[0129] To verify the sample condition and tumor cell content of the samples, all samples were divided into serial frozen sections. To this end, frozen tissue samples were embedded in Tissue-Tek OCT compound (Sakura Finetek GmbH) and frozen to fix them on a metal indenter. Frozen sections were prepared using a freezing microtome (Leica) equipped with a microtome blade C35 (FEATHER) cooled to -28°C. A total of 208 frozen sections were cut from all samples, four of which were HE-stained and evaluated by a pathologist for their tumor cell content. This resulted in three stacks of consecutive frozen sections, each bordered by an HE-stained section. Only stacks bordered on both sides by sections containing at least 60% tumor cells or sections containing up to 5% tumor cells were used as tumor or tumor-free samples, respectively. 50 frozen sections from the selected stacks were then subjected to RNA preparation.

[0130] RNA isolation Total RNA was isolated from cryopreserved tissues using Qiazol and miRNeasy Mini Kits on a QIAcube (all Qiagen), followed by manual digestion with DNase I. RNA concentration was determined using a Nanodrop 1000 (Peqlab). RNA integrity was verified with an Agilent Bioanalyzer 2100 (Agilent Technologies, Palo Alto, CA), and only RNA samples with an RNA integrity number (RIN) of at least 6 were further processed.

[0131] Genome-wide long RNA next-generation sequencing Genome-wide long-stranded RNA sequencing was performed on a subset of a retrospective PCa cohort consisting of eight prostate tissue samples from benign prostatic hyperplasia (BPH) as controls and 56 samples from prostate cancer patients (including tumor-free paired tissues from samples with a Gleason score of >7). Ribosomal RNA was depleted from 1 μg of total RNA using the Ribo-Zero rRNA Removal Kit (Epicentre). Sequencing libraries were prepared from 50 ng of rRNA-depleted RNA using the ScriptSeq v2 RNA-Seq Library Preparation Kit (Epicentre). Dual-tagged cDNA was purified using the Agencourt AMPure XP System Kit (Beckman Coulter). PCR was performed for 10 cycles to incorporate index barcodes for sample multiplexing and to amplify the cDNA library. The quality and concentration of the amplified libraries were determined using a DNA High Sensitivity Kit on an Agilent Bioanalyzer (Agilent Technologies). Four nanograms of each of the eight samples was pooled and size-selected on a 2% agarose gel using agarose gel electrophoresis. Samples ranging from 150 bp to 600 bp were excised from the gel and purified using the MinElute Gel Extraction Kit (Qiagen) according to the manufacturer's instructions. The purified libraries were quantified on the Agilent Bioanalyzer using a DNA High Sensitivity Chip (Agilent Technologies). The entire purified and size-selected library pool was then loaded onto an Illumina HiSeq2000 flow cell and distributed across all lanes. Cluster generation was performed using the TruSeq PE Cluster Kit v3 (Illumina Inc.) on an Illumina cBOT instrument according to the manufacturer's protocol. Sequencing was performed on an Illumina HiSeq2000 sequencing machine (Illumina, Inc.).The sequencing run details were as follows: paired-end sequencing strategy, 101 cycles for read 1, 7 cycles for the index sequence, and 101 cycles for read 2.

[0132] Analysis of sequencing data: preparation of raw data Raw sequencing data, including base call files (BCL files), were processed using CASAVA v1.8.1 (Illumina) to generate FASTQ files. The FASTQ files contained all sequenced RNA fragments (hereafter referred to as "reads") for each clinical sample. Specific adapter sequences were removed using cutadapt v1.6 (http: / / code.google.com / p / cutadapt / ).

[0133] Analysis of sequencing data: Genome mapping and transcript assembly: Reads were mapped to the human genome (assembly GRCh37 / hg19) using segemehl v0.1.4-382 and TopHat v2.0.87. Novel transcripts, i.e., transcripts not annotated in Gencode v17, were assembled using Cufflinks v2.1.18 and Cuffmerge v2.1.1. All novel transcripts and all known Gencode v17 transcripts were combined into a comprehensive annotation set.

[0134] Analysis of sequencing data: statistical analysis Htseq-count v0.6.0 (http: / / www-huber.embl.de / users / anders / HTSeq / doc / count.html) was used to calculate read counts per transcript and gene within a comprehensive annotation set of novel and known transcripts. Differentially expressed transcripts and genes were identified using R and the Bioconductor library DeSeq2 v1.81. Raw read counts were normalized and variance-stabilized (VST transformation). A negative binomial log-linear model was fitted to the read counts for each transcript or gene, and coefficients different from zero were identified using a likelihood ratio test. False discovery rates were controlled using the Benjamini-Hochberg adjustment.

[0135] Validation with custom microarrays Based on the sequencing results, we designed a custom microarray (Agilent SurePrint G3 Custom Exon Array, 4x180K, design ID 058029) with 180,000 probes, including mRNA, long non-coding RNA (gencode v15), novel transcripts, and all transcripts found to be differentially expressed between tumor and control tissue samples by RNA sequencing. Probe design was performed using Agilent's custom design tool, eArray.

[0136] Microarray screening was performed using a retrospective PCa cohort consisting of 40 prostate tissue samples and 8 control samples from patients with benign prostatic hyperplasia (BPH), as well as 164 tumor and 52 tumor-free tissue samples from prostate cancer patients after radical prostatectomy. cRNA was synthesized from 200 ng of total RNA using the Quick Amp Labeling Kit (Agilent), and 1650 ng of cRNA was hybridized on the array (Agilent Gene Expression Hybridization Kit).

[0137] Analysis of RNA custom microarray data: Differential expression analysis in the validation cohort was performed only on patients not included in the discovery cohort (164 PCa patients and 39 BPH patients). Differentially expressed probes were identified using R v3.2.2 and the Bioconductor package limma v3.24.15. Array quality control was performed by checking the distribution of "bright corner" and "dark corner" probes and spike-in concentrations relative to the normalized signal. BLAT v35 was used with the parameter -minIdentity=93 to obtain a set of probes mapping to unique genomic locations within hg19. All probes mapping to multiple distinct genomic regions were discarded. Quantile normalization was performed for inter-array normalization. Nonspecific filtering was applied to reduce the number of tests: probe expression must be greater than background expression on at least 20 arrays. Background expression is defined as the mean intensity plus three times the standard deviation of the negative control spots (Agilent 3xSLv spots). Additionally, probes were required to exhibit nonspecific expression changes with an IQR of at least 0.5. Finally, linear models were fitted using the R package limma, and reliable variance estimates were obtained using empirical Bayes adjusted t-statistics. False discovery rates were controlled using the Benjamini-Hochberg adjustment.

[0138] Validation by quantitative real-time PCR To validate the results obtained by next-generation sequencing and microarray screening, quantitative real-time PCR was used. cDNA was synthesized from 100 ng of total RNA using a high-capacity reverse transcription kit (Applied Biosystems) and random primers according to the manufacturer's instructions. Subsequent PCR assays were performed using 4 μl of diluted cDNA. Quantitative real-time PCR was performed on the reference and biomarker transcripts using custom-designed and predesigned TaqMan gene expression assays (Applied Biosystems) on an Applied Biosystems 7900HT real-time PCR system. The biomarker and reference genes tested are listed in Table 2, and the primers and probes used are listed in Table 3.

[0139] All samples were measured in triplicate, and the average of these measurements was used for subsequent calculations. All samples were checked for sufficient quality using the QC filters described above. Only samples that passed the QC filters were valid and suitable for inclusion in the final analysis. While the first QC filter checked for sufficient cell counts in the sample, the second QC filter confirmed sufficient prostate cell counts in the sample. This was further combined with an assessment of amplification quality to ensure that sample replicate measurements were reproducible and consistent, and that amplification met the minimum standards of the MIQE guidelines.

[0140] Validation in DRE urine samples: DRE urine sample collection and RNA isolation Urine samples were collected after digital rectal examination (DRE) of the prostate (DRE urine). This routine examination method allows for the collection of urine samples containing a consistent amount of prostate cells. DRE urine samples were centrifuged and washed twice with PBS. The resulting cell pellet was resuspended in 700 μl of Qiazol. Total RNA was isolated using the miRNeasy Mini Kit on a QIAcube (all Qiagen products) and then manually digested with DNase I. RNA concentration was determined using a Nanodrop 1000 (Peqlab). RNA integrity was verified using an Agilent Bioanalyzer 2100 (Agilent Technologies, Palo Alto, CA).

[0141] Quantitative real-time PCR screening of DRE urine samples cDNA was synthesized from 2 x 50 ng of total RNA using Superscript III reverse transcriptase (Applied Biosystems) and random primers according to the manufacturer's instructions. Subsequent PCR assays were performed using 4 μl of cDNA. Quantitative real-time PCR was performed on the Applied Biosystems 7900HT Real-Time PCR System using custom-designed and predesigned TaqMan Gene Expression Assays (Applied Biosystems) for housekeeping (PSA) and target transcripts. All samples were measured in duplicate, and the average of these measurements was used for subsequent calculations.

[0142] Genome-wide long RNA next-generation sequencing of DRE urine samples For genome-wide long-stranded RNA sequencing, total RNA from 11 DRE urine samples was concentrated by ethanol precipitation and resuspended in 10 μl of RNase-free water. rRNA removal was performed using 4 ng of total RNA using the Low Input Ribo-Zero rRNA Removal Kit (Epicentre, modified by Clontech) to obtain 10 μl of rRNA-depleted RNA. Sequencing libraries were prepared from 8 μl of rRNA-depleted RNA using the SMARTER stranded RNAseq Kit (Clontech). Dual-tagged cDNA was purified using the Agencourt AMPure XP System Kit (Beckman Coulter). PCR was performed for 18 cycles to incorporate index barcodes for sample multiplexing and to amplify the cDNA library. The quality and concentration of the amplified libraries were determined using a DNA high-sensitivity kit on an Agilent Bioanalyzer (Agilent Technologies). Samples were pooled, and cluster generation was performed using a 15 pmol / L pooled library and the TruSeq PE Cluster Kit v4 (Illumina Inc.) on an Illumina cBOT instrument according to the manufacturer's protocol. Sequencing was performed using HiSeq SBS v4 sequencing reagents (250 cycles) on an Illumina HiSeq2500 sequencing machine (Illumina, Inc.). Sequencing run details were as follows: paired-end sequencing strategy, 126 cycles for read 1, 7 cycles for the index sequence, and 126 cycles for read 2.

[0143] Statistical analysis of qRT-PCR results from DRE urine For relative quantification, gene expression changes in each sample were analyzed relative to the average expression level of genes within the same run. Data normalization was performed against the common reference genes listed in Table 4: EMC7 (SEQ ID NOs: 74-77), RAB7A (SEQ ID NOs: 78-95), and VCP (SEQ ID NOs: 96-131). Measurements between different runs were calibrated using calibrator samples derived from positive control pooled samples of patients diagnosed with benign prostatic hyperplasia (BPH) or prostate cancer.

[0144] After normalization and calibration, the relative expression levels of biomarkers were compared between tumor and control samples using the Mann-Whitney U test. Receiver operating characteristic (ROC) curves, representing the area under the curve (AUC), as a measure of the diagnostic power of each marker, were calculated using the pROC package. Variable selection was performed using variable importance with random forests and stepwise model selection (logistic regression: top-down with AIC). Final models were trained using logistic regression and ridge logistic regression. All statistical analyses were performed using R statistical software.

[0145] result The transcriptomes of 40 PCa tumor and 16 tumor-free samples obtained during RPE, as well as eight BPH prostate tissue samples serving as benign non-tumor controls, were analyzed using strand-specific paired-end long RNA next-generation sequencing (NGS). Approximately 150 frozen sections were prepared, with at least three segments per sample, for optimal data quality and analytical robustness. Upon pathological evaluation, only segments meeting a maximum tumor cell count of 60% were retained for subsequent analysis in tumor samples, and only segments meeting a minimum tumor cell count of 5% were retained for subsequent analysis in tumor-free samples. The transcriptome sequencing (RNA-seq) approach aimed to comprehensively identify and quantify RNAs expressed in normal and cancerous prostate tissues. All classes of coding and long non-coding transcripts were sequenced, regardless of polyadenylation status. High RNA input was used to ensure high library complexity. Furthermore, by sequencing an average of 200 million paired-end reads (2 × 100 nt) per library, high coverage enabled the assembly of novel, low-expression transcripts. This approach outperformed most comparable published studies that analyzed a larger number of samples. In total, we assembled approximately 3,000 novel transcripts that showed no exon overlap with transcripts annotated in Gencode v17. At a false discovery rate of 0.01, we observed 6,442 differentially expressed genes across all contrasts.

[0146] We successfully replicated most of the transcripts previously reported to be differentially expressed between prostate tumor and normal tissues. In addition, we identified several novel PCa-associated transcripts that could be used to develop diagnostic assays for PCa. We selected the most promising transcripts for validation in a test cohort of PCa tumor and BPH control samples by qRT-PCR.

[0147] Several of these novel biomarker candidates significantly exceed the specificity and sensitivity of PCA3, a biomarker already used for PCa diagnosis. In the sequencing cohort, PCA3 was clearly associated with PCa, but tended to be reduced in high-risk groups (Figure 2).

[0148] The experimental results demonstrate the high specificity and sensitivity of the novel biomarkers for PCa detection. As shown in Figure 3, which compares the differences in expression levels of various PCa biomarkers between PCa samples and BPH control samples, the results are presented for both the newly identified biomarker (ProstaCheck) and previously established biomarkers (SelectMDx and PCA3). Figure 3 clearly demonstrates that, in contrast to previously known biomarkers, more pronounced differences in expression levels can be observed between PCa and control samples for the newly identified biomarkers. Therefore, assays based on measuring these newly discovered biomarkers, alone or in combination (or in combination with other markers), in all materials that may contain prostate tumor cells or their parts (including vesicles such as exosomes, microvesicles, and other vesicles, as well as free or protein-bound RNA molecules derived from prostate tumor cells) can be designed and used for the diagnosis of PCa. As shown in Table 1, these new assays provide excellent diagnostic reliability.

[0149] The advantages of diagnostic assays based on these biomarkers include dramatically reduced false positive rates compared to current assays, and measuring the expression levels of these biomarkers in urine samples avoids the need for unnecessary invasive prostate biopsies.

[0150] Sequence Listing The sequences in the Sequence Listing are submitted separately. Table 6 below discloses the Ensembl transcript IDs or Ensembl gene IDs, represented by individual SEQ ID NOs, when no known transcript variants are available.

[0151]

Table 6

[0152] References Aubin SMJ et al. J Urol. 2010;184:1947-52. Yoav Benjamini and Yosef Hochberg. Controlling false discovery rate: a practical and powerful approach to multiple testing: Journal of the Royal Statistical Society 1995. Cucchiara, V., Cooperberg, M. R., Dall'Era, M., Lin, D. W., Montorsi, F., Schalken, J. A. u. Evans, C. P.: Genomic Markers in Prostate Cancer Decision Making. European urology (2(17) Duffy, Michael J. "Biomarkers for prostate cancer: prostate-specific antigen and beyond" Clinical Chemistry and Laboratory Medicine (CCLM), (2020) Gleason, D. F. (1977). "The Veteran's Administration Cooperative Urologic Research Group: histologic grading and clinical staging of prostatic carcinoma". In Tannenbaum, M. Urologic Pathology: The Prostate. Philadelphia: Lea and Febiger. pp. 171-198. Guescini M, Sisti D, Rocchi MB, Stocchi L, Stocchi V. A new real-time PCR method to overcome significant quantitative inaccuracy due to slight amplification inhibition. BMC Bioinformatics. 2008 Jul 30;9:326. doi: 10.1186 / 1471-2105-9-326 Haese Aet al. Eur Urol. 2008;54:1081-8. Irshad S et al. Sci Transl Med 2013;5:202ra122. Jansen FH et al. Eur Urol. 2009;55:563-74. Jansen FH, van Schaik RH, Kurstjens J, Horninger W, Klocker H, Bektic J, et al. Prostate-specific antigen (PSA) isoform p2PSA in combination with total PSA and free PSA improves diagnostic accuracy in prostate cancer detection. Eur Urol 2010;57:921-7. Lendinez-Cano G. Prospective study of diagnostic accuracy in the detection of high-grade prostate cancer in biopsy-naive patients with clinical suspicion of prostate cancer who underwent the Select MDx test. Prostate. 2021 Sep;81(12):857-865 Moyer VA et al. Ann Intern Med. 2012;157:120-134. Salagierski Met al. J Urol 2012;187: 795-801. Thompson IM et al. J Am Med Assoc. 2005;294:66-70. Van Gils MPet al. Clin Cancer Res. 2007;13:939-43. Vickers AJ, Cronin AM, Aus G, Pihl CG, Becker C, Pettersson K, et al. A panel of kallikrein markers can reduce unnecessary biopsy for prostate cancer: data from the European Randomized Study of Prostate Cancer Screening in Goteborg, Sweden. BMC Med 2008;6:19.

Claims

1. a) providing a sample from a patient suspected of having prostate cancer; b) analyzing the expression level of at least the nucleic acids set forth in SEQ ID NO: 2 and SEQ ID NO: 20 in said sample; Including, A method for ex vivo diagnosis of prostate cancer, wherein the sample is designated as prostate cancer positive when the expression level of the nucleic acid is above a threshold value.

2. a) providing a sample from a patient suspected of having prostate cancer; b) analyzing the expression level of at least the nucleic acids set forth in SEQ ID NO: 2, SEQ ID NO: 19-20, and SEQ ID NO: 35-37 in said sample; Including, A method for ex vivo diagnosis of prostate cancer, wherein the sample is designated as prostate cancer positive when the expression level of the nucleic acid is above a threshold value.

3. a) providing a sample from a patient suspected of having prostate cancer; b) analyzing the expression level of at least the nucleic acids set forth in SEQ ID NOs: 1-3, 20-34, and 38-50 in said sample; Including, A method for ex vivo diagnosis of prostate cancer, wherein the sample is designated as prostate cancer positive when the expression level of the nucleic acid is above a threshold value.

4. 4. The method of any one of claims 1 to 3, wherein the sample is selected from the group comprising prostate tissue, biopsy material, lymph nodes, urine, ejaculate, blood, serum, plasma, circulating tumor cells in the blood or lymph, any tissue suspected of containing metastasis, and any material that may contain prostate tumor cells or parts thereof, such as vesicles, including exosomes, microvesicles, and other vesicles, and free or protein-bound RNA molecules derived from prostate tumor cells.

5. 5. The method according to claim 1, wherein the analysis of the expression level is carried out by measuring the fluorescence of a labeled primer, a labeled probe, or a fluorescent detection agent, preferably wherein the analysis of the expression level is carried out by qRT-PCR.

6. The method of any one of claims 1 to 5, wherein the expression levels are normalized to a general reference gene and a prostate cell-specific reference gene.

7. 7. The method of any one of claims 1 to 6, wherein the general reference genes and the prostate cell-specific reference genes are used to assess the quality of the sample, and the expression levels of the general reference genes and the prostate cell-specific reference genes are used to provide at least one quality control filter, and only samples that pass the at least one quality control filter are included in the analysis.

8. A nucleic acid that hybridizes under stringent conditions with one of the nucleic acids of claims 1, 2, or 3.

9. 9. The nucleic acid of claim 8, which is about 10 to 100 nucleotides in length.

10. 10. The nucleic acid of claim 8 or 9, comprising a detectable label.

11. Use of a nucleic acid according to any one of claims 8 to 10 for the diagnosis of prostate cancer by a method according to any one of claims 1 to 7.

12. A nucleic acid having one of the sequences of SEQ ID NOs: 1-131, or the reverse complement thereof or any part thereof, or a nucleic acid sharing preferably at least 85%, 90%, 95%, or 99% sequence identity with one of the sequences of SEQ ID NOs: 1-131, or the reverse complement thereof or any part thereof, preferably selected from the group comprising SEQ ID NOs: 136-174.

13. Use of the nucleic acid according to claim 12 for the diagnosis of prostate cancer by the method according to any one of claims 1 to 7.

14. A kit for diagnosing prostate cancer, comprising the nucleic acid according to any one of claims 8 to 10 or the nucleic acid according to claim 12, and reagents for amplifying and / or quantifying and / or detecting the nucleic acid.

15. a) analyzing the expression level of a nucleic acid according to any one of claims 1, 2 or 3 in a sample from a patient; b) designating the sample as prostate cancer positive when the expression level of the nucleic acid is above a threshold; c) administering to said patient one or more prostate cancer therapeutic agents; 10. A method for the treatment and diagnosis of prostate cancer, comprising: