Kit and method for prognosis, diagnosis and treatment of prostate cancer

By measuring the expression levels of specific genes in urine, the MyProstateScore 2.0 model was developed, which solves the problem that existing technologies cannot accurately identify high-grade prostate cancer, enabling more precise diagnosis and treatment guidance and reducing unnecessary treatments.

CN120917152APending Publication Date: 2025-11-07THE RGT UNIV OF MICHIGAN
View PDF 27 Cites 0 Cited by

Patent Information

Application Number
CN202480009773.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-17
Filing Date
2024-01-29
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing methods for diagnosing prostate cancer cannot accurately identify high-grade prostate cancer, leading to unnecessary or overtreatment. Furthermore, there is a lack of reliable non-invasive screening tools. Existing tools such as PSA and MRI are costly, resource-intensive, and subjective in their interpretation, making them unsuitable for effective application at the population level.

Method used

The MyProstateScore 2.0 (MPS2) model was developed by measuring the expression levels of the genes TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6 in the urine of prostate cancer patients. This model is used to identify the likelihood of high-grade prostate cancer and provide corresponding treatment options.

Benefits of technology

It improves the diagnostic accuracy of high-grade prostate cancer, reduces unnecessary invasive biopsies, lowers the treatment burden on healthy subjects, and provides more accurate prognostic and treatment guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120917152A_ABST
    Figure CN120917152A_ABST
Patent Text Reader

Abstract

Provided herein are kits and methods useful for cancer diagnosis, prognosis, research, and therapy. In particular, provided herein are methods for diagnosis, prognosis and / or treatment of prostate cancer based on expression levels of cancer markers.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Statement as to Federally Sponsored Research or Development

[0002] This invention was made with government support under U01 CA214170, P50 CA186786, and R35 CA231996 awarded by the National Cancer Institute. The government has certain rights in the invention.

[0003] Cross Reference to Related Applications

[0004] This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 442,045, filed January 30, 2023, and U.S. Provisional Patent Application No. 63 / 446,596, filed February 17, 2023, each of which is incorporated by reference herein in its entirety.

[0005] Reference to Electronic Sequence Listing

[0006] The contents of the electronic sequence listing (LXDX-002-001WO.xml; size 50,183 bytes; and date created: January 28, 2024) are incorporated herein by reference in their entirety. TECHNICAL FIELD

[0007] Provided herein are kits and methods useful for cancer diagnosis, prognosis, research, and therapy. In particular, provided herein are methods of diagnosing, prognosticating, and / or treating prostate cancer based on the expression level of cancer markers. BACKGROUND

[0008] Prostate cancer is the third most common urological malignancy and can originate from the prostate parenchyma or the urinary collecting system. Prostate cell carcinoma originates from the prostate parenchyma, which is the most common malignant prostate tumor, with 64,000 cases of incidence and approximately 14,000 deaths per year in the United States. From the urinary collecting system, urothelial cell carcinoma is the most common malignancy, accounting for approximately 10-15% of all prostate tumors. The overall incidence of malignant prostate tumors is increasing and is currently the third most common form of urogenital cancer. With the use of advanced cross-sectional imaging, malignant and benign prostate tumors are increasingly diagnosed in an incidental manner. Due to the lack of accurate diagnosis of benign versus malignant tumor types, patients can receive unnecessary or over-treatment. Furthermore, there are currently no diagnostic tests performed by needle biopsy, urine, or blood that accurately characterize prostate tumors or identify patients at risk for prostate tumors. Due to the existence of multiple benign prostate tumor types and the fact that many small malignant prostate parenchymal tumors can only be observed and not definitively treated, the diagnosis and treatment of prostate tumors is complex.

[0009] Early detection and treatment of aggressive prostate cancer is critical to reduce its harm, but current diagnostic tests do not reliably identify clinically significant (e.g., classified as Grade Group [GG] > 2) prostate cancer. Serum prostate specific antigen (PSA) is poorly specific for cancer, its harm as a stand-alone diagnostic tool is well established, and several cancer-specific biomarkers have been proposed to augment PSA. These tools have shown incremental benefit, potentially avoiding 15-30% of biopsies performed for PSA, but at the cost of failing to diagnose 8-15% of GG > 2 prostate cancers. MRI is also used in multiple academic centers to play this role. In addition to growing evidence that some GG > 2 cancers are not discovered by MRI, MRI is costly, resource intensive, and subjective to interpret, making it less practical as a population-level diagnostic tool. Thus, there remains an urgent need for practical (affordable, repeatable, standardizable) non-invasive tests to reliably detect aggressive prostate cancer in a locally curable state.

[0010] While multiple molecular mechanisms contribute to the biology of aggressive prostate cancer, most patients harbor tumors that reflect only a limited number of these mechanisms. Currently, there are no accurate, user-friendly, and widely available tissue, blood, or urine level screening tools to enable the ideal clinical management of prostate tumors. SUMMARY

[0011] Provided herein are methods of treating prostate cancer, the methods comprising: a) determining the expression level of one or more genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6 in a sample from a subject diagnosed with prostate cancer; and b) administering a prostate cancer treatment to a subject identified as having an altered expression level of the one or more genes relative to a subject not having prostate cancer or a subject having low-grade prostate cancer.

[0012] Also provided are methods of characterizing, prognosticating, or recommending treatment of prostate cancer, the methods comprising: a) determining the expression level of one or more genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6 in a sample from a subject diagnosed with prostate cancer; and b) identifying the subject as having high-grade prostate cancer when the subject is identified as having an altered expression level of the gene relative to a subject not having prostate cancer or a subject having low-grade prostate cancer.

[0013] Also provided are methods for informing the survival outcome of prostate cancer, the methods comprising: (i) detecting the expression amount of at least three genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6, wherein the expression amount is present in the urine of a subject; (ii) determining a score based on the expression amount, wherein the score is correlated with or informs the likelihood of the subject having or developing prostate cancer of a grade group >2; and (iii) generating a report comprising the score.

[0014] Also provided are methods for identifying a subject having a high likelihood of having or developing prostate cancer of a grade group >2, the methods comprising detecting the expression amount of at least three genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6, wherein the expression amount is present in the urine of a subject, and indicates whether the subject has a high likelihood of having prostate cancer of a grade group >2 with a diagnostic accuracy (AUC) of >0.75.

[0015] Also provided is a method for identifying a likelihood of detecting a prostate cancer of a grade group >2 from a prostate biopsy of a subject, the method comprising detecting an expression amount of at least three genes selected from TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6, wherein the expression amount is present in urine of the subject and indicates the likelihood of detecting a prostate cancer of a grade group >2 from a prostate biopsy of the subject with a diagnostic accuracy (AUC) >0.75.

[0016] Also provided is a method for screening an expression amount of at least three genes, the method comprising: (a) reacting a urine sample from a human subject with a reagent for detecting an expression amount of at least three genes, wherein the at least three genes are selected from TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6; and (b) detecting the expression amount of the at least three genes, wherein the expression amount is present in the sample and the detecting comprises using an in vitro assay.

[0017] Also provided is a method for detecting an amount of mRNA expressed by at least three genes, the method comprising: (a) synthesizing cDNA from mRNA expressed by the at least three genes and present in a urine sample from a human subject, wherein the at least three genes are selected from TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6; b) amplifying the cDNA to provide amplified cDNA; and c) detecting the amplified cDNA, wherein the amplified cDNA indicates the amount of mRNA expressed by the at least three genes.

[0018] Further provided is a method for detecting an amount of mRNA expressed by at least three genes, the method comprising: a) isolating nucleic acids from a first composition comprising urine from a human subject to provide isolated nucleic acids; b) reacting the isolated nucleic acids with a second composition comprising reagents for detecting an amount of mRNA present in the first composition and expressed by the at least three genes, wherein the at least three genes are selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6; and c) detecting the amount of mRNA expressed by the at least three genes.

[0019] Further provided is a kit comprising: a container comprising a reagent composition for detecting an amount of expression of at least three genes; and instructions for detecting the amount of expression, wherein the amount of expression is present in urine of a subject and the at least three genes are selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6. BRIEF DESCRIPTION OF DRAWINGS

[0020] Figure 1A and 1B . Biomarker discovery Figure 1A ) and MyProstateScore 2.0 (MPS2) training cohort (from University of Michigan, UM) and validation cohort (from NCI-Early Detection Research Network, EDRN) sample inclusion Figure 1B ) flowchart.

[0021] Figures 2A-2C . Model development procedure. Redundant variables (highly correlated) Figure 2A ) were removed if Variance Inflation Factor (VIF) > 5. To select a robust gene detection panel, training of elastic net regression model was performed on 40 subsampled data Figure 2B ). Clinically meaningful prostate cancer calibration curves of MPS2 and MPS2+ in external validation cohort Figure 2C ).

[0022] Figures 3A-3EPerformance evaluation of the MPS2 model on the training cohort. Receiver operating characteristic (ROC) of the original MyProstateScore (MPS) (TMPRSS2-ERG + PCA3) Figure 3A ROC of the MPS2 gene panel Figure 3B ROC of the MPS2 plus clinical variables (MPS2c) Figure 3C ROC of the MPS2 plus clinical variables and prostate volume (MPS2cv) Figure 3D Calibration analysis of the model after calibration by adjusting the slope and intercept Figure 3E

[0023] Performance evaluation of the MPS2 model on the validation cohort. ROC and area under the curve (AUC) of the MPS2 model Figures 4A-4D Calibration curve of the calibrated risk probabilities Figure 4A Decision curve analysis showing the net benefit of the MPS2 model relative to "treat all" or "treat none" at different probability thresholds Figure 4B Interventions (biopsies) avoided at different probability thresholds Figure 4C Figure 4D

[0024] Figure 5 Association of selected genes with high-grade prostate cancer in the TCGA PRAD (Cancer Genome Atlas Prostate Adenocarcinoma) cohort. The 17 genes used in the final MPS2 model are: TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6.

[0025] Figure 6 Evaluation of feature selection methods. Boxplot showing the area under the curve (AUC) of each feature selection method obtained by repeated cross-validation. RFE, recursive feature elimination.

[0026] Figure 7 Calibration curve of the MPS2 model on the validation cohort without calibration. In the validation cohort, the prediction of risk is overestimated without correction for imbalanced classes.

[0027] Figure 8A and 8B MPS2 values calculated by biopsy pathology in the external validation cohort. Boxplot and dotplot showing the MPS2 of men with biopsy negative, biopsy with GG1 cancer, and biopsy with GG≥2 cancer in the external validation cohort Figure 8A ​​) and MPS2+( Figure 8B ) values. P-values for the paired comparison of GG≥2 cancers (biopsy-negative) and GG1 cancers were both <0.001 for both MPS2 models.

[0028] Figure 9 . Area under the curve of MPS2 and MPS2+ ROC curves for clinically significant prostate cancer in the external validation cohort. Receiver operating characteristic curves and area under the curve (AUC) for PSA (grey), PCPTrc (PCa Prevention Trial Risk Calculator, yellow), Prostate Health Index (PHI, purple), dmx2 (derived multiplex 2 gene model (HOXC6, DLX1), pink), dmx3 (derived multiplex 3 gene model (PCA3, ERG, SPDEF), maroon), MPS (MyProstateScore, orange), MPS2 (green), and MPS2+ (blue) in the external validation cohort.

[0029] Figure 10A and 10B . Decision curve analysis for clinically significant prostate cancer in the external validation cohort. Figure 10A Decision curve analysis (DCA) plots showing the net clinical benefit of pre-biopsy testing using PSA (grey), PCPTrc (yellow), PHI (purple), dmx2 (pink), dmx3 (maroon), MPS (orange), MPS2 (green), and MPS2+ (blue) compared to the baseline method of "all biopsies" (black) or "no biopsies" (dark green). Figure 10B DCA plots showing the net reduction in biopsies per 100 patients compared to the baseline method of biopsying all patients based on pre-biopsy testing using PSA (grey), PCPTrc (yellow), PHI (purple), dmx2 (pink), dmx3 (maroon), MPS (orange), MPS2 (green), and MPS2+ (blue), without missing any GG≥2 cancer diagnoses.

[0030] Figure 11 . Flow diagram of the NCI-EDRN external validation cohort. The external validation cohort consisting of men who underwent prostate biopsy in the National Cancer Institute (NCI) - Early Detection Research Network (NCI-EDRN) PCA3 trial is shown.

[0031] Definitions

[0032] To facilitate understanding of the present disclosure, the following terms and phrases are defined:

[0033] As used herein, the term "detect," "detecting," or "detection" can describe the general act of finding or discerning or a specific observation of a composition. Detecting a composition can include determining whether the composition is present. Detecting can include quantifying the composition. For example, detecting includes determining the expression level of a composition. The composition can include a nucleic acid molecule. For example, the composition can include at least a portion of a cancer marker disclosed herein. Optionally, or in addition, the composition can be a detectably labeled composition.

[0034] As used herein, the term "subject" refers to any organism that is screened using the diagnostic methods described herein. Such organisms preferably include, but are not limited to, mammals (e.g., murines, simians, equines, bovines, porcines, canines, felines, etc.), and most preferably include humans. In some embodiments, the subject is a mammal having a prostate. In some embodiments, the subject is a human having a prostate.

[0035] As used herein, the term "diagnose" refers to identifying a disease by signs and symptoms of the disease, or by genetic analysis, pathological analysis, histological analysis, etc.

[0036] As used herein, the phrase "characterizing a cancer of a subject" refers to identifying one or more properties of a cancer sample of a subject, including but not limited to the presence of benign, pre-cancerous, or cancerous tissue, the stage of the cancer, and the prognosis of the subject. The cancer can be characterized by identifying the expression of one or more cancer marker genes, including but not limited to the cancer markers disclosed herein.

[0037] As used herein, the phrase "cancer staging" refers to a qualitative or quantitative assessment of the degree of progression of a cancer. Criteria for determining cancer staging include, but are not limited to, the size of the tumor and the extent of metastasis (e.g., local or distant).

[0038] As used herein, the term "high likelihood" when used, for example, to refer to the likelihood of having or developing prostate cancer (e.g., prostate cancer with a Gleason score > 2) refers to an increased likelihood of developing prostate cancer with a Gleason score > 2, or an absolute likelihood of developing prostate cancer with a Gleason score > 2 that is high relative to a low risk subject. In some embodiments, the likelihood of developing prostate cancer with a Gleason score > 2 is determined to be high based on the expression level of one or more genes described herein. In some embodiments, "high likelihood" is an increased likelihood of 50%, 100%, 200%, 500% or more relative to a healthy subject or a subject without the changes in gene expression described herein. In some embodiments, high likelihood refers to an absolute likelihood of developing prostate cancer with a Gleason score > 2. In some embodiments, "high likelihood" refers to a likelihood of 50% or more that a prostate biopsy of a subject will detect prostate cancer with a Gleason score > 2 in the subject. In some embodiments, "high likelihood" refers to a likelihood of 60% or more that a prostate biopsy of a subject will detect prostate cancer with a Gleason score > 2 in the subject. In some embodiments, "high likelihood" refers to a likelihood of 70% or more that a prostate biopsy of a subject will detect prostate cancer with a Gleason score > 2 in the subject. In some embodiments, "high likelihood" refers to a likelihood of 80% or more that a prostate biopsy of a subject will detect prostate cancer with a Gleason score > 2 in the subject. In some embodiments, "high likelihood" refers to a likelihood of 90% or more that a prostate biopsy of a subject will detect prostate cancer with a Gleason score > 2 in the subject. In some embodiments, "high likelihood" refers to a likelihood of 95% or more that a prostate biopsy of a subject will detect prostate cancer with a Gleason score > 2 in the subject. In some embodiments, "high likelihood" refers to a likelihood of 96% or more that a prostate biopsy of a subject will detect prostate cancer with a Gleason score > 2 in the subject. In some embodiments, "high likelihood" refers to a likelihood of 97% or more that a prostate biopsy of a subject will detect prostate cancer with a Gleason score > 2 in the subject. In some embodiments, "high likelihood" refers to a likelihood of 98% or more that a prostate biopsy of a subject will detect prostate cancer with a Gleason score > 2 in the subject. In some embodiments, "high likelihood" refers to a likelihood of 99% or more that a prostate biopsy of a subject will detect prostate cancer with a Gleason score > 2 in the subject. In some embodiments, "high likelihood" refers to a likelihood of 100% that a prostate biopsy of a subject will detect prostate cancer with a Gleason score > 2 in the subject.

[0039] As used herein, the term "low likelihood" when used, for example, in reference to the likelihood of having or developing prostate cancer (e.g., prostate cancer with a Gleason grouping of >2) means a reduced likelihood of developing prostate cancer (e.g., prostate cancer with a Gleason grouping of >2) or an absolute low likelihood of developing prostate cancer (e.g., prostate cancer with a Gleason grouping of >2) relative to an average risk subject. In some embodiments, a low likelihood of developing prostate cancer with a Gleason grouping of >2 is determined based on the expression level of one or more genes described herein. In some embodiments, "low likelihood" is a 50%, 100%, 200%, 500% or more reduction in likelihood relative to a healthy subject or a subject without altered gene expression described herein. In some embodiments, low likelihood refers to an absolute likelihood of developing prostate cancer with a Gleason grouping of >2. In some embodiments, "low likelihood" means that there is less than a 50% chance that a prostate biopsy of the subject will detect prostate cancer with a Gleason grouping of >2 in the subject. In some embodiments, "low likelihood" means that there is a 40% or less chance that a prostate biopsy of the subject will detect prostate cancer with a Gleason grouping of >2 in the subject. In some embodiments, "low likelihood" means that there is a 30% or less chance that a prostate biopsy of the subject will detect prostate cancer with a Gleason grouping of >2 in the subject. In some embodiments, "low likelihood" means that there is a 20% or less chance that a prostate biopsy of the subject will detect prostate cancer with a Gleason grouping of >2 in the subject. In some embodiments, "low likelihood" means that there is a 10% or less chance that a prostate biopsy of the subject will detect prostate cancer with a Gleason grouping of >2 in the subject. In some embodiments, "low likelihood" means that there is a 5% or less chance that a prostate biopsy of the subject will detect prostate cancer with a Gleason grouping of >2 in the subject. In some embodiments, "low likelihood" means that there is a 4% or less chance that a prostate biopsy of the subject will detect prostate cancer with a Gleason grouping of >2 in the subject. In some embodiments, "low likelihood" means that there is a 3% or less chance that a prostate biopsy of the subject will detect prostate cancer with a Gleason grouping of >2 in the subject. In some embodiments, "low likelihood" means that there is a 2% or less chance that a prostate biopsy of the subject will detect prostate cancer with a Gleason grouping of >2 in the subject. In some embodiments, "low likelihood" means that there is a 1% or less chance that a prostate biopsy of the subject will detect prostate cancer with a Gleason grouping of >2 in the subject. In some embodiments, "low likelihood" means that there is a 0% chance that a prostate biopsy of the subject will detect prostate cancer with a Gleason grouping of >2 in the subject.

[0040] As used herein, the term "nucleic acid molecule" refers to any molecule comprising nucleic acid, including but not limited to DNA or RNA. A nucleic acid molecule can comprise one or more nucleotides. The term can include nucleotide polymers in which the nucleotides and the linkages between them include non-naturally occurring synthetic analogs, such as, for example and without limitation, phosphorothioates, phosphoramidates, methyl phosphonates, chiral-methyl phosphonates, 2-O-methyl ribonucleotides, peptide-nucleic acids (PNAs), and the like. The term further encompasses sequences that can include any known base analogs of DNA and RNA including, but not limited to, 4-acetylcytosine, 8-hydroxy-N6-methyladenosine, aziridinylcytosine, pseudoisocytosine, 5-(carboxyhydroxylmethyl) uracil, 5- fluorouracil, 5-bromouracil, 5-carboxymethylaminomethyl-2-thiouracil, 5- carboxymethylaminomethyluracil, dihydrouracil, inosine, N6-isopentenyladenine, 1- methyladenine, 1-methylpseudouracil, 1-methylguanine, 1-methylinosine, 2,2-dimethyl guanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5- methylcytosine, N6-methyladenine, 7-methylguanine, 5-methylaminomethyluracil, 5- methoxyaminomethyl-2-thiouracil, beta-D-mannosylqueosine, 5'-methoxycarbonylmethyl uracil, 5-methoxyuracil, 2-methylthio-N6-isopentenyladenine, uracil-5- oxyacetic acid methylester, uracil-5-oxyacetic acid, oxybutoxosine, pseudouracil, queosine, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, N-uracil-5- oxyacetic acid methylester, uracil-5-oxyacetic acid, pseudouracil, queosine, 2- thiocytosine, and 2,6-diaminopurine. It should be understood that when a nucleotide sequence is represented by a DNA sequence (i.e., A, T, G, C), this sequence is also indicative of an RNA sequence (i.e., A, U, G, C) wherein "U" replaces "T".

[0041] The term "gene" refers to a nucleic acid (e.g., DNA) sequence that comprises coding sequences necessary for the production of a polypeptide, precursor, or RNA (e.g., rRNA, tRNA). The polypeptide can be encoded by a full-length coding sequence or by any portion of the coding sequence so long as the desired activity or functional properties (e.g., enzymatic activity, ligand binding, signal transduction, immunogenicity, etc.) of the full-length or fragment are retained. The term also encompasses the coding region of a structural gene, and plus and minus terminal sequences adjacent to the coding region on both ends, up to about 1 kb or more of each end such that the "gene" corresponds to the entire mRNA encoded by a full-length DNA. The sequences that are present on the 5' and 3' ends of a coding region and are not translated into a polypeptide are referred to as 5' and 3' non- translated or untranslated sequences. The term "gene" encompasses both cDNA and genomic forms of a gene. A genomic form or clone of a gene contains the coding region interrupted with non-coding sequences termed "introns" or "intervening sequences" or "intervening regions." The introns are removed by enzymatic cleavage from the transcript of the gene (called "pre-mRNA") during the production of the mature mRNA through a process called "splicing." The mature mRNA is translated into a polypeptide in the cytoplasm of the cell. The mRNA of some genes does not require "splicing" and thus contains no introns.

[0042] As used herein, the term "oligonucleotide" refers to a single-stranded polynucleotide chain of short length. Oligonucleotides are typically less than 200 nucleotide residues in length (e.g., between 15 and 100), although, as used herein, the term is also intended to encompass longer polynucleotide chains. Oligonucleotides are often named by their length. For example, a 24-residue oligonucleotide is referred to as a "24-mer." Oligonucleotides can form secondary and tertiary structures through self-hybridization or through hybridization with other polynucleotides. Such structures can include, but are not limited to, duplexes, hairpins, cruciforms, bent structures, and triplexes.

[0043] As used herein, the term "label" refers to any atom or molecule that can be used to provide a detectable (preferably quantifiable) effect and that can be attached to a nucleic acid or protein. Labels include, but are not limited to: dyes; radioactive labels such as 32P; a binding moiety such as biotin; a hapten such as digoxigenin; a luminescent, phosphorescent, or fluorescent moiety; and a fluorescent dye alone or in combination with a moiety that can inhibit or shift the emission spectrum by fluorescence resonance energy transfer (FRET). The label can provide a signal that is detectable by fluorescence, radioactivity, colorimetry, gravimetry, X-ray diffraction or absorption, magnetism, enzymatic activity, and the like. The label can be a charged moiety (e.g., positive or negative charge), or alternatively, can be electrically neutral. The label can comprise or consist of a nucleic acid or protein sequence, so long as the sequence comprising the label is detectable. In some embodiments, nucleic acids can be detected directly (e.g., sequence read directly) without a label.

[0044] As used herein, the term“sample” includes specimens or cultures obtained from any source, as well as biological and environmental samples. Biological samples can be obtained from animals (including humans) and encompass fluids (e.g., blood, urine), solids, tissues, and gases. Biological samples can include urine, urine supernatant and urine cell pellet, as well as blood products, such as plasma, serum, and the like. However, such examples should not be understood to limit the types of samples applicable to the present disclosure.

[0045] As used herein,“high-grade prostate cancer” means prostate cancer that is graded ≥ 2. In some embodiments, high-grade prostate cancer is GG≥3 prostate cancer.

[0046] As used herein,“low-grade prostate cancer” means prostate cancer that is graded < 2.

[0047] As used herein, a“score” is the likelihood that a subject’s prostate biopsy will detect prostate cancer that is graded ≥ 2 in the subject, i.e., the likelihood that a subject’s prostate biopsy will be positive for prostate cancer. The score is based on the expression level or amount of expression of one or more genes described herein present in a subject’s sample. In some embodiments, the score is a numerical value in the range of 0% to 100%. In some embodiments, the numerical value is expressed as a decimal in the range of 0.0 to 100.0. In some embodiments, the score is a qualitative readout of“low risk” or“elevated risk.”

[0048] As used herein, the term“altered,” e.g., in the context of“altered expression level of one or more genes,” means that the expression level of the gene is different (e.g., increased or decreased) from, e.g., a subject who does not have prostate cancer or a subject who has low-grade prostate cancer.

[0049] As used herein, the term "variant" (e.g., genetic variant) refers to sequence changes that do not affect gene identity. Such sequence changes are readily understood by one of skill in the art. In some embodiments, variants include mutations, substitutions, and / or deletions. In some embodiments, variants include polymorphisms. In some embodiments, variants include splice variants.

[0050] As used herein, unless otherwise indicated or inferred, the term "about" means ± 10% of the nominal value. When the term "about" is used in reference to a number, the disclosure also includes the specific number itself, unless explicitly stated otherwise. DETAILED DESCRIPTION

[0051] Provided herein are kits and methods useful for the diagnosis, prognosis, research, and therapy of cancer. In particular, provided herein are methods of diagnosing, prognosticating, and / or treating prostate cancer based on the expression level of cancer markers.

[0052] The present disclosure is based, at least in part, on the discovery that the likelihood of a subject having a prostate cancer with a grade group > 2 can be determined based on the amount of expression of one or more genes described herein.

[0053] Described herein are methods and kits that include one or more of 17 markers useful for the prognosis, diagnosis, or treatment of prostate cancer. Importantly, the detection of PSA (prostate specific antigen) is a routine method of prognosing and / or diagnosing prostate cancer, but it is not a necessary step of the methods described herein. The high proportion of men without cancer who undergo invasive and unnecessary biopsies as a result of an elevated PSA identified during PSA screening, and the frequent over-diagnosis of low-grade, indolent cancers (grade group 1 (GG1)), are well-documented. The kits and methods of the present disclosure provide a more accurate prognosis or diagnosis of prostate cancer and help to identify those subjects who can benefit from early, aggressive therapeutic intervention, while avoiding invasive procedures, such as biopsies, for those subjects with indolent disease. Thus, the present methods provide a new and unconventional set of prostate cancer biomarkers, and in particular, high-grade (e.g., GG > 2) prostate cancer biomarkers, that are not dependent on PSA.

[0054] Accordingly, provided herein are methods and kits useful for prognosing, diagnosing, or treating a subject having prostate cancer (in some embodiments, a prostate cancer that is stratified >2). For example, in some embodiments, provided herein are methods of treating prostate cancer, the method comprising: a) determining the expression level of one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or all 17) genes selected from, for example, TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6 in a sample from a subject being prognosed or diagnosed for prostate cancer; and b) administering a prostate cancer treatment to a subject identified as having an altered expression level of the one or more genes relative to a subject not having prostate cancer or a subject having a low-grade prostate cancer. In some embodiments, the subject has a high-grade prostate cancer.

[0055] Further embodiments provide methods of characterizing, prognosing, or recommending a prostate cancer treatment, the method comprising: a) determining the expression level of one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or all 17) genes selected from, for example, TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6 in a sample from a subject being prognosed or diagnosed for prostate cancer; and b) identifying the subject as having a high-grade prostate cancer when the subject is identified as having an altered expression level of the genes relative to a subject not having prostate cancer or a subject having a low-grade prostate cancer. In some embodiments, the method further comprises administering a prostate cancer treatment to the subject. In some embodiments, the method further comprises recommending to the subject or a healthcare provider of the subject that the subject undergo a prostate biopsy. In some embodiments, the prostate biopsy indicates that the subject has a prostate cancer that is stratified >2. In some embodiments, the prostate biopsy indicates that the subject does not have a prostate cancer that is stratified >2.

[0056] In some embodiments, the method further comprises performing a prostate biopsy on the subject. In some embodiments, the method further comprises recommending to the subject or a healthcare provider of the subject that the subject undergo a prostate biopsy. In some embodiments, the prostate biopsy indicates that the subject has a prostate cancer that is stratified >2. In some embodiments, the prostate biopsy indicates that the subject does not have a prostate cancer that is stratified >2.

[0057] In some embodiments, the method further comprises advising the subject or a healthcare provider of the subject that the subject does not undergo a prostate biopsy.

[0058] In some embodiments, the method does not comprise performing a prostate biopsy on the subject.

[0059] The methods described herein can be used to identify subjects having high-grade prostate cancer for treatment and allow those subjects identified as not having high-grade prostate cancer to avoid a biopsy or treatment, and thus the associated side effects thereof. The methods as provided herein are used to reduce the number of unnecessary prostate biopsies, sparing healthy subjects from an expensive and invasive procedure.

[0060] I. Methods of determining marker expression

[0061] As described herein, embodiments of the disclosure provide methods for prognosis, diagnosis, or treatment that utilize detecting an amount of expression or expression level of one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or all 17) genes selected from, for example, TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6. Exemplary, non-limiting methods are described herein.

[0062] Detecting a gene

[0063] In some embodiments, the expression level or amount of expression of one or more genes is determined. In some embodiments, the expression level or amount is the level or amount of mRNA or protein expressed by the gene.

[0064] In some embodiments, the one or more genes are selected from TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6.

[0065] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of a TMPRSS2-ERG gene. The TMPRSS2-ERG gene fusion overexpresses the transcription factor ERG, which is present in both early and late stage prostate cancer. Multiple variants of the TMPRSS2-ERG fusion have been identified, with the most common variant including exon 1 of TMPRSS2 and exons 4-11 of ERG. In some embodiments, the TMPRSS2-ERG gene fusion includes a fusion of the nucleotide sequences of Ensembl Gene Identifiers ENSG00000184012 and ENSG00000157554. In some embodiments, the TMPRSS2-ERG gene fusion includes the nucleotide sequence of SEQ ID NO: 1, or a variant thereof.

[0066] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of a SCHLAP1 gene. SCHLAP1 is a long non-coding RNA that is overexpressed in a subset of prostate cancers. SCHLAP1 antagonizes the genome-wide localization and regulatory functions of the SWI / SNF chromatin-modifying complex. In some embodiments, the SCHLAP1 gene includes the nucleotide sequence provided by the HUGO Gene Nomenclature Committee (HGNC). In some embodiments, the HGNC identifier for SCHLAP1 is 48603. In some embodiments, the SCHLAP1 gene is located at chromosome location 2q31.3. In some embodiments, the SCHLAP1 gene includes the nucleotide sequence of Ensembl Gene Identifier ENSG00000281131. In some embodiments, the SCHLAP1 gene includes the nucleotide sequence of SEQ ID NO: 2, or a variant thereof.

[0067] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of the OR51E2 gene. OR51E2 is an olfactory receptor (OR) that represents the largest family of G protein-coupled receptors (GPCRs) in the human genome. Activation of human ORs can affect cell proliferation. Specifically, OR51E2 has been identified to be involved in regulating cell growth, migration, and invasion of melanocytes, melanoma cells, and prostate cancer cells. In some embodiments, the OR51E2 gene comprises the nucleotide sequence provided by HGNC. In some embodiments, the HGNC identifier for OR51E2 is 15195. In some embodiments, the OR51E2 gene is located at chromosome location 11p15.4. In some embodiments, the OR51E2 gene comprises the nucleotide sequence of Ensembl gene identifier ENSG00000167332. In some embodiments, the OR51E2 gene comprises the nucleotide sequence of SEQ ID NO: 3, or a variant thereof.

[0068] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of the APOC1 gene. APOC1 is a minor apolipoprotein and is a component of triglyceride-rich lipoproteins and high-density lipoproteins. APOC1 is involved in various biological processes and is associated with the progression of various diseases, such as diabetic nephropathy, Alzheimer’s disease, and glomerulosclerosis. Recent studies have shown that APOC1 can be associated with the development of cancer, including breast cancer, pancreatic cancer, lung cancer, and prostate cancer. In some embodiments, the APOC1 gene comprises the nucleotide sequence provided by HGNC. In some embodiments, the HGNC identifier for APOC1 is 607. In some embodiments, the APOC1 gene is located at chromosome location 19q13.32. In some embodiments, the APOC1 gene comprises the nucleotide sequence of Ensembl gene identifier ENSG00000130208. In some embodiments, the APOC1 gene comprises the nucleotide sequence of SEQ ID NO: 4, or a variant thereof.

[0069] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of PCAT14 gene. PCAT14 is a long non-coding RNA that exhibits both cancer and lineage specificity. PCAT14 is transcriptionally regulated by the androgen receptor (AR), and endogenous PCAT14 overexpression suppresses cell invasion. In some embodiments, the PCAT14 gene comprises the nucleotide sequence provided by HGNC. In some embodiments, the HGNC identifier for PCAT14 is 48977. In some embodiments, the PCAT14 gene is located at chromosome location 22ql l.23. In some embodiments, the PCAT14 gene comprises the nucleotide sequence of Ensembl gene identifier ENSG00000280623. In some embodiments, the PCAT14 gene comprises the nucleotide sequence of SEQ ID NO: 5, or a variant thereof.

[0070] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of CAMKK2 gene. CAMKK2 is a direct target of AR, and its regulation can vary by disease stage. CAMKK2 has been identified as a driver of prostate cancer progression. In some embodiments, the CAMKK2 gene comprises the nucleotide sequence provided by HGNC. In some embodiments, the HGNC identifier for CAMKK2 is 1470. In some embodiments, the CAMKK2 gene is located at chromosome location 12q24.31. In some embodiments, the CAMKK2 gene comprises the nucleotide sequence of Ensembl gene ENSG00000110931. In some embodiments, the CAMKK2 gene comprises the nucleotide sequence of SEQ ID NO: 6, or a variant thereof.

[0071] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of PCA3 gene. PCA3 is a non-coding gene associated with prostate cancer. In some embodiments, the PCA3 gene comprises the nucleotide sequence provided by HGNC. In some embodiments, the HGNC identifier for PCA3 is 8637. In some embodiments, the PCA3 gene is located at chromosome location 9q21.2. In some embodiments, the PCA3 gene comprises the nucleotide sequence of Ensembl gene identifier ENSG00000225937. In some embodiments, the PCA3 gene comprises the nucleotide sequence of SEQ ID NO: 7, or a variant thereof.

[0072] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of NKAIN1 gene. NKAIN1 is a sodium / potassium transporting ATPase. In some embodiments, the NKAIN1 gene comprises the nucleotide sequence provided by HGNC. In some embodiments, the HGNC identifier for NKAIN1 is 25743. In some embodiments, the NKAIN1 gene is located at chromosome location lp35.2. In some embodiments, the NKAIN1 gene comprises the nucleotide sequence of Ensembl Gene Identifier ENSG00000084628. In some embodiments, the NKAIN1 gene comprises the nucleotide sequence of SEQ ID NO: 8, or a variant thereof.

[0073] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of B3GNT6 gene. B3GNT6 is a member of the O-GlcNAc transferase (OGT) family and is responsible for the production of core 3 structures of O-glycans. In some embodiments, the B3GNT6 gene comprises the nucleotide sequence provided by HGNC. In some embodiments, the HGNC identifier for B3GNT6 is 24141. In some embodiments, the B3GNT6 gene is located at chromosome location 11q13.5. In some embodiments, the B3GNT6 gene comprises the nucleotide sequence of Ensembl Gene Identifier ENSG00000198488. In some embodiments, the B3GNT6 gene comprises the nucleotide sequence of SEQ ID NO: 9, or a variant thereof.

[0074] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of TFF3 gene. TFF3 is a trefoil factor, which is a secreted peptide produced by normal intestinal mucosa. Trefoil factor family members are overexpressed in a variety of cancers and are associated with tumor invasion, anti-apoptosis, and metastasis. In some embodiments, the TFF3 gene comprises the nucleotide sequence provided by HGNC. In some embodiments, the HGNC identifier for TFF3 is 11757. In some embodiments, the TFF3 gene is located at chromosome location 21q22.3. In some embodiments, the TFF3 gene comprises the nucleotide sequence of Ensembl Gene Identifier ENSG00000160180. In some embodiments, the TFF3 gene comprises the nucleotide sequence of SEQ ID NO: 10, or a variant thereof.

[0075] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of the SPON2 gene. SPON2 belongs to the F-spondin family of secreted extracellular matrix proteins and is dysregulated in some tumors, including prostate cancer. In some embodiments, the SPON2 gene comprises the nucleotide sequence provided by HGNC. In some embodiments, the HGNC identifier for SPON2 is 11253. In some embodiments, the SPON2 gene is located at chromosome location 4pl6.3. In some embodiments, the SPON2 gene comprises the nucleotide sequence of Ensembl Gene Identifier ENSG00000159674. In some embodiments, the SPON2 gene comprises the nucleotide sequence of SEQ ID NO: 11, or a variant thereof.

[0076] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of the PCGEM1 gene. PCGEM1 is a long non-coding RNA that is a prostate-specific transcript. In some embodiments, the PCGEM1 gene comprises the nucleotide sequence provided by HGNC. In some embodiments, the HGNC identifier for PCGEM1 is 30145. In some embodiments, the PCGEM1 gene is located at chromosome location 2q32.3. In some embodiments, the PCGEM1 gene comprises the nucleotide sequence of Ensembl Gene Identifier ENSG00000227418. In some embodiments, the PCGEM1 gene comprises the nucleotide sequence of SEQ ID NO: 12, or a variant thereof.

[0077] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of the TRGV9 gene. TRGV9 is encoded by the TRG locus, which is rearranged to encode TCR gamma chains comprising 14 variable genes, only 6 of which are functional, including TRGV9. In some embodiments, the TRGV9 gene comprises the nucleotide sequence provided by HGNC. In some embodiments, the HGNC identifier for TRGV9 is 12295. In some embodiments, the TRGV9 gene is located at chromosome location 7pl4.1. In some embodiments, the TRGV9 gene comprises the nucleotide sequence of Ensembl Gene Identifier ENSG00000211695. In some embodiments, the TRGV9 gene comprises the nucleotide sequence of SEQ ID NO: 13, or a variant thereof.

[0078] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of the TMSB15A gene. TMSB15A is an isoform of human thymosin beta 15, which is an actin-binding protein. TMSB15A is expressed in normal human prostate and prostate cancer tissues. In some embodiments, the TMSB15A gene comprises the nucleotide sequence provided by HGNC. In some embodiments, the HGNC identifier for TMSB15A is 30744. In some embodiments, the TMSB15A gene is located at chromosome location Xq22.1. In some embodiments, the TMSB15A gene comprises the nucleotide sequence of Ensembl Gene Identifier ENSG00000158164. In some embodiments, the TMSB15A gene comprises the nucleotide sequence of SEQ ID NO: 14, or a variant thereof.

[0079] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of the ERG gene. ERG is a transcriptional regulator that is overexpressed in prostate cancer. In some embodiments, the ERG gene comprises the nucleotide sequence provided by HGNC. In some embodiments, the HGNC identifier for ERG is 3446. In some embodiments, the ERG gene is located at chromosome location 21q22.2. In some embodiments, the ERG gene comprises the nucleotide sequence of Ensembl Gene Identifier ENSG00000157554. In some embodiments, the ERG gene comprises the nucleotide sequence of SEQ ID NO: 15, or a variant thereof.

[0080] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of the KLK4 gene. KLK4 is a member of the kallikrein (KLK) family of highly conserved serine proteases that play key roles in a variety of physiological and pathological processes. KLKs are secreted proteins with extracellular substrates and functions. KLK4 is overexpressed in prostate cancer. In some embodiments, the KLK4 gene comprises the nucleotide sequence provided by HGNC. In some embodiments, the HGNC identifier for KLK4 is 6365. In some embodiments, the KLK4 gene is located at chromosome location 19q13.41. In some embodiments, the KLK4 gene comprises the nucleotide sequence of Ensembl Gene Identifier ENSG00000167749. In some embodiments, the KLK4 gene comprises the nucleotide sequence of SEQ ID NO: 16, or a variant thereof.

[0081] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of HOXC6 gene. HOXC6 is a homeobox (HOX) gene. HOX genes are involved in organ development and homeostasis and have been shown to be involved in the development of normal prostate and prostate cancer. HOXC6 is overexpressed in prostate cancer. In some embodiments, the HOXC6 gene comprises the nucleotide sequence provided by HGNC. In some embodiments, the HGNC identifier for HOXC6 is 5128. In some embodiments, the HOXC6 gene is located at chromosome location 12q13.13. In some embodiments, the HOXC6 gene comprises the nucleotide sequence of Ensembl Gene Identifier ENSG00000197757. In some embodiments, the HOXC6 gene comprises the nucleotide sequence of SEQ ID NO: 17, or a variant thereof.

[0082] Table A provides exemplary nucleotide sequences for TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6.

[0083] Table A. Exemplary nucleotide sequences for genes of the present disclosure.

[0084]

[0085]

[0086]

[0087]

[0088]

[0089]

[0090]

[0091]

[0092]

[0093]

[0094]

[0095]

[0096]

[0097]

[0098]

[0099]

[0100]

[0101]

[0102]

[0103]

[0104]

[0105]

[0106]

[0107]

[0108] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or 17 genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6.

[0109] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of at least three genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6. In some embodiments, the at least three genes are TMPRSS2-ERG, PCA3, and PCAT14.

[0110] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of at least four genes selected from TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6. In some embodiments, the at least four genes are TMPRSS2-ERG, PCA3, PCAT14, and OR51E2. In some embodiments, the at least four genes are TMPRSS2-ERG, PCA3, PCAT14, and TRGV9. In some embodiments, the at least four genes are TMPRSS2-ERG, PCA3, PCAT14, and ERG. In some embodiments, the at least four genes are TMPRSS2-ERG, PCA3, PCAT14, and TFF3. In some embodiments, the at least four genes are TMPRSS2-ERG, PCA3, PCAT14, and SCHLAP1. In some embodiments, the at least four genes are TMPRSS2-ERG, PCA3, PCAT14, and HOXC6. In some embodiments, the at least four genes are TMPRSS2-ERG, PCA3, PCAT14, and SPON2. In some embodiments, the at least four genes are TMPRSS2-ERG, PCA3, PCAT14, and TMSB15A. In some embodiments, the at least four genes are TMPRSS2-ERG, PCA3, PCAT14, and APOC1. In some embodiments, the at least four genes are TMPRSS2-ERG, PCA3, PCAT14, and B3GNT6. In some embodiments, the at least four genes are TMPRSS2-ERG, PCA3, PCAT14, and KLK4. In some embodiments, the at least four genes are TMPRSS2-ERG, PCA3, PCAT14, and CAMKK2. In some embodiments, the at least four genes are TMPRSS2-ERG, PCA3, PCAT14, and NKAIN1. In some embodiments, the at least four genes are TMPRSS2-ERG, PCA3, PCAT14, and PCGEM1.

[0111] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of at least five genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, OR51E2, and TRGV9. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, and OR51E2. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, TFF3, and TRGV9. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, OR51E2, and TFF3. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, SCHLAP1, and TRGV9. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, OR51E2, and SCHLAP1. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, HOXC6, and TRGV9. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, OR51E2, and HOXC6. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, SPON2, and TRGV9. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, and SPON2. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, TMPSB15A, and TRGV9. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, and TMSB15A. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, APOC1, and TRGV9. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, and APOC1. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, B3GNT6, and TRGV9. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, and B3GNT6.In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, KLK4, and TRGV9. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, and KLK4. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, CAMKK2, and TRGV9. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, and CAMKK2. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, NKAIN1, and TRGV9. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, and NKAIN1. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, PCGEM1, and TRGV9. In some embodiments, the at least five genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, and PCGEM1.

[0112] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of at least six genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, and OR51E2. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, NKAIN1, and PCGEM1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, TRGV9, and PCGEM1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, CAMKK2, and PCGEM1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, and NKAIN1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, OR51E2, and TCGEM1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, and TFF3. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, OR51E2, and SCHLAP1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TFF3, and HOXC6. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, SCHLAP1, and SPON2. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, HOXC6, and TMSB15A. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, SPON2, and APOC1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TMSB15A, and APOC1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TMSB15A, and B3GNT6.In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, APOC1, and KLK4. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, B3GNT6, and CAMKK2. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, KLK4, and NKAIN1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, CAMKK2, and PCGEM1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, KLK4, and PCGEM1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, and CAMKK12. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, OR51E2, and NKAIN1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TFF3, and PCGEM1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, and SCHLAP1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, OR51E2, and HOXC6. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TFF3, and SPON2. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, SCHLAP1, and TMSB15A. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, HOXC6, and APOC1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, SPON2, and B2GNT6. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TMSB15A, and KLK4. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, APOC1, and CAMKK2. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, B3GNT6, and NKAIN1.In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, B3GNT6, and PCGEMl. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, and KLK4. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, OR51E2, and CAMKK2. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TFF3, and NKAIN1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, SCHLAP1, and PCGEMl. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, and HOXC6. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, OR51E2, and SPON2. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TFF3, and TMSB15A. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, SCHLAP1, and APOC1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, HOXC6, and B3GNT6. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, SPON2, and KLK4. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TMSB15A, and CAMKK2. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, APOC1, and NKAIN1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, APOC1, and PCGEMl. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, and B3GNT6. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, OR51E2, and KLK4. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TFF3, and CAMKK2.In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, SCHLAP1, and NKAIN1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, HOXC6, and PCGEM1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, HOXC6, and SPON2. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, OR51E2, and TMSB15A. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TMSB15A, and PCGEM1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, and APOC1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, OR51E2, and B3GNT6. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TFF3, and KLK4. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, SCHLAP1, and CAMKK2. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, HOXC6, and NKAIN1. In some embodiments, the at least six genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, SPON2, and PCGEM1.

[0113] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of at least seven genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, and TFF3. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, NKAIN1, and PCGEM1. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, and PCGEM1. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, CAMKK2, and PCGEM1. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, and NKAIN1. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, KLK4, and PCGEM1. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, and CAMKK2. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, SCHLAP1, and HOXC6. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, B3GNT6, and PCGEM1. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, and KLK4. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, HOXC6, and SPON2. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, APOC1, and PCGEM1. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, and B3GNT6.In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, TMSB15A, and PCGEMl. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, and APOCl. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, KLK4, NKAIN1, and PCGEMl. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, B3GNT6, NKAIN1, and PCGEMl. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, OR51E2, TFF3, and NKAIN1. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, OR51E2, HOXC6, and SPON2. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, OR51E2, TFF3, and CAMKK2. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, OR51E2, SPON2, and TMPSB15A. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, OR51E2, TFF3, and KLK4. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TFF3, SCHLAP1, and PCGEMl. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TFF3, SPON2, and TMSB15A. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TFF3, SCHLAP1, and NKAIN1. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TFF3, TMSB15A, and APOCl. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TFF3, SCHLAP1, and CAMKK2. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, SCHLAP1, TMSB15A, and APOCl.In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, SCHLAP1, HOXC6, and PCGEM1. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, SCHLAP1, APOC1, and B3GNT6. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, SCHLAP1, HOXC6, and NKAIN1. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, HOXC6, APOC1, and B3GNT6. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, HOXC6, B3GNT6, and KLK4. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, HOXC6, SPON2, and PCGEM1. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, SPON2, B3GNT6, and KLK4. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, SPON2, KLK4, and CAMKK2. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, SPON2, NKAIN1, and PCGEM1. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TMSB15A, KLK4, and CAMKK2. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TMSB15A, CAMKK2, and NKAIN1. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TMSB15A, NKAIN1, and PCGEM1. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, APOC1, NKAIN1, and CAMKK2. In some embodiments, the at least seven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, APOC1, NKAIN1, and PCGEM1.

[0114] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of at least eight genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, and SCHLAP1. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, NKAIN1, and PCGEM1. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, and PCGEM1. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, CAMKK2, and PCGEM1. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, and NKAIN1. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, SCHLAP1, and HOXC6. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, KLK4, and PCGEM1. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, and CAMKK2. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, HOXC6, and SPON2. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, B3GNT6, and PCGEM1. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, and KLK4. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, SPON2, and PCGEM1.In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, APOC1, and PCGEM1. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, and B2GNT6. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, KLK4, NKAIN1, and PCGEM1. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, TFF3, SCHLAP1, and PCGEM1. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, TFF3, HOXC6, and SPON2. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, TFF3, SCHLAP1, and NKAIN1. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, TFF3, SPON2, and TMSB15A. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, TFF3, SCHLAP1, and CAMKK2. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, TFF3, SCHLAP1, and KLK4. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, SCHLAP1, SPON2, and TMSB15A. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, SCHLAP1, HOXC6, and PCGEM1. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, SCHLAP1, TMSB15A, and APOC1. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, SCHLAP1, HOXC6, and NKAIN1.In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, SCHLAP1, HOXC6, and CAMKK2. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, HOXC6, TMSB15A, and APOC1. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, HOXC6, APOC1, and B3GNT6. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, HOXC6, SPON2, and NKAIN1. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, SPON2, APOC1, and B3GNT6. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, SPON2, B3GNT6, and KLK4. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, SPON2, TMSB15A, and PCGEM1. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, TMSB15A, B3GNT6, and KLK4. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, TMSB15A, KLK4, and CAMKK2. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, TMSB15A, NKAIN1, and PCGEM1. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, APOC1, KLK4, and CAMKK2. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, APOC1, CAMKK2, and NKAIN1. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, APOC1, NKAIN1, and PCGEM1. In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, B3GNT6, CAMKK2, and NKAIN1.In some embodiments, the at least eight genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, B3GNT6, NKAIN1, and PCGEM1.

[0115] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of at least nine genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, and HOXC6. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, and pCGEM1. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, and NKAIN1. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, and SPON2. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, and CAMKK2. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, and HOXC6. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, and TMSB15A. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, and KLK4. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, and B3GNT6. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, NKAIN1, and PCGEM1. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, NKAIN1, and KLK4.In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, NKAIN1, and B3GNT6. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, NKAIN1, and SPON2. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, NKAIN1, and TMSB15A. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, CAMKK2, and PCGEM1. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, CAMKK2, and B3GNT6. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, CAMKK2, and APOC1. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, CAMKK2, and HOXC6. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, CAMKK2, and SPON2. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, HOXC6, and PCGEM1. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, HOXC6, and TMSB15A. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, HOXC6, and APOC1. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, HOXC6, and KLK4. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SPON2, and APOC1.In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SPON2, and PCGEMl. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SPON2, and B3GNT6. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, TMSB15A, and B3GNT6. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, TMSB15A, and KLK4. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, TMSB15A, and PCGEMl. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, TMSB15A, and APOCl. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, APOCl, and KLK4. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, APOCl, and PCGEMl. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, KLK4, and PCGEMl. In some embodiments, the at least nine genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, B3GNT6, and PCGEMl.

[0116] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of at least ten genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, and SPON2. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, and PCGEM1. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, NKAIN1, and pCGEM1. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, CAMKK2, and PCGEM1. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, and NKAIN1. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, SPON2, and TMSB15A. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, KLK4, and PCGEM1. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, and CAMKK2. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, TMBS15A, and APOC1. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, B3GNT6, and PCGEM1.In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, and KLK4. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, APOC1, and B3GNT6. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, CAMKK2, TMSB15A, and KLK4. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, CAMKK2, APOC1, and NKAIN1. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, CAMKK2, HOXC6, and SPON2. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, KLK4, NKAIN1, and PCGEM1. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, KLK4, SPON2, and B3GNT6. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, HOXC6, SPON2, and PCGEM1. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, HOXC6, TMSB15A, and APOC1. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, HOXC6, SPON2, and NKAIN1. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, HOXC6, APOC1, and B3GNT6.In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SPON2, APOC1, and B3GNT6. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SPON2, TMSB15A, and PCGEM1. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SPON2, TMSB15A, and NKAIN1. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, B3GNT6, NKAIN1, and PCGEM1. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, B3GNT6, APOC1, and TMSB15A. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, APOC1, NKAIN1, and PCGEM1. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, APOC1, TMSB15A, and PCGEM1. In some embodiments, the at least ten genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, TMSB15A, NKAIN1, and PCGEM1.

[0117] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of at least eleven genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6. In some embodiments, the at least eleven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, and TMSB15A. In some embodiments, the at least eleven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, and PCGEM1. In some embodiments, the at least eleven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, NKAIN1, and PCGEM1. In some embodiments, the at least eleven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, B3GNT6, and PCGEM1. In some embodiments, the at least eleven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, and KLK4. In some embodiments, the at least eleven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least eleven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, CAMKK2, KLK4, and NKAIN1. In some embodiments, the at least eleven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, CAMKK2, B3GNT6, and KLK4. In some embodiments, the at least eleven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, CAMKK2, SPON2, and TMSB15A.In some embodiments, the at least eleven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, APOC1, B3GNT6, and KLK4. In some embodiments, the at least eleven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, APOC1, NKAIN1, and PCGEM1. In some embodiments, the at least eleven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, APOC1, TMSB15A, and B3GNT6. In some embodiments, the at least eleven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, APOC1, SPON2, and TMSB15A. In some embodiments, the at least eleven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, APOC1, TMSB15A, and NKAIN1. In some embodiments, the at least eleven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, APOC1, B3GNT6, and PCGEM1. In some embodiments, the at least eleven genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, TMSB15A, KLK4, and PCGEM1.

[0118] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of at least twelve genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6. In some embodiments, the at least twelve genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least twelve genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, CAMKK2, B3GNT6, and KLK4. In some embodiments, the at least twelve genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, CAMKK2, KLK$, and NKAIN1. In some embodiments, the at least twelve genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, CAMKK2, SPON2, and PCGEM1. In some embodiments, the at least twelve genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, CAMKK2, SPON2, and TMSB15A. In some embodiments, the at least twelve genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, NKAIN1, and PCGEM1. In some embodiments, the at least twelve genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, and PCGEM1. In some embodiments, the at least twelve genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, and APOC1.In some embodiments, the at least twelve genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, and NKAIN1. In some embodiments, the at least twelve genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, APOC1, and B3GNT6. In some embodiments, the at least twelve genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, KLK4, and PCGEM1. In some embodiments, the at least twelve genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, TMSB15A, APOC1, and B3GNT6. In some embodiments, the at least twelve genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, TMSB15A, APOC1, and PCGEM1. In some embodiments, the at least twelve genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, TMSB15A, B3GNT6, and KLK4. In some embodiments, the at least twelve genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, TMSB15A, APOC1, and NKAIN1. In some embodiments, the at least twelve genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, APOC1, B3GNT6, and KLK4. In some embodiments, the at least twelve genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, KLK4, NKAIN1, and PCGEM1. In some embodiments, the at least twelve genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, B3GNT6, NKAIN1, and PCGEM1.In some embodiments, the at least twelve genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, APOC1, B3GNT6, and PCGEM1.

[0119] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of at least thirteen genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6. In some embodiments, the at least thirteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, and B3GNT6. In some embodiments, the at least thirteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, NKAIN1, and PCGEM1. In some embodiments, the at least thirteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, KLK4, and PCGEM1. In some embodiments, the at least thirteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, and CAMKK2. In some embodiments, the at least thirteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, KLK4, and CAMKK2. In some embodiments, the at least thirteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, TMSB15A, APOC1, B3GNT6, and KLK4. In some embodiments, the at least thirteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, TMSB15A, APOC1, B3GNT6, and NKAIN1. In some embodiments, the at least thirteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, TMSB15A, APOC1, CAMKK2, and NKAIN1.In some embodiments, the at least thirteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, APOC1, B3GNT6, KLK4, and CAMKK2. In some embodiments, the at least thirteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, APOC1, B3GNT6, KLK4, and PGEM1. In some embodiments, the at least thirteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, APOC1, B3GNT6, NKAIN1, and PCGEM1. In some embodiments, the at least thirteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, B3GNT6, KLK4, NKAIN1, and CAMKK2. In some embodiments, the at least thirteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, KLK4, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least thirteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least thirteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, APOC1, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least thirteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, B3GNT6, NKAIN1, and PCGEM1. In some embodiments, the at least thirteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, B3GNT6, KLK4, and CAMKK2.In some embodiments, the at least thirteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, B3GNT6, KLK4, and PCGEM1. In some embodiments, the at least thirteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, APOC1, KLK4, and NKAIN1. In some embodiments, the at least thirteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, TMSB15A, B3GNT6, CAMKK2, and PCGEM1.

[0120] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of at least fourteen genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, B3GNT6, KLK4, CAMKK2, and NKAIN1. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, SPON2, KLK4, CAMKK2, and NKAIN1. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, SPON2, TMSB15A, CAMKK2, and NKAIN1. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, SPON2, TMSB15A, APOC1, and NKAIN1. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, SPON2, TMSB15A, APOC1, and B3GNT6. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, NKAIN1, CAMKK2, KLK4, and APOC1. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, NKAIN1, CAMKK2, B3GNT6, and SPON2.In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, NKAIN1, KLK4, TMSB15A, and SPON2. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, CAMKK2, APOC1, TMSB15A, and SPON2. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, NKAIN1, CAMKK2, B3GNT6, and APOC1. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, NKAIN1, KLK4, B3GNT6, and SPON2. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, CAMKK2, KLK4, TMSB15A, and SPON1. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, NKAIN1, B3GNT6, APOC1, and TMSB15A. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, KLK4, B3GNT6, APOC1, and SPON2. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, NKAIN1, KLK4, B3GNT6, and APOC1. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, CAMKK2, KLK4, B3GNT6, and SPON2.In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, NKAIN1, CAMKK2, APOC1, and TMSB15A. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, NKIAN1, B3GNT6, APOC1, and SPON2. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, KLK4, B3GNT6, TMSB15A, and SPON2. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, CAMKK2, KLK4, B3GNT6, and APOC1. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, NKAIN1, CAMKK2, KLK4, and TMSB15A. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, NKAIN1, CAMKK2, APOC1, and SPON2. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, NKAIN1, B3GNT6, TMSB15A, and SPON2. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, NKAIN1, CAMKK2, B3GNT6, and TMSB15A. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, NKAIN1, KLK4, APOC1, and SPON2.In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, CAMKK2, B3GNT6, TMSB15A, and SPON2. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, CAMKK2, B3GNT6, APOC1, and TMSB15A. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, CAMKK2, KLK4, B3GNT6, and TMSB15A. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, PCGEM1, KLK4, APOC1, TMSB15A, and SPON2. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, NKAIN1, B3GNT6, APOC1, TMSB15A, and SPON2. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, NKAIN1, CAMKK2, APOC1, TMSB15A, and SPON2. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, NKAIN1, CAMKK2, KLK4, TMSB15A, and SPON2. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, NKAIN1, CAMKK2, KLK4, B3GNT6, and SPON2. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, NKAIN1, CAMKK2, KLK4, APOC1, and SPON2.In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, NKAIN1, KLK4, B3GNT6, APOC1, and SPON2. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, NKAIN1, KLK4, APOC1, TMSB15A, and SPON2. In some embodiments, the at least fourteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, KLK4, B3GNT6, APOC1, TMSB15A, and SPON2.

[0121] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of at least fifteen genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, B3GNT6, KLK4, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, KLK4, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, B3GNT6, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, B3GNT6, KLK4, and PCEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, B3GNT6, KLK4, and CAMKK2. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, B3GNT6, CAMKK2, NKAIN1, and PCGEM1.In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, KLK4, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, B3GNT6, CAMKK2, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, B3GNT6, KLK4, and NKAIN1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, APOC1, KLK4, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, APOC1, B3GNT6, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, B3GNT6, KLK4, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, KLK4, CAMKK2, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, B3GNT6, CAMKK2, and NKAIN1.In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, B3GNT6, KLK4, CAMKK2, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, KLK4, CAMKK2, and NKAIN1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, B3GNT6, KLK4, CAMKK2, and NKAIN1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, TMSB15A, APOC1, B3GNT6, KLK4, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, APOC1, B3GNT6, KLK4, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, SPON2, APOC1, B3GNT6, KLK4, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, TMSB15A, B3GNT6, KLK4, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, SPON2, TMSB15A, APOC1, B3GNT6, KLK4, CAMKK2, and NKAIN1.In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, SPON2, TMSB15A, B3GNT6, KLK4, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, SPON2, TMSB15A, APOC1, B3GNT6, KLK4, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, SPON2, TMSB15A, APOC1, KLK4, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, TMSB15A, APOC1, KLK4, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, TMSB15A, APOC1, B3GNT6, KLK4, CAMKK2, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, TMSB15A, APOC1, B3GNT6, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, APOC1, B3GNT6, KLK4, CAMKK2, and NKAIN1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, APOC1, B3GNT6, KLK4, NKAIN1, and PCGEM1.In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, APOC1, B3GNT6, KLK4, CAMKK2, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, TMSB15A, APOC1, B3GNT6, KLK4, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, TMSB15A, APOC1, B3GNT6, KLK4, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, TMSB15A, APOC1, B3GNT6, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, TMSB15A, APOC1, B3GNT6, KLK$, CAMKK2, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, TMSB15A, APOC1, KLK4, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, TMSB15A, APOC1, B3GNT6, KLK4, CAMKK2, and NKAIN1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, APOC1, B3GNT6, KLK4, CAMKK2, and PCGEM1.In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, APOC1, B3GNT6, KLK4, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, B3GNT6, KLK4, CAMKK2, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, APOC1, B3GNT6, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, B3GNT6, KLK4, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, KLK4, CAMKK2, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, B3GNT6, CAMKK2, and NKAIN1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, SPON2, TMSB15A, APOC1, B3GNT6, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, SPON2, TMSB15A, APOC1, KLK4, CAMKK2, NKAIN1, and PCGEM1.In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, SPON2, TMSB15A, APOC1, B3GNT6, KLK4, NKAIN1, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, SPON2, TMSB15A, APOC1, B3GNT6, KLK4, CAMKK2, and PCGEM1. In some embodiments, the at least fifteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, SPON2, TMSB15A, B3GNT6, KLK4, CAMKK2, NKAIN1, and PCGEM1.

[0122] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of at least sixteen genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6. In some embodiments, the at least sixteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, B3GNT6, KLK4, CAMKK2, and NKAIN1. In some embodiments, the at least sixteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, B3GNT6, KLK4, CAMKK2, and PCGEM1. In some embodiments, the at least sixteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, B3GNT6, KLK4, NKAIN1, and PCGEM1. In some embodiments, the at least sixteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, B3GNT6, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least sixteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, KLK4, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least sixteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, B3NGT6, KLK4, CAMKK2, NKAIN1, and PCGEM1.In some embodiments, the at least sixteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, APOC1, B3GNT6, KLK4, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least sixteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, TMSB15A, APOC1, B3GNT6, KLK4, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least sixteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, SPON2, TMSB15A, APOC1, B3GNT6, KLK4, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least sixteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, TFF3, HOXC6, SPON2, TMSB15A, APOC1, B3GNT6, KLK4, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least sixteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, OR51E2, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, B3GNT6, KLK4, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least sixteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, TRGV9, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, B3GNT6, KLK4, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least sixteen genes are TMPRSS2-ERG, PCA3, PCAT14, ERG, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, B3GNT6, KLK4, CAMKK2, NKAIN1, and PCGEM1. In some embodiments, the at least sixteen genes are TMPRSS2-ERG, PCA3, PCAT14, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, B3GNT6, KLK4, CAMKK2, NKAIN1, and PCGEM1.In some embodiments, the at least sixteen genes are TMPRSS2-ERG, PCA3, ERG, TRGV9, OR51E2, TFF3, SCHLAP1, HOXC6, SPON2, TMSB15A, APOC1, B3GNT6, KLK4, CAMKK2, NKAIN1, and PCGEM1.

[0123] In some embodiments, the methods and kits described herein can be used to detect the expression level or amount of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6.

[0124] In some embodiments, the expression level or expression amount of at least three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, sixteen, or seventeen genes described herein is higher in a subject having or at risk of developing a prostate cancer with a Gleason score > 2 relative to a subject having a prostate cancer with a Gleason score < 2 or no prostate cancer. In some embodiments, the expression level or expression amount of at least one of TMPRSS2-ERG, SCHLAP1, OR51E2, PCAT14, PCA3, B3GNT6, TFF3, SPON2, TRGV9, TMSB15A, ERG, KLK4, and HOXC6 is higher in a subject at risk of having a prostate cancer with a Gleason score > 2 relative to a subject having or at risk of developing a prostate cancer with a Gleason score < 2 or a subject with no prostate cancer. In some embodiments, the expression level or expression amount of each of TMPRSS2-ERG, SCHLAP1, OR51E2, PCAT14, PCA3, B3GNT6, TFF3, SPON2, TRGV9, TMSB15A, ERG, KLK4, and HOXC6 is higher in a subject at risk of having a prostate cancer with a Gleason score > 2 relative to a subject having or at risk of developing a prostate cancer with a Gleason score < 2 or a subject with no prostate cancer. In some embodiments, the expression level or expression amount of TMPRSS2-ERG, SCHLAP1, OR51E2, PCAT14, PCA3, B3GNT6, TFF3, SPON2, TRGV9, TMSB15A, ERG, KLK4, and HOXC6 is higher in a subject at risk of having a prostate cancer with a Gleason score > 2 relative to a subject having or at risk of developing a prostate cancer with a Gleason score < 2 or a subject with no prostate cancer.

[0125] In some embodiments, the expression level or amount of at least three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, sixteen, or seventeen genes described herein is lower in a subject having or at risk of developing prostate cancer with a Gleason score > 2 relative to a subject having or at risk of developing prostate cancer with a Gleason score < 2 or a subject not having prostate cancer. In some embodiments, the expression level or amount of at least one of APOCl, CAMKK2, NKAIN1, and PCGEM1 is lower in a subject at risk of developing prostate cancer with a Gleason score > 2 relative to a subject having or at risk of developing prostate cancer with a Gleason score < 2 or a subject not having prostate cancer. In some embodiments, the expression level or amount of each of APOCl, CAMKK2, NKAIN1, and PCGEM1 is lower in a subject at risk of developing prostate cancer with a Gleason score > 2 relative to a subject having or at risk of developing prostate cancer with a Gleason score < 2 or a subject not having prostate cancer. In some embodiments, the total expression level or amount of APOCl, CAMKK2, NKAIN1, and PCGEM1 is lower in a subject at risk of developing prostate cancer with a Gleason score > 2 relative to a subject having or at risk of developing prostate cancer with a Gleason score < 2 or a subject not having prostate cancer.

[0126] Methods for detecting gene expression

[0127] The expression level or amount of one or more genes of the present disclosure can be detected using any of a variety of nucleic acid techniques, including but not limited to: nucleic acid sequencing; nucleic acid hybridization; and nucleic acid amplification.

[0128] In some embodiments, nucleic acid sequencing methods are utilized (e.g., for detecting amplified nucleic acids). In some embodiments, the technologies provided herein find use with second generation (i.e., Next-Gen), third generation (i.e., Next-Next-Gen), or fourth generation (i.e., N3-Gen) sequencing technologies, including but not limited to, pyrosequencing, ligation sequencing, single molecule sequencing, sequencing by synthesis (SBS), semiconductor sequencing, massively parallel clonal, massively parallel single molecule SBS, massively parallel single molecule real-time, massively parallel single molecule real-time nanopore technology, and the like. A review of some such technologies is provided by Morozova and Marra in Genomics, 92:255 (2008), which is incorporated by reference herein in its entirety. Those skilled in the art will recognize that, because RNA is less stable in cells and experimentally more susceptible to nuclease attack, RNA can be reverse transcribed into DNA prior to sequencing.

[0129] Many DNA sequencing techniques are suitable for use in the present methods, including fluorescence-based sequencing methods (see, e.g., Birren et al., Genome Analysis: Analyzing DNA, 1, Cold Spring Harbor, N.Y.; which is incorporated by reference herein in its entirety). In some embodiments, sequencing is an automated sequencing technique as understood in the art. In some embodiments, sequencing is parallel sequencing of partitioned amplicons (PCT Publication No. WO 2006084132 to Kevin McKernan et al., which is incorporated by reference herein in its entirety). In some embodiments, sequencing is DNA sequencing by parallel oligonucleotide extension (see, e.g., U.S. Patent No. 5,750,341 to Macevicz et al. and U.S. Patent No. 6,306,597 to Macevicz et al., both of which are incorporated by reference herein in their entirety). Additional examples of sequencing techniques include the Church polony technique (Mitra et al., 2003, Analytical Biochemistry 320, 55-65; Shendure et al., 2005 Science 309, 1728-1732; U.S. Patent No. 6,432,360; U.S. Patent No. 6,485,944; U.S. Patent No. 6,511,803; which are incorporated by reference herein in their entirety), the 454 picotiter pyrosequencing technique (Margulies et al., 2005 Nature 437, 376-380; US20050130173; which are incorporated by reference herein in their entirety), the Solexa single base addition technique (Bennett et al., 2005, Pharmacogenomics, 6, 373-382; U.S. Patent No. 6,787,308; U.S. Patent No. 6,833,246; which are incorporated by reference herein in their entirety), the Lynx massively parallel signature sequencing technique (Brenner et al. (2000). Nat. Biotechnol. 18:630-634; U.S. Patent No. 5,695,934; U.S. Patent No. 5,714,330; which are incorporated by reference herein in their entirety), and the Adessi PCR colony technique (Adessi et al. (2000). Nucleic Acid Res. 28, E87; WO 00018957; which are incorporated by reference herein in their entirety).

[0130] Exemplary, non-limiting examples of nucleic acid hybridization techniques include, but are not limited to, in situ hybridization (ISH), microarray, and Southern or Northern blotting.

[0131] In situ hybridization (ISH) is a type of hybridization that uses labeled complementary DNA or RNA strands as probes to locate specific DNA or RNA sequences in a portion of a tissue or section of a tissue (in situ), or, if the tissue is small enough, the entire tissue (whole tissue ISH). DNA ISH can be used to determine chromosome structure. RNA ISH can be used to measure and locate mRNA and other transcripts (e.g., cancer markers) in a tissue section or whole tissue. Sample cells and tissues can be treated to fix the target transcripts in place and increase the accessibility of the probes. The probes hybridize to the target sequences at high temperature and then excess probes are washed away. The probes labeled with radioactivity, fluorescence, or antigen are located and quantified in the tissue using autoradiography, fluorescence microscopy, or immunohistochemistry, respectively. ISH can also use two or more probes labeled with radioactivity or other non-radioactive labels to detect two or more transcripts simultaneously.

[0132] One or more cancer markers in the methods described herein can be detected by performing one or more hybridization reactions. The one or more hybridization reactions can comprise one or more hybridization arrays, hybridization reactions, hybridization chain reactions, isothermal hybridization reactions, nucleic acid hybridization reactions, or combinations thereof. The one or more hybridization arrays can comprise hybridization array genotyping, hybridization array ratio sensing, DNA hybridization arrays, macroarrays, microarrays, high-density oligonucleotide arrays, genomic hybridization arrays, comparative hybridization arrays, or combinations thereof.

[0133] Different kinds of biological assays are referred to as microarrays, including but not limited to: DNA microarrays (e.g., cDNA microarrays and oligonucleotide microarrays); protein microarrays; tissue microarrays; transfection or cell microarrays; chemical compound microarrays; and antibody microarrays. DNA microarrays, often referred to as gene chips, DNA chips, or biochips, are collections of tiny spots of DNA attached to a solid surface (e.g., a glass, plastic, or silicon chip) that form an array for the purpose of simultaneously analyzing expression profiles or monitoring expression levels of thousands of genes. The immobilized DNA segments are called probes, and thousands of probes can be used in a single DNA microarray. Microarrays can be used to identify disease genes or transcripts (e.g., cancer markers) by comparing gene expression in disease cells and normal cells. Microarrays can be fabricated using a variety of techniques, including but not limited to: printing on glass slides using fine-tipped pens; photolithography using a pre-fabricated mask; photolithography using a dynamic micro-mirror device; inkjet printing; or, electrochemical reactions on a microelectrode array.

[0134] The methods disclosed herein can include performing one or more amplification reactions. Nucleic acids (e.g., cancer markers) can be amplified prior to or concurrently with detection. Performing one or more amplification reactions can include one or more PCR-based amplifications, non-PCR-based amplifications, or combinations thereof. Exemplary, non-limiting examples of nucleic acid amplification techniques include, but are not limited to, polymerase chain reaction (PCR), reverse transcription polymerase chain reaction (RT-PCR), nested PCR, linear amplification, multiple displacement amplification (MDA), real-time SDA, rolling circle amplification, circle-to-circle amplification, transcription-mediated amplification (TMA), ligase chain reaction (LCR), strand displacement amplification (SDA), and nucleic acid sequence-based amplification (NASBA). Those of ordinary skill in the art will recognize that certain amplification techniques (e.g., PCR) require reverse transcription of RNA into DNA prior to amplification (e.g., RT-PCR), while other amplification techniques directly amplify RNA (e.g., TMA and NASBA).

[0135] Polymerase chain reaction (U.S. Patent Nos. 4,683,195, 4,683,202, 4,800,159, and 4,965,188, each incorporated herein by reference in its entirety) generally referred to as PCR, uses multiple cycles of denaturation, annealing of primer pairs to complementary strands, and primer extension to exponentially increase the copy number of a target nucleic acid sequence. In a variant referred to as RT-PCR, reverse transcriptase (RT) is used to prepare complementary DNA (cDNA) from mRNA, and the cDNA is then amplified by PCR to produce multiple copies of DNA. For other various permutations of PCR, see, e.g., U.S. Patent Nos. 4,683,195, 4,683,202, and 4,800,159; Mullis et al., Meth. Enzymol. 155:335 (1987); and Murakawa et al., DNA 7:287 (1988), each incorporated herein by reference in its entirety.

[0136] Transcription-mediated amplification (U.S. Patent Nos. 5,480,784 and 5,399,491, each incorporated herein by reference in its entirety), generally referred to as TMA, autocatalytically synthesizes multiple copies of a target nucleic acid sequence under conditions of essentially constant temperature, ionic strength, and pH, wherein multiple RNA copies of the target sequence autocatalytically generate additional copies. See, e.g., U.S. Patent Nos. 5,399,491 and 5,824,518, each incorporated herein by reference in its entirety. In a variant described in U.S. Publication No. 20060046265 (incorporated herein by reference in its entirety), TMA is optionally combined with blocking, terminating, and other modifying moieties to improve the sensitivity and accuracy of the TMA process.

[0137] The ligase chain reaction (Weiss, R., Science 254: 1292 (1991), which is incorporated by reference herein in its entirety), commonly referred to as LCR, uses two sets of complementary DNA oligonucleotides that hybridize to adjacent regions of the target nucleic acid. The DNA oligonucleotides are covalently linked by DNA ligase in repeated cycles of heat denaturation, hybridization, and ligation to produce a detectable double-stranded ligated oligonucleotide product.

[0138] Strand displacement amplification (Walker, G. et al., Proc. Natl. Acad. Sci. USA 89: 392-396 (1992); U.S. Patent Nos. 5,270,184 and 5,455,166, each of which is incorporated by reference herein in its entirety), commonly referred to as SDA, uses the following cycle to obtain exponential amplification of a product: annealing of a primer sequence pair to complementary strands of a target sequence, primer extension in the presence of dNTPs to produce double-stranded hemi-thiophosphate primer extension products, endonuclease-mediated cleavage of hemi-modified restriction endonuclease recognition sites, and polymerase-mediated primer extension from the cleaved 3' end to displace the existing strand and produce a strand for the next round of primer annealing, cleavage, and strand displacement. Thermophilic SDA (tSDA) uses thermophilic endonucleases and polymerases at higher temperatures in essentially the same method (European Patent No. 0684 315).

[0139] Other amplification methods include, for example: nucleic acid sequence-based amplification (U.S. Patent No. 5,130,238, incorporated by reference in its entirety), commonly referred to as NASBA; one that uses an RNA replicase to amplify the probe molecule itself (Lizardi et al., BioTechnol. 6: 1197 (1988), incorporated by reference in its entirety), commonly referred to as Qβ replicase; a transcription-based amplification method (Kwoh et al., Proc. Natl. Acad. Sci. USA 86: 1173 (1989)); and self-sustained sequence replication (Guatelli et al., Proc. Natl. Acad. Sci. USA 87: 1874 (1990), each of which is incorporated by reference in its entirety). For further discussion of known amplification methods, see Persing, David H., “In Vitro Nucleic Acid Amplification Techniques” in Diagnostic Medical Microbiology: Principles and Applications (Persing et al., Eds.), pp. 51-87 (American Society for Microbiology, Washington, DC (1993)).

[0140] In some embodiments, the amplification method is a real-time quantitative PCR method (QPCR). Real-time polymerase chain reaction (real-time PCR or qPCR) is a molecular biology laboratory technique based on the polymerase chain reaction (PCR). It monitors the amplification of a targeted DNA molecule during (i.e., in real-time) PCR, rather than at the end of amplification as with traditional PCR. Real-time PCR can be used for quantification (quantitative real-time PCR) and semi-quantification (i.e., above / below a certain amount of DNA molecules) (semi-quantitative real-time PCR). Two commonly used methods for detecting PCR products in real-time PCR are (1) a non-specific fluorescent dye that intercalates into any double-stranded DNA and (2) a sequence-specific DNA probe consisting of an oligonucleotide labeled with a fluorescent reporter that can only be detected after the probe has hybridized to its complementary sequence.

[0141] Exemplary, non-limiting examples of immunoassays include, but are not limited to: immunoprecipitation; Western blot; ELISA; immunohistochemistry; immunocytochemistry; flow cytometry; and immuno-PCR. Detectably-labeled polyclonal or monoclonal antibodies using various techniques known to one of ordinary skill in the art (e.g., colorimetric, fluorescent, chemiluminescent, or radioactive) are suitable for use in immunoassays.

[0142] Immunoprecipitation is a technique that uses antigen-specific antibodies to precipitate antigens from solution. This process can identify protein complexes present in a cell extract by targeting proteins believed to be present in a complex. These complexes are precipitated from solution by insoluble antibody-binding proteins originally isolated from bacteria, such as protein A and protein G. Antibodies can also be coupled to agarose beads, which can be easily separated from solution. After washing, the precipitate can be analyzed using mass spectrometry, Western blot, or any other method for identifying components in the complex.

[0143] Western blot or immunoblotting is a method of detecting proteins in a given sample of tissue homogenate or extract. It uses gel electrophoresis to separate denatured proteins by mass. The proteins are then transferred from the gel and onto a membrane, typically polyvinylidene difluoride or nitrocellulose, where probing with antibodies specific to the protein of interest occurs. Thus, researchers can examine the amount of protein in a given sample and compare levels between groups.

[0144] ELISA is an acronym for Enzyme-Linked ImmunoSorbent Assay, which is a biochemical technique that detects the presence of antibodies or antigens in a sample. It utilizes at least two antibodies, one of which has specificity for an antigen, and the other of which is coupled to an enzyme. The second antibody causes a colorimetric or fluorescent substrate to produce a signal. Variants of ELISA include sandwich ELISA, competitive ELISA, and ELISPOT. Since an ELISA can be performed to evaluate the presence of an antigen or the presence of an antibody in a sample, it is a useful tool for both determining serum antibody concentrations as well as detecting the presence of an antigen.

[0145] Immunohistochemistry and immunocytochemistry refer to the process of localizing proteins in tissue sections or cells, respectively, via the principle of antigen binding to their respective antibodies in the tissue or cell. Visualization is achieved by labeling the antibodies with a color or fluorescent tag. Typical examples of color tags include, but are not limited to, horseradish peroxidase and alkaline phosphatase. Typical examples of fluorophore tags include, but are not limited to, fluorescein isothiocyanate (FITC) or phycoerythrin (PE).

[0146] Immuno-polymerase chain reaction (IPCR) utilizes nucleic acid amplification technology to increase signal generation in antibody-based immunoassays. Since there is no protein equivalent of PCR, i.e., proteins cannot be replicated in the same way that nucleic acids are replicated in the PCR process, the only way to improve detection sensitivity is through signal amplification. The target protein is bound to an antibody conjugated directly or indirectly to an oligonucleotide. Unbound antibody is washed away, and the remaining bound antibody has its amplified oligonucleotide. Protein detection is achieved via detection of the amplified oligonucleotide using standard nucleic acid detection methods, including real-time methods.

[0147] In some embodiments, the level or amount of mRNA is detected using an RT-qPCR analysis, which provides a Ct (cycle threshold) for each mRNA detected. In real-time PCR assays, positive reactions are detected by the accumulation of a fluorescent signal. The Ct value is defined as the number of cycles required for the fluorescent signal to reach a threshold value (i.e., to exceed background levels). Ct levels are inversely proportional to the amount of target nucleic acid in the sample (i.e., the lower the Ct value, the greater the amount of mRNA in the sample).

[0148] In some embodiments, the expression level or amount of expression of any one of the genes described herein is normalized to the expression level or amount of expression of a reference gene. In some embodiments, the amount of expression of mRNA is normalized to the mRNA expression level or amount of expression of a reference gene. Suitable reference genes for normalization are known to those of skill in the art and include, but are not limited to, KLK3, CYPB561A3, EEFlA2, GAPDH, HPN, KLK2, KLK4, LBH, NUDT8, SPDEF, or TRGV. In some embodiments, the reference gene is KLK3.

[0149] Compositions, such as reagent compositions, for use in the methods described herein include, but are not limited to, antibodies, probes, amplification oligonucleotides, and the like.

[0150] Compositions and kits can include one or more, two or more, three or more, or four or more antibodies, probes, probe pairs, amplification oligonucleotide pairs, or sequencing primers.

[0151] A probe or primer can hybridize to one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, eleven or more, twelve or more, thirteen or more, fourteen or more, fifteen or more, twenty or more, or twenty-one or more target molecules. A target molecule can be an RNA, a DNA, a cDNA, an mRNA, a portion or fragment thereof, or a combination thereof. In some cases, at least a portion of the target molecule is a cancer marker. A probe can hybridize to one or more, or two or more, cancer markers disclosed herein.

[0152] Generally, a probe or primer includes a target-specific sequence. A target-specific sequence can be complementary to at least a portion of a target molecule. A target-specific sequence can be at least about 50% or more, 55% or more, 60% or more, 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, or 100% complementary to at least a portion of a target molecule.

[0153] The length of the target-specific sequence can be at least about 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, 20 or more nucleotides. In some cases, the length of the target-specific sequence is about 8 to about 20 nucleotides, 10 to about 18 nucleotides, or 12 to about 16 nucleotides.

[0154] The composition and kit can include a plurality of probes or primers, wherein two or more probes of the plurality of probes include the same target-specific sequence. The composition and kit can include a plurality of probes, wherein two or more probes of the plurality of probes include different target-specific sequences.

[0155] The probe can further include a unique sequence. The unique sequence is not complementary to the cancer marker. The unique sequence can include a label, a barcode, or a unique identifier. The unique sequence can include a random sequence, a non-random sequence, or a combination thereof. The length of the unique sequence can be at least about 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, 20 or more, 22 or more, 24 or more, 26 or more, 28 or more, 30 or more nucleotides. In some cases, the length of the unique sequence is about 8 to about 20 nucleotides, 10 to about 18 nucleotides, or 12 to about 16 nucleotides.

[0156] The probe can further include a universal sequence. The universal sequence can include a primer binding site. The universal sequence can enable detection of the target sequence. The universal sequence can enable amplification of the target sequence. The universal sequence can enable transcription or reverse transcription of the target sequence. The universal sequence can enable sequencing of the target sequence.

[0157] The probe or primer composition of the present disclosure can be provided on a solid support. The solid support can include one or more of a bead, a plate, a solid surface, a well, a chip, or a combination thereof. The bead can be magnetic, antibody-coated, protein A cross-linked, protein G cross-linked, streptavidin-coated, oligonucleotide-conjugated, silica-coated, or a combination thereof. Examples of the bead include, but are not limited to, Ampure beads, AMPure XP beads, streptavidin beads, agarose beads, magnetic beads, Microbeads, antibody-conjugated beads (e.g., anti-immunoglobulin microbeads), protein A-conjugated beads, protein G-conjugated beads, protein A / G-conjugated beads, protein L-conjugated beads, oligo dT-conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorochrome microbeads, and BcMag TM Carboxyl-terminated magnetic beads.

[0158] Compositions and kits can include primers and primer pairs capable of amplifying a target molecule or fragment or subsequence or complement thereof. Nucleotide sequences of target molecules can be provided in computer readable media for computer applications and as a basis for designing appropriate primers for amplification of one or more target molecules.

[0159] Primers based on nucleotide sequences of target molecules can be designed for amplification of target molecules. For use in amplification reactions (such as PCR), primer pairs can be used. The exact composition of primer sequences is not important to the present disclosure, but for most applications, primers can hybridize to specific sequences of target molecules or to universal sequences of probes under stringent conditions, particularly under high stringency conditions, as known in the art. Primer pairs are typically selected to generate amplification products of at least about 15 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 450 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more nucleotides. Selection algorithms for primer sequences are generally known and can be available in commercial software packages. These primers can be used in standard quantitative or qualitative PCR-based assays to assess transcriptional expression levels of target molecules. Optionally, these primers can be used in combination with probes, such as molecular beacons, in real-time PCR amplification.

[0160] The nucleotide sequence of a full-length primer need not be derived from the target sequence. Thus, for example, a primer can include at the 5' and / or 3' end a nucleotide sequence that is not derived from the target molecule. The nucleotide sequence of the non-derivative of the target molecule nucleotide sequence can provide additional functionality to the primer. For example, they can provide a restriction enzyme recognition sequence or a "tag" that facilitates detection, isolation, purification, or immobilization on a solid support. Optionally, the additional nucleotides can provide a self-complementary sequence, causing the primer to adopt a hairpin configuration. Such a configuration can be necessary for certain primers, for example, molecular beacons and Scorpion primers, which are useful in solution hybridization techniques.

[0161] If desired, the probe or primer can incorporate a moiety that facilitates detection, isolation, purification, or immobilization. Such moieties are well known in the art (see, e.g., Ausubel et al., (1997 & updates) Current Protocols in Molecular Biology, Wiley & Sons, New York), and such moieties are selected so as not to interfere with the ability of the probe to hybridize to its target molecule.

[0162] Examples of suitable moieties are detectable labels, such as radioisotopes, fluorophores, chemilumophores, enzymes, colloidal particles and fluorescent microparticles, as well as antigens, antibodies, haptens, avidin / streptavidin, biotin, haptens, enzyme cofactors / substrates, enzymes, and the like.

[0163] The label can be attached to or incorporated into the probe or primer, optionally, to enable detection and / or quantification of the target polynucleotide that is representative of the target molecule of interest. The target polynucleotide can be the expressed target molecule RNA itself, a cDNA copy thereof, or an amplification product derived therefrom, and can be the plus or minus strand, so long as it can be specifically detected in the detection method used. Similarly, the antibody can be labeled.

[0164] In certain multiplexed formats, the labels used to detect different target molecules can be distinguishable. The labels can be attached directly (e.g., via a covalent bond) or indirectly, e.g., via a bridging molecule or series of molecules (e.g., a molecule or complex that can bind to an assay component, or via a member of a binding pair that can be incorporated into an assay component, e.g., biotin-avidin or streptavidin). Many labels are commercially available in active form that can be readily used in such conjugation (e.g., by amine acylation), or the labels can be attached by known or determinable conjugation protocols, many of which are known in the art.

[0165] Labels useful in the disclosures described herein include any substance that can be detected when bound to or incorporated into a target molecule. Any effective detection method can be used, including optical, spectroscopic, electrical, piezoelectric, magnetic, Raman scattering, surface plasmon resonance, colorimetric, calorimetric, and the like. Labels are typically selected from the group consisting of a chromophore, a luminophore, a fluorophore, a member of a quenching system, a chromogen, a hapten, an antigen, a magnetic particle, a material exhibiting nonlinear optics, a semiconductor nanocrystal, a metal nanoparticle, an enzyme, an antibody or binding portion or equivalent thereof, an aptamer, and a member of a binding pair, and combinations thereof. A quenching scheme can be used, in which a quencher and a fluorophore can be used on a probe as members of a quenching pair, such that upon binding to a target a change in an optical parameter occurs, introducing or quenching a signal from the fluorophore. One example of such a system is a molecular beacon. Suitable quencher / fluorophore systems are known in the art. Labels can be bound through a variety of intermediate linkages. For example, a target polynucleotide can include a biotin binding substance, and an optically detectable label can be conjugated to biotin, which then binds to the labeled target polynucleotide. Similarly, a polynucleotide sensor can include an immunological species, such as an antibody or fragment, and a second antibody containing an optically detectable label can be added.

[0166] Chromophores useful in the methods described herein include any substance capable of absorbing energy and emitting light. For multiplexed assays, a plurality of different signaling chromophores can be used, with different detectable emission spectra. The chromophore can be a luminophore or a fluorophore. Typical fluorophores include fluorescent dyes, semiconductor nanocrystals, lanthanide chelates, polynucleotide specific dyes, and green fluorescent protein.

[0167] Encoding schemes can optionally be used, including encoding particles and / or encoding tags associated with different polynucleotides of the disclosure. A variety of different encoding schemes are known in the art, including fluorophores (including SCNCs), deposited metals, and RF tags.

[0168] Subjects and Samples

[0169] The methods and kits described herein are suitable for detecting the expression level or amount of expression of one or more genes described herein in a sample from a subject. In some embodiments, the subject from which the sample is obtained can be selected by a skilled practitioner. In some embodiments, the selection of the subject is based on consideration or analysis of one or more factors. These considerations include, but are not limited to, a family history of a particular disease, a genetic predisposition to a disease, an increased risk of developing a disease, physical symptoms indicative of a disease, or environmental causes. Environmental causes can include, but are not limited to, lifestyle or exposure to factors that cause or contribute to a particular disease. In some embodiments, the selection of the subject is based on the subject’s prior medical history, a positive diagnosis prior to therapy or after therapy, treatment of a disease, or remission or recovery from a disease.

[0170] In some embodiments, the sample used with the kits and for the methods of the disclosure comprises nucleic acids suitable to provide RNA expression information. In principle, the biological sample from which the expressed RNA is obtained and from which the expression of the target molecules is analyzed can be any material suspected to comprise cancerous tissue or cells. The sample can be a biological sample used directly in the methods of the disclosure. Optionally, the sample can be a sample prepared from a biological sample.

[0171] In some embodiments, the sample or sample portion comprising or suspected to comprise cancerous tissue or cells can be any biological material source, including cells, tissue, secretions, or fluids, including bodily fluids. Non-limiting examples of sample sources include aspirates, needle biopsies, cytological sediment, bulk tissue preparation or sections thereof (e.g., obtained by surgery or autopsy), lymphatic fluid, blood, plasma, serum, tumors, and organs. Optionally, or in addition, the sample source can be urine, bile, fecal matter, sweat, tears, spinal fluid, and fecal matter. In some embodiments, the sample source is a secretion. In some embodiments, the secretion is an exosome. In some embodiments, the sample is a urine sample. In some embodiments, the urine sample is obtained after the subject has performed a digital rectal exam (DRE). In some embodiments, the urine sample is obtained within 30 minutes after the subject has performed a DRE. In some embodiments, the urine sample is obtained between 30 minutes and 60 minutes after the subject has performed a DRE. In some embodiments, the urine sample is obtained between 30 minutes and 180 minutes after the subject has performed a DRE. In some embodiments, the urine sample is obtained within one hour after the subject has performed a DRE. In some embodiments, the urine sample is obtained within two hours after the subject has performed a DRE. In some embodiments, the urine sample is obtained within three hours after the subject has performed a DRE. In some embodiments, the urine sample is obtained from a subject who has not performed a DRE.

[0172] Without wishing to be bound by theory, it is believed that a DRE increases the concentration of mRNA or protein expressed by one or more of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6 in a sample (e.g., a urine sample). This increase in concentration facilitates the detection of mRNA or protein expressed by one or more genes.

[0173] In some embodiments, the sample is combined with a buffer, e.g., for processing. In some embodiments, the amount of expression of one or more genes described herein is determined from a composition (e.g., a solution or suspension) comprising the sample and the buffer. Buffers suitable for use with a sample are known to those of skill in the art and can be determined based on the type of sample collected. In some embodiments, the composition further comprises a preservative to provide sufficient stability to the sample. In some embodiments, the ratio of buffer to sample is 2:5. In some embodiments, the ratio of buffer to sample is 1 :5, 2:5, 3:5, or 4:5.

[0174] The sample can be an archival sample with a known and recorded medical outcome, or can be a sample from a current patient for whom the final medical outcome is not yet known.

[0175] In some embodiments, the sample can be dissected prior to molecular analysis. The sample can be prepared via macrodissection of a bulk tumor specimen or portion thereof, or can be processed via microdissection, such as via laser capture microdissection (LCM).

[0176] The sample can be initially provided in a variety of states, such as fresh tissue, fresh-frozen tissue, fine needle aspirate, and can be fixed or unfixed. Typically, medical laboratories routinely prepare medical samples in a fixed state to facilitate tissue storage. A variety of fixatives can be used to fix tissue to stabilize cell morphology, and can be used alone or in combination with other reagents. Exemplary fixatives include cross-linking agents, alcohols, acetone, Bouin's solution, Zenker's solution, Hely's solution, osmic acid solution, and Carnoy's solution.

[0177] Cross-linking fixatives can include any reagent suitable for forming two or more covalent bonds, such as an aldehyde. Commonly used sources of aldehyde for fixation include formaldehyde, paraformaldehyde, glutaraldehyde, or formalin. Preferably, the cross-linking agent includes formaldehyde, which can be included in its native form or in the form of paraformaldehyde or formalin. Those of skill in the art will appreciate that for samples that have been fixed using a cross-linking fixative, special preparation steps can be required, including, for example, heating steps and proteinase-k digestion.

[0178] One or more alcohols can be used to fix tissue, alone or in combination with other fixatives. Exemplary alcohols for fixation include methanol, ethanol, and isopropanol.

[0179] Formalin fixation is frequently used in medical laboratories. Formalin includes an alcohol (typically methanol) and formaldehyde, both of which can act to fix a biological sample.

[0180] Regardless of fixation or non-fixation, a biological sample can optionally be embedded in an embedding agent. Exemplary embedding agents used in histology include paraffin, V.I.P. (TM), Paramat, Paramat Extra, Paraplast, Paraplast X-tra, Paraplast Plus, Peel Away Paraffin Embedding Wax, Polyester Wax, Carbowax Polyethylene Glycol, Polyfin TM , Tissue Freezing Medium TFMFM, Cryo-Gef TM and OCT Compound (Electron Microscopy Sciences, Hatfield, PA). Prior to performing molecular analysis, the embedding material can be removed by any suitable technique known in the art. For example, when the sample is embedded in wax, the embedding material can be removed by extraction with one or more organic solvents (e.g., xylene). Kits are commercially available for removing embedding agents from tissue. The sample or section thereof can be subjected to further processing steps as desired, such as successive hydration or dehydration steps.

[0181] In some embodiments, the sample is a fixed, wax-embedded biological sample. Samples provided by medical laboratories are typically fixed, wax-embedded samples, most commonly formalin-fixed paraffin-embedded (FFPE) tissue.

[0182] In some embodiments, the subject is undergoing a first prostate biopsy, i.e., the subject has not undergone a prostate biopsy. In some embodiments, the subject has a prior negative prostate biopsy result. In some embodiments, the prostate biopsy result is negative for prostate cancer with a grade group > 2. In some embodiments, one or more additional clinical variables are associated with the subject. In some embodiments, the method comprises determining one or more additional clinical variables (e.g., including but not limited to prostate volume, PSA level or amount, PSA density, biopsy Gleason score, race, family history of prostate cancer, prior negative prostate biopsy, or abnormal DRE of the subject). In some embodiments, one or more additional clinical variables are associated with the subject having a prior negative prostate biopsy result.

[0183] Determining the likelihood of having or developing prostate cancer with a grade group > 2

[0184] In some embodiments, the expression level or amount of expression of one or more genes described herein determines the likelihood of detecting prostate cancer in a subject. In some embodiments, the likelihood of detecting prostate cancer in a subject is based on a prostate biopsy of the subject. In some embodiments, the expression level or amount of expression of one or more genes described herein determines the likelihood of detecting a prostate cancer that is graded and grouped >2 in a subject. In some embodiments, the likelihood is represented in a score based on the amount or level of expression of one or more genes described herein present in a sample of the subject. In some embodiments, the likelihood of detecting a prostate cancer that is graded and grouped >2 is represented in a score ranging from 0% to 100%. In some embodiments, the likelihood of detecting a prostate cancer that is graded and grouped >2 is represented in a score ranging from 0.0 to 100.0. In some embodiments, a subject undergoing an initial biopsy that obtains a score of 0-7.5% means a low risk or low likelihood of detecting a prostate cancer that is graded and grouped >2 from a prostate biopsy of the subject. In some embodiments, a subject with a prior negative prostate biopsy result that obtains a score of 0-5.4% means a low risk or low likelihood of detecting a prostate cancer that is graded and grouped >2 from a prostate biopsy of the subject. In some embodiments, a subject undergoing an initial biopsy that obtains a score of >7.6% means a high risk or high likelihood of detecting a prostate cancer that is graded and grouped >2 from a prostate biopsy of the subject. In some embodiments, a subject with a prior negative prostate biopsy result that obtains a score of >5.5% has a high risk or high likelihood of detecting a prostate cancer that is graded and grouped >2 from a prostate biopsy of the subject.

[0185] In some embodiments, a computer-based analysis program is used to convert the raw data generated by the detection assay (e.g., the presence, absence, or amount of one or more markers given) into data that has predictive value to the clinician, subject, or healthcare provider of the subject. The clinician, subject, or healthcare provider of the subject can access the raw data using any suitable means. Thus, in some embodiments, the present disclosure provides a further benefit in that the clinician, subject, or healthcare provider of the subject, who can not have received training in genetics or molecular biology, does not need to understand the raw data. The data can be presented directly to the clinician or healthcare provider of the subject in its most useful form. This allows the clinician or healthcare provider to immediately utilize this information to optimize the care of the subject.

[0186] Information can be received, processed, or transmitted to one or more laboratories performing the assays, information providers, medical personnel, or subjects using any suitable method. For example, in some embodiments of the disclosure, a sample (e.g., a biopsy or serum or urine sample) is taken from a subject and submitted to an analysis service (e.g., a clinical laboratory of a medical facility, a genomic analysis business, etc.) located anywhere in the world (e.g., a different country than the subject resides or the country where the information is ultimately used) to generate raw data. When the sample includes a tissue or other biological sample, the subject can visit a medical center to obtain the sample and send it to the analysis center, or the subject can collect the sample (e.g., a urine sample) themselves and send it directly to the analysis center. When the sample includes previously determined biological information, the information can be sent directly to the analysis service by the subject (e.g., an information card including the information can be scanned by a computer and the data transmitted to a computer of the analysis center using an electronic communication system). Once the analysis service receives the sample, the sample can be processed and an analysis (i.e., expression data) generated that is useful for diagnostic or prognostic information desired by the subject.

[0187] The analysis data is then prepared in a format suitable for interpretation by one or more medical personnel (e.g., a treating clinician, physician assistant, nurse, or pharmacist). For example, rather than providing raw expression data, the prepared format can be a diagnosis or risk assessment for the subject (e.g., levels of the cancer markers described herein), and recommendations regarding particular treatment options. The data can be displayed to the medical personnel by any suitable method. For example, in some embodiments, the analysis service generates a report that can be printed for the medical personnel (e.g., at the point of care) or displayed on a computer display for the medical personnel.

[0188] In some embodiments, the information is first analyzed at the point of care or regional facility. The raw data is then sent to a central processing facility for further analysis and / or to convert the raw data into information useful to medical personnel or subjects. The central processing facility provides the advantages of privacy (all data is stored in a central facility with uniform security protocols), speed, and uniformity of data analysis. The central processing facility can then control the fate of the data after the subject is treated. For example, using an electronic communication system, the central facility can provide the data to medical personnel, subjects, or researchers.

[0189] In some embodiments, the subject or the subject's health care provider has direct access to the data using an electronic communication system. The subject can select further intervention or consultation based on the results.

[0190] In some embodiments, the data is used for research purposes. For example, the data can be used to further optimize the inclusion or elimination of markers as useful indicators of a particular disease condition or stage, or as a companion diagnostic to determine a treatment regimen.

[0191] In some embodiments, the expression level or amount of expression of one or more genes described herein is used to determine a score. In some embodiments, determining a score comprises performing an algorithm that generates a score. In some embodiments, the score is related to or informs the likelihood that the subject has or will develop a prostate cancer that is a grade group >2. In some embodiments, the score indicates the likelihood that a prostate cancer that is a grade group >2 will be detected from a prostate biopsy of the subject. Algorithms that determine a score with acceptable diagnostic accuracy can be derived based on, for example, but not limited to, logistic regression with stepwise feature selection, logistic regression with recursive feature elimination, and regularized logistic regression with elastic net. In some embodiments, performing an algorithm comprises using a processor.

[0192] In some embodiments, the algorithm is Equation 1, 2, 3, or 4 below.

[0193] In some embodiments, the subject has a prior negative prostate biopsy result, and determining a score comprises performing Equation 1:

[0194] Equation 1:

[0195] where x = Intercept + Slope((-1)((a)(CRT mean APOC1 - CRT mean Ref) + (b)(CRT mean B3GNT6 - CRT mean Ref) + (c)(CRT mean CAMKK2 - CRT mean Ref) + (d)(CRT mean ERG - CRT mean Ref) + (e)(CRT mean HOXC6 - CRT mean Ref) + (f)(CRT mean KLK4 - CRT mean Ref) + (g)(CRT mean NKAIN1 - CRT mean Ref) + (h)(CRT mean OR51E2 - CRT mean Ref) + (i)(CRT mean PCA3 - CRT mean Ref) + (j)(CRT mean PCAT14 - CRT mean Ref) + (k)(CRT mean PCGEM1 - CRT mean Ref) + (1)(CRT mean SCHLAP1 - CRT mean Ref) + (m)(CRT mean SPON2 - CRT mean Ref) + (n)(CRT mean TFF3 - CRT mean Ref) + (o)(CRT mean T2:ERG - CRT mean Ref) + (p)(CRT mean TMSB15A - CRT mean Ref) + (q)(CRT mean TRGV9 - CRT mean Ref)) + ((r)(Age) + (s)(Family Hx) + (t)(Abnormal DRE) + (u)(Bx Prior Neg) + (v)(PSA) + (w)(Prostate Volume))).

[0196] In some embodiments, the subject is undergoing a first prostate biopsy (i.e., has not had a prior prostate biopsy), and determining the score comprises performing Equation 2:

[0197] Equation 2:

[0198] where x = intercept + slope((-1)((a)(CRT mean APOC1 - CRT mean reference) + (b)(CRT mean B3GNT6 - CRT mean reference) + (c)(CRT mean CAMKK2 - CRT mean reference) + (d)(CRT mean ERG - CRT mean reference) + (e)(CRT mean HOXC6 - CRT mean reference) + (f)(CRT mean KLK4 - CRT mean reference) + (g)(CRT mean NKAIN1 - CRT mean reference) + (h)(CRT mean OR51E2 - CRT mean reference) + (i)(CRT mean PCA3 - CRT mean reference) + (j)(CRT mean PCAT14 - CRT mean reference) + (k)(CRT mean PCGEM1 - CRT mean reference) + (1)(CRT mean SCHLAP1 - CRT mean reference) + (m)(CRT mean SPON2 - CRT mean reference) + (n)(CRT mean TFF3 - CRT mean reference) + (o)(CRT mean T2:ERG - CRT mean reference) + (p)(CRT mean TMSB15A - CRT mean reference) + (q)(CRT mean TRGV9 - CRT mean reference))).

[0199] In Equations 1 and 2, "reference" refers to a reference gene described herein (e.g., KLK3).

[0200] In some embodiments, the subject has a prior negative prostate biopsy result, and determining the score comprises performing Equation 3:

[0201] Equation 3:

[0202] where x = intercept + slope((-1)((a)(CRT mean APOC1 - CRT mean KLK3) + (b)(CRT mean B3GNT6 - CRT mean KLK3) + (c)(CRT mean CAMKK2 - CRT mean KLK3) + (d)(CRT mean ERG - CRT mean KLK3) + (e)(CRT mean HOXC6 - CRT mean KLK3) + (f)(CRT mean KLK4 - CRT mean KLK3) + (g)(CRT mean NKAIN1 - CRT mean KLK3) + (h)(CRT mean OR51E2 - CRT mean KLK3) + (i)(CRT mean PCA3 - CRT mean KLK3) + (j)(CRT mean PCAT14 - CRT mean KLK3) + (k)(CRT mean PCGEM1 - CRT mean KLK3) + (1)(CRT mean SCHLAP1 - CRT mean KLK3) + (m)(CRT mean SPON2 - CRT mean KLK3) + (n)(CRT mean TFF3 - CRT mean KLK3) + (o)(CRT mean T2:ERG - CRT mean KLK3) + (p)(CRT mean TMSB15A - CRT mean KLK3) + (q)(CRT mean TRGV9 - CRT mean KLK3))).

[0203] In some embodiments, the subject is undergoing a first prostate biopsy (i.e., has not undergone a prostate biopsy in the past), and determining the score comprises performing Equation 4:

[0204] Equation 4:

[0205] where x = Intercept + Slope((-1)((a)(CRT mean APOC1 - CRT mean KLK3) + (b)(CRT mean B3GNT6 - CRT mean KLK3) + (c)(CRT mean CAMKK2 - CRT mean KLK3) + (d)(CRT mean ERG - CRT mean KLK3) + (e)(CRT mean HOXC6 - CRT mean KLK3) + (f)(CRT mean KLK4 - CRT mean KLK3) + (g)(CRT mean NKAIN1 - CRT mean KLK3) + (h)(CRT mean OR51E2 - CRT mean KLK3) + (i)(CRT mean PCA3 - CRT mean KLK3) + (j)(CRT mean PCAT14 - CRT mean KLK3) + (k)(CRT mean PCGEM1 - CRT mean KLK3) + (1)(CRT mean SCHLAP1 - CRT mean KLK3) + (m)(CRT mean SPON2 - CRT mean KLK3) + (n)(CRT mean TFF3 - CRT mean KLK3) + (o)(CRT mean T2:ERG - CRT mean KLK3) + (p)(CRT mean TMSB15A - CRT mean KLK3) + (q)(CRT mean TRGV9 - CRT mean KLK3)) + ((r)(Age) + (s)(Family Hx) + (t)(Abnormal DRE) + (u)(Bx Prior Neg) + (v)(PSA) + (w)(Prostate Volume))).

[0206] In each of Equations 1-4: (i)“CRT” refers to the cycle threshold identified by the methods described herein for determining the amount of gene expression, (ii)“e” is Euler’s number, and (iii)“MPS2” is a score representing the likelihood of detecting a Gleason grade >2 prostate cancer from a prostate biopsy of the subject. In Equations 1 and 3,“Family Hx,”“Abnormal DRE,” and“Bx Prior Neg” are binary values of 1 or 0, where 1 = yes and 0 = no. Specifically, the value is 1 if the subject has a family history of prostate cancer, 1 if the subject has an abnormal DRE, and 1 if the subject has a prior negative prostate biopsy. In each of Equations 1-4, (a)-(w) are independent coefficients based on the selected model, e.g., logistic regression with stepwise feature selection, logistic regression with recursive feature elimination, or regularized logistic regression with elastic net. Illustrative coefficients are provided in Table B.

[0207] Table B. Illustrative coefficients used in Equations 1-4

[0208] where x = Intercept + Slope((-1)((a)(CRT mean APOC1 - CRT mean KLK3) + (b)(CRT mean B3GNT6 - CRT mean KLK3) + (c)(CRT mean CAMKK2 - CRT mean KLK3) + (d)(CRT mean ERG - CRT mean KLK3) + (e)(CRT mean HOXC6 - CRT mean KLK3) + (f)(CRT mean KLK4 - CRT mean KLK3) + (g)(CRT mean NKAIN1 - CRT mean KLK3) + (h)(CRT mean OR51E2 - CRT mean KLK3) + (i)(CRT mean PCA3 - CRT mean KLK3) + (j)(CRT mean PCAT14 - CRT mean KLK3) + (k)(CRT mean PCGEM1 - CRT mean KLK3) + (1)(CRT mean SCHLAP1 - CRT mean KLK3) + (m)(CRT mean SPON2 - CRT mean KLK3) + (n)(CRT mean TFF3 - CRT mean KLK3) + (o)(CRT mean T2:ERG - CRT mean KLK3) + (p)(CRT mean TMSB15A - CRT mean KLK3) + (q)(CRT mean TRGV9 - CRT mean KLK3)) + ((r)(Age) + (s)(Family Hx) + (t)(Abnormal DRE) + (u)(Bx Prior Neg) + (v)(PSA) + (w)(Prostate Volume))).

[0206] In each of Equations 1-4: (i)“CRT” refers to the cycle threshold identified by the methods described herein for determining the amount of gene expression, (ii)“e” is Euler’s number, and (iii)“MPS2” is a score representing the likelihood of detecting a Gleason grade >2 prostate cancer from a prostate biopsy of the subject. In Equations 1 and 3,“Family Hx,”“Abnormal DRE,” and“Bx Prior Neg” are binary values of 1 or 0, where 1 = yes and 0 = no. Specifically, the value is 1 if the subject has a family history of prostate cancer, 1 if the subject has an abnormal DRE, and 1 if the subject has a prior negative prostate biopsy. In each of Equations 1-4, (a)-(w) are independent coefficients based on the selected model, e.g., logistic regression with stepwise feature selection, logistic regression with recursive feature elimination, or regularized logistic regression with elastic net. Illustrative coefficients are provided in Table B.

[0207] Table B. Illustrative coefficients used in Equations 1-4

[0208]

[0209]

[0210] N / A = not applicable.

[0211] - = not present in any of equations 1-4.

[0212] The methods disclosed herein can also include the transmission of data / information. For example, data / information derived from the detection and / or quantification of a target can be transmitted to another device and / or instrument. In some cases, information obtained from an algorithm can also be transmitted to another device and / or instrument. The transmission of data / information can include the transmission of data / information from a first source to a second source. The first source and the second source can be located at approximately the same location (e.g., within the same room, building, block, campus). Optionally, the first source and the second source can be located at multiple locations (e.g., multiple cities, states, countries, continents, etc.).

[0213] The transmission of data / information can include digital transmission or analog transmission. Digital transmission can include the physical transmission of data (a stream of digital bits) over a point-to-point or point-to-multipoint communication channel. Examples of such channels are copper wires, optical fibers, wireless communication channels, and storage media. The data can be represented as electromagnetic signals, such as voltages, radio waves, microwaves, or infrared signals.

[0214] Analog transmission can include the transmission of continuously varying analog signals. Messages can be represented by means of a sequence of pulses of a line code (baseband transmission), or also by using a digital modulation method from a finite set of continuously varying waveforms (passband transmission). Passband modulation and the corresponding demodulation (also called detection) can be performed by modem devices. According to the most common definition of digital signals, both baseband and passband signals representing a bit stream are considered as digital transmissions, while another definition considers only baseband signals as digital, while passband transmission of digital data is considered as a form of digital-to-analog conversion.

[0215] In some embodiments, a report comprising a score is generated. In some embodiments, the score indicates the likelihood that a grade group >2 prostate cancer is detected from a prostate biopsy of a subject. In some embodiments, the report is accessible or provided to a subject’s healthcare provider. In some embodiments, the report is available or provided in digital or paper copy. In some embodiments, the report is delivered to a subject’s healthcare provider in a digital format as described herein (e.g., via email), or via courier if the report is a paper copy.

[0216] In some embodiments, the report comprises a treatment selection. In some embodiments, the report comprises a treatment selection for a grade group >2 prostate cancer.

[0217] Diagnostic accuracy

[0218] Diagnostic accuracy of the methods or kits described herein can be determined by analyzing the area under the curve (AUC) derived from a receiver operating characteristic (ROC) curve. An ROC curve is a graphical plot that shows the ability of a binary classifier system as its decision threshold varies. The ROC curve is plotted with the true positive rate versus the false positive rate, with the true positive rate on the y-axis and the false positive rate on the x-axis. The true positive rate (also known as sensitivity) is calculated by dividing the number of true positives by the sum of true positives and false negatives. The false positive rate is calculated by either (1) dividing the number of false positives by the sum of true negatives and false positives, or (2) subtracting the specificity from one, where the specificity is calculated by dividing the number of true negatives by the sum of true negatives and false positives. In some embodiments, an ROC curve is generated based on the individual expression amounts of each gene. In some embodiments, an ROC curve is generated based on a combination of the expression amounts of each gene.

[0219] In some embodiments, the AUC value of the method or kit described herein is greater than 0.50. In some embodiments, the AUC value of the method or kit described herein is at least 0.60. In some embodiments, the AUC value of the method or kit described herein is at least 0.70. In some embodiments, the AUC value of the method or kit described herein is at least 0.71. In some embodiments, the AUC value of the method or kit described herein is at least 0.72. In some embodiments, the AUC value of the method or kit described herein is at least 0.73. In some embodiments, the AUC value of the method or kit described herein is at least 0.74. In some embodiments, the AUC value of the method or kit described herein is at least 0.75. In some embodiments, the AUC value of the method or kit described herein is at least 0.76. In some embodiments, the AUC value of the method or kit described herein is at least 0.77. In some embodiments, the AUC value of the method or kit described herein is at least 0.78. In some embodiments, the AUC value of the method or kit described herein is at least 0.79. In some embodiments, the AUC value of the method or kit described herein is at least 0.80. In some embodiments, the AUC value of the method or kit described herein is at least 0.81. In some embodiments, the AUC value of the method or kit described herein is at least 0.82. In some embodiments, the AUC value of the method or kit described herein is at least 0.83. In some embodiments, the AUC value of the method or kit described herein is at least 0.84. In some embodiments, the AUC value of the method or kit described herein is at least 0.85. In some embodiments, the AUC value of the method or kit described herein is at least 0.86. In some embodiments, the AUC value of the method or kit described herein is at least 0.87. In some embodiments, the AUC value of the method or kit described herein is at least 0.88. In some embodiments, the AUC value of the method or kit described herein is at least 0.89. In some embodiments, the AUC value of the method or kit described herein is at least 0.90.

[0220] By implementing critical analysis, diagnostic accuracy of individual gene expression or combinations of specific gene expression can be maximized, taking into account sensitivity, specificity, negative predictive value (NPV), positive predictive value (PPV), positive likelihood ratio (PLR), and negative likelihood ratio (NLR) necessary for clinical utility. Results of expression are analyzed in a variety of ways. In some embodiments, results are analyzed using univariate or single variable analysis (SV). In some embodiments, results are analyzed using multivariate analysis (MV).

[0221] Generation of ROC curves and analysis of sample populations can be used to establish cut-off values that are directly useful in distinguishing different subgroups of subjects. For example, a cut-off value can be used to distinguish between a high likelihood of detecting a prostate cancer with a Gleason score > 2 from a prostate biopsy of a subject and a low likelihood of detecting a prostate cancer with a Gleason score > 2 from a prostate biopsy of a subject. In some embodiments, a cut-off value can distinguish between these subjects. In some embodiments, a cut-off value can distinguish between subjects with non-invasive cancer and invasive cancer.

[0222] In some embodiments, the methods or kits described herein provide a score that indicates a likelihood of detecting a prostate cancer with a Gleason score > 2 from a prostate biopsy of a subject with a diagnostic accuracy of at least 0.70. In some embodiments, the methods or kits described herein provide a score that indicates a likelihood of detecting a prostate cancer with a Gleason score > 2 from a prostate biopsy of a subject with a diagnostic accuracy of at least 0.75. In some embodiments, the methods or kits described herein provide a score that indicates a likelihood of detecting a prostate cancer with a Gleason score > 2 from a prostate biopsy of a subject with a diagnostic accuracy of at least 0.80.

[0223] In some embodiments, the methods or kits described herein provide a score that indicates a likelihood of detecting a prostate cancer with a Gleason score > 2 from a prostate biopsy of a subject who is undergoing a prostate biopsy for the first time with a diagnostic accuracy of at least 0.70. In some embodiments, the methods or kits described herein provide a score that indicates a likelihood of detecting a prostate cancer with a Gleason score > 2 from a prostate biopsy of a subject who is undergoing a prostate biopsy for the first time with a diagnostic accuracy of at least 0.75. In some embodiments, the methods or kits described herein provide a score that indicates a likelihood of detecting a prostate cancer with a Gleason score > 2 from a prostate biopsy of a subject who is undergoing a prostate biopsy for the first time with a diagnostic accuracy of at least 0.80.

[0224] In some embodiments, the methods or kits described herein provide a score that indicates the likelihood of detecting a grade group >2 prostate cancer from a prostate biopsy of a subject with a prior negative prostate biopsy with a diagnostic accuracy of at least 0.70. In some embodiments, the methods or kits described herein provide a score that indicates the likelihood of detecting a grade group >2 prostate cancer from a prostate biopsy of a subject with a prior negative prostate biopsy with a diagnostic accuracy of at least 0.75. In some embodiments, the methods or kits described herein provide a score that indicates the likelihood of detecting a grade group >2 prostate cancer from a prostate biopsy of a subject with a prior negative prostate biopsy with a diagnostic accuracy of at least 0.80. In some embodiments, the methods or kits described herein provide a score that indicates the likelihood of detecting a grade group >2 prostate cancer from a prostate biopsy of a subject with a prior negative prostate biopsy with a diagnostic accuracy of at least 0.81. In some embodiments, the methods or kits described herein provide a score that indicates the likelihood of detecting a grade group >2 prostate cancer from a prostate biopsy of a subject with a prior negative prostate biopsy with a diagnostic accuracy of at least 0.82.

[0225] In some embodiments, each reference diagnostic accuracy can be achieved where the urine sample is obtained within one hour of the subject undergoing a digital rectal exam (DRE). In some embodiments, each reference diagnostic accuracy can be achieved where the urine sample is obtained within 30 minutes to 60 minutes of the subject undergoing a DRE. In some embodiments, the urine sample is obtained 30 minutes to 180 minutes after the subject undergoes a DRE. In some embodiments, the urine sample is obtained within one hour of the subject undergoing a DRE. In some embodiments, the urine sample is obtained within two hours of the subject undergoing a DRE. In some embodiments, the urine sample is obtained within three hours of the subject undergoing a DRE. In some embodiments, the urine sample is obtained from a subject who does not undergo a DRE.

[0226] Kits and devices

[0227] In some embodiments, the present disclosure provides a kit for analyzing a cancer, comprising (a) a probe set comprising a plurality of probes comprising a target-specific sequence complementary to one or more target molecules, wherein the one or more target molecules comprise one or more cancer markers; and (b) a computer model or algorithm for analyzing the expression level and / or expression profile of the one or more target molecules in a sample. The target molecules can comprise one or more or a combination thereof described herein.

[0228] In some embodiments, the disclosure provides a kit for analyzing cancer comprising (a) a probe set comprising a plurality of probes comprising a target-specific sequence complementary to one or more target molecules in a biomarker library; and (b) a computer model or algorithm for analyzing the expression level and / or expression profile of the one or more target molecules in a sample. Control samples and / or nucleic acids can optionally be provided in the kit. The control samples can include tissue and / or nucleic acids obtained from or representative of a tumor sample from a healthy subject, and tissue and / or nucleic acids obtained from or representative of a tumor sample from a subject diagnosed with cancer.

[0229] Instructions (instructions) for using the kit to perform one or more methods of the disclosure can be provided and can be provided in any fixed medium. The instructions can be located on or inside the container or packaging, and / or can be printed on or inside the packaging. The kit can be multiplexed for simultaneous detection and / or quantification of one or more different target polynucleotides representative of expression.

[0230] In some embodiments, the disclosure provides a kit comprising a container containing a reagent composition for detecting the amount of expression of at least three genes described herein; and instructions for detecting the amount of expression. In some embodiments, the reagent composition comprises polynucleotide reagents for detecting the amount of mRNA expressed by the at least three genes. In some embodiments, the reagent composition comprises polynucleotide reagents for detecting the amount of expression of a reference gene, and the instructions are further for normalizing the amount of expression of the at least three genes to the amount of expression of the reference gene. In some embodiments, the instructions are further for generating a report comprising a score determined from the amount of expression of the at least three genes, wherein the score indicates a likelihood of detecting a grade group >2 prostate cancer from a prostate biopsy of a subject.

[0231] Devices useful for performing the methods of the disclosure are also provided. The devices can include means for characterizing the expression levels of the target molecules of the disclosure, such as components for performing one or more methods of nucleic acid extraction, amplification, and / or detection. Such components can include one or more of the following: an amplification chamber (e.g., a thermal cycler), a plate reader, a spectrophotometer, a capillary electrophoresis apparatus, a chip reader, and / or a robotic sample processing component. These components can ultimately obtain data reflecting the expression levels of the target molecules used in the employed assays.

[0232] The devices can include excitation and / or detection means. Any instrument that can excite a wavelength of the target species and shorter than the one or more emission wavelengths to be detected can be used for excitation. Commercially available devices can provide suitable excitation wavelengths as well as suitable detection components.

[0233] Exemplary excitation sources include a broadband UV light source such as a deuterium lamp with appropriate filters, the output of a white light source such as a xenon or deuterium lamp after extraction of the desired wavelength or wavelengths by a monochromator, a continuous wave (cw) gas laser, a solid state diode laser, or any pulsed laser. Emitted light can be detected by any suitable means or technique; many suitable methods are known in the art. For example, a fluorometer or spectrophotometer can be used to detect whether a test sample emits light at a wavelength characteristic of the label used in the assay.

[0234] The device can include means for identifying a given sample and associating the results obtained with that sample. Such means can include manual labeling, bar codes, and other indicators that can be associated with the sample container, and / or can optionally be included in the sample itself, for example in the case of encoded particles added to the sample. The results can be associated with the sample, for example in a computer memory containing a record of the sample identification and expression levels obtained from the sample. The association of the results with the sample can also include association with a particular sample container in the device, which is also associated with the sample identity.

[0235] The device can also include means for associating the expression levels of the target molecules being studied with a prognosis of a disease outcome. Such means can include one or more of a variety of correlation techniques, including look-up tables, algorithms, multivariate models, and linear or non-linear combinations of expression models or algorithms. The expression levels can be converted to one or more likelihood scores reflecting the likelihood that a subject providing the sample will exhibit a particular disease outcome. These models and / or algorithms can be provided in machine-readable form, and can optionally further specify a treatment modality for the subject or a class of subjects.

[0236] The device can also include output means for outputting the disease status, prognosis, and / or treatment modality. This output can take any form of communication of the results to the subject and / or to a health care provider, and can include a monitor, a printout, or both. The device can use a computer system to perform one or more of the steps provided.

[0237] II. Prognosis, Diagnosis, or Treatment

[0238] The methods, compositions, and kits disclosed herein can be used to prognose, diagnose, predict, monitor, and / or treat cancer (e.g., prostate cancer, and in some embodiments, prostate cancer that is stratified as >2) in a subject. In some embodiments, predicting and / or monitoring the status or outcome of cancer comprises assessing the presence or risk of high-grade prostate cancer (i.e., prostate cancer that is stratified as >2). In some embodiments, predicting and / or monitoring the status or outcome of cancer comprises determining the efficacy of a treatment. In some embodiments, the methods and kits disclosed herein can be used to indicate the likelihood of detecting prostate cancer that is stratified as >2 from a prostate biopsy of a subject.

[0239] In some embodiments, the method comprises determining, recommending, or administering a treatment regimen. In some embodiments, the treatment regimen is an anti-cancer therapy. In some embodiments, the method comprises modifying the treatment regimen. Modifying the treatment regimen can comprise increasing the dosage of the treatment, decreasing the dosage of the treatment, or terminating the treatment regimen.

[0240] For example, in some embodiments, the methods described herein can be used to identify a subject having high-grade prostate cancer. In some embodiments, the methods described herein can be used to identify a subject having a high likelihood of having high-grade prostate cancer that can be detected from a prostate biopsy. Such a subject can be administered a prostate cancer therapy (e.g., one or more of surgery, radiation therapy, hormone therapy, targeted therapy, chemotherapy, immunotherapy, radiopharmaceutical, or bone-modifying drug).

[0241] In contrast, in some embodiments, a subject identified as having low-grade prostate cancer or as having a low likelihood of having high-grade prostate cancer (e.g., based on the expression level of the marker) can be given the option to avoid biopsy or treatment and opt for close observation or minimal treatment.

[0242] In some embodiments, the prostate cancer therapy comprises administration of a chemotherapeutic agent. Examples of chemotherapeutic agents include alkylating agents, antimetabolites, plant alkaloids and terpenoids, vinca alkaloids, podophyllotoxins, taxoids, topoisomerase inhibitors, and cytotoxic antibiotics. Cisplatin, carboplatin, and oxaliplatin are all examples of alkylating agents. Other alkylating agents include nitrogen mustards, cyclophosphamide, chlorambucil, ifosfamide. Alkylating agents can impair cellular function by forming covalent bonds with amino, carboxyl, sulfhydryl, and phosphate groups in biologically important molecules. Optionally, alkylating agents can chemically modify the DNA of a cell.

[0243] Biotherapies (sometimes called immunotherapies, biologic therapies, or biologic response modifier (BRM) therapies) directly or indirectly harness the body’s immune system to fight cancer or to reduce side effects that can be caused by some cancer treatments. Biotherapies include interferons, interleukins, colony-stimulating factors, monoclonal antibodies, vaccines, gene therapies, and nonspecific immunomodulators.

[0244] In some embodiments, the biotherapy is an immune checkpoint therapy. Immune checkpoint inhibitors target CTLA-4, PD-1, or PD-L1. Examples include, but are not limited to, ipilimumab, nivolumab, cemiplimab, avelumab, durvalumab, tremelimumab, dostarlimab, pembrolizumab, spartalizumab, and atezolizumab.

[0245] In some embodiments, the prostate cancer therapy is an FDA-approved for the treatment of prostate cancer. In some embodiments, the prostate cancer therapy is: abiraterone acetate, apulutamide, bicalutamide, cabazitaxel, casodex, darolutamide, degarelix, docetaxel, eligard, enzalutamide, erleada, firmagon, flutamide, goserelin acetate, jevtana, leuprolide acetate, Lupron depot, lutetium lu 177 vipivotide tetraxetan, Lynparza, mitoxantrone hydrochloride, nilandron, nilutamide, nubeqa, Olaparib, orgovyx, pluvicto, provege, radium dichloride 223, relugolix, rubraca, rucaparib camsylate, sipuleucel-t, taxotere, xofigo, xtandi, yonsa, zoladex, xytiga, or any combination thereof.

[0246] Experiments

[0247] The following examples are provided to demonstrate and further illustrate certain embodiments and aspects of the present disclosure, and are not to be construed as limiting the scope thereof.

[0248] Example 1

[0249] Methods

[0250] Initial genetic screening

[0251] RNA-seq data from The Cancer Genome Atlas (TCGA) Prostate Adenocarcinoma (PRAD) cohort was used to select potential grade-associated genes (The Cancer Genome Atlas Research Network. The Molecular Taxonomy of Primary Prostate Cancer. Cell. 2015; 163(4): 1011-25). Differential analysis between high-grade (Gleason > 6) and low-grade (Gleason = 6) and between high-grade and benign was performed following the limma+voom procedure (Law CW, et al., voom: precision weights unlock linear model analysis tools for RNA-seq read counts. Genome Biol. 2014; 15(2): R29). Candidate genes were hand-picked by evaluating the logFC and p-value of both comparisons. Several genes of known prostate cancer biomarkers were also included (Table 5). A total of 53 genes and 1 gene fusion (TMPRSS2-ERG) were selected. Figure 5

[0252] Patient cohort

[0253] Figure 1B Table 1 gives information on the training (University of Michigan) and validation (NCI-EDRN) cohorts.

[0254] Of the 815 participants, qPCR produced valid results for 761 (93%). The median age was 63 years (IQR 58-68), the median PSA was 5.6 ng / mL (IQR 4.6-7.2), and 163 patients (21%) had a prior negative biopsy result (Table 1). Based on the study biopsy, 293 men (39%) had GG > 2 cancer. The accuracy of each candidate gene was quantified by elastic net mathematical models (Table 4).

[0255] The final MPS2 model included standard clinical variables and the 17 most informative markers, including 13 from the discovery analysis (four high-grade specific [APOC1, B3GNT6, NKAIN1, SCHLAP1] and nine prostate cancer specific [PCGEM1, SPON2, TRGV9, PCA3, OR51E2, CAMKK2, TFF3, PCAT14, TMSB15A]), four curated markers (HOXC6, ERG, TMPRSS2:ERG, KLK4), and the reference gene KLK3. Model coefficients were determined across the entire cohort ​

[0256] 10 (Table 6). Calibration and internal cross-validation were performed Figure 2C and Figures 3C-3D and the MPS2 model was locked for external validation.

[0257] Urine RNA extraction and cDNA synthesis

[0258] RNA isolation for MPS2 analysis was performed using the mirVana Total RNA Isolation Kit (ThermoFisher ) according to the manufacturer’s instructions. Briefly, 500 pL of urine / Hologic urine transport medium 1 : 1 mixture was mixed with the lysis binding mix (component of the mirVana Total RNA Isolation Kit from ThermoFisher ). Then, the binding bead mix (component of the mirVana Total RNA Isolation Kit from ThermoFisher ) was added to enrich nucleic acids in the urine sample, followed by TURBO digestion and washing. Finally, RNA was eluted. For high-throughput extraction of urine RNA, urine samples were processed by semi-automated KingFisher Flex (ThermoFisher ). After RNA extraction, cDNA was synthesized by using 16 pL of RNA using the SuperScript IV Master Mix (ThermoFisher ) followed by pre-amplification by using the Master Mix (ThermoFisher ). Analysis

[0259]

[0260] The TaqMan® Array CGH 2.0 Kit (ThermoFisher ) is a high-throughput real-time PCR genotyping method that allows rapid screening of multiple assays in multiple samples. This real-time method involves the use of an array composed of 3072 wells run on a TaqMan® 12K Flex Real-Time PCR System with a TaqMan® Fast 480 Module. For each sample, 2.5 pL of pre-amplified cDNA and 2.5 pL of 2x TaqMan® Universal PCR Master Mix were manually mixed and according to the manufacturer’s instructions (ThermoFisher

[0261] ​​​​​​The instruction manual was inserted into the 384-well plate. 12K Flex The system transfers the previously generated mixture to OpenArray board. Use 12K Flex Real-Time PCR System (ThermoFisher) Amplification was performed using the instrument and QuantStudio 12K Flex software (ThermoFisher). The expression was analyzed using the ΔΔCt method.

[0262] Data preprocessing

[0263] In order to To remove obvious outliers and resolve undetected data points in the three technical replicates of qPCR, the following data preprocessing steps were established: 1) If Ct = "Undetermined" or Amp.Status = "Undetermined / No Amp", Ct was set to 35; 2) The standard deviation (SD) of the three replicates was calculated; 3) If SD >= 1, the replicate with the largest difference from the mean was removed; if SD < 1, all three replicates were retained; 4) The mean Ct was calculated from the remaining two or three replicates. All Ct means were normalized using KLK3, using the equation - [mean Ct of gene X - mean Ct of KLK3]. The normalized Ct was used for downstream model construction. Since no exponential operation was applied, the normalized data were scaled logarithmically. Because samples with low KLK3 may indicate invalid DRE and may be unreliable, a 95th percentile threshold was set to remove samples with high KLK3 mean Ct.

[0264] Mathematical model construction

[0265] To avoid multicollinearity in the regression models, a stepwise procedure was used to identify and remove highly correlated variables. Specifically, the variance inflation factor (VIF) was calculated for all variables (54 probe gene expression + clinical variables including PSA density and prostate volume) and the variable with the highest VIF was removed; the VIF was recalculated with the remaining variables and this step was repeated until no variable had a VIF > 5. Nine probes (including probes targeting COL9A2, PLA2G7, HPN, CYB561A3, PDLIM5, MYO6, GAPDH, GDF15, and one of the two probes targeting PCA3) and PSA density were removed in the pre-filtering step. To select important genes, 3 model construction strategies were evaluated, including logistic regression with stepwise feature selection (James DA, et al., Modern applied statistics with S-PLUS. Technometrics. 1996; 38(1): 77), logistic regression with recursive feature elimination, and regularized logistic regression with elastic net (Friedman J, et al., Regularization Paths for Generalized Linear Models via Coordinate Descent. J Stat Softw. 2010; 33(1): 1-22). The model construction steps were implemented using the glmStepAIC (forward and backward), rfe, and glmnet functions in the R package "caret" (Kuhn M, et al., caret: Classification and Regression Training. R package version 6.0-86. Astrophysics Source Code). Elastic net was considered a mathematical model with built-in feature selection because unimportant variables were assigned zero importance. Repeated cross-validation (10-fold repeated 3 times) was used to evaluate the performance of the mathematical models. During training, the small class (high grade, 39%, Table 1) was up-sampled to create a balanced class. For stepwise and RFE, feature selection was considered part of the mathematical model construction and was encapsulated within each fold Figure 2A ).

[0266] Elastic net was chosen for the final mathematical model construction because it showed the best performance in terms of median AUC across the 30 resamplings. To select a robust gene panel, information from multiple elastic net regression models built by resampling was integrated to develop a final mathematical model in an ensemble approach Figure 2B). Specifically, the training data were first randomly split into 4 partitions and this step was repeated 4 times to generate a total of 40 resamples. Elastic net regression models were fitted to the 40 subsamples of the training data; then the frequency and importance of each gene were added together across all resamples. Genes were subsequently ranked by selecting the frequency and total importance and the top 17 genes were selected to be included in the final model. Based on the preliminary analysis of the optimal feature size using RFE and the number of genes was determined by the optimal design of the plate. The coefficients of each of the 17 genes were estimated by fitting the entire training data using an elastic net regression model (referred to as "MPS2"). In addition, an enhanced model was constructed by incorporating the 17 genes and clinical variables (age, race, family history, abnormal DRE, and prior negative biopsy) (referred to as "MPS2c"). In addition to the above variables, prostate volume was added to construct a third model that can be used when this information is available (referred to as "MPS2cv" or "MPS2+").

[0267] Mathematical model calibration

[0268] Calibration curves were used to assess the agreement between the predicted probability and the observed prevalence of disease in each bin. Logistic regression is considered a well-calibrated classifier. However, calibration is necessary when there is a distributional bias between the training and validation samples. In this study, the classes were balanced during training, while the validation cohort was a consecutive cohort and its 20% high-grade prevalence reflects an unbalanced real-world distribution. Without calibration, the calibration curve ( Figure 7 ) shows that the overall risk is overestimated. Therefore, it is critical to calibrate the model so that the prediction for each patient reflects the true risk of the real population. Two calibration techniques were tested in the resampled UM cohort - large-scale re-scaling (re-estimating the model intercept) and re-scaling (re-estimating the intercept and slope) (Vergouwe Y, et al., A closed testing procedure to select an appropriate method for updating prediction models. Stat Med. 2017;36(28):4529-39) to match the 20% high-grade in the validation cohort. Calibration curves were generated from the raw probability and the calibrated probability. In the high-probability bins, re-scaling was better and all subsequent calibrations were performed using this method.

[0269] Model validation

[0270] The raw data of the validation cohort were pre-processed in the same way as the training cohort. Using the normalized Ct values of the validation cohort, predictions were made for the three models (MPS2, MPS2c, and MPS2cv) by the “predict” function of caret (Kuhn M, et al., caret: Classification and Regression Training. R package version 6.0-86. Astrophysics Source Code.). Then, the model predictions were calibrated using the intercept and slope of each model, respectively, estimated as described in the “Model Calibration” section above.

[0271] Statistical analysis

[0272] Statistical analysis was performed using R version 4.1 (R Core Team. R: A language and environment for statistical computing. Vienna, Austria: R Foundation for Statistical Computing; 2013 2013). Kruskal-Wallis test was used for comparison between groups, and p-value < 0.05 was considered statistically significant. Regularized logistic regression with elastic net was used to build the model to predict high-grade prostate cancer, and the “glmnet” wrapper function provided by caret was used (Kuhn M, et al., caret: Classification and Regression Training. R package version 6.0-86. Astrophysics Source Code.). The diagnostic potential was visualized by receiver operating characteristic (ROC) curves and quantified by the area under the curve (AUC) using the R package pROC (Robin X, et al., pROC: an open-source package for R and S+ to analyze and compare ROC curves. BMC Bioinformatics. 2011; 12: 77). Calibration analysis was performed using the calibration function of caret by setting cuts = 8 (number of bins). Decision curve analysis (DCA) was performed using the dca function of dcurves (Sjoberg DD. dcurves: decision curve analysis for model evaluation, 2021).

[0273] The MPS2 model provided the highest net clinical benefit across all tests in the range of clinically relevant thresholds of about 4% to 20% Figure 10A ). The threshold probability (x-axis) reflects the patient's and clinician's assessment of the potential clinical outcome. For example, if a patient has a risk of about 4% or greater of having clinically significant prostate cancer, the threshold probability for selecting a biopsy is about 4%. A threshold probability of about 4% for clinically significant prostate cancer represents the opposite of the population at risk, such as a young man with a long life expectancy. In practical terms, this means that the clinician is willing to perform up to 20 biopsies to detect an additional case of clinically significant prostate cancer. On the other hand, a threshold probability of 20% applies to a patient who selects a biopsy only when the risk of clinically significant prostate cancer is >20%. This population values avoiding a biopsy very highly and is willing to accept a higher risk of delaying detection of clinically significant prostate cancer. The units of net benefit (y-axis) are true positives. A net benefit of 0.15 corresponds to a situation in which 15 out of 100 patients are biopsied based on the use of the test, and all 15 patients are found to have clinically significant prostate cancer. The plots were made with ggplot2 (Wickham H. Springer; New York: 2009. Ggplot2: elegant graphics for data analysis).

[0274] Results

[0275] Results include urinary transcript markers of high-grade prostate cancer.

[0276] Additional cancer and high-grade prostate cancer-specific transcripts were identified and added to the MPS test (Tomlins SA, et al., Urine TMPRSS2: ERG fusion transcript stratifies prostate cancer risk in men with elevated serum PSA. Sci Transl Med. 2011; 3(94):94ra72; Tomlins SA, et al., Urine TMPRSS2: ERG Plus PCA3 for Individualized Prostate Cancer Risk Assessment. Eur Urol. 2016; 70(1):45-53). For transcript 4 nominations, RNA-seq data from the Cancer Genome Atlas (TCGA) prostate adenocarcinoma (PRAD) cohort were analyzed (see Methods above). In short, biomarker discovery was performed using RNA sequencing (RNA-seq) data from 220 benign prostates, 71 GG1, and 484 GG≥2 cancers, which were available from the Cancer Genome Atlas (TCGA), the Genotype Tissue Expression (GTEx) portal, and the University of Michigan (UM). Forty-four transcripts that met the predetermined nomination criteria were supplemented with 10 selected cancer-related genes. The analysis yielded a panel of 54 biomarkers, including two reference genes and PCA3 and T2:ERG gene fusions from the original MPS assay (Table 5). Figures 4A-4D The study showed the relationship between the expression of these genes and an increase in Gleason scores.

[0277] Then, develop based on customization A qPCR platform (see Methods) was used to detect these transcripts in urine samples collected immediately after a digital rectal examination (DRE). Training cohort patients ( Figure 1B Table 1) comprises men who underwent prostate biopsies at the University of Michigan (UM). Of the initial 921 patients included in the UM training cohort, 761 men had available clinical variables from the Prostate Cancer Prevention Trial (PCPT), PSA <10 ng / mL, available prostate volume, and a threshold cycle (Ct) of KLK (PSA) transcripts <27 in their urinary biopsies. Figure 1B Of these 761 patients, 293 (38.5%) were found to have GG≥2 prostate cancer during biopsy (Table 1).

[0278] After QPCR analysis of 54 markers in the training cohort urine samples, MPS2 model development was performed Figure 2A ). An important aspect of model development was dimension reduction to reduce model complexity and avoid overfitting. After pre-filtering using variance inflation factor (VIF) > 5 to remove redundant variables, 46 genes remained. Three different model building algorithms with embedded feature selection were first evaluated on the training cohort using repeated cross-validation (CV), including logistic regression with stepwise feature selection (forward and backward), logistic regression with recursive feature elimination (RFE), and regularized logistic regression with elastic net. Elastic net has built-in feature selection because it provides feature importance that can be used for this purpose. The median area under the curve (AUC) was highest for the elastic net regression method Figure 6 ), so this method was chosen to build the final MPS2 mathematical model. Using an ensemble approach, i.e., by resampling, integrating data from multiple mathematical models, the development set was randomly divided into four partitions, and for each partition the model that produced the highest AUC was identified. This was repeated ten times using different random seeds, resulting in a total of 40 elastic net models. The frequency of inclusion for each model was tabulated along with the importance for the detection of clinically significant prostate cancer. Based on the analysis of the optimal feature size and OpenArray TM platform technical features, 17 biomarkers that had the best discriminative accuracy for GG > 2 prostate cancer were included in the MPS2 and MPS2+ (plus prostate volume) models along with the standard clinical variables and the normalizing gene KLK3. The models were calibrated and internally cross-validated (see Figure 1B ) before external validation was performed. Model performance was evaluated using repeated CV starting with the 54 MPS2 genes, as well as the MPS2 genes plus clinical variables (with and without prostate volume). By including more genes, the prediction of high-grade prostate cancer was improved compared to PCA3 and T2:ERG in terms of AUC only (0.784 vs 0.731) Figure 3A and 3B ). The inclusion of clinical variables (without prostate volume) increased the AUC to 0.802 Figure 3C ), and the inclusion of prostate volume increased the AUC to 0.820 Figure 3D .

[0279] It was observed that certain genes were frequently selected in different CV folds, while the selection of other genes was more random. To select a robust set of genes, the training data with all genes was partitioned into four subsamples, and the data partitioning was repeated 10 times using different random seeds, resulting in a total of 40 subsamples Figure 2BEach subsample was trained using an elastic network, and the top 17 genes were selected for the final mathematical model based on importance and frequency (Table 3). The final MPS2 model was constructed using the 17 genes from the entire training cohort, and the MPS2c and MPS2cv models were constructed by adding clinical variables with and without prostate volume, respectively. Calibration curves after class imbalance correction showed that the predicted risk was consistent with the observed risk. Figure 3E Next, the performance of the locked 17-transcript MPS2 model was tested on a validation cohort consisting of a blinded, multi-institutional National Cancer Institute-Early Detection Research Network (NCI-EDRN) prostate biopsy cohort. Figure 1B (Table 1). Of the 743 final patients included in the validation cohort, 20.3% had GG≥2 prostate cancer based on biopsy (Table 1). For the MPS2 model, although logistic regression is considered a well-calibrated classifier, calibration is necessary when there is a distribution shift between the training and validation populations. The classes in this study were balanced during training, while the validation cohort was a continuous cohort, and its 20% high-grade prevalence reflects an imbalanced true distribution. The MPS2 calibration curve (without calibration) Figure 7 This indicates that the overall risk is overestimated. Therefore, model calibration is crucial to ensure that predictions for each patient reflect the true risk in the real population, as described in the Methods section above (see also Tables 6, 7, and 8).

[0280] like Figure 4A As shown, the final MPS2 model outperforms the original MPS model, with values ​​similar to those obtained in the training cohort. Compared to the original MPS model's AUC of 0.730, the AUC values ​​for MPS2, MPS2c, and MPS2cv are 0.750, 0.807, and 0.818, respectively. The final calibration curves for each model show that the predicted risk is consistent with the risk observed in the validation cohort. Figure 4B ). Decision curve analysis also demonstrates the net benefit of the MPS2 model relative to "full treatment" or "no treatment" at different probability thresholds. Figure 4C Furthermore, it calculates the interventions (biopsy) avoided by each model at different probability thresholds. Figure 4D (Table 4).

[0281] Of the 859 men who participated in the PCA3 trial, 46 (5.4%) were ineligible for the current analysis due to insufficient urine volume or lack of clinical data. Of the 813 validation patients ( Figure 11), 743 (91%) were qPCR successful. Median PSA was 5.6 ng / mL (IQR 4.1-8.0) and 247 men (33%) had prior negative biopsy (Table 1). Based on study biopsy, 151 men (20%) had GG >2 PCa. Median MPS2 value was significantly higher in men with GG >2 PCa than in men with negative biopsy and men with GG1 PCa (0.44 vs. 0.08 and 0.20, respectively; both p<0.001) (Table 1, Figure 8A ). Similarly, median MPS2+ was significantly higher in men with GG >2 cancer than in men with negative or GG1 biopsy (0.54 vs. 0.08 and 0.25, respectively; p<0.001, Figure 8B ). AUC for GG >2 cancer was 0.60 for PSA, 0.66 for PCPTrc, 0.77 for PHI, 0.76 for dmx2, 0.72 for dmx3, and 0.74 for MPS, compared to 0.81 for MPS2 and 0.82 for MPS2+ ( Figure 9 ). Observed prevalence of GG >2 cancer was very close to the predicted probability by MPS2 and MPS2+ ( Figure 2C ), reflecting good calibration. Critically, MPS2 model was particularly well calibrated for predicted probability <30%.

[0282] Using a test threshold to detect 95% of GG >2 PCa (i.e., 95% sensitivity), the proportion of unnecessary biopsies avoided using each test was: 11% for PSA, 20% for PCPTrc, 26% for PHI, 27% for dmx2, 17% for dmx3, and 23% for MPS, compared to 37% for MPS2 and 41% for MPS2+. Table 2 lists the full performance metrics and unnecessary biopsies avoided. Critically, MPS2 and MPS2+ provided 99% sensitivity and 99% NPV for GG >3 PCa.

[0283] The initial biopsy population included 496 patients with median PSA of 5.0 ng / mL (IQR 3.8-6.6) (Table 7). Based on study biopsy, 133 patients (27%) had GG >2 cancer. Using a 95% sensitivity threshold, the proportion of unnecessary biopsies avoided was 15% for PSA, 27% for PCPTrc, 30% for PHI, 30% for dmx2, 17% for dmx3, and 27% for MPS, compared to 35% for MPS2 (Table 2). Although patients with initial biopsy can not have sufficient prostate volume, 42% of unnecessary biopsies would be avoided using MPS2+.

[0284] The repeat biopsy cohort included 247 men with a median PSA of 7.2 ng / mL (IQR 5.5-9.8), of which 18 (7.3%) were found to have GG > 2 prostate cancer (Table 7). The sensitivity for GG > 2 cancer was 95% with MPS2, compared to 46% for MPS2 and 51% for MPS2+, and the proportion of unnecessary biopsies avoided was 15% for PSA, 8.7% for PHI, 14% for dmx2, 16% for dmx3, and 15% for MPS (Table 2). Thus, the MPS2 test would avoid approximately half of the unnecessary biopsies while maintaining 95% detection of GG > 2 prostate cancer. Subgroups provided the performance of the MPS2 model with and without clinical factors (Tables 8-9). The MPS2 model provided the highest net clinical benefit across all tests within the 5% to 20% range of clinically relevant thresholds Figure 10A ). Expressing benefit as net reduction in unnecessary biopsies, MPS2 provided the greatest net reduction in unnecessary biopsies and did not miss a single patient with GG > 2 prostate cancer for biopsy Figure 10B

[0285] Converting the sequencing-based findings into a scalable qPCR platform, the test provided herein includes 17 Pca markers as well as markers uniquely overexpressed by high-grade cancer. Three MPS2 models were developed, including the 17 biomarkers alone or with clinical data (MPS2c) and prostate volume (MPS2cv). Validation of the MPS2 models in a blinded external cohort demonstrated that the models improved the diagnostic accuracy of the original MPS model as well as improved specificity. The MPS2 test had a sensitivity of 95% for GG > 2 cancer with NPV of 95-99% and specificity of 35-51% across subgroups. For individual patients, the near 100% NPV provides clear guidance to make confident decisions. For clinicians, uniform use of MPS2 can avoid up to half of the unnecessary biopsies while preserving the immediate detection of 95% of GG > 2 cancers diagnosed under the "total biopsy" approach. Critically, MPS2 provides 99% sensitivity and 99% NPV for GG > 3 cancer, meaning that the rare false negative MPS2 result is almost uniformly more favorable for the GG2 cancer that is least likely to metastasize.

[0286] In summary, the results demonstrate that the MPS2 test can improve the detection of clinically meaningful prostate cancer and can be used to identify the prostate cancer patients most likely to benefit from more aggressive treatment.

[0287] Table 1. Overall characteristics of the development and validation cohorts and characteristics stratified by prostate biopsy pathology.

[0288]

[0289] ​Table 2. Performance of PSA, PCPTrc, PHI, dmx2, dmx3, MPS, MPS2, and MPS2+ in the EDRN validation cohort: overall (N=743), initial biopsy (N=496), and repeat biopsy (N=247) subgroups.

[0290]

[0291] Abbreviations: MPS2, MyProstateScore 2.0; MPS2+, MyProstateScore 2.0 plus; dmx2, derived multiplex 2 gene model (HOXC6, DLX1); dmx3, derived multiplex 3 gene model (PCA3, ERG, SPDEF); PHI, Prostate Health Index; PSA, prostate specific antigen.

[0292] Table 3. Genes selected for the MPS2 predictive model of high-grade prostate cancer. Frequency of inclusion and cumulative importance of the 17 most informative markers among 40 elastic net models evaluated in the development.

[0293]

[0294] a Cumulative importance represents the total relative weight of marker importance across repeated samplings derived by elastic net models.

[0295] Table 4. Performance metrics for each threshold in the validation cohort. PPV, positive predictive value; NPV, negative predictive value.

[0296]

[0297] Table 5. Information of the 54 genes included in the OpenArray assay.

[0298]

[0299] Table 6. Model coefficients for the MPS2 and MPS2+ models.

[0300] Covariate MPS2 a ]]> MPS2+ b ]]> (intercept) 5.902430658 6.676363594 T2ERG 0.111906862 0.148584627 SCHLAP1 0.17335791 0.205268829 OR51E2 0.200676934 0.228791882 APOC1 -0.07916931 -0.08896388 PCAT14 0.14420976 0.16009867 CAMKK2 -0.26364401 -0.277941 927 PCA3.1 0.080881661 0.074209893 NKA1N1 -0.06946207 -0.093791082 B3GNT6 0.047475092 0.072524885 TFF3 0.186395669 0.2128103 SPON2 0.156664808 0.1740959 PCGEM1 -0.16940833 -0.149084289 TRGV9 0.096184103 0.177972309 TMSB15A 0.151071453 0.214870771 ERG 0.023544761 0.030085251 KLK4 0.149451849 0.214609188 HOXC6 0.05612131 0 Age 0.000134221 0.021446485 African American 0.828856591 1.232493234 Family history 0.148709502 0.292757369 Abnormal DRE 0.888379309 1.094432439 Prior biopsy -0.8505938 -0.61694213 PSA 0.073709982 0.092554335 Prostate volume N / A -0.024051593

[0301] a MPS2: calibrated logit = -1.453526 + logit * 1.302089

[0302] b MPS2+: calibrated logit = -1.41207 + logit * 1.077061

[0303] Table 7. Characteristics of the NCI-EDRN external validation population stratified by prior biopsy status.

[0304]

[0305] Abbreviations: DRE, digital rectal examination; GG, indicates Gleason grade group; IQR, interquartile range; MPS, MyProstateScore; MPS2, MyProstateScore 2.0; MPS2+, MyProstateScore 2.0 plus; PCA3, prostate cancer antigen 3; PHI, Prostate Health Index; PSA, prostate specific antigen

[0306] a Measured by transrectal ultrasound.

[0307] b PSA density is equal to serum PSA divided by prior prostate volume.

[0308] c MPS2 and MPS2+ values are reported on a continuous scale as the likelihood of having a clinically significant prostate cancer detected on biopsy.

[0309] Table 8. Clinical performance of high-sensitivity MPS2 thresholds in the initial biopsy subpopulation of the external validation cohort (N=496).

[0310] Threshold Sensitivity Specificity NPV PPV MPS2+ 0.05 g7% 21% 95% 31% 0.06 97% 25% 96% 32% 0.07 96% 29% 95% 33% 0.075 96% 31% 96% 34% 0.08 95% 32% 95% 34% 0.09 95% 35% 95% 35% 0.10 95% 38% 95% 36% 0.11 a ]] 95% 42% 96% 37% 0.12 92% 44% 94% 38% 0.13 92% 48% 95% 39% 0.14 92% 51% 94% 41% 0.15 90% 53% 94% 41% MPS2 0.05 96% 21% 94% 31% 0.06 96% 25% 95% 32% 0.07 96% 28% 95% 33% 0.075 96% 31% 96% 34% 0.08 95% 33% 94% 34% 0.087 a ]] 95% 35% 95% 35% 0.09 94% 36% 94% 35% 0.10 94% 39% 95% 36% 0.11 93% 41% 94% 37% 0.12 92% 46% 94% 38% 0.13 91% 49% 94% 39% 0.14 89% 52% 93% 41% 0.15 89% 54% 93% 41% Marker only 0.05 95% 20% 92% 30% 0.06 95% 28% 94% 33% 0.07 95% 33% 94% 34% 0.075 95% 35% 95% 35% 0.077 a ]] 95% 35% 95% 35% 0.08 94% 37% 94% 35% 0.09 92% 41% 93% 36% 0.10 90% 44% 92% 37%

[0311] The optimal threshold provided 95% sensitivity for GG2 or higher grade cancer.

[0312] Table 9. Clinical performance of high-sensitivity MPS2 thresholds in the repeat biopsy subpopulation of the external validation cohort (N=247).

[0313] Threshold Sensitivity Specificity NPV PPV MPS2+ 0.04 100% 40% 100% 12% 0.05 94% 47% 99% 12% 0.054 94% 49% 99% 13% 0.058 a ]] 94% 51% 99% 13% 0.06 88% 52% 98% 13% MPS2 0.038 100% 41% 100% 12% 0.04 94% 42% 99% 11% 0.044 a ]] 94% 46% 99% 12% 0.05 89% 48% 98% 12% 0.054 89% 49% 98% 12% 0.06 89% 52% 98% 13% Marker only 0.04 100% 22% 100% 9.1% 0.05 100% 26% 100% 9.6% 0.054 100% 29% 100% 10% 0.06 100% 32% 100% 10% 0.07 100% 35% 100% 11% 0.08 a ]]> 94% 42% 99% 11%

[0314] a The optimal threshold provided 94.4% sensitivity for GG2 or higher grade cancer.

[0315] Table 10. MPS2 model coefficients redeveloped using the highest grade group in all biopsy and surgical specimens.

[0316]

[0317] Abbreviations: DRE, digital rectal examination; MPS2, MyProstateScore 2.0; MPS2+, MyProstateScore 2.0 plus; PSA, prostate specific antigen; RP, radical prostatectomy and / or repeat biopsy.

[0318] a Considering the potential misclassification of clinically significant prostate cancer due to under-sampling of the biopsy, we evaluated the MPS2 model including pathology data obtained after study urine collection (e.g., repeat biopsy, radical prostatectomy). Of the 761 patients in the development set, 382 (50%) had serious prostate cancer. The table includes model parameters based on the highest cancer grade examined (i.e., RP-derived model). Of the 17 informative markers in the MPS2 model, 13 were retained in the RP-derived model. The direction of the biomarker-outcome association for all markers remained unchanged. The AUC difference for the cross-validated models was 1% for MPS2 (0.802 vs. 0.792) and 0.1% for MPS2+ (0.821 vs. 0.822).

[0319] Example 2

[0320] Sample collection and mRNA detection

[0321] A urine sample is collected from a subject suspected of having prostate cancer or at risk of developing prostate cancer to determine the likelihood of detecting a prostate cancer with a Gleason grade group >2 from a prostate biopsy of the subject. The subject is subjected to a digital rectal exam (DRE) and a urine sample is collected within about 1 hour after the DRE. The urine is placed into a collection tube with a stabilizing buffer at a volume ratio of buffer to sample of about 2:5.

[0322] Positive and negative controls are prepared. The negative control is a sample from a subject in the "low risk category" previously reported. The positive control is a sample from a subject in the "high risk category" previously reported.

[0323] RNA is isolated from the test sample, negative control, and positive control using a commercially available RNA isolation kit. The extracted RNA is subjected to RT-PCR to generate cDNA, pre-amplification, and qPCR to determine the amount of mRNA expressed by each of the following genes: TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6. The amount of mRNA expressed by a reference gene, such as KLK3, can also be determined using RT-PCR, pre-amplification, and qPCR. The cDNA is amplified using target-specific primers and gene-specific probes that release a fluorophore are used to accurately and quantitatively measure the expression level of the above target genes.

[0324] Analysis and MPS2 score generation

[0325] Based on the qPCR performed above, the average Crt value (cycle threshold) for each target gene is determined and normalized to the average Crt value of the reference gene (e.g., KLK3) using the following equation: Crt average (target) - Crt average (reference). The normalized Crt is multiplied by a gene-specific coefficient. The following table provides exemplary gene-specific coefficients. The sum of normalized Crt x coefficient = logit value. The logit value is recalibrated using the intercept and slope. The logit value is converted to a score using the logit equation. The gene-specific coefficients, slope, and intercept are different for subjects who have a first biopsy performed or subjects who have a prior negative prostate biopsy.

[0326] Table 11

[0327]

[0328] An exemplary calculation for subjects who have a first biopsy performed is as follows:

[0329] 1. Normalize each Crt average for each of the 17 targets to KLK3 = (target gene Crt average - KLK3 Crt average)

[0330] 2. Normalize value for specified target * coefficient

[0331] 3. Sum of (2) for all targets) + intercept

[0332] 4. (intercept + (3)) * (slope) = logit

[0333] 5. MPS2 probability = (Exp(4)) / Exp(4) + 1)

[0334] An exemplary calculation for subjects who have a prior negative prostate biopsy is as follows:

[0335] 1. Normalize each Crt average for each of the 17 targets to KLK3 = (target gene Crt average - KLK3 Crt average)

[0336] 2. Add the following clinical variables and prostate volume to (1):

[0337] a. Age

[0338] b. African American (binary)

[0339] c. Negative biopsy = 1

[0340] d. Abnormal DRE (binary)

[0341] e. Family history (binary)

[0342] f. Serum PSA

[0343] g. PSA volume

[0344] 3. Normalized value * coefficient of specified target + sum of (2a-2g)

[0345] 4. Sum of (3) for all targets + intercept

[0346] 5. (Intercept + (3)) * (slope) = logit

[0347] 6. MPS2 probability = (Exp(4)) / Exp(4) + 1)

[0348] Thresholds for determining low risk or high risk are different for subjects undergoing initial biopsy and subjects with prior negative prostate biopsies:

[0349] Table 12.

[0350]

[0351] Based on the above results, a report is generated with a score and risk category and provided to the subject's healthcare provider (e.g., the subject's urologist) who requested the test.

[0352] All publications, patents, patent applications and accession numbers mentioned in the above specification are herein incorporated by reference in their entirety. Although the foregoing application has been described in some detail by way of illustration and example for purposes of clarity and understanding, it is readily apparent to those of ordinary skill in the art in light of the teachings of this application that various changes and modifications can be made thereto without departing from the spirit of the application. It is therefore intended that this application not be limited to the exact compositions and methods described above, but that described be susceptible to various modifications and embodiments within the scope of the following claims.

Claims

1. A method of treating prostate cancer, the method comprising: a) determining the expression level of one or more genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6 in a sample from a subject diagnosed with prostate cancer; and b) administering a prostate cancer treatment to a subject identified as having an altered expression level of the genes relative to a subject not having prostate cancer or a subject having low grade prostate cancer.

2. A method of characterizing, prognosticating, or recommending a prostate cancer treatment, the method comprising: a) determining the expression level of one or more genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6 in a sample from a subject diagnosed with prostate cancer; and b) identifying the subject as having high grade prostate cancer when the subject is identified as having an altered expression level of the genes relative to a subject not having prostate cancer or a subject having low grade prostate cancer.

3. A method for informing prostate cancer survival outcome, the method comprising: a) detecting an expression amount of at least three genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6, wherein the expression amount is present in a urine of a subject; b) determining a score based on the expression amount, wherein the score is associated with or informs a likelihood of the subject having or developing prostate cancer of a grade group > 2; and.

4. A method for identifying a subject having a high likelihood of having or developing a prostate cancer with a Gleason score > 2, the method comprising detecting the amount of expression of at least three genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6, wherein, the expression amount is present in the urine of the subject and indicates whether the subject has a high likelihood of having prostate cancer of a grade group > 2 with a diagnostic accuracy (AUC) of > 0.

75.

5. A method for identifying the likelihood of detecting a Gleason Score > 2 prostate cancer from a prostate biopsy of a subject, the method comprising detecting the amount of expression of at least three genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6, wherein, the expression amount is present in the urine of the subject and indicates a likelihood of detecting prostate cancer of a grade group > 2 from a prostate biopsy of the subject with a diagnostic accuracy (AUC) of > 0.

75.

6. The method of any one of claims 2 to 5, further comprising administering a prostate cancer treatment to the subject.

7. The method of claim 1, wherein, the subject has high grade prostate cancer.

8. The method of claim 7, wherein, the high grade prostate cancer is prostate cancer of a grade group > 2.

9. The method of any of the preceding claims, wherein, The method further comprises determining a score based on the expression level or amount, wherein the score indicates a likelihood that the subject has or develops prostate cancer that is graded and / or staged ≥2.

10. The method of claim 9, further comprising generating a report comprising the score.

11. The method of claim 9 or 10, wherein, prostate cancer that is graded and / or staged ≥2 is determined by a prostate biopsy of the subject.

12. The method of any one of claims 3, 9, 10, or 11, wherein, The score has a diagnostic accuracy (AUC) of >0.

75.

13. The method of any one of claims 9-12, further comprising forwarding the report to the subject or a healthcare provider of the subject.

14. The method of any one of claims 1 or 8-14, wherein, The one or more genes are two or more genes.

15. The method of any of the preceding claims, wherein, The one or more genes or the at least three genes are five or more genes.

16. The method of any of the preceding claims, wherein, The one or more genes or the at least three genes are 10 or more genes.

17. The method of any of the preceding claims, wherein, The one or more genes or the at least three genes are TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6.

18. The method of any of the preceding claims, wherein, The method comprises determining the expression level of one to 20 additional genes.

19. The method of any of the preceding claims, wherein, The subject has no prior prostate biopsy.

20. The method of any of the preceding claims, wherein, The subject has a prior negative prostate biopsy result.

21. The method of any one of claims 9-20, wherein, One or more clinical variables are associated with the subject, and the method further comprises identifying at least one of the one or more clinical variables and determining the score based on at least one of the one or more clinical variables.

22. The method of claim 21, wherein, At least one of the one or more clinical variables is an age, race, family history of prostate cancer, digital rectal exam (DRE) result, prostate biopsy result, prostate specific antigen (PSA) expression value based on a serum sample, multi-parametric MRI (mpMRI) result, or any combination thereof of the subject.

23. The method of claim 22, wherein, The DRE or prostate biopsy of the subject is performed within 30 days of obtaining the urine sample.

24. The method of any one of claims 21-23, wherein, The subject has a prior negative prostate biopsy result and determining the score comprises performing: a) Equation 1 where x = Intercept + Slope((-1)((a)(CRT mean APOC1 - CRT mean Ref) + (b)(CRT mean B3GNT6 - CRT mean Ref) + (c)(CRT mean CAMKK2 - CRT mean Ref) + (d)(CRT mean ERG - CRT mean Ref) + (e)(CRT mean HOXC6 - CRT mean Ref) + (f)(CRT mean KLK4 - CRT mean Ref) + (g)(CRT mean NKAIN1 - CRT mean Ref) + (h)(CRT mean OR51E2 - CRT mean Ref) + (i)(CRT mean PCA3 - CRT mean Ref) + (j)(CRT mean PCAT14 - CRT mean Ref) + (k)(CRT mean PCGEM1 - CRT mean Ref) + (1)(CRT mean SCHLAP1 - CRT mean Ref) + (m)(CRT mean SPON2 - CRT mean Ref) + (n)(CRT mean TFF3 - CRT mean Ref) + (o)(CRT mean T2:ERG - CRT mean Ref) + (p)(CRT mean TMSB15A - CRT mean Ref) + (q)(CRT mean TRGV9 - CRT mean Ref)) + ((r)(Age) + (s)(Family Hx) + (t)(Abnormal DRE) + (u)(Bx Prior Neg) + (v)(PSA) + (w)(Prostate Volume))); or b) Equation 3 where x = intercept + slope((-1)((a)(CRT mean APOC1 - CRT mean KLK3) + (b)(CRT mean B3GNT6 - CRT mean KLK3) + (c)(CRT mean CAMKK2 - CRT mean KLK3) + (d)(CRT mean ERG - CRT mean KLK3) + (e)(CRT mean HOXC6 - CRT mean KLK3) + (f)(CRT mean KLK4 - CRT mean KLK3) + (g)(CRT mean NKAIN1 - CRT mean KLK3) + (h)(CRT mean OR51E2 - CRT mean KLK3) + (i)(CRT mean PCA3 - CRT mean KLK3) + (j)(CRT mean PCAT14 - CRT mean KLK3) + (k)(CRT mean PCGEM1 - CRT mean KLK3) + (1)(CRT mean SCHLAP1 - CRT mean KLK3) + (m)(CRT mean SPON2 - CRT mean KLK3) + (n)(CRT mean TFF3 - CRT mean KLK3) + (o)(CRT mean T2:ERG - CRT mean KLK3) + (p)(CRT mean TMSB15A - CRT mean KLK3) + (q)(CRT mean TRGV9 - CRT mean KLK3)).

25. The method of any one of claims 21-23, wherein, The subject has a first-time prostate biopsy and determining the score comprises performing: a) Equation 2 wherein x = intercept + slope((-1)((a)(CRT mean APOC1 - CRT mean reference) + (b)(CRT mean B3GNT6 - CRT mean reference) + (c)(CRT mean CAMKK2 - CRT mean reference) + (d)(CRT mean ERG - CRT mean reference) + (e)(CRT mean HOXC6 - CRT mean reference) + (f)(CRT mean KLK4 - CRT mean reference) + (g)(CRT mean NKAIN1 - CRT mean reference) + (h)(CRT mean OR51E2 - CRT mean reference) + (i)(CRT mean PCA3 - CRT mean reference) + (j)(CRT mean PCAT14 - CRT mean reference) + (k)(CRT mean PCGEM1 - CRT mean reference) + (1)(CRT mean SCHLAP1 - CRT mean reference) + (m)(CRT mean SPON2 - CRT mean reference) + (n)(CRT mean TFF3 - CRT mean reference) + (o)(CRT mean T2:ERG - CRT mean reference) + (p)(CRT mean TMSB15A - CRT mean reference) + (q)(CRT mean TRGV9 - CRT mean reference))); or b) Equation 4 where x = Intercept + Slope((-1)((a)(CRT Mean APOC1 - CRT Mean KLK3) + (b)(CRT Mean B3GNT6 - CRT Mean KLK3) + (c)(CRT Mean CAMKK2 - CRT Mean KLK3) + (d)(CRT Mean ERG - CRT Mean KLK3) + (e)(CRT Mean HOXC6 - CRT Mean KLK3) + (f)(CRT Mean KLK4 - CRT Mean KLK3) + (g)(CRT Mean NKAIN1 - CRT Mean KLK3) + (h)(CRT Mean OR51E2 - CRT Mean KLK3) + (i)(CRT Mean PCA3 - CRT Mean KLK3) + (j)(CRT Mean PCAT14 - CRT Mean KLK3) + (k)(CRT Mean PCGEM1 - CRT Mean KLK3) + (1)(CRT Mean SCHLAP1 - CRT Mean KLK3) + (m)(CRT Mean SPON2 - CRT Mean KLK3) + (n)(CRT Mean TFF3 - CRT Mean KLK3) + (o)(CRT Mean T2:ERG - CRT Mean KLK3) + (p)(CRT Mean TMSB15A - CRT Mean KLK3) + (q)(CRT Mean TRGV9 - CRT Mean KLK3)) + ((r)(Age) + (s)(Family Hx) + (t)(Abnormal DRE) + (u)(Bx Prior Neg) + (v)(PSA) + (w)(Prostate Volume))).

26. The method of claim 24 or 25, wherein, The performing comprises using a processor.

27. The method of any one of the preceding claims, wherein, The score has a diagnostic accuracy (AUC) of >0.

80.

28. The method of any one of claims 1 or 6-27, wherein, The prostate cancer treatment is one or more of: surgery, radiation therapy, hormone therapy, targeted therapy, chemotherapy, immunotherapy, radiopharmaceutical, and bone-modifying drugs.

29. The method of any one of the preceding claims, wherein, The expression level or amount is an amount of mRNA or protein expressed by the gene.

30. The method of any one of claims 1-2 or 7-29, wherein, The sample is selected from the group consisting of tissue, blood, plasma, serum, urine, prostate secretions, and prostate cancer cells.

31. The method of claim 30, wherein, The sample is urine, and the urine is obtained within 30 minutes after a DRE of the subject.

32. The method of any one of claims 9-31, further comprising determining a prostate volume of the subject and determining the score based on the prostate volume of the subject.

33. The method of claim 32, wherein, the score has a diagnostic accuracy (AUC) of > 0.

81.

34. The method of claim 32, wherein, The diagnostic accuracy of the score is 1-10% higher than a score determined by the expression level or amount of only PCA3 and TMPRSS2-ERG.

35. The method of any one of the preceding claims, wherein, Detecting the expression level or amount of the gene comprises detecting the mRNA expression amount of the gene.

36. The method of claim 35, wherein, Detecting the mRNA expression level or amount comprises reacting the sample or urine with a reagent composition comprising a polynucleotide reagent.

37. The method of claim 35, wherein, Detecting the mRNA expression level or amount comprises synthesizing a cDNA complementary to the mRNA expressed by the gene, amplifying the cDNA, and detecting the cDNA.

38. The method of any one of the preceding claims, further comprising detecting an expression level or amount of a reference gene, and normalizing the expression amount of the one or more genes or the at least three genes to the expression amount of the reference gene.

39. The method of claim 38, wherein, Detecting the expression level or amount of the reference gene comprises detecting the level or amount of mRNA expressed by the reference gene.

40. The method of claim 38 or 39, wherein, The reference gene is KLK3, CYPB561A3, EEFlA2, GAPDH, HPN, KLK2, KLK4, LBH, NUDT8, SPDEF, or TRGV9.

41. The method of claim 40, wherein, The reference gene is KLK3.

42. The method of any one of the preceding claims, wherein, The expression level or amount of the gene is different than the expression amount of the gene in a subject having or at risk of developing a prostate cancer of a grade group < 2 or a subject not having prostate cancer.

43. The method of any one of the preceding claims, comprising detecting the expression level or amount of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or 17 genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6.

44. The method of any one of the preceding claims, comprising detecting the expression level or amount of 1-10 genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6.

45. The method of any one of the preceding claims, comprising detecting the expression level or amount of 5-10 genes selected from TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6.

46. The method of any one of the preceding claims, comprising detecting the expression level or amount of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6.

47. The method of claim 46, wherein, The expression level or amount of each of APOC1, CAMKK2, NKAIN1, and PCGEM1 is lower in a subject at risk of having prostate cancer with a grade group of >2 than in a subject at risk of having or developing prostate cancer with a grade group of <2 or in a subject not having prostate cancer.

48. The method of claim 46 or 47, wherein, The expression level or amount of each of TMPRSS2-ERG, SCHLAP1, OR51E2, PCAT14, PCA3, B3GNT6, TFF3, SPON2, TRGV9, TMSB15A, ERG, KLK4, and HOXC6 is higher in a subject at risk of having prostate cancer with a grade group of >2 than in a subject at risk of having or developing prostate cancer with a grade group of <2 or in a subject not having prostate cancer.

49. A method for screening the amount of expression of at least three genes, the method comprising: a) reacting a urine sample from a human subject with a reagent for detecting the amount of expression of at least three genes, wherein the at least three genes are selected from TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6; and b) detecting the amount of expression of the at least three genes, wherein the amount of expression is present in the sample and the detecting comprises using an in vitro assay.

50. The method of claim 49, wherein, The in vitro assay is a nucleic acid amplification assay.

51. The method of claim 50, wherein, The nucleic acid amplification assay comprises performing a reverse transcription polymerase chain reaction.

52. A method for detecting the amount of mRNA expressed by at least three genes, the method comprising: a) synthesizing cDNA from mRNA expressed by the at least three genes and present in a urine sample from a human subject, wherein the at least three genes are selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6; b) amplifying the cDNA to provide amplified cDNA; and c) detecting the amplified cDNA, wherein the amplified cDNA is indicative of the amount of mRNA expressed by the at least three genes.

53. A method for detecting the amount of mRNA expressed by at least three genes, the method comprising: a) isolating nucleic acid from a first composition comprising urine from a human subject to provide isolated nucleic acid; b) reacting the isolated nucleic acid with a second composition comprising reagents for detecting the amount of mRNA present in the first composition and expressed by the at least three genes, wherein the at least three genes are selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6; and c) detecting the amount of mRNA expressed by the at least three genes.

54. The method of any one of claims 52-53, comprising detecting the amount of expression of 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or 17 genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6.

55. The method of any one of claims 52-54, comprising detecting the amount of expression of 3-10 genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6.

56. The method of any one of claims 53-54, comprising detecting the amount of expression of 5-10 genes selected from TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6.

57. The method of any one of claims 53-54, comprising detecting the amount of expression of each of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6.

58. The method of any one of the preceding claims, further comprising informing the subject or a healthcare provider of the subject about treatment options for prostate cancer stratified into a group of >2.

59. The method of any one of claims 10-58, wherein, the report comprises information about treatment options for prostate cancer stratified into a group of >2.

60. The method of any one of the preceding claims, further comprising providing the subject or a healthcare provider of the subject with instructions for administering a prostate cancer treatment stratified into a group of >2 to the subject.

61. A kit comprising: a container comprising a reagent composition for detecting the amount of expression of at least three genes; and instructions for detecting the amount of expression, wherein the amount of expression is present in a urine of a subject and the at least three genes are selected from TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6.

62. The kit of claim 61, wherein, the reagent composition comprises polynucleotide reagents for detecting the amount of mRNA expressed by the at least three genes.

63. The kit of claim 61 or 62, wherein, the reagent composition comprises polynucleotide reagents for detecting the amount of expression of a reference gene and the instructions are further for normalizing the amount of expression of the at least three genes to the amount of expression of the reference gene.

64. The kit of claim 63, wherein, the reference gene is KLK3, CYPB561A3, EEFlA2, GAPDH, HPN, KLK2, KLK4, LBH, NUDT8, SPDEF, or TRGV9.

65. The kit of claim 64, wherein, the reference gene is KLK3.

66. The kit of any one of claims 61-65, wherein, the instructions are further for generating a report comprising a score determined by the amount of expression of the at least three genes, wherein the score indicates a likelihood that the subject has or develops prostate cancer stratified into a group of >2.

67. The kit of claim 66, wherein, prostate cancer stratified into a group of >2 is determined by a prostate biopsy of the subject.

68. The kit of any one of claims 61-67, wherein, the subject has no prior prostate biopsy.

69. The kit of any one of claims 61-68, wherein, the subject has a prior negative prostate biopsy result.

70. The kit of any one of claims 61-69, wherein, one or more clinical variables are associated with the subject, and the instructions are further to determine a score based on at least one of the one or more clinical variables.

71. The kit of claim 70, wherein, at least one of the one or more clinical variables is the subject's age, race, family history of prostate cancer, digital rectal exam (DRE) results, prostate biopsy results, prostate specific antigen (PSA) expression values based on serum samples, multi-parametric MRI (mpMRI) results, or any combination thereof.

72. The kit of any one of claims 61-71, wherein, the instructions are further to determine a score based on the subject's prostate volume.

73. The kit of any one of claims 61-72, wherein, the genes are 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or 17 genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6.

74. The kit of any one of claims 61-72, wherein, the genes are 3-10 genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6.

75. The kit of any one of claims 61-72, wherein, the genes are 5-10 genes selected from the group consisting of TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6.

76. The kit of any one of claims 61-72, wherein, the genes are TMPRSS2-ERG, SCHLAP1, OR51E2, APOC1, PCAT14, CAMKK2, PCA3, NKAIN1, B3GNT6, TFF3, SPON2, PCGEM1, TRGV9, TMSB15A, ERG, KLK4, and HOXC6.

77. The kit of any one of claims 61-76, wherein, the instructions are further to inform the subject about treatment options for prostate cancer stratified into a group ≥ 2.

78. The kit of any one of claims 66-77, wherein, the report includes treatment options for prostate cancer stratified into a group ≥ 2.

79. The kit of any one of claims 66-78, wherein, the instructions are further to administer to the subject a treatment for prostate cancer stratified into a group ≥ 2.

80. The method of any one of claims 1-60, wherein, the method does not include performing a prostate biopsy on the subject.

81. The method of any one of claims 2-5, 9-10, 12-27, 29-58, or 81, wherein, the subject is spared an unnecessary prostate biopsy.

82. The method of any one of claims 1-60, wherein, the method further includes performing a prostate biopsy on the subject.

83. The method of any one of claims 1-60, wherein, the method further includes advising the subject or a health care provider of the subject that the subject undergo a prostate biopsy.

84. The method of claim 81 or 82, wherein, the prostate biopsy indicates that the subject has prostate cancer stratified into a group ≥ 2.

85. The method of claim 81 or 82, wherein, the prostate biopsy indicates that the subject does not have prostate cancer stratified into a group ≥ 2.

86. The method of any one of claims 2-60 and 80-83, wherein, The method further comprises administering to the subject a treatment for staging prostate cancer of > 2.

Citation Information

Patent Citations

  • radio (st222)

    CN3229253D

  • Strand displacement amplification using thermophilic enzymes

    EP0684315A1

  • Methods of amplifying and sequencing nucleic acids

    US20050130173A1

  • Single-primer nucleic acid amplification methods

    US20060046265A1

  • Process for amplifying, detecting, and / or-cloning nucleic acid sequences

    US4683195A