Method for providing cancer risk information using cox regression analysis algorithm
By employing a Cox regression analysis algorithm with a combination of predictive genes and signaling pathway scores, the method addresses the limitations of current prostate cancer diagnostic tools, providing enhanced accuracy in metastasis and recurrence prediction and improving treatment decisions.
Patent Information
- Application Number
- PCT/KR2024/018426
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-20
- Filing Date
- 2024-11-20
- Publication Date
- 2025-05-30
AI Technical Summary
Current diagnostic methods for prostate cancer, such as the PSA test and RNA expression level-based tools, have limited accuracy for early detection, metastasis prediction, and recurrence assessment, leading to difficulties in determining appropriate treatment strategies.
A method using a combination of predictive genes and signaling pathway scores applied to a Cox regression analysis algorithm to provide more precise information on cancer risk, specifically the possibility of metastasis or recurrence, thereby aiding in treatment decision-making.
This approach enhances the accuracy of prostate cancer prognosis prediction, allowing for early identification of metastasis and recurrence possibilities, which can lead to improved patient outcomes and survival rates.
Smart Images

Figure KR2024018426_30052025_PF_FP_ABST
Abstract
Description
A method for providing information on cancer risk using the Cox regression analysis algorithm.
[0001] The present invention relates to a method for predicting cancer risk by applying a Cox regression analysis algorithm.
[0002] Cancer is one of the most common causes of death worldwide. Approximately 10 million new cases occur each year, accounting for approximately 12% of all deaths, making it the third leading cause of death. Among various types of cancer, prostate cancer is the most common cancer among men worldwide and the second leading cause of death. It occurs primarily in men over 50 years of age and is characterized by a rapid increase in the number of patients with the disease with increasing age. While it typically progresses slowly, once it becomes malignant and metastasizes, it is extremely difficult to treat. Metastases typically begin in the lymph nodes, pelvic bones, spine, and bladder surrounding the prostate cancer, gradually spreading throughout the body.
[0003] Currently, primary methods for diagnosing prostate cancer include the prostate-specific antigen (PSA) test and digital rectal examination. Imaging modalities include transrectal ultrasound, CT, MRI, and WBBS (Whole Body Bone Scan). Biopsy is also performed. However, most of these methods have low diagnostic accuracy, make early diagnosis difficult, and struggle to detect metastases. Furthermore, they often struggle to distinguish between benign conditions like benign prostatic hyperplasia and prostatitis, making it difficult to differentiate between the two. Consequently, diagnosing prostate cancer by type is challenging.
[0004] Although prognostic tools based on RNA expression levels (e.g., Decipher, Prolaris, and oncotypeDX) have recently entered the market, they are technologies that utilize classical methods such as PCR and DNA microarray, and thus the number of markers that can be predicted simultaneously is limited to about 30. In addition, the prognostic accuracy of factors currently used for prognostic prediction, including the prostate cancer stage (TNM stage), is significantly lower than that of other cancer types. Furthermore, among prognostic prediction tools, there is a lack of tools that can precisely predict metastasis and the possibility of recurrence.
[0005] Accordingly, the inventors of the present invention have discovered a combination of biomarkers that can analyze whether prostate cancer has metastasized and enable more precise diagnosis, and by applying these combined markers to the Cox regression analysis algorithm, they aim to provide information on the cancer risk of each patient more easily when deciding on a treatment method.
[0006] Although various diagnostic technologies for prostate cancer have been developed, the accuracy of prognosis prediction is significantly lower than that of other cancers, and there is currently no diagnostic technology that accurately predicts the possibility of metastasis and recurrence.
[0007] The purpose of the present invention is to provide a method for providing improved information on the risk of cancer compared to existing diagnostic methods by applying a combination marker consisting of predictive genes and signaling pathway scores that can predict the possibility of metastasis or recurrence at an early stage and quickly determine a treatment method to a statistical algorithm.
[0008] Hereinafter, various embodiments described herein will be described with reference to the drawings. In the following description, various specific details, such as specific configurations, compositions, and processes, are set forth to provide a thorough understanding of the present invention. However, certain embodiments may be practiced without one or more of these specific details, or in conjunction with other known methods and configurations. In other instances, well-known processes and manufacturing techniques have not been described in specific detail so as not to unnecessarily obscure the present invention. Reference throughout this specification to "one embodiment" or "an embodiment" means that a particular feature, configuration, composition, or characteristic described in connection with the embodiment is included in one or more embodiments of the present invention. Thus, the appearances of "in one embodiment" or "an embodiment" in various places throughout this specification do not necessarily refer to the same embodiment of the present invention. Additionally, the particular features, configurations, compositions, or characteristics may be combined in any suitable manner in one or more embodiments.
[0009] Unless otherwise specifically defined herein, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention belongs.
[0010]
[0011] In this specification, "Next Generation Sequencing (NGS)" refers to a method that divides the genome into countless fragments, decodes and reassembles the genetic information of each fragment, and then analyzes the entire base sequence. It has the advantage of being able to analyze the base sequence of the genome at high speed, and is also called high-throughput sequencing, massive parallel sequencing, or second-generation sequencing. Compared to NGS, Sanger sequencing can also read the entire human genome, but it can only target known genes, which limits the scope of testing, and requires repeated experiments to examine multiple genes. For example, compared to Sanger sequencing, which requires approximately 3 million segments, next-generation sequencing is a significantly improved analysis method in terms of time and cost. Massively parallel base sequence analysis enabled by next-generation sequencing (NGS) technology is another way to access the enumeration of RNA transcripts in tissue samples, and RNA-sequencing utilizes this. It is currently the most powerful analytical tool used for transcriptome analysis, including differences in gene expression levels between different physiological conditions or changes occurring during development or disease progression. Specifically, RNA-sequencing can be used to study phenomena such as gene expression changes, alternative splicing events, allele-specific gene expression, gene fusion events, de novo transcripts, and chimeric transcripts, including RNA editing.
[0012] As used herein, the term "prognosis" refers to a broad concept that includes the course of a disease, such as cancer migration and invasion into tissues, metastasis to other tissues, recurrence, and death due to the disease, as well as the possibility of complete recovery. For the purposes of the present invention, prognosis refers to the course of a disease or survival prognosis of a patient with solid cancer, specifically prostate cancer. Using the method of the present invention, the survival prognosis according to prostate cancer can be easily determined, thereby easily determining whether to use additional treatment methods. Ultimately, the survival rate after the onset of prostate cancer can be improved. More specifically, the prognosis to be predicted in the present invention includes the possibility of recurrence, metastasis, or a combination thereof after surgical treatment.
[0013] As used herein, the term "metastasis" refers to the process by which cancer cells detach from the primary tumor (the initial cancer) and migrate to other parts of the body, forming new tumors there. Metastasis is a critical stage that increases the lethality of cancer, and it is known that more than 90% of cancer deaths are due to metastasis. The present invention enables a more precise diagnosis of cancer metastasis, particularly in prostate cancer.
[0014] As used herein, the term "recurrence" refers to the recurrence of cancer that appeared to be cured for a certain period of time after treatment. This occurs when cancer cells remaining after treatment regrow over time and form a new tumor, making prior diagnosis crucial. Depending on the location of the cancer recurrence, it is categorized into local recurrence, regional recurrence, and distant recurrence. Specifically, recurrence in the present invention may be, but is not limited to, biochemical recurrence (BCR).
[0015] The term "biochemical recurrence" in this specification generally refers to a state in which biochemical markers detect signs of cancer reactivation after prostate cancer treatment. Biochemical recurrence is considered when the PSA level is measured as 0.2 ng / mL or higher twice consecutively after radical prostatectomy, and biochemical recurrence is defined as a PSA level that rises by 2 ng / mL or more from the post-treatment nadir after radiation therapy. BCR can be an early warning sign of cancer recurrence, but it is not necessarily the same as clinically apparent recurrence (e.g., symptoms, lesions on imaging tests). If BCR is detected, the treatment plan is determined after comprehensively considering additional tests and the patient's condition.
[0016] As used herein, the term "prediction" refers to the act of predicting the course and outcome of a disease. More specifically, prognostic prediction can be interpreted as any act of predicting the course of a disease after treatment, taking into account the patient's physiological and environmental conditions, which may vary depending on the patient's condition.
[0017]
[0018] According to one embodiment of the present invention, the present invention relates to a composition for predicting the prognosis of cancer and a composition for predicting the possibility of metastasis or recurrence of cancer.
[0019] In the present invention, the composition includes ALDH1A3, ANO7, CRACR2B, ENO1, EPS8L1, HSPA1A, ITGB4, KLK3, KLK4, MALAT1, NPM1, PLA2G2A, SLC45A3, SPINT2, ZFP36, ACTA2, ACTB, ACTG1, ALDOA, ARID1A, B2M, BCAM, C3, CD63, CHCHD10, CKB, CLU, COL1A2, COL6A1, COX6A1, CST3, DDT, DENND2B, EEF1A1, EEF1G, EEF2, EGR1, EHMT1, FAM193B, FAU, FOS, FOSB, FTH1, FTL, GAPDH, GBF1, GFAP, GSTP1, H2AJ, H2BC12, H3-3A, H3-3B, H4C12, HLA-DRB1, HNRNPA1, HNRNPC, HNRNPH1, HSF4, HSPA8, HSPB1, HSPB6, ITM2B, JUNB, KLF13, LENG8, LTBP4, MAN2C1, MBP, MGP, MIB2, MIF, MYL6, MYO15B, NCOR2, NDRG1, NDUFA1, NPDC1, NPIPA6, NPIPA9, NR4A1, OAZ1, PEBP1, PER1, PKD1, PKM, PPIA, PSAP, RACK1, RERE, SH2B3, SPARC, SPON2, SREBF1, TCEAL4, TMSB10, TMSB4X, TNK2, TNS2, TPT1, TRIM3, TUBA1B, WDR1 and ZNF414 It may include a formulation that measures the expression level of at least one gene selected from the group or a protein encoded thereby.
[0020] In the present invention, the composition comprises ACLY, ACOX1, ACSL4, ACVR1, AKT1, AXIN1, AXIN2, BCL2L1, BMP4, BMPR1A, BRCA1, BTK, CASP3, CASP8, CASP9, CCNA2, CCNB2, CCND1, CCR5, CD40, CDK1, CDK2, CDK4, CDKN1A, CDKN2A, CHEK1, CS, CTNNB1, CXCR4, DVL1, DVL2, EP300, FYN, GADD45A, GNB1, GPX4, GRB2, GSK3B, H3C12, H3C13, H4C6, HMOX1, HRAS, IKBKB, IKBKG, IL1B, IL6, INS, JAG1, JAK1, JAK2, JAK3, JUN, The composition may further include an agent for measuring the expression level of at least one gene selected from the group consisting of KRAS, LCK, LEF1, LPCAT3, LYN, MAML1, MAML2, MAML3, MAPK1, MAPK3, MCM7, MLST8, MTOR, NCOA4, NFKB1, NFKBIA, NOTCH1, NOTCH2, NOTCH3, NOTCH4, NRAS, PIK3CA, PIK3R1, PLCG2, PTEN, RBPJ, RELA, RHEB, RHOA, RPTOR, RUNX1, SDHB, SMAD2, SMAD3, SMAD4, SOS1, SRC, STAT1, STAT3, SYK, TFRC, TGFB1, TNF, TNFRSF1A, TP53, TRAF6, TSC2 and TYK2 or a protein encoded by the same.
[0021] Among the genes in the present invention, ACTB, H3-3B, ACLY, ACOX1, ACSL4, ACVR1, AKT1, AXIN1, AXIN2, BCL2L1, BMP4, BMPR1A, BRCA1, BTK, CASP3, CASP8, CASP9, CCNA2, CCNB2, CCND1, CCR5, CD40, CDK1, CDK2, CDK4, CDKN1A, CDKN2A, CHEK1, CS, CTNNB1, CXCR4, DVL1, DVL2, EP300, FYN, GADD45A, GNB1, GPX4, GRB2, GSK3B, H3C12, H3C13, H4C6, HMOX1, HRAS, IKBKB, IKBKG, IL1B, IL6, INS, JAG1, JAK1, JAK2, The genes JAK3, JUN, KRAS, LCK, LEF1, LPCAT3, LYN, MAML1, MAML2, MAML3, MAPK1, MAPK3, MCM7, MLST8, MTOR, NCOA4, NFKB1, NFKBIA, NOTCH1, NOTCH2, NOTCH3, NOTCH4, NRAS, PIK3CA, PIK3R1, PLCG2, PTEN, RBPJ, RELA, RHEB, RHOA, RPTOR, RUNX1, SDHB, SMAD2, SMAD3, SMAD4, SOS1, SRC, STAT1, STAT3, SYK, TFRC, TGFB1, TNF, TNFRSF1A, TP53, TRAF6, TSC2, and TYK2 are characterized as signal transduction pathway-related genes.
[0022] In the present invention, the signal transduction pathway may be a signal transduction pathway related to cancer death, cancer proliferation inhibition, cancer metastasis or recurrence, and may be specifically a signal transduction pathway related to prostate cancer death, proliferation inhibition, metastasis or recurrence, but is not limited thereto.
[0023] In the present invention, the "signal transduction pathway-related gene" refers to a gene that is involved in the cancer-related signal transduction pathway described above or can obtain information on activation or inactivation of the pathway through its expression pattern.
[0024] In the present invention, the signal transduction pathway is an adipogenesis pathway, an apoptosis pathway Ⅱ, an E2F targets pathway, an estrogen response early pathway, an estrogen response late pathway, a NOTCH signaling pathway, a WNT-Beta Catenin signaling pathway, a MAPK signaling pathway, an ErbB signaling pathway, a Ras signaling pathway, a Rap1 signaling pathway, a chemokine signaling pathway, an NF-kappa B signaling pathway, a cell cycle pathway, a p53 signaling pathway, an mTOR signaling pathway. PI3K-Akt signaling pathway, Apoptosis pathway Ⅰ, Ferroptosis pathway, Necroptosis pathway, Cellular senescence pathway, Wnt signaling pathway, Notch signaling pathway, TGF-beta signaling pathway, JAK-STAT signaling pathway, T cell receptor signaling pathwayIt may be at least one signaling pathway selected from the group consisting of, but not limited to, B cell receptor signaling pathway, TNF signaling pathway, Transcriptional misregulation pathway in cancer, Prostate cancer pathway, ESR Mediated signaling pathway, Estrogen Dependent gene expression pathway, and Androgen Receptor Network pathway in Prostate cancer.
[0025] In the composition of the present invention, the agent for measuring the expression level of the gene may include, but is not limited to, one or more selected from the group consisting of a primer, a probe, and an antisense nucleotide that specifically bind to the gene.
[0026] In the composition of the present invention, the agent for measuring the expression level of the protein encoded by the gene may include, but is not limited to, one or more selected from the group consisting of antibodies, oligopeptides, ligands, PNA (peptide nucleic acid), and aptamers that specifically bind to the protein.
[0027] In the present invention, the cancer may be at least one selected from the group consisting of prostate cancer, breast cancer, mammary cancer, glioma, thyroid cancer, lung cancer, liver cancer, pancreatic cancer, head and neck cancer, stomach cancer, colon cancer, urothelial cancer, kidney cancer, testicular cancer, penile cancer, uterine cancer, cervical cancer, endometrial cancer, cervical cancer, fallopian tube cancer, vaginal cancer, ovarian cancer, melanoma, skin cancer, blood cancer, bone cancer, skin cancer, brain cancer, endocrine cancer, parathyroid cancer, ureteral cancer, urethral cancer, bronchial cancer, bladder cancer, bone marrow cancer, leukemia, brain tumor, intestinal cancer, esophageal cancer, Ewing's sarcoma, tongue cancer, lymphoma, kaposi sarcoma, mesothelioma, multiple myeloma, neuroblastoma, osteosarcoma, and retinoblastoma, and may be specifically prostate cancer, but is not limited thereto.
[0028]
[0029] According to another embodiment of the present invention, the present invention relates to a kit for predicting cancer prognosis, a kit for predicting the possibility of cancer metastasis or recurrence, comprising the composition.
[0030] In the present invention, the kit may be, but is not limited to, an NGS kit, an RT-PCR kit, a DNA chip kit, an RNA sequencing kit, an ELISA kit, a protein chip kit, or a rapid kit.
[0031] The kit of the present invention may further comprise one or more other component compositions, solutions or devices suitable for the analysis method.
[0032] For example, in the present invention, the kit may further include essential elements necessary for performing a reverse transcription polymerase reaction. The reverse transcription polymerase reaction kit includes a pair of primers specific for a gene encoding a marker protein. The primers are nucleotides having a sequence specific for the nucleic acid sequence of the gene, and may have a length of about 7 bp to 50 bp, more preferably about 10 bp to 30 bp. It may also include a primer specific for the nucleic acid sequence of a control gene. In addition, the reverse transcription polymerase reaction kit may include a test tube or other appropriate container, a reaction buffer (with various pH and magnesium concentrations), deoxynucleotides (dNTPs), an enzyme such as Taq polymerase and reverse transcriptase, DNase, RNase inhibitor DEPC-water, sterile water, etc.
[0033] Additionally, the kit of the present invention may include essential elements necessary for performing a DNA chip. The DNA chip kit may include a substrate to which cDNA or oligonucleotides corresponding to a gene or fragment thereof are attached, and reagents, preparations, enzymes, etc. for producing a fluorescently labeled probe. The substrate may also include cDNA or oligonucleotides corresponding to a control gene or fragment thereof.
[0034] Additionally, the kit of the present invention may include essential components necessary for performing an ELISA. The ELISA kit includes an antibody specific for the protein. The antibody is an antibody with high specificity and affinity for the marker protein and little cross-reactivity with other proteins, and may be a monoclonal antibody, polyclonal antibody, or recombinant antibody. The ELISA kit may also include an antibody specific for a control protein. In addition, the ELISA kit may include reagents capable of detecting bound antibodies, such as labeled secondary antibodies, chromophores, enzymes (e.g., conjugated to antibodies), and their substrates or other substances capable of binding to antibodies.
[0035] In the kit of the present invention, a fixative for the antigen-antibody binding reaction may be a nitrocellulose membrane, a PVDF membrane, a well plate synthesized from polyvinyl resin or polystyrene resin, a glass slide glass, etc., but is not limited thereto.
[0036] In addition, in the kit of the present invention, the label of the secondary antibody is preferably a conventional chromogen that undergoes a color development reaction, and labels such as fluorescein and dyes such as HRP (horseradish peroxidase), alkaline phosphatase, colloid gold, FITC (poly L-lysine-fluorescein isothiocyanate), and RITC (rhodamine-B-isothiocyanate) can be used, but are not limited thereto.
[0037] In addition, in the kit of the present invention, it is preferable to use a chromogenic substrate for inducing color development according to a marker that undergoes a color development reaction, and TMB (3,3',5,5'-tetramethyl bezidine), ABTS [2,2'-azino-bis(3-ethylbenzothiazoline-6-sulfonic acid)], OPD (o-phenylenediamine), etc. can be used. At this time, it is more preferable that the chromogenic substrate is provided in a state dissolved in a buffer solution (0.1 M NaAc, pH 5.5). A chromogenic substrate such as TMB is decomposed by HRP used as a marker of a secondary antibody conjugate to generate a chromogenic precipitate, and the presence or absence of the marker proteins is detected by visually confirming the degree of deposition of this chromogenic precipitate.
[0038] In the kit of the present invention, the washing solution preferably contains phosphate buffer, NaCl, and Tween 20, and a buffer solution (PBST) composed of 0.02 M phosphate buffer, 0.13 M NaCl, and 0.05% Tween 20 is more preferred. After the antigen-antibody binding reaction, the washing solution reacts the antigen-antibody complex with a secondary antibody, and then adds an appropriate amount to the fixative and washes 3 to 6 times. The reaction stopping solution can preferably be a sulfuric acid solution (H2SO4).
[0039]
[0040] According to another embodiment of the present invention, there is provided a method for providing information on predicting cancer prognosis, and a method for providing information on predicting the possibility of cancer metastasis or recurrence.
[0041] In the present invention, the method comprises the steps of: ALDH1A3, ANO7, CRACR2B, ENO1, EPS8L1, HSPA1A, ITGB4, KLK3, KLK4, MALAT1, NPM1, PLA2G2A, SLC45A3, SPINT2, ZFP36, ACTA2, ACTB, ACTG1, ALDOA, ARID1A, B2M, BCAM, C3, CD63, CHCHD10, CKB, CLU, COL1A2, COL6A1, COX6A1, CST3, DDT, DENND2B, EEF1A1, EEF1G, EEF2, EGR1, EHMT1, FAM193B, FAU, FOS, FOSB, FTH1, FTL, GAPDH, GBF1, GFAP, GSTP1, H2AJ, H2BC12, H3-3A, H3-3B, H4C12, HLA-DRB1, HNRNPA1, HNRNPC, HNRNPH1, HSF4, HSPA8, HSPB1, HSPB6, ITM2B, JUNB, KLF13, LENG8, LTBP4, MAN2C1, MBP, MGP, MIB2, MIF, MYL6, MYO15B, NCOR2, NDRG1, NDUFA1, NPDC1, NPIPA6, NPIPA9, NR4A1, OAZ1, PEBP1, PER1, PKD1, PKM, PPIA, PSAP, RACK1, RERE, SH2B3, SPARC, SPON2, SREBF1, TCEAL4, TMSB10, TMSB4X, TNK2, TNS2, TPT1, TRIM3, TUBA1B, WDR1 and It may include a step of measuring the expression level of at least one gene selected from the group consisting of ZNF414 or a protein encoded by the same.
[0042] In the present invention, the method comprises the steps of: ACLY, ACOX1, ACSL4, ACVR1, AKT1, AXIN1, AXIN2, BCL2L1, BMP4, BMPR1A, BRCA1, BTK, CASP3, CASP8, CASP9, CCNA2, CCNB2, CCND1, CCR5, CD40, CDK1, CDK2, CDK4, CDKN1A, CDKN2A, CHEK1, CS, CTNNB1, CXCR4, DVL1, DVL2, EP300, FYN, GADD45A, GNB1, GPX4, GRB2, GSK3B, H3C12, H3C13, H4C6, HMOX1, HRAS, IKBKB, IKBKG, IL1B, IL6, INS, JAG1, JAK1, JAK2, The method may further include a step of measuring the expression level of at least one gene selected from the group consisting of JAK3, JUN, KRAS, LCK, LEF1, LPCAT3, LYN, MAML1, MAML2, MAML3, MAPK1, MAPK3, MCM7, MLST8, MTOR, NCOA4, NFKB1, NFKBIA, NOTCH1, NOTCH2, NOTCH3, NOTCH4, NRAS, PIK3CA, PIK3R1, PLCG2, PTEN, RBPJ, RELA, RHEB, RHOA, RPTOR, RUNX1, SDHB, SMAD2, SMAD3, SMAD4, SOS1, SRC, STAT1, STAT3, SYK, TFRC, TGFB1, TNF, TNFRSF1A, TP53, TRAF6, TSC2 and TYK2 or a protein encoded by the same.
[0043] The measurement of the expression level according to the present invention is a process of confirming the presence and degree of expression of mRNA of the gene group, and can be performed by measuring the expression amount of the corresponding gene from mRNA extracted from a sample of a subject. Analysis methods for measuring the expression level include, but are not limited to, RT-PCR, competitive RT-PCR, real-time RT-PCR, RNase protection assay (RPA), northern blotting, DNA microarray chips, etc., and can be performed using any appropriate method commonly used in the art.
[0044] Preferably, the agent for measuring the expression level of mRNA according to the present invention is an antisense oligonucleotide, primer, or probe. Based on the base sequence of the gene group, a primer or probe that specifically amplifies a specific region of these genes can be designed. Since the base sequence of the gene group according to the present invention is registered in GenBank and is known in the art, those skilled in the art can design an antisense oligonucleotide, primer, or probe that can specifically amplify a specific region of these genes based on the base sequence.
[0045] According to the present invention, the expression level of each gene or transcript of each sample can be confirmed by using the number of reads mapped through RNA sequencing as a method for calculating the differential expression level of genes. However, there may be an error for each sample in defining the expression level by the number of mapped reads, and it is difficult to view it as an objective value. Therefore, a normalization process was performed as a method for deriving a more objective value.
[0046] In the present invention, the cancer may be at least one selected from the group consisting of prostate cancer, breast cancer, mammary cancer, glioma, thyroid cancer, lung cancer, liver cancer, pancreatic cancer, head and neck cancer, stomach cancer, colon cancer, urothelial cancer, kidney cancer, testicular cancer, penile cancer, uterine cancer, cervical cancer, endometrial cancer, cervical cancer, fallopian tube cancer, vaginal cancer, ovarian cancer, melanoma, skin cancer, blood cancer, bone cancer, skin cancer, brain cancer, endocrine cancer, parathyroid cancer, ureteral cancer, urethral cancer, bronchial cancer, bladder cancer, bone marrow cancer, leukemia, brain tumor, intestinal cancer, esophageal cancer, Ewing's sarcoma, tongue cancer, lymphoma, kaposi sarcoma, mesothelioma, multiple myeloma, neuroblastoma, osteosarcoma, and retinoblastoma, and specifically, may be prostate cancer.
[0047] The method for providing information on cancer prognosis of the present invention can not only predict the cancer prognosis of a subject, but can also be used to predict the possibility of cancer metastasis and, further, the possibility of cancer recurrence of the subject.
[0048]
[0049] According to another embodiment of the present invention, a cancer prognosis diagnostic device is provided, comprising: (a) a detection unit that measures and normalizes the expression level of a group of genes or a group of proteins encoded by the same in a biological sample obtained from a subject; and (b) an output unit that predicts and outputs a cancer prognosis.
[0050] In the present invention, the detection unit detects ALDH1A3, ANO7, CRACR2B, ENO1, EPS8L1, HSPA1A, ITGB4, KLK3, KLK4, MALAT1, NPM1, PLA2G2A, SLC45A3, SPINT2, ZFP36, ACTA2, ACTB, ACTG1, ALDOA, ARID1A, B2M, BCAM, C3, CD63, CHCHD10, CKB, CLU, COL1A2, COL6A1, COX6A1, CST3, DDT, DENND2B, EEF1A1, EEF1G, EEF2, EGR1, EHMT1, FAM193B, FAU, FOS, FOSB, FTH1, FTL, GAPDH, GBF1, GFAP, GSTP1, H2AJ, H2BC12, H3-3A, H3-3B, H4C12, HLA-DRB1, HNRNPA1, HNRNPC, HNRNPH1, HSF4, HSPA8, HSPB1, HSPB6, ITM2B, JUNB, KLF13, LENG8, LTBP4, MAN2C1, MBP, MGP, MIB2, MIF, MYL6, MYO15B, NCOR2, NDRG1, NDUFA1, NPDC1, NPIPA6, NPIPA9, NR4A1, OAZ1, PEBP1, PER1, PKD1, PKM, PPIA, PSAP, RACK1, RERE, SH2B3, SPARC, SPON2, SREBF1, TCEAL4, TMSB10, TMSB4X, TNK2, TNS2, TPT1, TRIM3, TUBA1B, The expression level of at least one gene selected from the group consisting of WDR1 and ZNF414 or a protein encoded by the same can be measured.
[0051] In the present invention, the detection unit can additionally measure the expression level of a signal transduction pathway-related gene or a protein encoded by the same from the biological sample obtained from the subject.
[0052] In the present invention, the signal transduction pathway may be a signal transduction pathway related to cancer death, cancer proliferation inhibition, cancer metastasis or recurrence, and may be specifically a signal transduction pathway related to prostate cancer death, proliferation inhibition, metastasis or recurrence, but is not limited thereto.
[0053] In the present invention, the "signal transduction pathway-related gene" refers to a gene that is involved in the cancer-related signal transduction pathway described above or can obtain information on activation or inactivation of the pathway through its expression pattern.
[0054] In the present invention, the signal transduction pathway is an adipogenesis pathway, an apoptosis pathway Ⅱ, an E2F targets pathway, an estrogen response early pathway, an estrogen response late pathway, a NOTCH signaling pathway, a WNT-Beta Catenin signaling pathway, a MAPK signaling pathway, an ErbB signaling pathway, a Ras signaling pathway, a Rap1 signaling pathway, a chemokine signaling pathway, an NF-kappa B signaling pathway, a cell cycle pathway, a p53 signaling pathway, an mTOR signaling pathway. PI3K-Akt signaling pathway, Apoptosis pathway Ⅰ, Ferroptosis pathway, Necroptosis pathway, Cellular senescence pathway, Wnt signaling pathway, Notch signaling pathway, TGF-beta signaling pathway, JAK-STAT signaling pathway, T cell receptor signaling pathwayIt may be at least one signaling pathway selected from the group consisting of, but not limited to, B cell receptor signaling pathway, TNF signaling pathway, Transcriptional misregulation pathway in cancer, Prostate cancer pathway, ESR Mediated signaling pathway, Estrogen Dependent gene expression pathway, and Androgen Receptor Network pathway in Prostate cancer.
[0055] In the present invention, the cancer may be at least one selected from the group consisting of prostate cancer, breast cancer, mammary cancer, glioma, thyroid cancer, lung cancer, liver cancer, pancreatic cancer, head and neck cancer, stomach cancer, colon cancer, urothelial cancer, kidney cancer, testicular cancer, penile cancer, uterine cancer, cervical cancer, endometrial cancer, cervical cancer, fallopian tube cancer, vaginal cancer, ovarian cancer, melanoma, skin cancer, blood cancer, bone cancer, skin cancer, brain cancer, endocrine cancer, parathyroid cancer, ureteral cancer, urethral cancer, bronchial cancer, bladder cancer, bone marrow cancer, leukemia, brain tumor, intestinal cancer, esophageal cancer, Ewing's sarcoma, tongue cancer, lymphoma, kaposi sarcoma, mesothelioma, multiple myeloma, neuroblastoma, osteosarcoma, and retinoblastoma, and specifically, may be prostate cancer.
[0056]
[0057] According to another embodiment of the present invention, a method for providing information on cancer risk is provided.
[0058] In the present invention, the above method may be based on a Cox regression analysis algorithm.
[0059] In the present invention, the method may include a step of calculating gene expression levels and signal transduction pathway scores for a biological sample obtained from a subject and classifying them as explanatory variables.
[0060] In the present invention, the term "Cox regression analysis" refers to a statistical technique, first developed by Cox in 1972, used in survival analysis. It analyzes the relationship between survival time and one or more predictor variables, making it a crucial tool for survival analysis. It can analyze the risk of an event over time by linking it to various factors. It is widely used, particularly in the medical field, to evaluate patient survival rates and treatment effectiveness, and is a highly useful tool when handling data that satisfies the proportional hazards assumption.
[0061] In the present invention, the term "cancer" generally refers to a physiological condition in mammals characterized by uncontrolled cell proliferation. Examples of cancer include, but are not limited to, prostate cancer, breast cancer, mammary cancer, glioma, thyroid cancer, lung cancer, liver cancer, pancreatic cancer, head and neck cancer, stomach cancer, colon cancer, urothelial cancer, kidney cancer, testicular cancer, penile cancer, uterine cancer, cervical cancer, endometrial cancer, cervical cancer, fallopian tube cancer, vaginal cancer, ovarian cancer, melanoma, skin cancer, blood cancer, bone cancer, skin cancer, brain cancer, endocrine cancer, parathyroid cancer, ureter cancer, urethral cancer, bronchial cancer, bladder cancer, bone marrow cancer, leukemia, brain tumor, bowel cancer, esophageal cancer, Ewing's sarcoma, tongue cancer, lymphoma, kaposi sarcoma, mesothelioma, multiple myeloma, neuroblastoma, osteosarcoma, and retinoblastoma, and may be specifically prostate cancer.
[0062] In the present invention, the term "cancer risk" may have a broad meaning encompassing the possibility that a specific individual will develop cancer, the possibility that a specific individual will metastasize or recur from cancer, or the possibility that a specific individual will not show a beneficial response to anticancer treatment, etc., through a genetic risk assessment excluding environmental factors and lifestyle factors, and may specifically have a probabilistic meaning regarding the possibility of cancer metastasis or recurrence, but is not limited thereto.
[0063] In the present invention, the term "subject" refers to an individual that has developed a disease or is likely to develop a disease, and may be a mammal including a human, and may be selected from the group consisting of, for example, a human, a rat, a mouse, a guinea pig, a hamster, a rabbit, a monkey, a dog, a cat, a cow, a horse, a pig, a sheep, and a goat, and may be, but is not limited to, a human.
[0064] In the present invention, the term "biological sample" means any sample capable of confirming a patient's genetic information, including, but not limited to, a solid tissue sample, a tissue culture, a liquid tissue sample, cells, or cell fragments. Also, non-limiting examples of biological samples include whole blood, leukocytes, peripheral blood mononuclear cells, buffy coat, plasma, serum, sputum, tears, mucus, nasal washes, nasal aspirate, breath, urine, semen, saliva, peritoneal washings, ascites, cystic fluid, meningeal fluid, amniotic fluid, glandular fluid, pancreatic fluid, lymph fluid, pleural fluid, nipple aspirate, bronchial aspirate, synovial fluid, joint fluid. It may include at least one selected from the group consisting of joint aspirate, organ secretions, cells, cell extracts, and cerebrospinal fluid, and is not limited thereto, as long as it corresponds to a sample from which nucleic acids can be specifically extracted.
[0065] In the present invention, the term "gene expression level" refers to the amount of DNA gene sequence information converted into transcript RNA (either the initial non-spliced RNA transcript or mature mRNA) or encoded protein product. Gene expression level can be monitored by measuring the level of either the total RNA or protein product of the gene, or their subsequent levels.
[0066] In the present invention, the term "signal transduction pathway" refers to a signal transduction pathway related to cancer death, cancer proliferation inhibition, cancer metastasis or recurrence, and may specifically be a signal transduction pathway related to prostate cancer death, proliferation inhibition, metastasis or recurrence, but is not limited thereto.
[0067] In the present invention, the signal transduction pathway is an adipogenesis pathway, an apoptosis pathway Ⅱ, an E2F targets pathway, an estrogen response early pathway, an estrogen response late pathway, a NOTCH signaling pathway, a WNT-Beta Catenin signaling pathway, a MAPK signaling pathway, an ErbB signaling pathway, a Ras signaling pathway, a Rap1 signaling pathway, a chemokine signaling pathway, an NF-kappa B signaling pathway, a cell cycle pathway, a p53 signaling pathway, an mTOR signaling pathway. PI3K-Akt signaling pathway, Apoptosis pathway Ⅰ, Ferroptosis pathway, Necroptosis pathway, Cellular senescence pathway, Wnt signaling pathway, Notch signaling pathway, TGF-beta signaling pathway, JAK-STAT signaling pathway, T cell receptor signaling pathwayIt may be at least one signaling pathway selected from the group consisting of, but not limited to, B cell receptor signaling pathway, TNF signaling pathway, Transcriptional misregulation pathway in cancer, Prostate cancer pathway, ESR Mediated signaling pathway, Estrogen Dependent gene expression pathway, and Androgen Receptor Network pathway in Prostate cancer.
[0068] In the present invention, the term "explanatory variable" refers to the expression level of a differentially expressed gene (DEG gene) or a signal transduction pathway score, as defined herein. Examples can be readily understood by referring to the explanatory variable items listed in Tables 4 and 5 of the Examples section of this specification.
[0069] In the present invention, the gene expression level refers to the expression level of a differentially expressed gene, and as a method of calculating this, the expression level of each gene or transcript of each sample can be confirmed by using the number of reads mapped through RNA sequencing, but there may be an error for each sample when defining the expression level with the number of mapped reads, and it is difficult to view it as an objective value, so it may be derived by performing a normalization process as a method for deriving a more objective value.
[0070] In the present invention, the signal transduction pathway score is characterized in that it is calculated from the signal transduction pathway-related genes of Table 2 included in each signal transduction pathway using the GSVA (Gene Set Variation Analysis) program.
[0071] In the present invention, the method may further include a step of calculating a risk score of the following equation 1 using a Cox regression analysis algorithm by combining at least one gene described in Table 1 of the present specification and at least one signal transduction pathway described in Table 2 of the present specification.
[0072] [Formula 1]
[0073]
[0074] In the above equation 1,
[0075] The above p represents the number of explanatory variables,
[0076] The above explanatory variables represent gene expression levels or signal transduction pathway scores.
[0077] In the present invention, the regression coefficient in the above equation 1 may include a value within a ±5% error range of each regression coefficient value matching each explanatory variable in Table 4 or Table 5 of the present specification. More specifically, it may include a value within a ±4% error range of each regression coefficient value matching each explanatory variable in Table 4 or Table 5, even more specifically, it may include a value within a ±3% error range, even more specifically, it may include a value within a ±2% error range, and most specifically, it may include a value within a ±1% error range. As a non-limiting example, the regression coefficient value matching the EEF1A1 gene in Table 4 of the present specification is -24.5157, and when the error range of the above value is applied, the regression coefficient value may include a value in a range of -25.7415 or more and -23.2899 or less. As another non-limiting example, the regression coefficient value matching the p53 signaling pathway (hsa04115_p53_signaling_pathway) in Table 5 of the present specification is -12.2718, and when the error range of the above value is applied, the regression coefficient value can include a value in the range of -11.6582 or more and -12.8854 or less. More specifically, it can include a value within the error range of ±4% of each regression coefficient value matching each explanatory variable in Table 4 or Table 5, even more specifically, it can include a value within the error range of ±3%, even more specifically, it can include a value within the error range of ±2%, and most specifically, it can include a value within the error range of ±1%.
[0078] In the present invention, the gene may be a DEG gene according to one embodiment of the present invention, and specifically, ALDH1A3, ANO7, CRACR2B, ENO1, EPS8L1, HSPA1A, ITGB4, KLK3, KLK4, MALAT1, NPM1, PLA2G2A, SLC45A3, SPINT2, ZFP36, ACTA2, ACTB, ACTG1, ALDOA, ARID1A, B2M, BCAM, C3, CD63, CHCHD10, CKB, CLU, COL1A2, COL6A1, COX6A1, CST3, DDT, DENND2B, EEF1A1, EEF1G, EEF2, EGR1, EHMT1, FAM193B, FAU, FOS, FOSB, FTH1, FTL, GAPDH, GBF1, GFAP, GSTP1, H2AJ, H2BC12, H3-3A, H3-3B, H4C12, HLA-DRB1, HNRNPA1, HNRNPC, HNRNPH1, HSF4, HSPA8, HSPB1, HSPB6, ITM2B, JUNB, KLF13, LENG8, LTBP4, MAN2C1, MBP, MGP, MIB2, MIF, MYL6, MYO15B, NCOR2, NDRG1, NDUFA1, NPDC1, NPIPA6, NPIPA9, NR4A1, OAZ1, PEBP1, PER1, PKD1, PKM, PPIA, PSAP, RACK1, RERE, SH2B3, SPARC, SPON2, SREBF1, TCEAL4, TMSB10, TMSB4X, TNK2, TNS2, TPT1, TRIM3, TUBA1B, It is characterized by at least one gene selected from the group consisting of WDR1 and ZNF414.
[0079] In the present invention, the signal transduction pathway-related gene refers to a gene that is involved in the cancer-related signal transduction pathway described above or can obtain information on activation or inactivation of the pathway through its expression pattern. Specifically, ACTB, H3-3B, ACLY, ACOX1, ACSL4, ACVR1, AKT1, AXIN1, AXIN2, BCL2L1, BMP4, BMPR1A, BRCA1, BTK, CASP3, CASP8, CASP9, CCNA2, CCNB2, CCND1, CCR5, CD40, CDK1, CDK2, CDK4, CDKN1A, CDKN2A, CHEK1, CS, CTNNB1, CXCR4, DVL1, DVL2, EP300, FYN, GADD45A, GNB1, GPX4, GRB2, GSK3B, H3C12, H3C13, H4C6, HMOX1, HRAS, IKBKB, IKBKG, IL1B, IL6, INS, JAG1, JAK1, JAK2, It is characterized by at least one selected from the group consisting of JAK3, JUN, KRAS, LCK, LEF1, LPCAT3, LYN, MAML1, MAML2, MAML3, MAPK1, MAPK3, MCM7, MLST8, MTOR, NCOA4, NFKB1, NFKBIA, NOTCH1, NOTCH2, NOTCH3, NOTCH4, NRAS, PIK3CA, PIK3R1, PLCG2, PTEN, RBPJ, RELA, RHEB, RHOA, RPTOR, RUNX1, SDHB, SMAD2, SMAD3, SMAD4, SOS1, SRC, STAT1, STAT3, SYK, TFRC, TGFB1, TNF, TNFRSF1A, TP53, TRAF6, TSC2, and TYK2.
[0080] The published Gene ID information (ensembl_gene_id) of the signal transduction pathway related genes of the present invention is as follows:
[0081] ACTB(ensembl_gene_id: ENSG00000075624), H3-3B(ensembl_gene_id: ENSG00000132475), ACLY(ensembl_gene_id: ENSG00000131473), ACOX1(ensembl_gene_id: ENSG00000161533), ACSL4(ensembl_gene_id: ENSG00000068366), ACVR1(ensembl_gene_id: ENSG00000115170), AKT1(ensembl_gene_id: ENSG00000142208), AXIN1(ensembl_gene_id: ENSG00000103126), AXIN2(ensembl_gene_id: ENSG00000168646), BCL2L1(ensembl_gene_id: ENSG00000171552), BMP4(ensembl_gene_id: ENSG00000125378), BMPR1A(ensembl_gene_id: ENSG00000107779), BRCA1(ensembl_gene_id: ENSG00000012048), BTK(ensembl_gene_id: ENSG00000010671), CASP3(ensembl_gene_id: ENSG00000164305), CASP8(ensembl_gene_id: ENSG00000064012), CASP9(ensembl_gene_id: ENSG00000132906), CCNA2(ensembl_gene_id: ENSG00000145386), CCNB2(ensembl_gene_id: ENSG00000157456), CCND1(ensembl_gene_id: ENSG00000110092), CCR5(ensembl_gene_id: ENSG00000160791), CD40(ensembl_gene_id: ENSG00000101017), CDK1(ensembl_gene_id: ENSG00000170312), CDK2(ensembl_gene_id: ENSG00000123374),CDK4(ensembl_gene_id: ENSG00000135446), CDKN1A(ensembl_gene_id: ENSG00000124762), CDKN2A(ensembl_gene_id: ENSG00000147889), CHEK1(ensembl_gene_id: ENSG00000149554), CS(ensembl_gene_id: ENSG00000062485), CTNNB1(ensembl_gene_id: ENSG00000168036), CXCR4(ensembl_gene_id: ENSG00000121966), DVL1(ensembl_gene_id: ENSG00000107404), DVL2(ensembl_gene_id: ENSG00000004975), EP300(ensembl_gene_id: ENSG00000100393), FYN(ensembl_gene_id: ENSG00000010810), GADD45A(ensembl_gene_id: ENSG00000116717), GNB1(ensembl_gene_id: ENSG00000078369), GPX4(ensembl_gene_id: ensembl_gene_id: ENSG00000167468), GRB2(ensembl_gene_id: ENSG00000177885), GSK3B(ensembl_gene_id: ENSG00000082701), H3C12(ensembl_gene_id: ENSG00000197153), H3C13(ensembl_gene_id: ENSG00000183598), H4C6(ensembl_gene_id: ENSG00000274618), HMOX1(ensembl_gene_id: ENSG00000100292), HRAS(ensembl_gene_id: ENSG00000174775), IKBKB(ensembl_gene_id: ENSG00000104365), IKBKG(ensembl_gene_id: ENSG00000269335), IL1B(ensembl_gene_id: ENSG00000125538),IL6(ensembl_gene_id: ENSG00000136244), INS(ensembl_gene_id: ENSG00000254647), JAG1(ensembl_gene_id: ENSG00000101384), JAK1(ensembl_gene_id: ENSG00000162434), JAK2(ensembl_gene_id: ENSG00000096968), JAK3(ensembl_gene_id: ENSG00000105639), JUN(ensembl_gene_id: ENSG00000177606), KRAS(ensembl_gene_id: ENSG00000133703), LCK(ensembl_gene_id: ENSG00000182866), LEF1(ensembl_gene_id: ENSG00000138795), LPCAT3(ensembl_gene_id: ENSG00000111684), LYN(ensembl_gene_id: ENSG00000254087), MAML1(ensembl_gene_id: ENSG00000161021), MAML2(ensembl_gene_id: ENSG00000184384), MAML3(ensembl_gene_id: ENSG00000196782), MAPK1(ensembl_gene_id: ENSG00000100030), MAPK3(ensembl_gene_id: ENSG00000102882), MCM7(ensembl_gene_id: ENSG00000166508), MLST8(ensembl_gene_id: ENSG00000167965), MTOR(ensembl_gene_id: ENSG00000198793), NCOA4(ensembl_gene_id: ENSG00000266412), NFKB1(ensembl_gene_id: ENSG00000109320), NFKBIA(ensembl_gene_id: ENSG00000100906), NOTCH1(ensembl_gene_id: ENSG00000148400),NOTCH2(ensembl_gene_id: ENSG00000134250), NOTCH3(ensembl_gene_id: ENSG00000074181), NOTCH4(ensembl_gene_id: ENSG00000204301), NRAS(ensembl_gene_id: ENSG00000213281), PIK3CA(ensembl_gene_id: ENSG00000121879), PIK3R1(ensembl_gene_id: ENSG00000145675), PLCG2(ensembl_gene_id: ENSG00000197943), PTEN(ensembl_gene_id: ENSG00000171862), RBPJ(ensembl_gene_id: ENSG00000168214), RELA(ensembl_gene_id: ENSG00000173039), RHEB(ensembl_gene_id: ENSG00000106615), RHOA(ensembl_gene_id: ENSG00000067560), RPTOR(ensembl_gene_id: ENSG00000141564), RUNX1(ensembl_gene_id: ENSG00000159216), SDHB(ensembl_gene_id: ENSG00000117118), SMAD2(ensembl_gene_id: ENSG00000175387), SMAD3(ensembl_gene_id: ENSG00000166949), SMAD4(ensembl_gene_id: ENSG00000141646), SOS1(ensembl_gene_id: ENSG00000115904), SRC(ensembl_gene_id: ENSG00000197122), STAT1(ensembl_gene_id: ENSG00000115415), STAT3(ensembl_gene_id: ENSG00000168610), SYK(ensembl_gene_id: ENSG00000165025), TFRC(ensembl_gene_id: ENSG00000072274),TGFB1 (ensembl_gene_id: ENSG00000105329), TNF (ensembl_gene_id: ENSG00000232810), TNFRSF1A (ensembl_gene_id: ENSG00000067182), TP53 (ensembl_gene_id: ENSG00000141510), TRAF6 (ensembl_gene_id: ENSG00000175104), TSC2 (ensembl_gene_id: ENSG00000103197) and TYK2 (ensembl_gene_id: ENSG00000105397);
[0082] In the present invention, the adipogenesis pathway-related gene among the signal transduction pathways may include at least one selected from the group consisting of ACLY, ACOX1, CS, GADD45A, GPX4, LPCAT3, and SDHB, but is not limited thereto.
[0083] In the present invention, the apoptosis pathway Ⅱ related gene among the signal transduction pathways may include at least one selected from the group consisting of BCL2L1, BRCA1, CASP3, CASP8, CASP9, CCND1, CDK2, CDKN1A, CTNNB1, GADD45A, GPX4, HMOX1, IL1B, IL6, JUN, LEF1, RELA, and TNF, but is not limited thereto.
[0084] In the present invention, genes related to the E2F targets pathway among the signal transduction pathways may include at least one selected from the group consisting of BRCA1, CCNB2, CDK1, CDK4, CDKN1A, CDKN2A, CHEK1, MCM7, TFRC, and TP53, but are not limited thereto.
[0085] In the present invention, the gene related to the Estrogen Response Early pathway among the signal transduction pathways may include at least one selected from the group consisting of CCND1 and JAK2, but is not limited thereto.
[0086] In the present invention, the gene related to the Estrogen Response Late pathway among the signal transduction pathways may include at least one selected from the group consisting of CCND1, JAK1, and JAK2, but is not limited thereto.
[0087] In the present invention, among the signal transduction pathways, the NOTCH signaling pathway-related genes may include at least one selected from the group consisting of CCND1, DVL1, DVL2, EP300, JAG1, MAML1, MAML2, MAML3, NOTCH1, NOTCH2, NOTCH3, NOTCH4, and RBPJ, but are not limited thereto.
[0088] In the present invention, among the signal transduction pathways, the WNT-Beta Catenin signaling pathway-related genes may include at least one selected from the group consisting of AXIN1, AXIN2, CTNNB1, DVL2, JAG1, LEF1, MAML1, NOTCH1, NOTCH4, RBPJ, and TP53, but are not limited thereto.
[0089] In the present invention, the MAPK signaling pathway-related genes among the signal transduction pathways may include at least one selected from the group consisting of AKT1, CASP3, GADD45A, GRB2, HRAS, IKBKB, IKBKG, IL1B, INS, JUN, KRAS, MAPK1, MAPK3, NFKB1, NRAS, RELA, SOS1, TGFB1, TNF, TNFRSF1A, TP53, and TRAF6, but are not limited thereto.
[0090] In the present invention, the ErbB signaling pathway-related genes among the signal transduction pathways may include at least one selected from the group consisting of AKT1, CDKN1A, GRB2, GSK3B, HRAS, JUN, KRAS, MAPK1, MAPK3, MTOR, NRAS, PIK3CA, PIK3R1, PLCG2, SOS1, and SRC, but are not limited thereto.
[0091] In the present invention, genes related to the Ras signaling pathway among the signal transduction pathways may include at least one selected from the group consisting of AKT1, BCL2L1, GNB1, GRB2, HRAS, IKBKB, IKBKG, INS, KRAS, MAPK1, MAPK3, NFKB1, NRAS, PIK3CA, PIK3R1, PLCG2, RELA, RHOA, and SOS1, but are not limited thereto.
[0092] In the present invention, genes related to the Rap1 signaling pathway among the signal transduction pathways may include at least one selected from the group consisting of ACTB, AKT1, CTNNB1, HRAS, INS, KRAS, MAPK1, MAPK3, NRAS, PIK3CA, PIK3R1, RHOA, and SRC, but are not limited thereto.
[0093] In the present invention, genes related to the chemokine signaling pathway among the signal transduction pathways may include at least one selected from the group consisting of AKT1, CCR5, CXCR4, GNB1, GRB2, GSK3B, HRAS, IKBKB, IKBKG, JAK2, JAK3, KRAS, LYN, MAPK1, MAPK3, NFKB1, NFKBIA, NRAS, PIK3CA, PIK3R1, PLCG2, RELA, RHOA, SOS1, SRC, STAT1, and STAT3, but are not limited thereto.
[0094] In the present invention, the NF-kappa B signaling pathway-related genes among the signal transduction pathways may include at least one selected from the group consisting of BCL2L1, BTK, CD40, GADD45A, IKBKB, IKBKG, IL1B, LCK, LYN, NFKB1, NFKBIA, PLCG2, RELA, SYK, TNF, TNFRSF1A, and TRAF6, but are not limited thereto.
[0095] In the present invention, the cell cycle pathway-related genes among the signal transduction pathways may include at least one selected from the group consisting of CCNA2, CCNB2, CCND1, CDK1, CDK2, CDK4, CDKN1A, CDKN2A, CHEK1, EP300, GADD45A, GSK3B, MCM7, SMAD2, SMAD3, SMAD4, TGFB1, and TP53, but are not limited thereto.
[0096] In the present invention, genes related to the p53 signaling pathway among the signal transduction pathways may include at least one selected from the group consisting of BCL2L1, CASP3, CASP8, CASP9, CCNB2, CCND1, CDK1, CDK2, CDK4, CDKN1A, CDKN2A, CHEK1, GADD45A, PTEN, TP53, and TSC2, but are not limited thereto.
[0097] In the present invention, among the signal transduction pathways, the mTOR signaling pathway-related genes may include at least one selected from the group consisting of AKT1, DVL1, DVL2, GRB2, GSK3B, HRAS, IKBKB, INS, KRAS, MAPK1, MAPK3, MLST8, MTOR, NRAS, PIK3CA, PIK3R1, PTEN, RHEB, RHOA, RPTOR, SOS1, TNF, TNFRSF1A, and TSC2, but are not limited thereto.
[0098] In the present invention, genes related to the PI3K-Akt signaling pathway among the signal transduction pathways may include at least one selected from the group consisting of AKT1, BCL2L1, BRCA1, CASP9, CCND1, CDK2, CDK4, CDKN1A, GNB1, GRB2, GSK3B, HRAS, IKBKB, IKBKG, IL6, INS, JAK1, JAK2, JAK3, KRAS, MAPK1, MAPK3, MLST8, MTOR, NFKB1, NRAS, PIK3CA, PIK3R1, PTEN, RELA, RHEB, RPTOR, SOS1, SYK, TP53, and TSC2, but are not limited thereto.
[0099] In the present invention, the gene related to the apoptosis pathway Ⅰ among the signal transduction pathways may include at least one selected from the group consisting of ACTB, AKT1, BCL2L1, CASP3, CASP8, CASP9, GADD45A, HRAS, IKBKB, IKBKG, JUN, KRAS, MAPK1, MAPK3, NFKB1, NFKBIA, NRAS, PIK3CA, PIK3R1, RELA, TNF, TNFRSF1A, and TP53, but is not limited thereto.
[0100] In the present invention, the ferroptosis pathway-related gene among the signal transduction pathways may include at least one selected from the group consisting of ACSL4, GPX4, HMOX1, LPCAT3, NCOA4, TFRC, and TP53, but is not limited thereto.
[0101] In the present invention, the necroptosis pathway-related gene among the signal transduction pathways may include at least one selected from the group consisting of CASP8, IL1B, JAK1, JAK2, JAK3, STAT1, STAT3, TNF, TNFRSF1A, and TYK23, but is not limited thereto.
[0102] In the present invention, the cellular senescence pathway-related genes among the signal transduction pathways may include at least one selected from the group consisting of AKT1, CCNA2, CCNB2, CCND1, CDK1, CDK2, CDK4, CDKN1A, CDKN2A, CHEK1, GADD45A, HRAS, IL6, KRAS, MAPK1, MAPK3, MTOR, NFKB1, NRAS, PIK3CA, PIK3R1, PTEN, RELA, RHEB, SMAD2, SMAD3, TGFB1, TP53, and TSC2, but are not limited thereto.
[0103] In the present invention, among the signal transduction pathways, the Wnt signaling pathway-related genes may include at least one selected from the group consisting of AXIN1, AXIN2, CCND1, CTNNB1, DVL1, DVL2, EP300, GSK3B, JUN, LEF1, RHOA, SMAD3, SMAD4, and TP53, but are not limited thereto.
[0104] In the present invention, the Notch signaling pathway-related genes among the signal transduction pathways may include at least one selected from the group consisting of DVL1, DVL2, EP300, JAG1, MAML1, MAML2, MAML3, NOTCH1, NOTCH2, NOTCH3, NOTCH4, and RBPJ, but are not limited thereto.
[0105] In the present invention, among the signal transduction pathways, the genes related to the TGF-beta signaling pathway may include at least one selected from the group consisting of ACVR1, BMP4, BMPR1A, EP300, MAPK1, MAPK3, RHOA, SMAD2, SMAD3, SMAD4, TFRC, TGFB1, and TNF, but are not limited thereto.
[0106] In the present invention, among the signal transduction pathways, the JAK-STAT signaling pathway-related genes may include at least one selected from the group consisting of AKT1, BCL2L1, CCND1, CDKN1A, EP300, GRB2, HRAS, IL6, JAK1, JAK2, JAK3, MTOR, PIK3CA, PIK3R1, SOS1, STAT1, STAT3, and TYK2, but are not limited thereto.
[0107] In the present invention, genes related to the T cell receptor signaling pathway among the signal transduction pathways may include at least one selected from the group consisting of AKT1, CDK4, FYN, GRB2, GSK3B, HRAS, IKBKB, IKBKG, JUN, KRAS, LCK, MAPK1, MAPK3, NFKB1, NFKBIA, NRAS, PIK3CA, PIK3R1, RELA, RHOA, SOS1, and TNF, but are not limited thereto.
[0108] In the present invention, among the signal transduction pathways, the B cell receptor signaling pathway-related genes may include at least one selected from the group consisting of AKT1, BTK, GRB2, GSK3B, HRAS, IKBKB, IKBKG, JUN, KRAS, LYN, MAPK1, MAPK3, NFKB1, NFKBIA, NRAS, PIK3CA, PIK3R1, PLCG2, RELA, SOS1, and SYK, but are not limited thereto.
[0109] In the present invention, genes related to the TNF signaling pathway among the signal transduction pathways may include at least one selected from the group consisting of AKT1, CASP3, CASP8, IKBKB, IKBKG, IL1B, IL6, JAG1, JUN, MAPK1, MAPK3, NFKB1, NFKBIA, PIK3CA, PIK3R1, RELA, TNF, and TNFRSF1A, but are not limited thereto.
[0110] In the present invention, genes related to the transcriptional misregulation pathway in cancer among the signal transduction pathways may include at least one selected from the group consisting of BCL2L1, CCNA2, CD40, CDKN1A, GADD45A, H3-3B, H3C12, H3C13, IL6, NFKB1, RELA, RUNX1, and TP53, but are not limited thereto.
[0111] In the present invention, among the signal transduction pathways, the prostate cancer pathway-related genes may include at least one selected from the group consisting of AKT1, CASP9, CCND1, CDK2, CDKN1A, CTNNB1, EP300, GRB2, GSK3B, HRAS, IKBKB, IKBKG, INS, KRAS, LEF1, MAPK1, MAPK3, MTOR, NFKB1, NFKBIA, NRAS, PIK3CA, PIK3R1, PTEN, RELA, SOS1, and TP53, but are not limited thereto.
[0112] In the present invention, genes related to the ESR mediated signaling pathway among the signal transduction pathways may include at least one selected from the group consisting of AKT1, AXIN1, CCND1, EP300, GNB1, H3-3B, H3C12, H3C13, H4C6, HRAS, JUN, KRAS, MAPK1, MAPK3, NRAS, PIK3CA, PIK3R1, RUNX1, and SRC, but are not limited thereto.
[0113] In the present invention, the gene related to the estrogen dependent gene expression pathway among the signal transduction pathways may include at least one selected from the group consisting of AXIN1, CCND1, EP300, H3-3B, H3C12, H3C13, H4C6, JUN, and RUNX1, but is not limited thereto.
[0114] In the present invention, among the signal transduction pathways, the genes related to the androgen receptor network pathway in prostate cancer may include at least one selected from the group consisting of AKT1, BRCA1, CASP3, CASP8, CASP9, CCND1, CDK1, CDK2, CDK4, CHEK1, GRB2, HRAS, JAK1, JUN, MAPK1, MAPK3, MTOR, PIK3CA, PTEN, RHEB, RPTOR, SMAD2, SMAD3, SOS1, STAT1, STAT3, TP53, and TSC2, but are not limited thereto.
[0115] In the present invention, the method may further include a step of grouping samples of biological samples based on the Risk score value calculated above.
[0116] In the present invention, the grouping step may be to divide the risk score obtained by multiplying the regression coefficient matching each explanatory variable by the variable's unique value into a High group and a Low group according to the cut-off criteria disclosed in the embodiment of the present invention.
[0117] In the present invention, the cut-off value, which serves as the cut-off criterion, may be applied as a value within a margin of error of ±5% based on -257.50477. Specifically, a value within a margin of error of ±4% may be applied, more specifically, a value within a margin of error of ±3% may be applied, even more specifically, a value within a margin of error of ±2% may be applied, and most specifically, a value within a margin of error of ±1% may be applied.
[0118] In the present invention, the unique value of the variable is characterized by being a signal transduction pathway score derived from the gene expression level of Table 1 or the signal transduction pathway-related genes of Table 2.
[0119] In the present invention, the High group may be determined as a patient with a high possibility of cancer metastasis or a high risk of biochemical recurrence due to a low possibility of showing a beneficial response to anticancer treatment, and specifically, may be determined as a patient with a high possibility of prostate cancer metastasis or a high risk of biochemical recurrence.
[0120] In the present invention, the Low group may be determined as a patient with a low possibility of cancer metastasis or a low risk of biochemical recurrence, as the patient is likely to exhibit a beneficial response to anticancer treatment, and specifically, may be determined as a patient with a low possibility of prostate cancer metastasis or a low risk of biochemical recurrence.
[0121] The term "beneficial response" as used herein means improvement in any measure of patient status, such as overall survival, long-term survival, recurrence-free survival, and distant recurrence-free survival, which are commonly used in the art. Recurrence-free survival (RFS) refers to the time (in months) from surgery to the first local recurrence, regional recurrence, or distant recurrence. Distant recurrence-free survival (DRFS) or distant metastasis-free survival (DMFS) refers to the time (in months) from surgery to the first distant recurrence. Recurrence refers to RFS and / or DFRS. The term "long-term" survival as used herein refers to survival of at least 3 years, or at least 5 years, or at least 8 years, or at least 10 years after surgery or other treatment.
[0122]
[0123] According to another embodiment of the present invention, the present invention relates to a diagnostic device for predicting cancer risk.
[0124] In the present invention, the diagnostic device may include (a) an input unit for inputting a gene expression level and a signal transduction pathway score for a biological sample obtained from a subject; (b) a calculation unit for calculating a risk score of Equation 1 below using a Cox regression analysis algorithm based on a combination of at least one gene listed in Table 1 and at least one signal transduction pathway listed in Table 2; and (c) an output unit for grouping biological sample samples based on the calculated Risk score value to determine and output a High group and a Low group.
[0125] [Formula 1]
[0126]
[0127] In the present invention, the term "diagnostic device" means equipment capable of diagnosing diseases externally based on substances produced in the human body, such as blood, saliva, urine, etc., and includes, for example, an input unit; a calculation unit; and an output unit, etc., and is not limited as long as it is in a form capable of analyzing the level of gene or protein expression from the substance.
[0128] In the above diagnostic device of the present invention, the description of Cox regression analysis, cancer, cancer risk, subject, biological sample, gene, signal transduction pathway, signal transduction pathway score, regression coefficient, explanatory variable, High group, and Low group is the same as that described in the method for providing information on cancer risk, and thus is omitted to avoid excessive complexity of the present specification.
[0129] By applying statistical algorithms to newly discovered marker combinations that can precisely predict cancer risk, particularly the likelihood of metastasis or recurrence in prostate cancer, we can provide information on cancer risk, particularly prostate cancer risk and the likelihood of metastasis / recurrence. By predicting the likelihood of metastasis or recurrence early, we can expedite treatment decisions and significantly improve patient survival rates.
[0130] FIG. 1 is a schematic diagram of an algorithm that firstly classifies the risk of each patient and secondly determines molecular subtypes to predict drug responsiveness by using the score of genes related to gene expression patterns and signal transduction pathways through target RNA sequencing according to one embodiment of the present invention.
[0131] FIG. 2 is a diagram showing an ROC curve that verifies the performance of a RandomForest model using only DEG markers according to one embodiment of the present invention using TestSet.
[0132] FIG. 3 is a diagram showing an ROC curve that confirms the performance of a RandomForest model using only DEG markers according to one embodiment of the present invention with a validation set that was not used for model development.
[0133] FIG. 4 is a diagram showing an ROC curve that verifies a RandomForest model using only signal transmission pathway-related markers according to one embodiment of the present invention using TestSet.
[0134] FIG. 5 is a diagram showing an ROC curve that verifies a RandomForest model using only signal transmission pathway-related markers according to one embodiment of the present invention using a validation set that was not used in model development.
[0135] FIG. 6 is a diagram showing an ROC curve that verifies a RandomForest model using a combination of DEG markers and signal transmission pathway-related markers according to one embodiment of the present invention using TestSet.
[0136] FIG. 7 is a diagram showing an ROC curve that verifies a RandomForest model using a combination of DEG markers and signal transduction pathway-related markers according to one embodiment of the present invention using a validation set that was not used for model development.
[0137] Figure 8a is a Kaplan-Meier graph showing the distribution of metastasis-free survival (MFS) over the entire follow-up observation period according to one embodiment of the present invention.
[0138] Figure 8b is a Kaplan-Meier graph showing the distribution of biochemical relapse-free survival (BFS) over the entire follow-up period according to one embodiment of the present invention.
[0139] Figure 8c is a Kaplan-Meier graph showing the distribution of survival (OS) according to the entire follow-up observation period, according to one embodiment of the present invention.
[0140] Figure 8d is a Kaplan-Meier graph showing the distribution of metastasis-free survival (MFS) over a 5-year follow-up period according to one embodiment of the present invention.
[0141] Figure 8e is a Kaplan-Meier graph showing the distribution of biochemical relapse-free survival (BFS) over a 5-year follow-up period according to one embodiment of the present invention.
[0142] Figure 8f is a diagram showing the distribution of survival (OS) according to a 5-year follow-up observation period using a Kaplan-Meier graph according to one embodiment of the present invention.
[0143] Figure 9 is a diagram showing the results of predicting whether or not metastasis will occur based on a risk score calculated in Cox regression analysis according to one embodiment of the present invention.
[0144] Example
[0145]
[0146] Selection of target prostate cancer patients and preparation of test tissues
[0147] To screen genes associated with prognosis in prostate cancer, tissue samples were obtained from patients who underwent surgery after a diagnosis of prostate cancer but did not develop clinical recurrence (BCR) or metastasis (METS; no evidence of disease) (NED), patients with biochemical recurrence (BCR), and patients with metastasis (METZ). RNA was extracted from representative formalin-fixed paraffin-embedded (FFPE) blocks. The goal was to obtain a list of genes whose expression levels were differentially expressed among patients who underwent surgery after a diagnosis of prostate cancer (NED), patients with biochemical recurrence (BCR), and patients with metastasis (METZ). In-house Korean prostate cancer patient data were utilized. The sample data from this in-house Korean prostate cancer patient data were broadly divided into three patient groups: NED, BCR, and METS. Patients who underwent surgery for localized or locally advanced prostate cancer and were followed up were selected if they met the following three criteria. The first group of patients is the BCR and no metastases / recurrence group (NED) for 7 years after radical prostatectomy. The second group of patients is the BCR group (BCR) group, which is a group of patients who have experienced biochemical recurrence (blood PSA level increased by 0.2 ng / ml or more on two consecutive occasions after radical prostatectomy) but have not developed metastases within the 5-year follow-up period. The third group of patients is the METS group, which is a group of patients who have developed metastases within the 5-year follow-up period.
[0148]
[0149] RNA-sequencing
[0150] Before performing NGS analysis, the DQ score (DCgen Quality score, DQ score) system, which is a system that selects high-quality NGS libraries suitable for NGS data production by synthesizing the QC information generated, was used to select high-quality NGS libraries suitable for NGS data production. If it was confirmed to have a certain level of RNA quality, a cDNA library was created by selecting high-quality NGS libraries. Only specific genes were detected using gene panel probes. All candidate genes were sequenced using next-generation sequencing (NGS) equipment and aligned to the human genome. They were mapped to the public human genome reference using an alignment algorithm called STAR. The expression level of each gene was measured from the information of the aligned analysis sequences.
[0151]
[0152] Screening of markers related to prognosis, metastasis, or recurrence of prostate cancer based on RNA-sequencing expression information.
[0153] To discover the optimal biomarker combination, a step-by-step screening process was performed through 1) in-house data sample matching, 2) extraction of metastasis-related markers, 3) extraction of key markers related to prostate cancer signal transduction, and 4) extraction of housekeeping genes.
[0154] The patient samples used in the development were collected by matching the high clinical risk of patients with confirmed recurrence (estimated by tumor size, PSA level, and Gleason score) with clinical information. This is because when samples with low clinical risk are used to build a model, clinical indicators contribute most to recurrence, so only gene expression levels can be used as variables when building a model. At this time, in order to differentiate it from models that use existing clinical information, sample matching was performed using a statistical methodology called propensity score matching (PSM) so that the distributions of 1) cancer stage, 2) pre-treatment blood PSA level, and 3) Gleason score were similar. PSM matching was performed using the MatchIT package in R. Before matching, ANOVA tests and Kruskal tests were performed, and the necessity of matching was confirmed when the difference in variables by institution was less than a p-value of 0.05. PSM matching was performed with the nearest criterion of 0.8 in the MatchIT package as the tolerance. To enable analysis of differences in data range on a common scale, z-score (standardized score) was applied to log2(TPM) values.
[0155] To obtain a list of genes whose expression levels are increased or decreased in patients with metastasis compared to NED patients, Limma, a differentially expressed gene (DEG) analysis tool, was used. The selection criteria were genes with a p-value of 0.05 or less, which indicates the significance of the difference and the actual difference in expression levels, and a difference in the absolute log2 (fold-change) of 0.585 or more. Genes whose expression levels differed depending on whether the patients had metastasis were selected. To verify these genes, additional genes showing differences in expression levels were selected by combining factors that influence cancer metastasis, such as ISUP grade and T-stage grade, and then an additional step of comparative analysis was performed. As shown in Table 1, 103 significant genes were identified, and the results of GSEA analysis showed that the selected genes had a higher interpretability related to biological mechanisms and a high correlation with prostate cancer metastasis than other factors. Ultimately, they were selected as the optimal combination of biomarkers for diagnosing prostate cancer metastasis.
[0156] gene_idgene_namelogFCP valueabs_logFCENSG00000004776HSPB6-0.8513781890.0020388780.851378189ENSG00000008710PKD1-0.5895362560.0079586420.589536256ENSG00000034510TMSB102.3188484155.0705E-112.318848415ENSG00000061938TNK2-0.6400866053.68972E-060.640086605ENSG00000067225PKM0.7070830491.53094E-050.707083049ENSG00000071127WDR10.597693092.06496E-070.59769309ENSG00000072310SREBF1-1.0386935450.000904511.038693545ENSG00000074800ENO10.8408507721.09025E-060.840850772ENSG00000075624ACTB3.4541195034.9876E-073.454119503ENSG00000084207GSTP10.7984585260.0025045180.798458526ENSG00000087086FTL0.9297589720.0229501870.929758972ENSG00000089220PEBP11.1517138243.422E-081.151713824ENSG00000090006LTBP4-0.8605539470.0097468520.860553947ENSG00000092199HNRNPC0.8912723791.91104E-070.891272379ENSG00000092841MYL60.8960464630.0028728480.896046463ENSG00000099977DDT0.8145820991.6297E-070.814582099ENSG00000101439CST30.7306515080.0002006040.730651508ENSG00000102878HSF4-0.7322745884.08602E-050.732274588ENSG00000104419NDRG10.9650266810.0055515870.965026681ENSG00000104904OAZ10.7386965961.20419E-050.738696596ENSG00000106211HSPB11.3872529130.000292921.387252913ENSG00000107281NPDC1-0.7818102260.0001577340.781810226ENSG00000107796ACTA21.1626671480.0001375471.162667148ENSG00000107862GBF1-0.7104547070.0174457170.710454707ENSG00000109971HSPA80.6750120373.33441E-110.675012037ENSG00000110171TRIM3-0.6976552940.0001224930.697655294ENSG00000111077TNS2-0.5904710790.0202271820.590471079ENSG00000111252SH2B3-0.9691366311.60892E-050.969136631ENSG00000111341MGP0.930021750.0006324780.93002175ENSG00000111640GAPDH1.1427423671.94777E-071.142742367ENSG00000111775COX6A10.6287831891.26053E-080.628783189ENSG00000113140SPARC0.8944042514.76476E-060.894404251ENSG00000117713ARID1A-1.0602781452.75276E-061.060278145ENSG00000120738EGR1-1.003498730.0063606881.00349873ENSG00000120885CLU2.7572438214.42145E-092.757243821ENSG00000123358NR4A1-5.0655581434.05657E-085.065558143ENSG00000123416TUBA1B0.8373781471.41634E-090.837378147ENSG00000125356NDUFA10.7995192431.89763E-070.799519243ENSG00000125730C31.0921543384.1011E-061.092154338ENSG00000125740FOSB-1.9059772740.0119815821.905977274ENSG00000128016ZFP36-2.1417986873.02412E-052.141798687ENSG00000131037EPS8L1-0.8789790930.0013747910.878979093ENSG00000131095GFAP3.7262478098.48121E-053.726247809ENSG00000132470ITGB4-0.6640210830.0025619270.664021083ENSG00000132475H3-3B1.0012056787.65756E-051.001205678ENSG00000133112TPT11.3075101660.000456711.307510166ENSG00000133142TCEAL40.8178104370.0054334950.817810437ENSG00000133250ZNF414-1.3950452160.0020668921.395045216ENSG00000135404CD630.5853294441.74459E-070.585329444ENSG00000135486HNRNPA11.5398965414.3155E-051.539896541ENSG00000136156ITM2B0.6296112820.0019876250.629611282ENSG00000140400MAN2C1-0.6150563320.0052688650.615056332ENSG00000142156COL6A1-0.6780234970.0017879110.678023497ENSG00000142515KLK3-10.477635060.00076448310.47763506ENSG00000142599RERE-0.8704199429.36718E-070.870419942ENSG00000146067FAM193B-0.6146478220.0129232320.614647822ENSG00000146205ANO7-1.6140528545.37859E-101.614052854ENSG00000149806FAU0.7753578053.30801E-080.775357805ENSG00000149925ALDOA1.0156746924.37283E-071.015674692ENSG00000156508EEF1A11.6551875941.78296E-061.655187594ENSG00000158715SLC45A3-1.8094760132.57772E-051.809476013ENSG00000159674SPON2-2.3728628330.0077123962.372862833ENSG00000163041H3-3A0.6358115140.0044394810.635811514ENSG00000164692COL1A20.723872762.36013E-100.72387276ENSG00000166165CKB1.0428677220.0013148761.042867722ENSG00000166444DENND2B-0.6566261110.0121911650.656626111ENSG00000166710B2M2.5397947210.0190058312.539794721ENSG00000167615LENG8-1.0788582911.90007E-071.078858291ENSG00000167642SPINT20.5895419910.0007815660.589541991ENSG00000167658EEF22.2063793460.0159636562.206379346ENSG00000167749KLK4-0.9506628670.0026706560.950662867ENSG00000167996FTH11.9426348294.94462E-081.942634829ENSG00000169045HNRNPH11.9304231837.37271E-051.930423183ENSG00000169926KLF13-0.6457839590.002370950.645783959ENSG00000170345FOS-2.5798835510.0029992962.579883551ENSG00000171223JUNB-0.7706724610.0035091790.770672461ENSG00000177685CRACR2B-0.7860919892.27557E-050.786091989ENSG00000179094PER1-1.0401991430.0003174631.040199143ENSG00000181090EHMT1-0.6611780770.0004329130.661178077ENSG00000181163NPM10.8078896680.0022681780.807889668ENSG00000183889NPIPA6-0.5937876460.0484294860.593787646ENSG00000184009ACTG11.1384500369.60393E-121.138450036ENSG00000184254ALDH1A3-1.4962454352.27668E-081.496245435ENSG00000187244BCAM-0.7292334430.0117935630.729233443ENSG00000188257PLA2G2A0.9364135590.0353842230.936413559ENSG00000196126HLA-DRB11.0292104189.16255E-091.029210418ENSG00000196262PPIA1.4535063756.82079E-141.453506375ENSG00000196498NCOR2-0.61240550.0292152590.6124055ENSG00000197530MIB2-0.754085940.0013834460.75408594ENSG00000197746PSAP0.6054375895.69857E-050.605437589ENSG00000197903H2BC121.0202134087.21446E-081.020213408ENSG00000197971MBP4.2092817670.00013514.209281767ENSG00000204389HSPA1A0.6518087150.0065809410.651808715ENSG00000204628RACK10.6854785770.0305744790.68547 8577ENSG00000205542TMSB4X1.5784523811.52324E-051.578452381ENSG0000023 3024NPIPA9-0.6435793950.0168807530.643579395ENSG00000240972MIF1.19512 90250.0026090811.195129025ENSG00000246705H2AJ-0.6075997250.0134350110. 607599725ENSG00000250479CHCHD100.6923582850.0075903320.692358285ENSG0 0000251562MALAT1-1.0694685120.0048273871.069468512ENSG00000254772EEF1 G0.9411978157.8788E-050.941197815ENSG00000266714MYO15B-1.564944639.48 967E-061.56494463ENSG00000273542H4C120.7734277858.7048E-070.773427785.
[0157]
[0158] Screening and analysis of biomarkers related to signaling pathways associated with prostate cancer metastasis
[0159] Pathway analysis was performed to determine the biological mechanisms associated with the biomarkers selected through the screening above. A list of signal transduction pathways containing keywords related to prostate cancer and its development was selected, and a list of pathways related to prostate cancer was selected from the results of a GSEA analysis based on genes showing differences in expression levels. The MSigDB database, a related database, was used to extract a list of genes included in the signal transduction pathways. To identify the central genes that are strongly connected to other signals and have the most interactions in the signal transduction pathway analysis, the STRING DB, a protein-protein interaction database, was used. Based on the protein-protein interaction database, the top 10 central genes with the highest centrality in each signal transduction pathway and the highest number of connections between two other genes were selected, resulting in a final selection of 103 central genes.
[0160] As a result of the analysis, as shown in Table 2 below, among the signaling pathways closely related to prostate cancer metastasis, 33 signaling pathways with the highest centrality were discovered, and among the biomarkers acting on the signaling pathways, the combination of biomarkers with the most close correlation was derived.
[0161] DB source genes (signaling pathway) genes (selected genes)MSigDBProstate cancerAKT1, CASP9, CCND1, CDK2, CDKN1A, CTNNB1, EP300, GRB2, GSK3B, HRAS, IKBKB, IKBKG, . INS, KRAS, LEF1, MAPK1, MAPK3, MTOR, NFKB1, NFKBIA, NRAS, PIK3CA, PIK3R1, PTEN, RELA, SOS1, TP53MSigDBTranscriptional misregulation in cancerBCL2L1, CCNA2, CD40, CDKN1A, GADD45A, H3-3B, H3C12, H3C13, IL6, NFKB1, RELA, RUNX1, TP53MSigDBTNF signaling pathwayAKT1, CASP3, CASP8, IKBKB, IKBKG, IL1B, IL6, JAG1, JUN, MAPK1, MAPK3, NFKB1, NFKBIA, PIK3CA, PIK3R1, RELA, TNF, . TNFRSF1AMSigDBB cell receptor signaling pathwayAKT1, BTK, GRB2, GSK3B, HRAS, IKBKB, IKBKG, JUN, KRAS, LYN, MAPK1, MAPK3, NFKB1, NFKBIA, NRAS, PIK3CA, PIK3R1, PLCG2, RELA, SOS1, SYKMSigDBT cell receptor signaling pathwayAKT1, CDK4, FYN, GRB2, GSK3B, HRAS, IKBKB, IKBKG, JUN, KRAS, LCK, MAPK1, MAPK3, NFKB1, NFKBIA, NRAS, PIK3CA, PIK3R1, RELA, RHOA, SOS1, TNFMSigDBJAK-STAT signaling pathwayAKT1, BCL2L1, CCND1, CDKN1A, EP300, GRB2, HRAS, IL6, JAK1, JAK2, JAK3, MTOR, PIK3CA,PIK3R1, SOS1, STAT1, STAT3, TYK2MSigDBTGF-beta signaling pathwayACVR1, BMP4, BMPR1A, EP300, MAPK1, MAPK3, RHOA, SMAD2, SMAD3, SMAD4, TFRC, TGFB1, TNFMSigDBNotch signaling pathwayDVL1, DVL2, EP300, JAG1, MAML1, MAML2, MAML3, NOTCH1, NOTCH2, NOTCH3, NOTCH4, RBPJMSigDBWnt signaling pathwayAXIN1, AXIN2, CCND1, CTNNB1, DVL1, DVL2, EP300, GSK3B, JUN, LEF1, RHOA, SMAD3, SMAD4, TP53MSigDBCellular senescenceAKT1, CCNA2, CCNB2, CCND1, CDK1, CDK2, CDK4, CDKN1A, CDKN2A, CHEK1, GADD45A, HRAS, IL6, KRAS, MAPK1, MAPK3, MTOR, NFKB1, NRAS, PIK3CA, PIK3R1, PTEN, RELA, RHEB, SMAD2, SMAD3, TGFB1, TP53, TSC2MSigDBNecroptosisCASP8, IL1B, JAK1, JAK2, JAK3, STAT1, STAT3, TNF, TNFRSF1A, TYK2MSigDBFerroptosisACSL4, GPX4, HMOX1, LPCAT3, NCOA4, TFRC, TP53MSigDBApoptosisⅠACTB, AKT1, BCL2L1, CASP3, CASP8, CASP9, GADD45A, HRAS, IKBKB, IKBKG, JUN, KRAS, MAPK1, MAPK3, NFKB1, NFKBIA, NRAS, PIK3CA, PIK3R1, RELA, TNF, TNFRSF1A, TP53MSigDBPI3K-Akt signaling pathwayAKT1, BCL2L1, BRCA1, CASP9, CCND1, CDK2, CDK4, CDKN1A,GNB1, GRB2, GSK3B, HRAS, IKBKB, IKBKG, IL6, INS, JAK1, JAK2, JAK3, KRAS, MAPK1, MAPK3, MLST8, MTOR, NFKB1, NRAS, PIK3CA, PIK3R1, PTEN, RELA, RHEB, RPTOR, SOS1, SYK, TP53, TSC2MSigDBmTOR signaling pathwayAKT1, DVL1, DVL2, GRB2, GSK3B, HRAS, IKBKB, INS, KRAS, MAPK1, MAPK3, MLST8, MTOR, NRAS, PIK3CA, PIK3R1, PTEN, RHEB, RHOA, RPTOR, SOS1, TNF, TNFRSF1A, TSC2MSigDBp53 signaling pathwayBCL2L1, CASP3, CASP8, CASP9, CCNB2, CCND1, CDK1, CDK2, CDK4, CDKN1A, CDKN2A, CHEK1, GADD45A, PTEN, TP53, TSC2MSigDBCell cycleCCNA2, CCNB2, CCND1, CDK1, CDK2, CDK4, CDKN1A, CDKN2A, CHEK1, EP300, GADD45A, GSK3B, MCM7, SMAD2, SMAD3, SMAD4, TGFB1, TP53MSigDBNF-kappa B signaling pathwayBCL2L1, BTK, CD40, GADD45A, IKBKB, IKBKG, IL1B, LCK, LYN, NFKB1, NFKBIA, PLCG2, RELA, SYK, TNF, TNFRSF1A, TRAF6MSigDBChemokine signaling pathwayAKT1, CCR5, CXCR4, GNB1, GRB2, GSK3B, HRAS, IKBKB, IKBKG, JAK2, JAK3, KRAS, LYN, MAPK1, MAPK3, NFKB1, NFKBIA, NRAS, PIK3CA, PIK3R1, PLCG2, RELA, RHOA, SOS1, SRC, STAT1,STAT3MSigDBRap1 signaling pathwayACTB, AKT1, CTNNB1, HRAS, INS, KRAS, MAPK1, MAPK3, NRAS, PIK3CA, PIK3R1, RHOA, SRCMSigDBRas signaling pathwayAKT1, BCL2L1, GNB1, GRB2, HRAS, IKBKB, IKBKG, INS, KRAS, MAPK1, MAPK3, NFKB1, NRAS, PIK3CA, PIK3R1, PLCG2, RELA, RHOA, SOS1MSigDBErbB signaling pathwayAKT1, CDKN1A, GRB2, GSK3B, HRAS, JUN, KRAS, MAPK1, MAPK3, MTOR, NRAS, PIK3CA, PIK3R1, PLCG2, SOS1, SRCMSigDBMAPK signaling pathwayAKT1, CASP3, GADD45A, GRB2, HRAS, IKBKB, IKBKG, IL1B, INS, JUN, KRAS, MAPK1, MAPK3, NFKB1, NRAS, RELA, SOS1, TGFB1, TNF, TNFRSF1A, TP53, TRAF6MSigDBWNT beta CATENIN signalingAXIN1, AXIN2, CTNNB1, DVL2, JAG1, LEF1, MAML1, NOTCH1, NOTCH4, RBPJ, TP53MSigDBNOTCH signalingCCND1, DVL1, DVL2, EP300, JAG1, MAML1, MAML2, MAML3, NOTCH1, NOTCH2, NOTCH3, NOTCH4, RBPJMSigDBEstrogen Response LATECCND1, JAK1, JAK2MSigDBEstrogen Response EARLYCCND1, JAK2MSigDBE2F TargetsBRCA1, CCNB2, CDK1, CDK4, CDKN1A, CDKN2A, CHEK1, MCM7, TFRC, TP53MSigDBApoptosisⅡBCL2L1, BRCA1, CASP3, CASP8, CASP9,CCND1, CDK2, CDKN1A, CTNNB1, GADD45A, GPX4, HMOX1, IL1B, IL6, JUN, LEF1, RELA, TNFMSigDBAdipogenesisACLY, ACOX1, CS, GADD45A, GPX4, LPCAT3, SDHBMSigDBESR mediated signalingAKT1, AXIN1, CCND1, EP300, GNB1, H3-3B, H3C12, H3C13, H4C6, HRAS, JUN, KRAS, MAPK1, MAPK3, NRAS, PIK3CA, PIK3R1, RUNX1, SRCMSigDBEstrogen dependent gene expressionAXIN1, CCND1, EP300, H3-3B, H3C12, H3C13, H4C6, JUN, RUNX1MSigDBAndrogen receptor network in prostate cancerAKT1, BRCA1, CASP3, CASP8, CASP9, CCND1, CDK1, CDK2, CDK4, CHEK1, GRB2, HRAS, JAK1, JUN, MAPK1, MAPK3, MTOR, PIK3CA, PTEN, RHEB, RPTOR, SMAD2, SMAD3, SOS1, STAT1, STAT3, TP53, TSC2,
[0162]
[0163] 수별 창도 방탄소년단
[0164] The housekeeping gene list was prepared using the HRT Atlas v1.0 database, and only overlapping gene lists from the external public data, TCGA, and in-house data were used. The list extraction step was performed under appropriate experimental conditions and to provide evidence of good sample quality. We aimed to prioritize the selection of a list of genes with high expression levels that could minimize the influence of experimental noise and other factors and ensure stable expression levels. In addition, we analyzed the variability of each gene to further select genes with low variability that ensure consistent expression patterns and reproducibility of experimental results. In each dataset, genes satisfying high expression and low variance were selected, and the top 10 reference genes were selected by assigning rank scores to 23 genes. The top 10 genes were derived by checking all statistical values, including minimum and maximum values of gene expression levels and coefficient of variation, and consist of PPP2R1A, LAMTOR1, RNF167, ENSA, AP2M1, RNF10, BANF1, SLC25A3, APH1A, and DNAJB2.
[0165] Finally, 103 genes exhibiting differential expression levels and 103 core genes selected from 33 signaling pathways were selected. Among the genes exhibiting differential expression levels and genes selected from signaling pathways, H3-3B and ACTB were identified as common biomarkers. Therefore, a total of 204 genes were derived as the final biomarker combination for diagnosing prostate cancer prognosis, metastasis, and / or recurrence.
[0166]
[0167] Validation of prostate cancer prognosis prediction
[0168] Clinical validation was conducted on 172 prostate cancer patients who had been followed for more than five years. For all patients, risk scores were calculated using the prognostic tool described above, and the product's performance was compared with that of existing products. In order to improve the prediction ability during the verification process, a random forest algorithm classification model was applied, and 103 DEG markers showing differences in gene expression levels (ALDH1A3, ANO7, CRACR2B, ENO1, EPS8L1, HSPA1A, ITGB4, KLK3, KLK4, MALAT1, NPM1, PLA2G2A, SLC45A3, SPINT2, ZFP36, ACTA2, ACTB, ACTG1, ALDOA, ARID1A, B2M, BCAM, C3, CD63, CHCHD10, CKB, CLU, COL1A2, COL6A1, COX6A1, CST3, DDT, DENND2B, EEF1A1, EEF1G, EEF2, EGR1, EHMT1, FAM193B, FAU, FOS, FOSB, FTH1, FTL, GAPDH, GBF1, GFAP, GSTP1, H2AJ, H2BC12, H3-3A, H3-3B, H4C12, HLA-DRB1, HNRNPA1, HNRNPC, HNRNPH1, HSF4, HSPA8, HSPB1, HSPB6, ITM2B, JUNB, KLF13, LENG8, LTBP4, MAN2C1, MBP, MGP, MIB2, MIF, MYL6, MYO15B, NCOR2, NDRG1, NDUFA1, NPDC1, NPIPA6, NPIPA9, NR4A1, OAZ1, PEBP1, PER1, PKD1, PKM, PPIA, PSAP, RACK1, RERE, SH2B3, SPARC, SPON2, SREBF1, TCEAL4, TMSB10, TMSB4X, TNK2, TNS2, Results using only markers related to signaling pathways (ACTB, H3-3B, ACLY, ACOX1, ACSL4, TPT1, TRIM3, TUBA1B, WDR1, and ZNF414) (see Figs. 2 and 3)ACVR1, AKT1, AXIN1, AXIN2, BCL2L1, BMP4, BMPR1A, BRCA1, BTK, CASP3, CASP8, CASP9, CCNA2, CCNB2, CCND1, CCR5, CD40, CDK1, CDK2, CDK4, CDKN1A, CDKN2A, CHEK1, CS, CTNNB1, CXCR4, DVL1, DVL2, EP300, FYN, GADD45A, GNB1, GPX4, GRB2, GSK3B, H3C12, H3C13, H4C6, HMOX1, HRAS, IKBKB, IKBKG, IL1B, IL6, INS, JAG1, JAK1, JAK2, JAK3, JUN, KRAS, LCK, LEF1, LPCAT3, The results using only LYN, MAML1, MAML2, MAML3, MAPK1, MAPK3, MCM7, MLST8, MTOR, NCOA4, NFKB1, NFKBIA, NOTCH1, NOTCH2, NOTCH3, NOTCH4, NRAS, PIK3CA, PIK3R1, PLCG2, PTEN, RBPJ, RELA, RHEB, RHOA, RPTOR, RUNX1, SDHB, SMAD2, SMAD3, SMAD4, SOS1, SRC, STAT1, STAT3, SYK, TFRC, TGFB1, TNF, TNFRSF1A, TP53, TRAF6, TSC2, and TYK2) (see Figs. 4 and 5) and the results using DEG markers and signal transduction pathway-related markers (see Figs. 6 and 7) were confirmed. Referring to Fig. 2, it was confirmed that when the DEG marker was used alone, the AUC value was 0.892, indicating excellent diagnostic performance. Referring to Fig. 6, when the signal transmission pathway-related marker was applied in an overlapping manner with the DEG marker, the AUC value was greatly improved to 0.904, indicating that the performance of the model was significantly improved (see Table 3 below).
[0169] Feature SetDatasetSensitivitySpecificityPPVNPVF1ACCAUCDEG + pathwayRF_testSet0.85290.81940.69050.92190.76320.83020.904RF_validSet0.56720.80950 .65520.74560.6080.71510.8203DEGRF_testSet0.67650.97220.920.86420.77970.87740.8922R F_validSet0.46270.95240.86110.73530.60190.76160.8183PathwayRF_testSet0.70590.61110 .46150.81480.55810.64150.6801RF_validSet0.44780.77140.55560.68640.49590.64530.6748
[0170] As can be seen from the above results, when using a combination of 103 DEG markers and 103 signaling pathway-related markers of the present invention, it was confirmed that it is possible to predict the prognosis of prostate cancer, which was a limitation of the existing prostate cancer diagnostic model, and further, to more precisely determine the possibility of metastasis and recurrence of prostate cancer. Therefore, it is expected that this will greatly contribute to improving the survival rate of prostate cancer patients.
[0171] Input data processing for prostate cancer risk and prognosis prediction algorithms
[0172] The 103 genes and 33 signal transduction pathways derived from the above-described DEG markers were used as explanatory variables, and the 103 gene markers listed in Table 1 and the 33 signal transduction pathway-related gene markers listed in Table 2 were used here. Each event information was used as a response variable, and the initial analysis data preprocessing was performed by excluding missing values. The list of genes showing differences in expression levels reflects the difference in gene expression levels between metastatic and non-metastatic patients by utilizing data converted to log2(TPM), and the list of central genes (signal transduction pathway-related genes) of the signal transduction pathway was used for ssGSVA analysis to calculate the degree of activation of each signal transduction pathway for each sample as a score. A positive number indicates that the pathway is highly activated, and a negative number indicates that the pathway is relatively inactivated. The generated data were normalized to the log2(TPM) scale to enable analysis on a common scale with gene expression levels. Based on the selected biomarkers, we explored the correlation between the f / u period and events (Meta, BCR, OS) and performed Cox regression analysis to estimate characteristics influencing the time to event occurrence.
[0173]
[0174] Cox regression analysis algorithm
[0175] The regression coefficient estimated by Cox regression analysis indicates the influence of each explanatory variable on survival time, and based on this, the product of the values of the explanatory variables (differentially expressed gene expression levels, signal transduction pathway scores) and the Cox regression coefficients of the corresponding variables was integrated to calculate the risk score for each patient. The regression coefficients of each explanatory variable estimated in the regression model and the variable-specific values were used to calculate the risk score for each patient, and the performance of determining whether the score was transferred was evaluated, and the risk and risk level of the variable were predicted according to the regression coefficient values. This can be expressed as a formula as shown in Equation 1 below, where p represents the number of explanatory variables.
[0176] [Formula 1]
[0177]
[0178] The regression coefficients for each gene estimated from the regression analysis are shown in Table 4 below, and the regression coefficients for each signal transduction pathway estimated from the regression analysis are shown in Table 5 below. The final risk score was calculated by summing the columns of "Each term of the risk score" in Tables 4 and 5 below. In "Each term of the risk score," the gene name indicates the expression level of each gene, and the signal transduction pathway name indicates the signal transduction pathway score. Here, the signal transduction pathway score is characterized in that it is calculated from the signal transduction pathway-related genes in Table 2 included in each signal transduction pathway using the GSVA (Gene Set Variation Analysis) program. Table 4 below shows the 103 DEG markers derived previously as explanatory variables, and Table 5 below shows 33 signal transduction pathway scores calculated from the signal transduction pathway-related genes as explanatory variables.
[0179] Description of variable 회귀 계수Risk Score의 그하EEF1A1-24.5157-24.5157 * EEF1A1PKM-20.6784-20.6784 * PKMCST3-20.6048-20.6048 * CST3RACK1-19.7729-19.7729 * RACK1HSPA8-16.1424-16.1424 * HSPA8FTH1-14.4929-14.4929 * FTH1ARID1A-10.9563-10.9563 * ARID1ASLC45A3-10.8547-10.8547 * SLC45A3TPT1-10.3515-10.3515* TPT1SPARC-9.7075-9.7075* SPARCNCOR2-9.6148-9.6148 * NCOR2RERE-8.7200-8.72 * REREALDH1A3-8.1487-8.1487 * ALDH1A3ACTA2-7.9225-7.9225 * ACTA2EGR1-7.5508-7.5508 * EGR1TCEAL4-7.0247-7.0247 * TCEAL4MYL6-6.4505-6.4505 * MYL6HSF4-6.1940-6.194 * HSF4EHMT1-5.9082-5.9082 * EHMT1MYO15B-5.7667-5.7667 * MYO15BCOX6A1-5.6728-5.6728 * COX6A1EEF1G-4.6997-4.6997 * EEF1GH3-3B-4.3425-4.3425 * H3-3BDENND2B-4.1949-4.1949 * DENND2BPLA2G2A-4.0087-4.0087 * PLA2G2AJUNB-3.8773-3.8773 * JUNBTNS2-3.8637-3.8637 * TNS2MAN2C1-3.8205-3.8205 * MAN2C1NDRG1-3.7061-3.7061 * NDRG1SH2B3-3.5983-3.5983 * SH2B3GSTP1-3.5252-3.5252 * GSTP1TNK2-3.1408-3.1408 * TNK2NR4A1-3.0147-3.0147 * NR4A1C3-3.0057-3.0057 * C3ITM2B-3.0036-3.0036 * ITM2BHNRNPH1-2.9322-2.9322 * HNRNPH1PER1-2.7527-2.7527 * PER1ZNF414-2.2967-2.2967 * ZNF414HLA-DRB1-2.2345-2.2345 * HLA-DRB1MBP-2.1829-2.1829 * MBPNPIPA9-2.1081-2.1081 * NPIPA9OAZ1-1.8119-1.8119 * OAZ1SPON2-1.5800-1.58 * SPON2ITGB4-1.2803-1.2803 * ITGB4SPINT2-1.1729-1.1729 * SPINT2MIF-1.0697-1.0697 * MIFMALAT1-0.8059-0.8059 * MALAT1ZFP36-0.7027-0.7027 * ZFP36PKD1-0.5050-0.505 * PKD1H4C12-0.2988-0.2988 * H4C12SREBF10.39010.3901 * SREBF1COL1A20.53080.5308 * COL1A2HSPB10.62780.6278 * HSPB1TMSB4X0.69090.6909 * TMSB4XKLF130.89070.8907 * KLF13FOSB0.98920.9892 * FOSBACTG11.44131.4413 * ACTG1GFAP1.48051.4805 * GFAPGBF11.73821.7382 * GBF1WDR11.83341.8334 * WDR1H2AJ1.85471.8547 * H2AJH2BC122.07122.0712 * H2BC12EPS8L12.07532.0753 * EPS8L1CLU2.37052.3705 * CLUGAPDH2.44152.4415 * GAPDHMIB22.84162.8416 * MIB2NPM13.05333.0533 * NPM1NDUFA13.09953.0995 * NDUFA1ACTB3.34353.3435 * ACTBTUBA1B3.36503.365 * TUBA1BCHCHD103.41103.411 * CHCHD10TRIM33.62633.6263 * TRIM3FTL4.07494.0749 * FTLCD634.11984.1198 * CD63KLK44.15714.<h2 style=";text-align:left;direction:ltr">1571 * KLK4HSPB64.27074.2707 * HSPB6EEF24.46454.4645 * EEF2TMSB104.51234.5123 * TMSB10DDT4.54374.5437 * DDTH3-3A4.57344.5734 * H3-3AFAM193B4.79824.7982 * FAM193BHSPA1A4.90584.9058 * HSPA1ANPDC15.39455.3945 * NPDC1PSAP5.47865.4786 * PSAPCKB5.63185.6318 * CKBCRACR2B5.79435.7943 * CRACR2BANO76.61796.6179 * ANO7KLK36.87426.8742 * KLK3LTBP46.91516.9151 * LTBP4NPIPA68.21848.2184 * NPIPA6PPIA8.61028.6102 * PPIALENG88.79828.7982 * LENG8MGP8.94088.9408 * MGPENO18.94588.9458 * ENO1PEBP19.26089.2608 * PEBP1B2M9.52079.5207 * B2MFOS11.219711.2197 * FOSHNRNPC12.275812.2758 * HNRNPCBCAM12.284312.2843 * BCAMFAU13.603613.6036 * FAUCOL6A113.755313.7553 * COL6A1HNRNPA113.791913.7919 * HNRNPA1ALDOA19.599119.5991 * ALDOA.<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0180] 설명 변수회귀 계수Risk Score의 각 항p53_signaling_pathway-12.2718-12.2718 * p53_signaling_pathwayChemokine_signaling_pathway-11.3817-11.3817 * Chemokine_signaling_pathwayESR_MEDIATED_SIGNALING-11.3383-11.3383 * ESR_MEDIATED_SIGNALINGB_cell_receptor_signaling_pathway-9.4816-9.4816 * B_cell_receptor_signaling_pathwaymTOR_signaling_pathway-8.1686-8.1686 * mTOR_signaling_pathwayWnt_signaling_pathway-7.7882-7.7882 * Wnt_signaling_pathwayTNF_signaling_pathway-5.8849-5.8849 * TNF_signaling_pathwayAPOPTOSIS Ⅱ-2.4220-2.422 * APOPTOSIS ⅡNOTCH_SIGNALING-1.5922-1.5922 * NOTCH_SIGNALINGProstate_cancer-1.4705-1.4705 * Prostate_cancerANDROGEN_RECEPTOR_NETWORK_IN_PROSTATE_CANCER-1.2326-1.2326 * ANDROGEN_RECEPTOR_NETWORK_IN_PROSTATE_CANCERFerroptosis-0.9126-0.9126 * FerroptosisE2F_TARGETS-0.4964-0.4964 * E2F_TARGETSWNT_BETA_CATENIN_SIGNALING-0.1242-0.1242 * WNT_BETA_CATENIN_SIGNALINGESTROGEN_RESPONSE_EARLY-0.0156-0.0156 * ESTROGEN_RESPONSE_EARLYMAPK_signaling_pathway0.49500.495 * MAPK_signaling_pathwayApoptosisⅠ1.32091.3209 * ApoptosisⅠESTROGEN_DEPENDENT_GENE_EXPRESSION1.35571.3557 * ESTROGEN_DEPENDENT_GENE_EXPRESSIONNF_kappa_B_signaling_pathway1.50111.5011 * NF_kappa_B_signaling_pathwayRap1_signaling_pathway1.59301.593 * Rap1_signaling_pathwayNecroptosis1.59481.5948 * NecroptosisCell_cycle1.63491.6349 * Cell_cycleTGF_beta_signaling_pathway1.75691.7569 * TGF_beta_signaling_pathwayJAK_STAT_signaling_pathway1.93881.9388 * JAK_STAT_signaling_pathwayESTROGEN_RESPONSE_LATE2.23482.2348 * ESTROGEN_RESPONSE_LATEADIPOGENESIS2.60372.6037 * ADIPOGENESISTranscriptional_misregulation_in_cancer2.67462.6746 * Transcriptional_misregulation_in_cancerPI3K_Akt_signaling_pathway3.60563.6056 * PI3K_Akt_signaling_pathwayNotch_signaling_pathway5.80375.8037 * Notch_signaling_pathwayCellular_senescence7.71217.7121 * Cellular_senescenceErbB_signaling_pathway9.39899.3989 * ErbB_signaling_pathwayRas_signaling_pathway10.008310.0083 * Ras_signaling_pathwayT_cell_receptor_signaling_pathway11.622011.622 * T_cell_receptor_signaling_pathway.
[0181] Risk score-based sample grouping
[0182] Based on the risk score calculated from the Cox regression analysis, a receiver operating characteristic (ROC) curve was constructed to evaluate the patient's likelihood of metastasis. The ROC curve is a method used to evaluate the performance of a model by visually representing the true positive rate and false positive rate at different thresholds. The closer the area under the ROC curve (AUC) is to 1, the higher the predictive ability. Therefore, the optimal cutoff value of -257.50477 was derived based on the point where the AUC was maximum. Based on the cutoff value, risk scores above -257.50477 were classified into the High group, and those below -257.50477 were classified into the Low group, indicating a low risk. The distribution differences in survival / metastasis / biochemical recurrence between the groups were verified using Kaplan-Meier survival analysis and the log-rank test, and are shown in Table 6 and Figures 8a to 8f.
[0183] As a result of the experiment, the distribution differences for metastasis and biochemical recurrence in the Cox model both showed p-value < 0.05 or less, which confirmed that the difference in the distributions for not only the entire follow-up period but also the 5-year metastasis-free survival and 5-year biochemical recurrence-free survival was significant. However, considering that the cutoff standard was derived from the model that determines whether or not there was metastasis, the difference in the distribution of survival was not significant regardless of the follow-up period with p-value > 0.05, but it was confirmed that the risk of HR was predicted to be higher in the High group than in the Low group.
[0184] Follow-upSurvival TypeP-valueHR(Hazard Ratio)Complete FUMFSp = 0.00152.691BFSp = 0.00411.956OSP = 0.11.8015-year FUMFSp = 0.00152.88BFSp = 0.00382.006OSP=0.0892.852
[0185] Validation of prostate cancer prognosis prediction based on algorithm application
[0186] We aimed to further verify whether the risk score of a Cox regression multivariate analysis model that combined the expression levels of 103 DEG genes and the activation scores of 33 signaling pathways could be used to precisely predict the metastasis of prostate cancer.
[0187] The Cox regression analysis model above showed an AUC of 0.632 and an ACC of 0.6792 for predicting the likelihood of prostate cancer metastasis. The ROC curve is shown in Fig. 3. Based on the results, the present invention was able to provide information on the risk or prognosis of prostate cancer through a risk score derived by applying DEG markers and signal transduction pathway-related markers. It is expected that this can significantly contribute to improving the survival rate of prostate cancer patients by predicting the likelihood of early metastasis.
[0188]
[0189] While specific aspects of the present invention have been described in detail above, it should be apparent to those skilled in the art that these specific descriptions are merely preferred embodiments and do not limit the scope of the present invention. Therefore, the substantial scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. In a method of providing information on cancer risk, (a) a step of calculating gene expression levels and signal transduction pathway scores for biological samples obtained from a subject and classifying them as explanatory variables; (b) calculating a risk score of the following formula 1 using a Cox regression analysis algorithm with a combination of at least one gene described in Table 1 and at least one signal transduction pathway described in Table 2; and (c) a method comprising a step of grouping samples of a biological specimen based on the Risk score value calculated above; [Formula 1] In the above equation 1, The above p represents the number of explanatory variables, The above explanatory variables represent gene expression levels or signal transduction pathway scores.
2. In paragraph 1, The genes are ALDH1A3, ANO7, CRACR2B, ENO1, EPS8L1, HSPA1A, ITGB4, KLK3, KLK4, MALAT1, NPM1, PLA2G2A, SLC45A3, SPINT2, ZFP36, ACTA2, ACTB, ACTG1, ALDOA, ARID1A, B2M, BCAM, C3, CD63, CHCHD10, CKB, CLU, COL1A2, COL6A1, COX6A1, CST3, DDT, DENND2B, EEF1A1, EEF1G, EEF2, EGR1, EHMT1, FAM193B, FAU, FOS, FOSB, FTH1, FTL, GAPDH, GBF1, GFAP, GSTP1, H2AJ, H2BC12, H3-3A, H3-3B, H4C12, HLA-DRB1, HNRNPA1, HNRNPC, HNRNPH1, HSF4, HSPA8, HSPB1, HSPB6, ITM2B, JUNB, KLF13, LENG8, LTBP4, MAN2C1, MBP, MGP, MIB2, MIF, MYL6, MYO15B, NCOR2, NDRG1, NDUFA1, NPDC1, NPIPA6, NPIPA9, NR4A1, OAZ1, PEBP1, PER1, PKD1, PKM, PPIA, PSAP, RACK1, RERE, SH2B3, SPARC, SPON2, SREBF1, TCEAL4, TMSB10, TMSB4X, TNK2, TNS2, TPT1, TRIM3, TUBA1B, WDR1 and ZNF414 A method characterized by at least one gene selected from the group.
3. In paragraph 1, The above signal transduction pathways include the Adipogenesis pathway, the Apoptosis pathway Ⅱ, the E2F targets pathway, the Estrogen Response Early pathway, the Estrogen Response Late pathway, the NOTCH signaling pathway, the WNT-Beta Catenin signaling pathway, the MAPK signaling pathway, the ErbB signaling pathway, the Ras signaling pathway, the Rap1 signaling pathway, the Chemokine signaling pathway, the NF-kappa B signaling pathway, the Cell cycle pathway, the p53 signaling pathway, the mTOR signaling pathway. PI3K-Akt signaling pathway, Apoptosis pathway Ⅰ, Ferroptosis pathway, Necroptosis pathway, Cellular senescence pathway, Wnt signaling pathway, Notch signaling pathway, TGF-beta signaling pathway, JAK-STAT signaling pathway, T cell receptor signaling pathway,A method characterized by at least one selected from the group consisting of B cell receptor signaling pathway, TNF signaling pathway, Transcriptional misregulation pathway in cancer, Prostate cancer pathway, ESR Mediated signaling pathway, Estrogen Dependent gene expression pathway and Androgen Receptor Network pathway in Prostate cancer.
4. In paragraph 1, A method characterized in that the signal transduction pathway score is calculated from the signal transduction pathway-related genes of Table 2 included in each signal transduction pathway using the GSVA (Gene Set Variation Analysis) program.
5. In paragraph 4, The genes related to the above signal transduction pathway are ACTB, H3-3B, ACLY, ACOX1, ACSL4, ACVR1, AKT1, AXIN1, AXIN2, BCL2L1, BMP4, BMPR1A, BRCA1, BTK, CASP3, CASP8, CASP9, CCNA2, CCNB2, CCND1, CCR5, CD40, CDK1, CDK2, CDK4, CDKN1A, CDKN2A, CHEK1, CS, CTNNB1, CXCR4, DVL1, DVL2, EP300, FYN, GADD45A, GNB1, GPX4, GRB2, GSK3B, H3C12, H3C13, H4C6, HMOX1, HRAS, IKBKB, IKBKG, IL1B, IL6, INS, JAG1, JAK1, A method characterized by at least one selected from the group consisting of JAK2, JAK3, JUN, KRAS, LCK, LEF1, LPCAT3, LYN, MAML1, MAML2, MAML3, MAPK1, MAPK3, MCM7, MLST8, MTOR, NCOA4, NFKB1, NFKBIA, NOTCH1, NOTCH2, NOTCH3, NOTCH4, NRAS, PIK3CA, PIK3R1, PLCG2, PTEN, RBPJ, RELA, RHEB, RHOA, RPTOR, RUNX1, SDHB, SMAD2, SMAD3, SMAD4, SOS1, SRC, STAT1, STAT3, SYK, TFRC, TGFB1, TNF, TNFRSF1A, TP53, TRAF6, TSC2 and TYK2.
6. In paragraph 1, A method characterized in that the regression coefficient in the above equation 1 applies a value within an error range of ±5% of each regression coefficient value matching each explanatory variable in Table 4 or Table 5.
7. In paragraph 1, The above grouping step is a method characterized in that the risk score obtained by multiplying the regression coefficient matching each explanatory variable by the variable's own value is divided into a high group and a low group based on a cutoff value.
8. In paragraph 7, A method, characterized in that the unique value of the above variable is a signal transduction pathway score derived from the gene expression level of Table 1 or the signal transduction pathway-related genes of Table 2.
9. In paragraph 7, A method wherein the above High group is determined to be a patient with a high possibility of cancer metastasis or recurrence.
10. In paragraph 7, A method wherein the above Low group is determined to be a patient with a low possibility of cancer metastasis or recurrence.
11. In paragraph 1, A method according to claim 1, wherein the cancer is at least one selected from the group consisting of prostate cancer, breast cancer, mammary cancer, glioma, thyroid cancer, lung cancer, liver cancer, pancreatic cancer, head and neck cancer, stomach cancer, colon cancer, urothelial cancer, kidney cancer, testicular cancer, penile cancer, uterine cancer, cervical cancer, endometrial cancer, cervical cancer, fallopian tube cancer, vaginal cancer, ovarian cancer, melanoma, skin cancer, blood cancer, bone cancer, skin cancer, brain cancer, endocrine cancer, parathyroid cancer, ureteral cancer, urethral cancer, bronchial cancer, bladder cancer, bone marrow cancer, leukemia, brain tumor, intestinal cancer, esophageal cancer, Ewing's sarcoma, tongue cancer, lymphoma, kaposi sarcoma, mesothelioma, multiple myeloma, neuroblastoma, osteosarcoma, and retinoblastoma.
12. Diagnostic devices for predicting cancer risk, including: (a) an input section for inputting gene expression levels and signal transduction pathway scores for biological samples obtained from a subject; (b) a computational unit that calculates a risk score of the following equation 1 using a Cox regression analysis algorithm by combining at least one gene described in Table 1 and at least one signal transduction pathway described in Table 2; and (c) A diagnostic device including an output unit that groups samples of biological samples based on the risk score value calculated above and determines and outputs a high group and a low group; [Formula 1] In the above equation 1, The above p represents the number of explanatory variables, The above explanatory variables represent gene expression levels or signal transduction pathway scores.
Citation Information
Patent Citations
Novel system for predicting prognosis of locally advanced gastric cancer
KR1020140121522A
Methods for predicting risk of recurrence of breast cancer patients
KR1020180059192A
Kernel function-based mobile communication performance analysis method, apparatus, and system
KR1020220059911A
Coffee grounds, panels and their production methods
KR1020240080613A
Method for identifying high-risk AML patients
US20190300956A1