Biomarker for predicting prognosis of colorectal cancer patients
A biomarker composition measuring mRNA levels of specific genes like PLAAT3 and NDUFA4L2 offers improved performance in predicting colon cancer prognosis, addressing the limitations of current biomarkers and enhancing clinical utility.
Patent Information
- Application Number
- PCT/KR2024/014392
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-29
- Filing Date
- 2024-09-24
- Publication Date
- 2025-06-05
AI Technical Summary
Current biomarkers for predicting the prognosis of colorectal cancer, such as the sidedness of the cancer location and the KRAS mutation, have low performance, limiting their clinical utility.
Development of a biomarker composition that measures the mRNA levels of specific genes like PLAAT3, NDUFA4L2, and GAL, or their combinations, to predict the prognosis of colon cancer patients.
The proposed biomarker achieves higher performance in predicting the prognosis of both metastatic and non-metastatic colon cancer compared to conventional markers, providing improved clinical utility and personalized treatment options.
Smart Images

Figure KR2024014392_05062025_PF_FP_ABST
Abstract
Description
Biomarkers for predicting prognosis in colorectal cancer patients
[0001] The present invention was made under the support of the Ministry of Science and ICT (MSIT) under Project No. 1711190121. The research management organization of the project is the National Research Foundation of Korea, the research project name is "Biomedical Technology Development", the research project title is "Development of Prognosis-Treatment Prediction Biomarkers Based on Single-Cell Transcriptome of Colon Cancer", the main organization is Samsung Medical Center, and the research period is from January 1, 2023 to December 31, 2023.
[0002] In addition, the present invention relates to a biomarker for predicting the prognosis of a colon cancer patient, and more specifically, to a composition for predicting the prognosis of colon cancer by measuring the level of mRNA transcribed from marker genes such as PLAAT3, NDUFA4L2, and GAL or a protein expressed therefrom in a sample obtained from a subject, a prognosis prediction kit including the same, and a method for predicting the prognosis of colon cancer.
[0003] Colorectal cancer is the most common cancer in South Korea (12.6% of cases in 2020). Currently, 5-FU (5-Fluorouracil)-based anticancer drugs are used as the standard treatment. However, according to South Korean literature, the five-year survival rate for metastatic colorectal cancer is only 18.9%. To overcome this low survival rate, various drugs, including immunotherapy, antibody-drug conjugates, and novel targeted therapies, are being developed, and clinical trials are actively underway. In this context, biomarkers that can predict the prognosis of colorectal cancer are needed to develop better treatment plans and design more sophisticated clinical trials.
[0004] Meanwhile, currently, prognostic markers for metastatic colorectal cancer consider sidedness (left / right), a clinical indicator of colorectal cancer location, and KRAS mutation, a DNA marker. However, these markers have low performance, limiting their clinical utility. Therefore, a biomarker for colorectal cancer prognosis that demonstrates improved performance than the two currently used markers is needed.
[0005] Accordingly, the inventors of the present invention have made extensive research efforts to develop biomarkers for predicting colon cancer prognosis with improved performance, and as a result, they have discovered that colon cancer prognosis can be predicted with high performance through 35 types of genetic markers or combinations thereof, including PLAAT3, NDUFA4L2, and GAL.
[0006] Accordingly, the purpose of the present invention is to provide a biomarker for predicting the prognosis of colon cancer patients.
[0007] Another object of the present invention relates to a composition for predicting the prognosis of colon cancer patients.
[0008] Another object of the present invention relates to a kit for predicting the prognosis of colon cancer patients.
[0009] Another object of the present invention is to provide a method for providing information for predicting the prognosis of colon cancer patients.
[0010] The present invention relates to a biomarker for predicting the prognosis of colon cancer patients, and the prognosis of colon cancer patients can be predicted with higher performance than conventional technologies through the biomarker of the present invention.
[0011] Hereinafter, the present invention will be described in more detail.
[0012] One aspect of the present invention relates to a biomarker for predicting the prognosis of colon cancer, comprising one or more genes selected from the group consisting of ACAA2, ANXA1, AREG, ASAP1, CHRM3, EIF3H, FHIT, FTH1, GAL, GAREM1, GAS5, KRT17, LARGE1, LINC01811, MACC1, MACROD2, MDK, NAP1L1, NDUFA4L2, NDUFB9, NPM3, NSMCE2, PABPC1, PCAT1, PDE10A, PDE4B, PLAAT3, PLCG2, PVT1, SMOC1, SNHG32, SNHG7, SOD3, TSHZ2, and YBX1.
[0013] The term "colon cancer" in this specification encompasses all malignant tumors arising from the mucosa of the colon, including adenocarcinoma, lymphoma, sarcoma, squamous cell carcinoma, and metastatic lesions. Metastatic colon cancer refers to cancer that originated in an organ other than the colon and spread to the colon, and non-metastatic colon cancer refers to cancer caused by a primary tumor arising in the colon rather than in another organ.
[0014] The term "prognosis" as used herein refers to the act of predicting the course of a disease and the outcome of death or survival. More specifically, prognosis refers to a patient's physiological or environmental condition that changes over time, and "predicting the prognosis" can be understood to mean any act of predicting the course of a disease before and after treatment by comprehensively considering the patient's condition.
[0015] Considering the purpose of the present invention, the "prognosis prediction" of the present invention can be understood as an act of predicting the disease-free survival rate or survival rate of a colon cancer patient for a period of one year, two years or more by predicting in advance the course of the disease and whether it is cured before / after treatment of colon cancer.
[0016] In the present invention, each gene used for predicting the prognosis of colon cancer may include a base sequence identical to or homologous to a gene sequence indicated by a number in the "Entrez ID" or "Reference Sequence Information" column of the gene list described in Table 3 of the present invention. Entrez ID may refer to a unique identification number assigned to each gene by Entrez, a search engine for a database stored on the website of the National Center for Biotechnology Information (NCBI) in the United States, and reference sequence information may refer to a unique identification number assigned to each gene in a database stored on the website of NCBI.
[0017] At this time, each gene may include a base sequence that exhibits 80% or more, 90% or more, 95% or more, or 95% or more homology with a gene sequence indicated by a number in the gene list described in Table 3 or the Entrez ID or reference sequence information field.
[0018] The term "biomarker" in this specification refers to organic biomolecules such as polypeptides, proteins, nucleic acids, genes, lipids, glycolipids, glycoproteins, sugars, etc., which are substances detectable in biological samples obtained from an individual, and whose detection enables detection of changes in the organism.
[0019] The term "subject" as used herein refers to any animal, including humans, that has developed or is likely to develop colorectal cancer, and may particularly refer to an animal that has already been diagnosed with colorectal cancer. Animals may include, but are not limited to, mammals such as cows, horses, sheep, pigs, goats, camels, antelopes, dogs, and cats that require treatment for similar symptoms, as well as humans.
[0020] The term "biological sample" in this specification means a subject isolated from an individual who has developed or is likely to develop colorectal cancer, and for which the level of genes or proteins, etc. is directly measured. The biological sample may be one or more selected from the group consisting of tissues, cells, whole blood, serum, and plasma isolated from the individual.
[0021] In one embodiment of the present invention, the biomarker gene for predicting the prognosis of colon cancer may include genes such as PLAAT3 and NDUFA4L2.
[0022] In one embodiment of the present invention, the biomarker gene for predicting the prognosis of colon cancer may further include one or more genes selected from the group consisting of EIF3H, GAL, GASS, NPM3, NSMCE2, PLCG2, PVT1, SNHG32, and YBX1.
[0023] In one embodiment of the present invention, the colon cancer may be metastatic colon cancer or non-metastatic colon cancer.
[0024] When predicting the prognosis of a colon cancer patient using biomarker genes such as PLAAT3, ACAA2, MDK, TSHZ2, PDE10A, NDUFA4L2, GAL, etc. or a combination thereof according to the present invention, the prognosis of metastatic colon cancer can be predicted with higher accuracy than the sidedness marker and KRAS mutation marker (see FIGS. 1 and 2), which are indicators used to predict metastatic colon cancer in the past (see FIGS. 5 to 11), and it was confirmed that the prognosis of not only metastatic colon cancer but also non-metastatic colon cancer can be predicted with high accuracy (see FIG. 12).
[0025] Another aspect of the present invention relates to a composition for predicting the prognosis of colon cancer, comprising an agent for measuring the level of mRNA transcribed from one or more genes selected from the group consisting of ACAA2, ANXA1, AREG, ASAP1, CHRM3, EIF3H, FHIT, FTH1, GAL, GAREM1, GAS5, KRT17, LARGE1, LINC01811, MACC1, MACROD2, MDK, NAP1L1, NDUFA4L2, NDUFB9, NPM3, NSMCE2, PABPC1, PCAT1, PDE10A, PDE4B, PLAAT3, PLCG2, PVT1, SMOC1, SNHG32, SNHG7, SOD3, TSHZ2 and YBX1, or a protein expressed therefrom.
[0026] The term "agent for measuring the level of mRNA" in this specification refers to an agent used in a method for measuring the level of mRNA transcribed from a target gene in order to confirm whether the target gene contained in a biological sample is expressed, and may include, but is not limited to, a primer or probe that can specifically bind to a target gene used in a method such as RT-PCR, quantified real time PCR, competitive RT-PCR, real time quantitative RT-PCR, RNase protection assay (RPA), Northern blotting, DNA chip analysis, etc.
[0027] The term "agent for measuring the level of a protein" in this specification means an agent used in a method for measuring the level of a target protein contained in a biological sample, and may include, but is not limited to, antibodies used in methods such as western blotting, enzyme linked immunosorbent assay (ELISA), radioimmunoassay (RIA), radioimmunodiffusion, Ouchterlony, rocket immunoelectrophoresis, immunohistochemical staining, immunoprecipitation assay, complement fixation assay, immunofluorescence, immunochromatography, fluorescence activated cell sorter analysis (FACS), or protein chip technology assay.
[0028] In one embodiment of the present invention, the gene may comprise PLAAT3 and NDUFA4L2.
[0029] In one embodiment of the present invention, the gene may further comprise one or more genes selected from the group consisting of EIF3H, GAL, GASS, NPM3, NSMCE2, PLCG2, PVT1, SNHG32 and YBX1.
[0030] In one embodiment of the present invention, the colon cancer may be metastatic colon cancer or non-metastatic colon cancer.
[0031] In one embodiment of the present invention, the agent for measuring the level of mRNA may comprise a primer or probe that specifically binds to a gene, or the agent for measuring the level of protein may comprise an antibody or aptamer that specifically binds to a protein.
[0032] The term "specific binding" as used herein means that adjacent target substances and labeling substances exhibit a high binding affinity compared to other substances, such that the presence of the target substance can be detected using conventional analytical methods. The target substance may be the mRNA or protein being measured, and the labeling substance may be, but is not limited to, a primer, a probe, an antibody, or an aptamer.
[0033] The term "primer" as used herein refers to a short nucleic acid sequence having a short free 3' hydroxyl group, which can form base pairs with a complementary template strand, and which functions as a starting point for copying the replicating strand. It may be a primer set including a forward and reverse primer, which is a short strand RNA or DNA sequence that recognizes a target gene sequence. Since the nucleic acid sequence of the primer is a sequence that does not match the non-target sequence present in the sample, it amplifies only the target gene sequence containing the complementary primer binding site, and can impart high specificity when it is a primer that does not cause non-specific amplification.
[0034] The term "probe" in this specification refers to a substance that can specifically bind to a target substance to be detected in a biological sample, and the probe can confirm the presence of the target substance in the sample through specific binding to the target substance. In this case, the probe may be any PNA (peptide nucleic acid), LNA (locked nucleic acid), peptide, polypeptide, protein, RNA, or DNA known in the art, but is not limited thereto.
[0035] The term "antibody" as used herein refers to a protein molecule capable of specifically binding to an antigenic site of a protein or peptide molecule. Antibodies can be produced by cloning each gene into an expression vector according to a conventional method to obtain a protein encoded by a marker gene, and then producing the obtained protein by a conventional method. The form of the antibody is not particularly limited, and any part thereof, such as a polyclonal antibody, a monoclonal antibody, an antibody fragment, or a recombinant antibody, that has antigen-binding properties is also included in the antibodies of the present invention. In the present invention, antibodies may include all immunoglobulin antibodies as well as specialized antibodies such as humanized antibodies. Furthermore, antibodies include not only complete forms having two full-length light chains and two full-length heavy chains, but also functional fragments of antibody molecules. Functional fragments of antibody molecules refer to fragments that possess at least an antigen-binding function, and may be Fab, F(ab'), F(ab') 2, and Fv.
[0036] The term "aptamer" as used herein refers to a nucleic acid molecule that is a single-stranded oligonucleotide and has binding activity to a given target molecule. Aptamers can have various three-dimensional structures depending on their base sequences and can have high affinity for a specific substance, such as an antigen-antibody reaction. Aptamers can inhibit the activity of a given target molecule by binding to the target molecule, and can be RNA, DNA, biochemically modified nucleic acids, or mixtures thereof, and can be in the form of a chain or a ring, but are not limited thereto.
[0037] Another aspect of the present invention relates to a kit for predicting the prognosis of colon cancer, comprising an agent for measuring the level of mRNA transcribed from one or more genes selected from the group consisting of ACAA2, ANXA1, AREG, ASAP1, CHRM3, EIF3H, FHIT, FTH1, GAL, GAREM1, GAS5, KRT17, LARGE1, LINC01811, MACC1, MACROD2, MDK, NAP1L1, NDUFA4L2, NDUFB9, NPM3, NSMCE2, PABPC1, PCAT1, PDE10A, PDE4B, PLAAT3, PLCG2, PVT1, SMOC1, SNHG32, SNHG7, SOD3, TSHZ2 and YBX1, or a protein expressed therefrom.
[0038] In the present invention, the kit for predicting the prognosis of colon cancer may be an RT-PCR kit, a DNA chip kit, or a protein chip kit.
[0039] The RT-PCR kit according to the present invention may be a kit containing essential elements necessary for measuring the mRNA expression level of a target gene through RT-PCR. In addition to each primer pair specific for the target gene, the RT-PCR kit may include a test tube or other appropriate container, a reaction buffer (with an appropriate pH and magnesium concentration), deoxynucleotides (dNTPs), enzymes such as Taq polymerase and reverse transcriptase, DNase, RNAse inhibitors, DEPC water, sterile water, and the like. In addition, the kit may include a primer pair specific for a gene used as a quantitative control.
[0040] The DNA chip kit according to the present invention may be a kit containing essential elements necessary for performing a DNA chip analysis method. The DNA chip analysis kit may include a substrate to which a cDNA corresponding to a gene or a fragment thereof is attached as a probe, and reagents, preparations, enzymes, etc. for producing a fluorescently labeled probe. In addition, the substrate may include a cDNA corresponding to a quantitative control gene or a fragment thereof.
[0041] The protein chip kit according to the present invention may be a kit containing the essential elements necessary for measuring the level of a target protein expressed from a target gene using a protein chip assay. The protein chip kit may include a substrate for immunological detection of antibodies, a suitable buffer solution, a secondary antibody labeled with a chromogenic enzyme or fluorescent substance, a chromogenic substrate, and the like. As the above-mentioned substrate, for example, a nitrocellulose membrane, a 96-well plate synthesized with polyvinyl resin, a 96-well plate synthesized with polystyrene resin, and a glass slide glass can be used, and as the chromogenic enzyme, for example, peroxidase and alkaline phosphatase can be used, and as the fluorescent substance, for example, FITC, RITC, etc. can be used, and as the chromogenic substrate liquid, for example, ABTS (2,2'-azino-bis(3-ethylbenzothiazoline-6-sulfonic acid)) or OPD (O-phenylenediamine), TMB (tetramethyl benzidine) can be used, but is not limited thereto.
[0042] In one embodiment of the present invention, the gene may comprise PLAAT3 and NDUFA4L2.
[0043] In one embodiment of the present invention, the gene may further comprise one or more genes selected from the group consisting of EIF3H, GAL, GASS, NPM3, NSMCE2, PLCG2, PVT1, SNHG32 and YBX1.
[0044] In one embodiment of the present invention, the colon cancer may be metastatic colon cancer or non-metastatic colon cancer.
[0045] In one embodiment of the present invention, the agent for measuring the level of mRNA may comprise a primer or probe that specifically binds to a gene, or the agent for measuring the level of protein may comprise an antibody or aptamer that specifically binds to a protein.
[0046] In the present invention, the kit for predicting the prognosis of colon cancer may further include a reagent necessary for predicting the prognosis of colon cancer, for example, a buffer solution, an indicator, a kit preservation solution, a sample stabilization solution, or a combination thereof.
[0047] Another aspect of the present invention relates to a method for providing information for predicting the prognosis of colon cancer, comprising a measuring step of quantifying the expression level of one or more genes by measuring the level of mRNA transcribed from or protein expressed therefrom, one or more genes selected from the group consisting of ACAA2, ANXA1, AREG, ASAP1, CHRM3, EIF3H, FHIT, FTH1, GAL, GAREM1, GAS5, KRT17, LARGE1, LINC01811, MACC1, MACROD2, MDK, NAP1L1, NDUFA4L2, NDUFB9, NPM3, NSMCE2, PABPC1, PCAT1, PDE10A, PDE4B, PLAAT3, PLCG2, PVT1, SMOC1, SNHG32, SNHG7, SOD3, TSHZ2 and YBX1, in a biological sample isolated from an individual.
[0048] In the present invention, quantifying the level of gene expression may be done by directly expressing the level of gene expression by quantifying RNA transcribed from the gene or expressed protein, or by indirectly expressing the level of gene expression by performing a predetermined statistical processing technique or data processing technique based on the result of quantifying RNA or protein.
[0049] Any technique known in the art may be utilized for quantification of RNA or protein, and any statistical processing technique or data processing technique known in the art may be utilized for processing the RNA or protein quantification results. For example, in the present invention, the expression level of a gene may be expressed as a GVSA score derived by applying the GVSA (gene set variation analysis) function of the GVSA package of the statistical program R to the whole transcriptome sequencing results of the gene, but is not limited thereto.
[0050] In one embodiment of the present invention, the measuring step may include a contacting step of contacting the sample with a primer or probe that specifically binds to mRNA or contacting the sample with an antibody or aptamer that specifically binds to a protein.
[0051] In one embodiment of the present invention, the information providing method may further include a comparison step of comparing the level of mRNA or protein measured in a sample isolated from an individual with the level of mRNA or protein measured in a sample of a normal control group.
[0052] In one embodiment of the present invention, the information providing method may further include a comparison step of comparing the expression level of one or more genes measured in a sample isolated from an individual with a high or low expression determination cutoff for each gene in [Table 6]. At this time, the high or low expression determination cutoff, which serves as a comparison measure for the gene expression level, may be the same as [Table 2] of Example 4-1 of the present invention, [Table 4] of Example 4-2 of the present invention, or [Table 6] of Example 4-3 of the present invention.
[0053] The term "cutoff" in this specification refers to a numerical value that serves as a standard for determining whether a specific gene is highly or lowly expressed in an individual based on the analysis results of a sample isolated from the individual, and may be expressed as a value of the same dimension as a numerical value indicating the level of gene expression. For example, in the present invention, when the level of gene expression is expressed as a GVSA score, the cutoff for determining high or low expression may also be expressed as a GVSA score. This cutoff may also be understood to refer to a threshold.
[0054] In one embodiment of the present invention, the information providing method determines that the prognosis of colon cancer is poor when the expression level of one or more genes selected from the group consisting of ANXA1, FTH1, GAL, GAS5, KRT17, MDK, NDUFA4L2, NSMCE2, PDE10A, PDE4B, PLAAT3, PLCG2, SMOC1, SNHG32, SNHG7, TSHZ2 and YBX1 is higher than the high expression determination cutoff for each gene, and determines that the prognosis of colon cancer is good when the expression level of one or more genes selected from the group consisting of ACAA2, CHRM3, FHIT, LARGE1, LINCO1811, MACC1, MACROD2, NAP1L1, NDUFB9, NPM3, PCAT1 and SOD3 is lower than the low expression determination cutoff for each gene, and determines that the prognosis of colon cancer is poor .... It may include an additional step of judging that the prognosis of colon cancer is good if it is higher than the cutoff.
[0055] In one embodiment of the present invention, the biological sample may be at least one selected from the group consisting of tissue, cells, whole blood, serum, and plasma isolated from an individual.
[0056] The biomarker of the present invention can predict the prognosis of colon cancer patients with higher performance than conventional technologies, and in particular, it can be used to predict the prognosis of not only metastatic colon cancer but also non-metastatic colon cancer, so the biomarker of the present invention has high clinical utility value, and a personalized treatment method can be proposed based on the prognosis prediction results for each patient.
[0057] Figure 1 is a diagram showing the results of classifying single cells into eight unique tumor cell clusters based on gene expression patterns through single-cell gene expression measurement and clustering analysis of patients with colon cancer.
[0058] Figure 2 is a diagram showing the results of analyzing differentially expressed genes (DEGs) in the malignant goblet cell cluster among the eight unique tumor cell clusters compared to the remaining cell clusters (malignant cells or non-malignant cells) as a volcano plot.
[0059] Figure 3 is a diagram showing the difference in survival curves by group when performing a survival analysis of patients (SMC cohort group) using sidedness (left / right) markers used in conventional metastatic colorectal cancer prognosis prediction, and the performance of the prediction model expressed as C-index.
[0060] Figure 4 is a diagram showing the difference in survival curves by group when performing a survival analysis of patients (SMC cohort group) using the KRAS mutation marker used in predicting the prognosis of conventional metastatic colorectal cancer, and the performance of the prediction model expressed as the C-index.
[0061] FIG. 5 is a diagram showing the difference in survival curves by group and the performance of the prediction model expressed as C-index when performing survival analysis of metastatic colorectal cancer patients (SMC cohort group) using the NDUFA4L2 single gene among biomarker genes according to one embodiment of the present invention.
[0062] FIG. 6 is a diagram showing the difference in survival curves by group and the performance of the prediction model expressed as C-index when performing survival analysis on metastatic colorectal cancer patients (SMC cohort group) using the ACAA2 single gene among biomarker genes according to one embodiment of the present invention.
[0063] FIG. 7 is a diagram showing the difference in survival curves by group and the performance of the prediction model expressed as C-index when performing survival analysis on metastatic colorectal cancer patients (SMC cohort group) using the PDE10A single gene among biomarker genes according to one embodiment of the present invention.
[0064] FIG. 8 is a diagram showing the difference in survival curves by group and the performance of the prediction model expressed as C-index when performing survival analysis on metastatic colorectal cancer patients (SMC cohort group) using a combination of PLAAT3, NDUFA4L2, and PLCG2 genes among biomarker genes according to one embodiment of the present invention.
[0065] FIG. 9 is a diagram showing the difference in survival curves by group and the performance of the prediction model expressed as C-index when performing survival analysis on metastatic colorectal cancer patients (TCGA cohort group) using a combination of PLAAT3, NDUFA4L2, and PLCG2 genes among biomarker genes according to one embodiment of the present invention.
[0066] FIG. 10 is a diagram showing the difference in survival curves by group and the performance of the prediction model expressed as C-index when a survival analysis of metastatic colorectal cancer patients (SMC cohort group) was performed using the PLAAT3 single gene among biomarker genes according to one embodiment of the present invention.
[0067] FIG. 11 is a diagram showing the difference in survival curves by group when performing survival analysis on metastatic colorectal cancer patients (GSE17536 cohort group) using the PDE10A single gene among biomarker genes according to one embodiment of the present invention, and the performance of the prediction model expressed as C-index.
[0068] FIG. 12 is a diagram showing the difference in survival curves by group and the performance of the prediction model expressed as C-index when a survival analysis of non-metastatic colorectal cancer patients (GSE17536 cohort group) was performed using the PDE10A single gene among biomarker genes according to one embodiment of the present invention.
[0069] A composition for predicting the prognosis of colon cancer, comprising an agent for measuring the level of mRNA transcribed from one or more genes selected from the group consisting of ACAA2, ANXA1, AREG, ASAP1, CHRM3, EIF3H, FHIT, FTH1, GAL, GAREM1, GAS5, KRT17, LARGE1, LINC01811, MACC1, MACROD2, MDK, NAP1L1, NDUFA4L2, NDUFB9, NPM3, NSMCE2, PABPC1, PCAT1, PDE10A, PDE4B, PLAAT3, PLCG2, PVT1, SMOC1, SNHG32, SNHG7, SOD3, TSHZ2, and YBX1.
[0070] Hereinafter, the present invention will be described in more detail with reference to the following examples. However, these examples are only intended to illustrate the present invention, and the scope of the present invention is not limited by these examples.
[0071] Throughout this specification, "%" used to indicate the concentration of a particular substance is (wt / wt)% for solid / solid, (wt / vol)% for solid / liquid, and (vol / vol)% for liquid / liquid, unless otherwise stated.
[0072] Unless otherwise specified, all numbers, values, or expressions expressing ingredients, reaction conditions, or quantities of ingredients used in this specification are to be understood as being modified in all instances by the term "about" because these numbers are approximations that inherently reflect, among other things, the various uncertainties of measurement in obtaining those values.
[0073] Additionally, when a numerical range is disclosed herein, such range is continuous and includes all values from the minimum value to the maximum value inclusive, unless otherwise specified.
[0074] Also, the term "or" in this specification is intended to mean an inclusive "or" rather than an exclusive "or." That is, where a connection or use between components is not otherwise specified or clear from context, i.e., if X includes A; X includes B; or X includes both A and B, "X includes A or B" can be applied to any of these cases.
[0075] Example 1: Selection of single cell types associated with predicting the prognosis of colon cancer.
[0076] To select cell types highly correlated with predicting survival prognosis in patients with metastatic colorectal cancer, single-cell clustering analysis was performed on patients with colorectal cancer.
[0077] Specifically, we performed single-cell gene expression analysis (single-cell RNA sequencing) and clustering analysis based on the analysis results of 58,826 colorectal cancer epithelial cells from 61 colorectal cancer patients. The Seurat package in R was used for the analysis process, and the NormalizeData, FindVariableFeatures, ScaleData, and RunPCA functions of the Seurat package were used to perform the basic analysis process. As a specific option, 2,000 genes showing highly variable features were selected, and the expression of 2,000 genes was scaled considering mitochondrial expression (percent.mt), and then dimensionality reduction was performed using the PCA method based on this.
[0078] At this time, considering that different features were shown for each patient, the patient-specific expression was corrected using the harmony package in R with batch correction to compensate for this. Afterwards, for clustering, the corrected expression values were used through the RunUMAP, FindNeighbors, and FindClusters functions of the Seurat package and the harmony package, and the resolution for clustering was set to 0.5, and clusters showing similar expression were considered as one cluster and analyzed thereafter. The results of the clustering analysis are shown in Figure 1.
[0079] As a result of the analysis, single cells were classified into eight unique tumor cell clusters based on their gene expression patterns. Among them, when we checked the cluster with the highest correlation with the intrinsic consensus molecular subtype (iCMS) type related to prognosis or the high risk score (HR score) related to recurrence in the analysis according to the existing molecular classification of colon cancer, we confirmed that malignant goblet cells can be used as an indicator that best represents the prognosis.
[0080] Therefore, among the eight unique tumor cell clusters, malignant goblet cells were thought to be highly related to predicting the survival prognosis of patients with metastatic colorectal cancer.
[0081] Example 2. Selection of biomarker candidate genes related to predicting the prognosis of colon cancer.
[0082] Through differential gene expression analysis of the malignant goblet cell clusters selected in Example 1, 43 candidate genes that can be used as biomarkers for predicting the prognosis of metastatic colorectal cancer were selected.
[0083] Specifically, among the single cells of 61 patients with colorectal cancer, 5,937 malignant goblet cells were analyzed for differentially expressed genes (DEGs) between 7 patients who developed colorectal cancer metastasis and died and 54 patients who developed local recurrence or metastasis but survived. The analysis was performed using the MAST (Model-based Analysis of Single Cell Transcriptomic) method of the FindMarker function of Seurat, one of the R packages, and 43 genes satisfying the conditions of adjusted p-value (Bonferroni correction) < 0.01 and log2 fold change > 0.7 were selected. The results of the DEG analysis for malignant goblet cells are presented in the form of a volcano plot in Fig. 2. (x-axis: log2 fold change, y-axis: p-value)
[0084] Example 3. Establishment of a colorectal cancer prognosis prediction model.
[0085] Single-cell gene expression analysis currently struggles to be widely applied in clinical practice due to analytical costs and sample management. Therefore, to enhance its applicability in future clinical practice, we developed a colorectal cancer prognosis prediction model based on biomarker candidate genes identified through the single-cell DEG analysis in Example 2, using tissue-level bulk gene expression (bulk RNA sequencing or bulk gene expression microarray) data.
[0086] Specifically, we classified patients with metastatic colorectal cancer into a high group (high-expression group) based on the top 25% expression levels and a low group (low-expression group) based on the bottom 25% expression levels for bulk gene expression data for each gene, and constructed a prognostic prediction model based on survival analysis of overall survival according to gene expression. Each model was evaluated as having excellent performance if the C-index > 0.7.
[0087] At this time, gene expression was analyzed by using the GSVA (gene set variation analysis) function of the R GSVA package for the expression level of a single gene or a combination of multiple genes to summarize the gene expression into a single value (GSVA score). Patients with the top 25% expression level were classified into the high group, and patients with the bottom 25% expression level were classified into the low group.
[0088] Each constructed prognostic prediction model was evaluated in three cohorts: SMC, GSE17356 (Moffitt Cancer Center, Affymetrix expression array-based gene expression profile data), and TCGA (The Cancer Genome Atlas, bulk RNA sequencing data). The SMC cohort consisted of 20 patients, the GSE17356 cohort consisted of 39 patients, and the TCGA cohort consisted of 26 patients.
[0089] For the SMC cohort, the gene expression levels used as input for calculating the GSVA score were based on whole-transcriptome analysis results through RNA sequencing. For the GSE17536 cohort, gene expression levels were obtained using public data from microarray experiments (Human Genome U133 Plus 2.0, Affymerix), while for the TCGA cohort, gene expression levels were obtained using public data from whole-transcriptome analysis using Illumina HiSeq 2000.
[0090] To construct whole-transcriptome data for the SMC cohort, we first performed whole-transcriptome analysis on RNA extracted from tumor tissue. For whole-transcriptome analysis, we used the TruSeq RNA Library Prep Kit v2 (Illumina Inc.), and constructed an RNA sequencing library according to the manufacturer's protocol. Paired-end sequencing was then performed on the Illumina HiSeq 2500 Sequencing Platform to convert the RNA sequencing library into sequencing reads, which were then used to generate FASTQ files. Finally, after removing quality defects from the FASTQ files, we aligned the sequencing reads to the human reference genome (hg19) using STAR (v2.5.2b) software, and measured the expression levels of all genes using RSEM software (v1.3).
[0091] As a result of the performance evaluation for each cohort group, prognostic prediction models that showed superior performance for one or more groups were included in the final list of prognostic prediction models.
[0092] Example 4. Selection of biomarker genes for predicting colon cancer prognosis through evaluation of a colon cancer prognosis prediction model.
[0093] 4-1. Selection of biomarker genes for predicting the prognosis of metastatic colorectal cancer.
[0094] Survival analysis was performed on the SMC, GSE17356, or TCGA cohorts of metastatic colorectal cancer using each prognostic prediction model constructed with a single gene or a combination of multiple genes of the sideness and KRAS mutation markers previously used to predict the prognosis of metastatic colorectal cancer, and the biomarker candidate genes discovered in the single-cell DEG analysis of Example 2, and the C-index of each model was derived based on the analysis results. The genes or combinations thereof used to construct each model, the cohort groups to be analyzed, and the C-index of each model are shown in Table 1 below. In addition, for each model, whether a gene is highly expressed (high group) or low expressed (low group) is evaluated as a high-risk group, and the cutoffs for classification of high and low gene expression for each model (cutoff, based on GSVA score) are shown in Table 2 below.
[0095] In addition, the specific results of the C-index performance evaluation for representative prognostic prediction models constructed with conventional sidedness and KRAS mutation markers, and single genes NDUFA4L2, ACAA2, PDE10A (all SMC cohort groups) or three-gene combination of PLAAT3 + NDUFA4L2 + PLCG2 (SMC cohort groups and TCGA cohort groups) discovered from single-cell DEG analysis are shown in Figs. 3 to 9, respectively.
[0096] Model numberBiomarkerDistinctionAnalysis target populationPerformance (C-index)Additional analysis population (C-index>0.7)Remarksc1SidednessExisting markerSMC0.540c2KRAS mutationExisting markerSMC0.490m1NDUFA4L21SMC0.842m2ACAA21SMC0.813TCGA (0.765)m3PLAAT31SMC0.750GSE17536 (0.717)m4NAP1L11SMC0.813m5AREG1SMC0.778m6FTH11SMC0.842m7PDE10A1SMC0.857Representative model, 1 markerMaximum performancem8MACC11SMC0.794m9PLCG21SMC0.778 m10MACROD21eaSMC0.700 m11YBX11eaSMC0.816 m12MDK1eaSMC0.833 m13CHRM31eaSMC0.700 m14TSHZ21eaSMC0.833 m15PDE4B1eaSMC0.789 m16SMOC11eaSMC0.700 m17PCAT11eaSMC0.706 m18NSMCE21eaSMC0.763 m19FHIT1eaSMC0.778 m20ANXA11eaGSE175360.735 m21SOD31eaGSE175360.702 m22SNHG71eaTCGA0.736 m23LINC018111eaTCGA0.739 m24LARGE11eaTCGA0.750 m25PLAAT3 + NDUFA4L2 + GAL3SMC0.839TCGA (0.804)m26PLAAT3 + NDUFA4L2 + NSMCE23SMC0.828GSE17536 (0.718), TCGA (0.755)m27PLAAT3 + NDUFA4L2 + GAS53SMC0.813GSE17536 (0.747), TCGA (0.783)m28PLAAT3 + NDUFA4L2 + YBX13SMC0.850GSE17536 (0.706), TCGA (0.788)m29PLAAT3 + NDUFA4L2 + NPM33SMC0.841GSE17536 (0.725), TCGA (0.782)m30PLAAT3 + NDUFA4L2 + 13 PVTSMC0.804GSE17536 (0.741), TCGA (0.804)m31PLAAT3 + NDUFA4L2 + EIF3H3ea SMC0.827GSE17536 (0.728), TCGA (0.729)m32PLAAT3 + NDUFA4L2 + PLCG2ea SMC0.868TCGA (0.721) Representative model, 3 markers Maximum performancem33PLAAT3 + NDUFA4L2 + GAL + PDE10A4ea SMC0.833TCGA (0.813)m34PLAAT3 + NDUFA4L2 + GAL + ANXA14ea SMC0.833TCGA (0.779)m35PLAAT3 + NDUFA4L2 + GAL + PDE4B4ea SMC0.818TCGA (0.787)m36PLAAT3 + NDUFA4L2 + GAL + FHIT4SMC0.813TCGA (0.791)m37PLAAT3 + NDUFA4L2 + GAL + GAREM14SMC0.800TCGA (0.704)m38PLAAT3 + NDUFA4L2 + GAL + PCAT14SMC0.841TCGA (0.746)m39PLAAT3 + NDUFA4L2 + GAL + PVT14SMC0.813TCGA (0.784)m40PLAAT3 + NDUFA4L2 + GAL + NDUFB94SMC0.813TCGA (0.763)m41PLAAT3 + NDUFA4L2 + GAL + NPM34SMC0.841TCGA (0.784)m42PLAAT3 + NDUFA4L2 + GAL + PABPC14SMC0.813TCGA (0.81)m43PLAAT3 + NDUFA4L2 + GAL + ASAP14SMC0.800TCGA (0.754)m44PLAAT3 + NDUFA4L2 + GAL + LARGE14SMC0.800TCGA (0.804)m45PLAAT3 + NDUFA4L2 + GAL + KRT174SMC0.818TCGA (0.792)m46PLAAT3 + NDUFA4L2 + GAL + SNHG324SMC0.804GSE17536 (0.701), TCGA (0.768)m47PLAAT3 + NDUFA4L2 + GAL + TSHZ24SMC0.810GSE17536 (0.719), TCGA (0.804).
[0097]
[0098] Model numberBiomarkerHigh-risk group judgment criteria (cutoff criteria)High expression judgment cutoff (GSVA score)Low expression judgment cutoff (GSVA score)m1NDUFA4L2High-risk group with high expression0.23-0.55m2ACAA2High-risk group with low expression0.64-0.55m3PLAAT3High-risk group with high expression-0.05-0.92m4NAP1L1High-risk group with low expression0.27-0.83m5AREGHigh-risk group with high expression0.26-0.80m6FTH1High-risk group with high expression0.54-0.66m7PDE10AHigh-risk group with high expression0.58-0.64m8MACC1High-risk group with low expression0.52-0.60m9PLCG2High-risk group with high expression0.710.00m10MACROD2Low expression High-risk group 0.49-0.34m11YBX1 High-risk group 0.33-0.58m12MDK High-risk group 0.67-0.40m13CHRM3 Low-risk group 0.48-0.27m14TSHZ2 High-risk group 0.64-0.48m15PDE4B High-risk group 0.39-0.45m16SMOC1 High-risk group 0.68-0.48m17PCAT1 Low-risk group 0.04-0.74m18NSMCE2 High-risk group 0.22-0.59m19FHIT Low-risk group 0.56-0.38m20ANXA1 High-risk group 0.25-0.52m21SOD3 Low-expression High risk group 0.70-0.72m22SNHG7High expression High risk group 0.50-0.51m23LINC01811Low expression High risk group 0.61-0.22m24LARGE1Low expression High risk group 0.09-0.55m25PLAAT3 + NDUFA4L2 + GALHigh expression High risk group 0.15-0.62m26PLAAT3 + NDUFA4L2 + NSMCE2High expression High risk group 0.20-0.73m27PLAAT3 + NDUFA4L2 + GAS5High expression High risk group 0.01-0.49m28PLAAT3 + NDUFA4L2 + YBX1High expression High risk group 0.10-0.66m29PLAAT3 + NDUFA4L2 + NPM3High expression High-risk group 0.17-0.50m30PLAAT3 + NDUFA4L2 + PVT1 High-risk group 0.16-0.64m31 High-risk group when PLAAT3 + NDUFA4L2 + EIF3H is highly expressed 0.30-0.69m32 High-risk group when PLAAT3 + NDUFA4L2 + PLCG2 is highly expressed 0.42-0.42m33 High-risk group when PLAAT3 + NDUFA4L2 + GAL + PDE10A is highly expressed 0.28-0.58m34 High-risk group when PLAAT3 + NDUFA4L2 + GAL + ANXA1 is highly expressed 0.03-0.58m35 High-risk group when PLAAT3 + NDUFA4L2 + GAL + PDE4B is highly expressed 0.08-0.46m36 High-risk group when PLAAT3 + NDUFA4L2 + GAL + FHIT is highly expressed 0.17-0.35m37 High-risk group when PLAAT3 + NDUFA4L2 + GAL + High-risk group when GAREM1 is highly expressed 0.17-0.37m38PLAAT3 + NDUFA4L2 + GAL + PCAT1 High-risk group when PLAAT3 + NDUFA4L2 + GAL + PVT1 High-risk group when PLAAT3 + NDUFA4L2 + GAL + PVT1 High-risk group when PLAAT3 + NDUFA4L2 + GAL + NDUFB9 High-risk group when PLAAT3 + NDUFA4L2 + GAL + NPM3 High-risk group when PLAAT3 + NDUFA4L2 + GAL + PABPC1 High-risk group when PLAAT3 + NDUFA4L2 + GAL + ASAP1 High-risk group when PLAAT3 + NDUFA4L2 + GAL + ASAP1 High-risk group 0.18-0.50m44PLAAT3 + NDUFA4L2 + GAL + LARGE1 High-risk group with high expression 0.24-0.43m45PLAAT3 + NDUFA4L2 + GAL + KRT17 High-risk group with high expression 0.18-0.56m46PLAAT3 + NDUFA4L2 + GAL + SNHG32 High-risk group with high expression 0.16-0.53m47PLAAT3 + NDUFA4L2 + GAL + TSHZ2 High-risk group with high expression 0.38-0.56.
[0099]
[0100] As a result of the performance evaluation of the metastatic colorectal cancer prognosis prediction model, the model (m7) using a single PDE10A gene as a marker had a C-index of 0.857, and it was impossible to predict the 1-year or 2-year survival probability. The model (m1) using a single NDUFA4L2 gene as a marker had a C-index of 0.842, and the 2-year survival probability was predicted to be 25% in the high group and 100% in the low group. The model (m32) using a combination of three genes, PLAAT3 + NDUFA4L2 + PLCG2, as markers had a C-index of 0.868, and the 1-year survival probability was predicted to be 43% in the high group and 86% in the low group.
[0101] From these results, for the PDE10A, NDUFA4L2, or PLAAT3 + NDUFA4L2 + PLCG2 combination, belonging to the high group showed high expression of each gene or gene combination and poor prognosis for metastatic colorectal cancer.
[0102] On the other hand, in the model (m2) using one ACAA2 gene as a marker, the prognosis of the high group was good and the prognosis of the low group was bad.
[0103] 4-2. Analysis of the Possibility of Expanding the Application of Biomarkers for Predicting Prognosis of Metastatic Colorectal Cancer to Nonmetastatic Colorectal Cancer
[0104] We analyzed the possibility of expanding the applicability of a prognostic prediction model built on data from metastatic colorectal cancer patients to the non-metastatic colorectal cancer patient group.
[0105] Specifically, similar to the process of constructing a prognostic prediction model according to Example 3, a metastatic colorectal cancer prognostic prediction model was applied to bulk gene expression data of a group of non-metastatic colorectal cancer patients to analyze whether there was a difference in overall survival time and derive a C-index, and the prognostic prediction performance of each model for non-metastatic colorectal cancer was evaluated based on the derived C-index.
[0106] The nonmetastatic colorectal cancer prognostic prediction model was evaluated in three cohorts: SMC, GSE17356, and TCGA. The SMC cohort consisted of 185 patients, the GSE17356 cohort consisted of 138 patients, and the TCGA cohort consisted of 149 patients.
[0107] The genes or their combinations used to construct each model, the cohort group to be analyzed, and the C-index of each model are shown in Table 3 below. In each model, whether the gene is highly expressed (high group) or low expressed (low group) is evaluated as a high-risk group, and the cutoffs for classification of high and low gene expression (based on GSVA score) for each model are shown in Table 4 below.
[0108] In addition, the results of C-index performance analysis for metastatic colorectal cancer (SMC cohort group and GSE17536 cohort group) and non-metastatic colorectal cancer (GSE17536 cohort group) of a representative prognostic prediction model constructed with the PLAAT3 single gene discovered from single-cell DEG analysis are shown in Figures 10 to 12, respectively.
[0109] Model numberBiomarkerDistinctionAnalysis target populationPerformance (C-index)Additional analysis population (C-index>0.7)Remarksn1PLAAT31ea GSE175360.722Representative model, 1 markern2PLAAT3 + NDUFA4L2 + GAS53ea SMC0.707 GSE17536 (0.708)n3PLAAT3 + NDUFA4L2 + EIF3H3ea SMC0.711n4PLAAT3 + NDUFA4L2 + GAL + PCAT14ea SMC0.702n5PLAAT3 + NDUFA4L2 + GAL + SNHG324ea SMC0.748Maximum performancen6PLAAT3 + NDUFA4L2 + GAL + GAREM14ea GSE175360.715n7PLAAT3 + NDUFA4L2 + GAL + TSHZ24GSE175360.721n8PLAAT3 + NDUFA4L2 + GAL + NPM34TCGA0.727
[0110]
[0111] Model numberBiomarkerHigh-risk group judgment criteria (cutoff criteria)High expression judgment cutoff (GSVA score)Low expression judgment cutoff (GSVA score)n1PLAAT3High-risk group with high expression0.46-0.59n2PLAAT3 + NDUFA4L2 + GAS5High-risk group with high expression0.33-0.53n3PLAAT3 + NDUFA4L2 + EIF3HHigh-risk group with high expression0.26-0.51n4PLAAT3 + NDUFA4L2 + GAL + PCAT1High-risk group with high expression0.29-0.37n5PLAAT3 + NDUFA4L2 + GAL + SNHG32High-risk group with high expression0.20-0.44n6PLAAT3 + NDUFA4L2 + GAL + GAREM1High-risk group with high expression High-risk group 0.33-0.28n7PLAAT3 + NDUFA4L2 + GAL + TSHZ2 High-risk group 0.27-0.33n8PLAAT3 + NDUFA4L2 + GAL + NPM3 High-risk group 0.25-0.47
[0112]
[0113] The performance evaluation of the metastatic colorectal cancer prognosis prediction model for nonmetastatic colorectal cancer showed that the PLAAT3 gene performed well in predicting the prognosis of metastatic colorectal cancer with a C-index of 0.75, and also showed excellent performance in predicting the prognosis of nonmetastatic colorectal cancer with a C-index of 0.722. In addition, the prognosis prediction performance of the multiple gene combination model for nonmetastatic colorectal cancer was also excellent with a C-index > 0.7.
[0114] Therefore, single biomarker genes or their combination models discovered from single-cell DEG analysis of malignant goblet cells can be used to predict the prognosis of not only metastatic colorectal cancer but also non-metastatic colorectal cancer, and thus are expected to have high clinical utility.
[0115] 4-3. Results of selecting biomarker genes for predicting colon cancer prognosis
[0116] Based on the experimental results above, it was confirmed that 35 genes among 43 candidate genes derived from single-cell DEG analysis can be used to predict the prognosis of metastatic or non-metastatic colorectal cancer. The information on each gene is shown in Table 5 below. For each gene, whether the gene is highly expressed (high group) or low expressed (low group) is evaluated as a high-risk group, and the cutoffs for classification of high and low expression for each model gene (cutoff, based on GSVA score) are shown in Table 6 below.
[0117] 순번유전자 기호Entrez ID유전자 명칭레퍼런스 서열 정보1ACAA210449acetyl-CoA acyltransferase 2NM_0061112ANXA1301annexin A1NM_0007003AREG374amphiregulinNM_0016574ASAP150807ArfGAP with SH3 domain, ankyrin repeat and PH domain 1NM_0184825CHRM31131cholinergic receptor muscarinic 3NM_0007406EIF3H8667eukaryotic translation initiation factor 3 subunit HNM_0037567FHIT2272fragile histidine triad diadenosine triphosphataseNM_0020128FTH12495ferritin heavy chain 1NM_0020329GAL51083galanin and GMAP prepropeptideNM_01597310GAREM164762GRB2 associated regulator of MAPK1 subtype 1NM_02275111GAS560674growth arrest specific 5NR_00257812KRT173872keratin 17NM_00042213LARGE19215LARGE xylosyl- and glucuronyltransferase 1NM_00473714LINC01811101928114long intergenic non-protein coding RNA 1811XR_00174063515MACC1346389MET transcriptional regulator MACC1NM_18276216MACROD2140733mono-ADP ribosylhydrolase 2NM_08067617MDK4192midkineNM_00239118NAP1L14673nucleosome assembly protein 1 like 1NM_00453719NDUFA4L256901NDUFA4 mitochondrialcomplex associated like 2NM_02014220NDUFB94715NADH:ubiquinone oxidoreductase subunit B9NM_00500521NPM310360nucleophosmin / nucleoplasmin 3NM_00699322NSMCE2286053NSE2 (MMS21) homolog, SMC5-SMC6 complex SUMO ligaseNR_14619123PABPC126986poly(A) binding protein cytoplasmic 1NM_00256824PCAT1100750225prostate cancer associated transcript 1NR_04526225PDE10A10846phosphodiesterase 10ANM_00666126PDE4B5142phosphodiesterase 4BNM_00260027PLAAT311145phospholipase A and acyltransferase 3NM_00706928PLCG25336phospholipase C gamma 2NM_00266129PVT15820Pvt1 oncogeneNR_00336730SMOC164093SPARC related modular calcium binding 1NM_02213731SNHG3250854small nucleolar RNA host gene 32NM_01694732SNHG784973small nucleolar RNA host gene 7NR_00367233SOD36649superoxide dismutase 3NM_00310234TSHZ2128553teashirt zinc finger homeobox 2NM_17348535YBX14904Y-box binding protein 1NM_004559
[0118]
[0119] Sequence Gene Symbol High-risk group judgment criteria High expression judgment cutoff (GSVA score) Low expression judgment cutoff (GSVA score) 1 ACAA2 High-risk group with low expression 0.64-0.55 2 ANXA1 High-risk group with high expression 0.25-0.52 3 AREG High-risk group with high expression 0.26-0.80 4 ASAP1 No significant difference at single gene level 0.30-0.51 5 CHRM3 High-risk group with low expression 0.48-0.27 6 EIF3H No significant difference at single gene level 0.44-0.60 7 FHIT High-risk group with low expression 0.56-0.38 8 FTH1 High-risk group with high expression 0.54-0.66 9 GAL High-risk group with high expression 0.51-0.59 10 GAREM1 Significant difference at single gene level Absence 0.58-0.40 11 GAS5 High risk group when high expression 0.42-0.59 12 KRT17 High risk group when high expression 0.48-0.67 13 LARGE1 High risk group when low expression 0.09-0.55 14 LINC018 11 High risk group when low expression 0.61-0.22 15 MACC1 High risk group when low expression 0.52-0.60 16 MACROD2 High risk group when low expression 0.49-0.34 17 MDK High risk group when high expression 0.67-0.40 18 NAP1L1 High risk group when low expression 0.27-0.83 19 NDUFA4L2 High risk group when high expression 0.23-0.55 20 NDUFB9 High risk group when low expression 0.28-0.67 21 NPM3 Low expression High-risk group 0.63-0.73 22 NSMCE2 High-risk group with high expression 0.22-0.59 23 PABPC1 No significant difference at single gene level 0.54-0.27 24 PCAT1 High-risk group with low expression 0.04-0.74 25 PDE10A High-risk group with high expression 0.58-0.64 26 PDE4B High-risk group with high expression 0.39-0.45 27 PLAAT3 High-risk group with high expression -0.05-0.92 28 PLCG2 High-risk group with high expression 0.71 0.00 29 PVT1 No significant difference at single gene level 0.39-0.27 30 SMOC1 High-risk group with high expression 0.68-0.48 31 SNHG32 High-risk group with high expression 0.47-0.58 32 SNHG7 High-risk group High-risk group 0.50-0.5133 High-risk group with low SOD3 expression 0.70-0.7234 High-risk group with high TSHZ2 expression 0.64-0.High-risk group with high expression of 4835YBX10.03-0.58.
[0120]
[0121] While the present invention has been described in detail through representative examples above, those skilled in the art will understand that various modifications can be made to the above-described embodiments without departing from the scope of the present invention. Therefore, the scope of the present invention should not be limited to the described embodiments, but should be determined by all changes or modifications derived from the claims and equivalent concepts.
[0122] The purpose of the present invention is to provide a biomarker for predicting the prognosis of colon cancer patients.
[0123] Another object of the present invention relates to a composition for predicting the prognosis of colon cancer patients.
[0124] Another object of the present invention relates to a kit for predicting the prognosis of colon cancer patients.
[0125] Another object of the present invention is to provide a method for providing information for predicting the prognosis of colon cancer patients.
Claims
A composition for predicting the prognosis of colon cancer, comprising an agent for measuring the level of mRNA transcribed from one or more genes selected from the group consisting of ACAA2, ANXA1, AREG, ASAP1, CHRM3, EIF3H, FHIT, FTH1, GAL, GAREM1, GAS5, KRT17, LARGE1, LINC01811, MACC1, MACROD2, MDK, NAP1L1, NDUFA4L2, NDUFB9, NPM3, NSMCE2, PABPC1, PCAT1, PDE10A, PDE4B, PLAAT3, PLCG2, PVT1, SMOC1, SNHG32, SNHG7, SOD3, TSHZ2 and YBX1.
2. In paragraph 1, A composition comprising the above genes PLAAT3 and NDUFA4L2.
3. In paragraph 2, A composition, wherein the gene further comprises at least one gene selected from the group consisting of EIF3H, GAL, GASS, NPM3, NSMCE2, PLCG2, PVT1, SNHG32 and YBX1.
4. In paragraph 1, A composition wherein the colon cancer is metastatic colon cancer or non-metastatic colon cancer.
5. In paragraph 1, A composition wherein the agent for measuring the level of said mRNA comprises a primer or probe that specifically binds to said gene, or the agent for measuring the level of said protein comprises an antibody or aptamer that specifically binds to said protein.
6. A kit for predicting the prognosis of colorectal cancer, comprising a formulation for measuring the level of mRNA transcribed from one or more genes selected from the group consisting of ACAA2, ANXA1, AREG, ASAP1, CHRM3, EIF3H, FHIT, FTH1, GAL, GAREM1, GAS5, KRT17, LARGE1, LINC01811, MACC1, MACROD2, MDK, NAP1L1, NDUFA4L2, NDUFB9, NPM3, NSMCE2, PABPC1, PCAT1, PDE10A, PDE4B, PLAAT3, PLCG2, PVT1, SMOC1, SNHG32, SNHG7, SOD3, TSHZ2 and YBX1.
7. In paragraph 6, A kit wherein the above genes include PLAAT3 and NDUFA4L2.
8. In paragraph 7, A kit wherein the above gene further comprises at least one gene selected from the group consisting of EIF3H, GAL, GASS, NPM3, NSMCE2, PLCG2, PVT1, SNHG32 and YBX1.
9. In paragraph 6, A kit wherein the above colon cancer is metastatic colon cancer or non-metastatic colon cancer.
10. In paragraph 6, A kit wherein the agent for measuring the level of said mRNA comprises a primer or probe that specifically binds to said gene, or the agent for measuring the level of said protein comprises an antibody or aptamer that specifically binds to said protein.
11. A method for providing information for predicting the prognosis of colon cancer, comprising a measuring step of quantifying the expression level of one or more genes by measuring the level of mRNA transcribed from or protein expressed therefrom at least one gene selected from the group consisting of ACAA2, ANXA1, AREG, ASAP1, CHRM3, EIF3H, FHIT, FTH1, GAL, GAREM1, GAS5, KRT17, LARGE1, LINC01811, MACC1, MACROD2, MDK, NAP1L1, NDUFA4L2, NDUFB9, NPM3, NSMCE2, PABPC1, PCAT1, PDE10A, PDE4B, PLAAT3, PLCG2, PVT1, SMOC1, SNHG32, SNHG7, SOD3, TSHZ2, and YBX1 in a biological sample isolated from an individual.
12. In paragraph 11, A method wherein the measuring step comprises a contacting step of contacting the sample with a primer or probe that specifically binds to the mRNA or contacting the sample with an antibody or aptamer that specifically binds to the protein.
13. In paragraph 11, The above information providing method further includes a comparison step of comparing the expression level of one or more genes measured in a sample separated from an individual with a high or low expression judgment cutoff for each gene in [Table 6].
14. In paragraph 13, The above information provision method determines that the prognosis of colon cancer is poor if the expression level of one or more genes selected from the group consisting of ANXA1, FTH1, GAL, GAS5, KRT17, MDK, NDUFA4L2, NSMCE2, PDE10A, PDE4B, PLAAT3, PLCG2, SMOC1, SNHG32, SNHG7, TSHZ2, and YBX1 is higher than the high expression cutoff for each gene, or that the prognosis of colon cancer is good if the expression level of one or more genes selected from the group consisting of ACAA2, CHRM3, FHIT, LARGE1, LINCO1811, MACC1, MACROD2, NAP1L1, NDUFB9, NPM3, PCAT1, and SOD3 is lower than the low expression cutoff for each gene, or that the prognosis of colon cancer is poor if the expression level of one or more genes selected from the group consisting of ACAA2, CHRM3, FHIT, LARGE1, LINCO1811, MACC1, MACROD2, NAP1L1, NDUFB9, NPM3, PCAT1, and SOD3 is higher than the high expression cutoff for each gene. A method further comprising a judgment step of judging that the prognosis is good.
15. In paragraph 11, A method wherein the biological sample is at least one selected from the group consisting of tissue, cells, whole blood, serum and plasma isolated from an individual.
Citation Information
Patent Citations
Biomarker for predicting prognosis of colon cancer patient and application of biomarker
CN117431319A
Biomarkers indicative of colon cancer and metastasis and diagnosis and screening therapeutics using the same
KR1020100121949A
Inflammatory genes and microrna-21 as biomarkers for colon cancer prognosis
US20110183859A1
Methods for identifying genes which predict disease outcome for patients with colon cancer
US20110257034A1
Methods for treating, preventing and detecting the prognosis of colorectal cancer
WO2019222618A1