Composition for diagnosing colorectal cancer and use thereof

The colon cancer diagnostic kit uses blood-based gene expression analysis and AI to accurately detect colon cancer and adenomas, addressing the limitations of current screening methods with high sensitivity and specificity.

WO2026005448A1PCT designated stage Publication Date: 2026-01-02INOGENIX CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/008820
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-27
Filing Date
2025-06-24
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Current colorectal cancer screening methods, such as colonoscopy and stool tests, face challenges in detecting both cancer and adenomas with high sensitivity and patient compliance is low, while existing blood-based tests lack accuracy for colon cancer prevention.

Method used

A colon cancer diagnostic kit using blood that measures the relative expression levels of specific genes and proteins, including CCR1, CD177, DLG5, EPCAM, ERBB2, and others, through RT-PCR and AI analysis to detect colon cancer and adenomas with high sensitivity and specificity.

Benefits of technology

The kit achieves a sensitivity of 93% for detecting colon cancer and 91% for advanced adenomas, with a specificity of 91%, effectively identifying early-stage cancer and asymptomatic cases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025008820_02012026_PF_FP_ABST
    Figure KR2025008820_02012026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to: a composition for diagnosing colorectal cancer, comprising a substance capable of measuring the relative expression levels of CCR1, CD177, DLG5, EPCAM, ERBB2, FCGR1A, FOXA2, GK, KRT19, KRT73, MKI67, MPO, MUC1, NME1, NPTN, NQO2, SH2D1B, SNAI2, TERT, TUG1, and VIM genes or proteins encoded by the genes; a kit for diagnosing colorectal cancer, comprising the composition; and a method for diagnosing colorectal cancer by using the kit.
Need to check novelty before this filing date? Find Prior Art

Description

Composition for diagnosing colon cancer and its use

[0001] The present invention relates to a colon cancer diagnostic kit and its use, and more particularly, to a colon cancer diagnostic composition suitable for diagnosing early colon cancer, a kit including the composition, and a colon cancer diagnostic method.

[0002] As of 2020, colorectal cancer was the third most common cancer worldwide and the second leading cause of death. South Korea has the highest incidence of colorectal cancer in the world, with 45 cases per 100,000 people, the highest among the countries studied. Furthermore, according to 2020 Statistics Korea data, colorectal cancer was the third leading cause of cancer death.

[0003] If diagnosed early (Stage 1), colon cancer has a survival rate of over 90%. However, if diagnosed late (Stage 4), treatment effectiveness plummets to less than 15%. Therefore, early diagnosis is crucial.

[0004] However, since most colon cancers arise from adenomas in the colon, removing them when they are in the adenoma stage can prevent colon cancer with a 100% probability. In other words, the detection and removal of adenomas are crucial for complete colon cancer prevention. However, currently, more than 50% of colon cancers are found in stages 3 and 4, when treatment is less effective. In particular, approximately 77% of colon cancers found in people under 50 years of age are found in stages 3 and 4, which is the cause of the high mortality rate of colon cancer in this age group.

[0005] For colon cancer prevention, the detection and removal of both cancer and adenomas are crucial. Currently, colonoscopy is the only screening method capable of detecting both cancer and adenomas with a sensitivity of over 90%. However, colonoscopy is often avoided due to concerns and fears surrounding the extreme discomfort of ingesting large amounts of bowel cleansing agents, the risks of the procedure requiring anesthesia, and the invasive nature of the procedure, which involves inserting an endoscope through the anus.

[0006] Meanwhile, from the perspective of medical professionals and hospitals, colonoscopy is a high-risk examination due to the hassle of the preparation process, the cost burden of requiring a large number of personnel and a long time, and the risk of accidents such as bleeding, infection, and perforation that may occur during the procedure.

[0007] For this reason, global companies have been developing colorectal cancer diagnostic tests using relatively simple stool samples. However, even molecular-based colorectal cancer diagnostic tests, which offer the highest sensitivity among stool tests, only detect adenomas, the primary cause of colorectal cancer, in less than 50% of cases.

[0008] Moreover, due to low compliance with stool tests due to aversion to the specimen, the participation rate in regular colorectal cancer prevention health screenings using colonoscopies or stool tests is the lowest among the five major cancers, reaching less than 50% of the target population worldwide. Therefore, there is a need to develop a new colorectal cancer diagnostic test that (1) can detect not only colorectal cancer but also adenomas with high sensitivity, and (2) has higher patient compliance than colonoscopies or stool tests.

[0009] Blood is a fundamental specimen used in regular checkups. It minimizes the examinee's resistance and is collected through a less invasive method than tissue testing, making it very useful in developing screening tests for colon cancer prevention. Furthermore, blood contains various biomarkers that reflect physical changes related to diseases such as cancer, making blood-based diagnosis very useful. Therefore, the development of a blood-based colon cancer diagnostic test is highly suitable for health screenings and is highly suitable for reducing the incidence and mortality of colon cancer by increasing the rate of health screenings. For this reason, the U.S. Insurance Association has announced that it will immediately adopt blood-based colon cancer diagnostic tests with a specificity of 90% and a sensitivity of 74% or higher for the detection of colon cancer and adenomas.

[0010] However, the products developed so far have lower sensitivity for detecting colon cancer than stool tests, and their sensitivity for detecting polyps for colon cancer prevention is also very low, so they are evaluated as lacking accuracy as a diagnostic test for colon cancer prevention.

[0011] Therefore, there is a need for a blood-based colon cancer diagnostic test product that has a sensitivity suitable for detecting not only colon cancer but also colon cancer and adenomas that cause colon cancer, as a diagnostic test for colon cancer prevention.

[0012] [Prior patent literature]

[0013] Republic of Korea Patent Publication No. 10-2023-0014990

[0014] The present invention was conceived in response to the above-mentioned necessity, and the purpose of the present invention is to provide a colon cancer diagnostic kit using blood that detects colon cancer, adenoma, and advanced adenoma with high sensitivity and specificity.

[0015] The purpose of the present invention is to provide a method for diagnosing colon cancer using blood that detects colon cancer, adenoma, and advanced adenoma with high sensitivity and specificity.

[0016] In order to achieve the above purpose, the present invention includes a substance capable of measuring the relative expression level of CCR1, CD177, DLG5, EPCAM, ERBB2, FCGR1A, FOXA2, GK, KRT19, KRT73, MKI67, MPO, MUC1, NME1, NPTN, NQO2, SH2D1B, SNAI2, TERT, TUG1, VIM, 18S rRNA and GAPDH genes or proteins encoded by the genes, wherein the relative expression level is 2 -ΔCq It is indicated as 2 -ΔCq can be obtained by the following relationship:

[0017] [Relationship]

[0018] 2 -ΔCq = 2-(target gene Cq - GAPDH gene Cq)

[0019] Here, the target genes are CCR1, CD177, DLG5, EPCAM, ERBB2, FCGR1A, FOXA2, GK, KRT19, KRT73, MKI67, MPO, MUC1, NME1, NPTN, NQO2, SH2D1B, SNAI2, TERT, TUG1, VIM, 18S rRNA is an internal control, GAPDH is a reference gene, and the Cq value is set as a threshold for a certain fluorescence value, and the composition for diagnosing colon cancer is provided.

[0020] In one embodiment of the present invention, the material capable of measuring the relative expression level of the gene or the relative expression level of the protein encoded by the gene is preferably, but is not limited to, a primer and probe set or an antibody.

[0021] In another embodiment of the present invention, the primer and probe sequences are preferably those described in SEQ ID NOs: 1 to 69, but are not limited thereto.

[0022] The present invention also provides a colon cancer diagnostic kit comprising the composition of the present invention.

[0023] In addition, the present invention provides a method for providing information for predicting the risk of developing colon cancer, characterized in that the risk of developing colon cancer is predicted to be high when the relative expression of the CD177, FOXA2, SNAI2, TERT, DLG5, FCGR1A, NPTN, NQO2 gene or protein increases or the relative expression of the MKI67, KRT73, EPCAM, NME1 gene or protein decreases by measuring the relative expression levels of the CCR1, CD177, DLG5, EPCAM, ERBB2, FCGR1A, FOXA2, GK, KRT19, KRT73, MKI67, MPO, MUC1, NME1, NPTN, NQO2, SH2D1B, SNAI2, TERT, TUG1 and VIM genes or proteins encoded by the genes for a genetic sample obtained from a subject.

[0024] Here, the relative expression level is 2 -ΔCq It is indicated as 2 -ΔCq can be obtained by the following relationship:

[0025] [Relationship]

[0026] 2 -ΔCq = 2-(target gene Cq - GAPDH gene Cq)

[0027] Here, the target genes are CCR1, CD177, DLG5, EPCAM, ERBB2, FCGR1A, FOXA2, GK, KRT19, KRT73, 18S rRNA, MKI67, MPO, MUC1, NME1, NPTN, NQO2, SH2D1B, SNAI2, TERT, TUG1, and VIM genes, and the Cq value refers to the number of cycles at which a certain fluorescence value is set as a threshold and this threshold is reached.

[0028]

[0029] The primers of the present invention can be chemically synthesized using the phosphoramidite solid support method or other well-known methods. These nucleic acid sequences can also be modified using various means known in the art.

[0030] Non-limiting examples of such modifications include methylation, "capping," substitution with one or more analogs of a natural nucleotide, and modifications between nucleotides, such as modification with uncharged linkers (e.g., methyl phosphonate, phosphotriester, phosphoroamidate, carbamate, etc.) or charged linkers (e.g., phosphorothioate, phosphorodithioate, etc.). Nucleic acids may contain one or more additional covalently linked moieties, such as proteins (e.g., nucleases, toxins, antibodies, signal peptides, lysine, etc.), intercalating agents (e.g., acridine, psoralen, etc.), chelating agents (e.g., metals, radioactive metals, iron, oxidizing metals, etc.), and alkylating agents.

[0031] The nucleic acid sequence of the present invention may also be modified with a label capable of directly or indirectly providing a detectable signal. Examples of labels include radioisotopes, fluorescent molecules, biotin, and the like.

[0032] In the method of the present invention, the amplified target sequence can be labeled with a detectable label. In one embodiment, the label may be a fluorescent, phosphorescent, chemiluminescent, or radioactive material, but is not limited thereto. Preferably, the label may be fluorescein, phycoerythrin, rhodamine, lissamine, Cy-5, or Cy-3. When performing RT-PCR by labeling the 5'-end and / or 3'-end of the primer during amplification of the target sequence with Cy-5 or Cy-3, the target sequence can be labeled with a detectable fluorescent label.

[0033] Additionally, labeling using radioactive materials is used when performing RT-PCR. 32 P or 35 Adding a radioactive isotope, such as S, to the PCR reaction mixture can cause the radioactivity to be incorporated into the amplified product as it is synthesized, thereby radioactively labeling the amplified product. One or more sets of oligonucleotide primers can be used to amplify the target sequence.

[0034] The label provides a signal that can be detected by fluorescence, radioactivity, chromogenicity, gravimetric measurement, X-ray diffraction or absorption, magnetic, enzymatic activity, mass analysis, binding affinity, hybridization radiofrequency, nanocrystal.

[0035] According to one aspect of the present invention, the expression level is measured at the mRNA level through RT-PCR. To this end, a novel primer pair and a fluorescently labeled probe that specifically bind to the biomarker gene of the present invention and the GAPDH gene, etc. are required. In the present invention, the primers and probes specified by a specific base sequence may be used, but are not limited thereto. Anything that specifically binds to these genes and provides a detectable signal to perform RT-PCR may be used without limitation. In the above, FAM and Quen (Quencher) refer to fluorescent dyes.

[0036] The RT-PCR method applied to the present invention can be performed through a known process commonly used in the art.

[0037] The step of measuring the mRNA expression level can be used without limitation as long as it is a method capable of measuring the conventional mRNA expression level, and can be performed through radiometric measurement, fluorescence measurement, or phosphorescence measurement depending on the type of probe label used, but is not limited thereto.

[0038] As one of the methods for detecting an amplification product, the fluorescence measurement method is performed by labeling the 5'-end of the primer with Cy-5 or Cy-3, and when real-time RT-PCR is performed, the target sequence is labeled with a detectable fluorescent label, and the fluorescence labeled in this way can be measured using a fluorometer.

[0039] In addition, the radiometric measurement method is used when performing RT-PCR. 32 P or 35 After labeling the amplified product by adding a radioactive isotope such as S to the PCR reaction solution, the radioactivity can be measured using a radiometric measuring device such as a Geiger counter or a liquid scintillation counter.

[0040] According to a preferred embodiment of the present invention, a fluorescently labeled probe is attached to a PCR product amplified through the RT-PCR to emit fluorescence of a specific wavelength, and simultaneously with the amplification, the mRNA expression level of the genes of the present invention is measured in real time by a fluorescence meter of a PCR device, and the measured value is calculated and visualized through a PC, so that an examiner can easily check the expression level.

[0041] According to another aspect of the present invention, the screening kit may be a kit for diagnosing colon cancer and colon polyps, characterized in that it includes essential elements necessary for performing a reverse transcription polymerase reaction. The reverse transcription polymerase reaction kit may include a pair of primers specific for each gene of the present invention. The primers are nucleotides having a sequence specific to the nucleic acid sequence of each marker gene, and may have a length of about 7 bp to 50 bp, more preferably about 10 bp to 30 bp.

[0042] Other kits may include test tubes or other appropriate containers, reaction buffers (with varying pH and magnesium concentrations), deoxynucleotides (dNTPs), enzymes such as Taq polymerase and reverse transcriptase, DNAse and RNAse inhibitors, DEPC-water, and sterile water.

[0043] Additionally, the kit of the present invention may further include a user guide describing optimal reaction performance conditions.

[0044] A guide is a printed document that explains how to use the kit, for example, how to prepare buffers, what reaction conditions are suggested, etc.

[0045] The guide may include instructions in the form of a pamphlet or leaflet, a label attached to the kit, or on the surface of the package containing the kit. The guide may also include information disclosed or provided through electronic media, such as the Internet.

[0046] In the present invention, the term "method for providing colon cancer information" refers to providing objective basic information necessary for diagnosing cancer as a preliminary step for diagnosis, and excludes a doctor's clinical judgment or opinion.

[0047] The term "primer" refers to a short nucleic acid sequence having a short free three-terminal hydroxyl group, capable of forming base pairs with a complementary template and serving as an initiation point for copying the template strand. The primer can initiate DNA synthesis in the presence of a polymerization reagent (i.e., DNA polymerase or reverse transcriptase) and four different nucleoside triphosphates in an appropriate buffer and temperature. The primers of the present invention are sense and antisense nucleic acids having a sequence of 7 to 50 nucleotides, each of which is specific to a marker gene. The primers may incorporate additional features that do not alter the basic property of the primers as initiators of DNA synthesis.

[0048] The term "probe" refers to a single-stranded nucleic acid molecule, comprising a sequence complementary to a target nucleic acid sequence.

[0049] The term "real-time reverse transcription polymerase reaction (real-time RT-PCR)" refers to a molecular biological polymerization method that uses reverse transcriptase to reverse transcribe RNA into complementary DNA (cDNA), then uses the produced cDNA as a template to amplify the target using a target primer and a target probe containing a label, and simultaneously quantitatively detects the signal generated from the label of the target probe on the amplified target.

[0050] The colon cancer prediction method of the present invention can utilize data mining methods that can diagnose colon cancer through information learning, and can be effectively improved through AI analysis. Therefore, the method for diagnosing or predicting colon cancer of the present invention preferably utilizes a method capable of measuring the relative expression level of colon cancer diagnostic markers and / or an AI analysis method.

[0051] In the present invention, when AI analysis is used for a colon cancer prediction model, various interpretable models can be used without limitation, and models such as linear regression, logistic regression, neural network analysis, decision trees, decision rules, rule fit, and support vector machines can be applied without limitation, and in a preferred embodiment of the present invention, logistic regression analysis, decision trees, neural network analysis, and support vector machines are used in particular.

[0052] Meanwhile, the prediction model of the present invention may include a colon cancer diagnosis unit, a classification unit, and a weight assignment unit, wherein the colon cancer diagnosis unit uses relative expression level information received from a relative expression level information reception unit of a patient's disease-related gene marker as input information, the colon-related disease classification unit may perform a process of classifying colon cancer using a neural network as a classifier, and the weight assignment unit may select colon cancer by assigning weights to the classification results.

[0053] Neural network analysis according to embodiments of the present invention refers to a system that constructs one or more layers and performs judgment based on multiple data. For example, in neural network analysis, the input layer is a layer that inputs relative expression level information of genetic markers as data into a neural network analysis model, and the output layer is a layer that outputs a result that can determine the presence or absence of a colon cancer patient based on various input information. The hidden layer is a layer that performs a process that can determine the presence or absence of a patient by assigning weights to various judgment criteria (gene mutation information).

[0054] A method for predicting colon cancer using an AI analysis technique according to an embodiment of the present invention estimates a neural network analysis model having the number of hidden nodes using an MLP neural network. In addition, among several neural network models constructed through various variable transformations of input and output variables, the neural network model with the highest estimated accuracy from each model is determined as the final neural network model for predicting skin diseases. The AI ​​analysis may be composed of an input layer, a hidden layer, and an output layer, and the neural network analysis model through the neural network analysis step may be a neural network model having several hidden layers and several hidden nodes.

[0055] The present invention is described below.

[0056] The present inventors selected blood mRNA markers that appear in the process of colon cancer and adenoma formation and developed a colon cancer diagnostic kit and diagnostic method (ONCOCHECK) that can detect colon cancer and adenoma through an artificial intelligence learning process using these markers. TM Colorectal Cancer) was developed.

[0057] As can be seen through the present invention, the kit and diagnostic method of the present invention were confirmed to have a sensitivity of 93% (45 / 48) for detecting colon cancer, a sensitivity of 91% (87 / 95) for detecting advanced adenoma, and a specificity of 91% (147 / 161) as a result of testing 413 test clinical samples, and in particular, a sensitivity of 96% was confirmed for early colon cancer (stage 1-2), and a specificity of 93% was confirmed for asymptomatic and symptomatic patients with negative colonoscopy findings.

[0058] Figure 1 is a scatter plot, where the X-axis represents the fold change value of the NGS data set, and the Y-axis represents the fold change value of the GEO-PAD data set. Seven transcripts were identified through GEO-PAD-NGS log2(FC) analysis. Two transcripts (ADAMTS1 and SLC26A8) were removed because their patterns were different between the GEO-PAD log2(FC) and NGS log2(FC), and five transcripts were identified to have the same up- or down-regulated patterns in the GEO-PAD log2(FC) and NGS log2(FC) results. CD177 and NQO2 were upregulated in red from GEO-PAD log2(FC) and NGS log2(FC), and SH2D1B, DLG5, and KRT73 were downregulated in blue from GEO-PAD log2(FC) and NGS log2(FC), {log2(FC) = 0: HC and CRC expressed genes are the same, log2(FC) > 0: CRC expressed genes are higher than HC, CRC, CRC, log2(FC) < 0: HC expressed genes are higher than CRC}

[0059] Figure 2 is a graph showing the t-test clinical results and characteristics of clinical samples for biomarker candidates. DLG5 and CD177 were statistically significant with p-values ​​< 0.0001. SH2D1B and NQO2 were statistically significant with p-values ​​< 0.05. KRT73 was not statistically significant with p-values ​​> 0.05 (ns = p-value > 0.05, * = p-value < 0.05, ** = p-value < 0.01, *** = p-value < 0.001, **** = p-value < 0.0001). In the figure, DLG5 represents Discs Large MAGUK Scaffold Protein 5; CD177 represents CD177 molecule; SH2D1B represents SH2 Domain Containing 1B; NQO2 represents N-Ribosyldihydronicotinamide: Quinone Reductase 2; KRT73 is an abbreviation for Keratin 73.

[0060] Figure 3 shows the RT-qPCR results for FCGR1A and MPO expression in clinical samples. MPO was upregulated in both the NA and AA groups compared to the HC group, with statistical significance. FCGR was slightly upregulated in the NA group compared to the HC group, but without statistical significance. FCGR1A was upregulated in the AA group and was statistically significant compared to the HC group. P values ​​(*, **, and *** indicate P < 0.05, P < 0.01, and P < 0.001, respectively) were calculated using a two-tailed Student's test.

[0061] Figures 4 and 5 are drawings showing RT-qPCR results for the expression of the biomarker of the present invention.

[0062] Figure 6 is a diagram showing the ROC curve in the test set.

[0063] The present invention is described in more detail below through non-limiting examples. However, the following examples are intended to illustrate the present invention, and the scope of the present invention is not to be construed as being limited by the following examples.

[0064] The development process of the present invention is summarized below.

[0065] Each step is explained below.

[0066] Example 1: Selection of circulating markers

[0067] The applicant collected and conducted a study over the past eight years, based on more than 2,100 whole blood samples with colonoscopy and biopsy results.

[0068] To select mRNA markers related to the development of colon cancer and adenoma, research was conducted using three major strategies.

[0069] (1) Exploration of useful markers through validation of known blood cancer cell-associated mRNA markers and selection of useful markers through qPCR validation.

[0070] (2) Exploration of new markers through RNA-Seq and public data analysis, and discovery of new colon cancer and adenoma markers through qPCR validation.

[0071] (3) Search for blood mRNA markers associated with colon cancer and adenoma identified by reference to literature and selection of useful markers through qPCR verification.

[0072] 1) Obtaining research subjects and blood samples

[0073] The applicant conducted a study using specimens from a total of 2,100 examinees with colonoscopy results for the development of a colorectal cancer screening test. The specimens were (a) examinees (n = 1,100) who visited the departments of gastroenterology at three tertiary hospitals in Seoul and underwent colonoscopy (IRB #4-2017-0148, #3-2017-0024, #2017-02-022-009) and (b) examinees (n = 1,000) who visited a university hospital in Wonju, Gangwon-do for the purpose of health checkups (IRB #CR319115, #CR222015). The study was conducted after securing IRB approval through two routes.

[0074] In the present invention, the results of colonoscopy and histopathological examination were classified into four groups.

[0075] A. Control group (NC): Asymptomatic individuals scheduled to undergo colonoscopy for health screening purposes and symptomatic individuals who visited the gastroenterology department of a tertiary hospital due to digestive symptoms and were recommended to undergo colonoscopy by a doctor. Among these individuals, those who did not find any abnormalities in the colon during colonoscopy (NC2) and those who did not have adenomas such as inflammatory or hyperplastic polyps but had various other symptoms (NC1)

[0076] B. Non-advanced adenoma group (NA): Patients with tubular adenoma (TA) or sessile serrated adenoma (SSA) less than 10 mm in size, regardless of the number.

[0077] C. Advanced adenoma group (AA): Group in which low-grade TA or SSA measuring 10 mm or more was found regardless of the number (AA4), group in which high-grade dysplasia (AA3), villous adenoma or tubulovillous adenoma (AA2), and group in which carcinoma in situ was found (AA1)

[0078] D. Colon cancer group: Group in which colon cancer (stages 1 to 4) was found

[0079] Major classificationSubclassificationHistopathological classificationColorectal cancer (CRC)CRC1Colorectal cancer stage 1CRC2Colorectal cancer stage 2CRC3Colorectal cancer stage 3CRC4Colorectal cancer stage 4Advanced adenoma group (AA)AA1Carcinoma in situAA2Villous adenoma (≥75%) / Tubular villous adenoma (≥25%)AA3High-grade dysplasiaAA4Tubular adenoma / serrated adenoma (≥10 mm)Non-advanced adenoma group (NA)NA1Tubular adenoma / serrated adenoma (<10 mm, n≥3)NA2Tubular adenoma / serrated adenoma (<10 mm, n<3)Control group (NC)NC1Non-neoplastic findings (hyperplastic polyp, inflammatory polyp, chronic non-specific inflammation, etc.)NC2Asymptomatic and symptomatic patients with negative colonoscopy findings

[0080] Table 1 shows the classification according to histopathological classification.

[0081] 2) Verification of CTC-derived markers

[0082] We selected and evaluated 10 circulating tumor cell-derived mRNA markers that have been previously shown to be useful for breast cancer diagnosis (Park S, et al. Blood Test for Breast Cancer Screening through the Detection of Tumor-Associated Circulating Transcripts, Int J Mol Sci. 2022 Aug 15;23(16):9140; Wang, HY et al. Detection of circulating tumor cell-specific markers in breast cancer patients using the quantitative RT-PCR assay. Int. J. Clin. Oncol. 2015, 20, 878-890). Circulating tumor cells are cancer cells with epithelial cell characteristics, and therefore contain epithelial gene markers, proliferation gene markers related to cell proliferation, and epithelial-to-mesenchymal transition-related gene markers that are necessary for epithelial cells to exist and survive in the blood. Using approximately 700 samples, qPCR was performed, and the usefulness was confirmed through a t-test. Among them, nine CTC-associated mRNA markers were selected as useful for the diagnosis of colorectal cancer. These markers were named Tumor-Associated Circulating Transcripts (TACT).

[0083] 3) Selection of CTC-derived mRNA markers

[0084] It is well known that the blood of colorectal cancer patients contains circulating cancer cells (CBCs) derived from the cancerous tissue. Diagnostic products that isolate CBCs from colorectal cancer using patient blood have been commercialized. Molecular diagnostic products utilizing CBC-associated mRNA have also been commercialized.

[0085] Therefore, the applicant developed ONCOCHECK™ Colorectal Cancer, a blood-based screening test for colorectal cancer and polyps. We have identified CTCs found in the blood of individuals with colorectal cancer and polyps, non-coding RNAs expressed by colorectal cancer and polyps, and numerous RNAs expressed by immune cells in response to the development of colorectal cancer and polyps, and have elucidated their diagnostic utility.

[0086] The applicant developed an in vitro diagnostic medical device (ONCOCHECK™Colorectal Cancer) for the purpose of screening for colorectal cancer using blood by performing qPCR targeting TACTs with confirmed diagnostic utility and diagnosing the expression profile of each gene obtained as a result using an artificial intelligence-based algorithm.

[0087] 4) Discovery of new markers associated with colon cancer and adenomas

[0088] Discovery of CD177, NQO2, DLG5, SH2D1B, and KRT73 markers

[0089] In the present invention, 956 transcripts were mapped to the PPI network in PPI analysis, and 132 transcripts were selected as GEO-PPIs with 10 or more interactions.

[0090] We analyzed the transcriptome using the CRC expression dataset ( / humantfs.ccbr.utoronto.ca / index.php) by searching the current list of human TFs and their motifs (v1.01) on the Human Transcription Factor site. 2,765 TF transcripts were included in the TF file name and their motifs (v1.01). The dataset consisting of these 2,765 TF transcripts was named the TF dataset. We performed an intersection between the 956 transcripts in GEO and the 2,765 transcripts in the TF dataset. The intersection resulted in the identification of 127 transcripts. The resulting set of 127 transcripts was named GEO-TF.

[0091] Furthermore, the GWAS search results were 124 summary data. After unifying the 124 summary data, 3,793 cancer patients and 410,350 healthy controls were used to identify CRC-related transcripts based on GWAS. Specifically, there were 15,281 SNPs using the useMart and getBM packages in the biomaRt package. A total of 1,592 transcripts met the p-value criterion and were selected as GWAS transcripts for comparison with the GEO dataset. The dataset of GWAS results was named the GWAS dataset. There were 53 transcripts overlapping between the GEO and GWAS datasets. The 53 transcripts from the GEO and GWAS analyses were named the GEO-GWAS dataset.

[0092] Additionally, the Z-score was used to determine whether genetic variation was associated with the number of SNPs. In the present invention, the Z-score was determined to be 1.15 or higher, which was applied to select transcripts with a higher correlation between the frequency of genetic variation and the number of SNPs in the eQTL list genes. After applying the elbow point of 1.15 to the eQTL data set, 1,289 transcripts remained. These 1,289 transcripts were designated as the eQTL data set. Common transcripts were identified by crossing the GEO data set with the eQTL data set. 194 transcripts were identified as common transcripts and designated as the GEO-eQTL data set.

[0093] Meanwhile, common transcripts were identified by intersecting 2,056 transcripts from the CoReCG dataset with 956 transcripts from the GEO dataset. Ninety-seven transcripts were identified as common transcripts and were named the GEO-CoReCG dataset. The DigSeE dataset contains 8,866 CRC-associated transcripts. 2,486 common transcripts between the statistically selected transcripts and the CRC-associated transcripts were selected from the DigSeE dataset and named the DigSeE dataset. The 2,486 transcripts from DigSeE were intersected with the 956 transcripts from the GEO dataset, and 48 transcripts were intersected with the GEO-DigSeE dataset.

[0094] GEO-PPI identified 45 upregulated and 87 downregulated transcripts. GEO-TF identified 24 upregulated and 103 downregulated transcripts. GEO-GWAS data set identified 11 upregulated and 42 downregulated transcripts. GEO-eQTL data set identified 43 upregulated and 151 downregulated transcripts. GEO-DiGSeE data set identified 21 upregulated and 27 downregulated transcripts. GEO-CoReCG data set identified 33 upregulated and 64 downregulated transcripts. These transcripts totaled 440 transcripts, and six PAD results were analyzed with GEO data set. GEO data set was bioinformatically analyzed with PPI, and 132 transcripts were selected as GEO-PPI transcripts. The GEO dataset was bioinformatically analyzed using the DiGSeE dataset, and 48 transcripts were selected for GEO-DiGSeE. The GEO dataset was bioinformatically analyzed using the TF dataset, and 127 transcripts were selected as GEO-TF. The GEO dataset was bioinformatically analyzed using the eQTL dataset, and 194 transcripts were selected as GEO-eQTL dataset. The GEO dataset was bioinformatically analyzed using the GWAS dataset, and 53 transcripts were selected for GEO-GWAS. The GEO dataset was bioinformatically analyzed using the CoreCG dataset, and 97 transcripts were selected for GEO-CoReCG. The total number of up- and down-regulated transcripts in GEO-PPI, GEO-TF, GEO-GWAS, GEO-eQTL, GEO-DiGSeE, and GEO-CoReCG was 132, 48, 127, 194, 53, and 97, respectively. The total number of CRC-related transcripts in GEO-PAD was 440.A cutoff of 6 for average expression was set in the MA plot. Sixty-five transcripts with an average expression lower than 6 were removed from the MA plot. Based on the selection of transcripts with higher average expression, reliable transcripts for CRC differentiation were selected for clinical validation. Therefore, 65 of the 440 transcripts were removed, leaving 375 transcripts in the GEO-PAD dataset.

[0095] We also used the p-value and gene expression values ​​from the NGS fold change as a reference for the GEO dataset. Characteristics of the CRC and HC groups for the NGS dataset. The mean age in the CRC group was 54.0 years, and the mean age in the HC group was 46.8 years. The gender ratio in the CRC and HC groups was 5:5, and the fold change value was approximately 1.5-fold or greater. We applied the Korean NGS dataset to identify CRC-related transcripts using the criteria of p-value <0.05 and gene expression value >6. Based on the fold change values, 92 up-regulated transcripts and 125 down-regulated transcripts were selected. A total of 217 transcripts were identified based on the criteria of a fold change value of approximately 1.5-fold or greater and a p-value <0.05. Subsequently, 194 transcripts with expression values ​​greater than 6 were identified. Therefore, 194 transcripts were identified and named the NGS dataset.

[0096] Additionally, the GEO-PAD dataset contained 375 transcripts, while the NGS dataset contained 194 transcripts. The GEO and NGS transcripts were intersected to identify common transcripts associated with CRC. Intersecting these two datasets revealed that seven transcripts from both GEO-PAD and NGS showed significant differences in CRC compared to HC.

[0097] Among the seven transcripts, two were included in all positive change regions of the scatterplot. Three of the seven transcripts were included in all negative change regions of the scatterplot. The remaining two transcripts were identified as having differential expression patterns compared to the NGS and GEO datasets (Figure 1). Therefore, two transcripts with different expression patterns in the GEO and NGS datasets were excluded. After verification of the GEO and NGS datasets, two upregulated transcripts (CD177, NQO2) were selected. Three downregulated transcripts (DLG5, SH2D1B, KRT73) were selected. Five transcripts were identified by the GEO-PAD and NGS datasets. The fold changes of DLG5, SH2D1B, and KRT73 were identified as negative fold changes using the DEG method. A negative fold change indicated a negative change in gene expression compared to HC. The fold changes of CD177 and NQO2 were identified as positive fold changes using the DEG method, which indicated a positive change in gene expression compared to HC. The average of all five transcripts was greater than 6 and the p-value was less than 0.05. SH2D1B and KRT73 were identified by PPI among the six biological PADs. CD177 and NQO2 were identified by eQTL among the six biological PADs. DLG5 was identified by DiGSee among the six biological PADs (Table 2).

[0098] No. Gene symbol fold change (log function) Average (gene expression) p-value PADGEONGSGEONGSGEONGS1DLG5-0.59-0.637.047.560.00050.007DiGsEc2CD1771.091.347.938.440.000010.03eQTL3SH2D1B-0.58-0.869.258.850.0000010.01PPI4NQO20.570.7710.3611.040.0000020.01eQTL5KRT73-0.76-1.006.237.070.00010.001PPI

[0099] Table 2 shows the logFold change, mean expression, p-value, and biomarker candidate results for PAD.

[0100] RT-qPCR was performed for clinical validation of five circulating transcripts, including DLG5, CD177, SH2D1B, NQO2, and KRT73. RT-qPCR was performed using 106 samples from the CRC group and 123 samples from the HC group. To construct a validation dataset for testing RT-qPCR using total RNA from the CRC and HC groups, the mean age in the CRC group was 63.2 years, and the mean age in the HC group was 48.9 years. The sex ratio in the CRC group was 59.4:40.6.

[0101] And in the HC group, it was 57.7:42.3. The RT-qPCR results of five circulating transcripts, DLG5, CD177, SH2D1B, and NQO2, were significantly upregulated in the CRC group. DLG5, CD177, SH2D1B, and NQO2 were confirmed to have statistical significance in the t-test. KRT73 was downregulated in the CRC group (Fig. 2).

[0102] FCGR1A and MPO marker discovery

[0103] Using KEGG pathway analysis, 87 pathways for 187 DEGs were identified. In this study, we focused specifically on the NET formation pathway due to its association with neutrophil-related genes, particularly CD177, MPO, and DEFA, which are prominently expressed during the ACS stage-specific transformation. Among the 187 DEGs, FCGR1A and MPO were included among transcripts associated with the NET pathway. They were located between IFI27, a representative transcript of the NDC group, and DEFA4, DEFA3, and DEFA1A, representative transcripts of the NA group. Furthermore, MPO binds to CD177, a representative gene for CRC stages, through the protein-protein interaction pathway. Changes in FCGR1A and MPO patterns indicate changes in neutrophils during ACS. FCGR1A, present in the neutrophil membrane, was activated when transitioning from the HC group to the NDC group. However, MPO, present in neutrophil granules, was stabilized. FCGR1A and MPO were found to be activated when transitioning from the NDC group to the NA group. Furthermore, transitioning from the NA group to the AA group stabilizes FCGR1A and activates MPO, whereas transitioning from AA to CRC reactivates FCGR1A and stabilizes MPO. This pattern of changes in NET-associated genes across the ACS stages highlights their association with immune processes, particularly NETs (187 DEGs).

[0104] Based on the ACS model, three GO analyses and three network analyses of circulating transcripts from the HC group to the CRC group were combined to reveal that these transcripts were related to the immune response, particularly neutrophils. Therefore, NET formation was selected among the biological mechanisms related to neutrophils, and the pathways of FCGR1 and MPO could be determined as representative transcripts. For clinical validation of FCGR1A and MPO, RT-qPCR was performed using 20 AA, 20 NA, and 20 HC groups. FCGR1A was upregulated in the NA and AA groups compared to the HC group, and was statistically significant in the AA group, but there was no statistical difference in the NA group. MPO was upregulated in both the NA and AA groups and was statistically significant.

[0105] Example 4: Validation of markers associated with colorectal cancer and advanced adenoma.

[0106] (1) Exploration of mRNA markers associated with colon cancer and advanced adenoma

[0107] In the present invention, in order to explore circulating transcript candidates related to CRC or AA in blood, five markers capable of distinguishing CRC or AA were selected from circulating transcripts reported in the existing literature (Nichita C, Ciarloni L, Monnier-Benoit S, Hosseinian S, Dorta G, Regg C. A novel gene expression signature in peripheral blood mononuclear cells for early detection of colorectal cancer. Aliment Pharmacol Ther. 2014 Mar;39(5):507-17).

[0108] Among the genes reported in the literature, the genes with the highest distinguishing power between AP and CRC were selected and validated using clinical samples to determine the final colorectal cancer and advanced adenoma-associated markers.

[0109] (2) qPCR data generation

[0110] qPCR was performed to generate expression values ​​for 38 selected mRNA markers using 1,375 clinical specimens classified by tissue examination results (Table 3).

[0111] Classification Gastroenterology Department (persons) Health Screening Center (persons) Number of specimens by classification Colon cancer 1630163 Advanced adenoma group 3068314 Non-advanced adenoma group 256107363 Control group 305230535 Total 10303451375

[0112] Table 3 shows the number of specimens by classification in each institution.

[0113] The process of generating mRNA marker expression levels through qPCR was as follows.

[0114] 1) Blood sample collection

[0115] Blood is Tempus containing RNA stabilizing solution TM 3 mL of blood was collected using a Blood RNA Tube (ThermoFisher, USA).

[0116] 2) RNA isolation and cDNA synthesis

[0117] RNA isolation was performed using Tempus TM RNA was isolated according to the instructions of the Spin RNA Isolation Kit (ThermoFisher, USA). The isolated RNA was synthesized into cDNA according to the instructions of the High-Capacity cDNA Reverse Transcription Kit (ThermoFisher, USA) and used as the final test sample.

[0118] 3) qPCR reaction

[0119] Using the synthesized cDNA, 38 mRNA markers and reference markers were selected and subjected to qPCR according to the qPCR reaction conditions in (Table 4) to generate Cq values.

[0120] Step Temperature TimeCycleUNG-Incubation50°C2 min1 cyclePre-Denaturation95°C20 sec1 cycleDenaturation95°C1 sec40 cyclesAnnealing& Extension60°C20 sec

[0121] Table 4 shows the qPCR conditions.

[0122] 4) Calculating relative expression level

[0123] The Cq value of each generated mRNA marker was calculated as a relative expression level using the Cq value of the reference marker according to the following formula.

[0124] i. Calculation of relative expression level: 2 -ΔCq

[0125] ii.ΔCq = (Cq of circulating marker) - (Cq of reference marker)

[0126] (3) Development of diagnostic algorithms

[0127] 1) Feature Selection & Development of a 21-Marker Based Diagnostic Algorithm

[0128] Gene expression patterns of 38 markers selected from 1,375 samples classified by colonoscopy examination and histological data were used to select 23 markers that distinguish colon cancer and precancerous lesions as benign using an artificial intelligence algorithm (Principal Component Analysis, Decision Tree, Logistic Regression, Neural Network, Support Vector Machine). Through training using 673 clinical samples and validation using 289 clinical samples, ONCOCHECK TM We developed an IgxAI model for colorectal cancer. The AI ​​algorithm uses Monte Carlo cross-validation to prevent overfitting on specific datasets and generate generalized models, thereby increasing accuracy.

[0129] 2) Characteristics of the final selected biomarkers for ONCOCHECK™ Colorectal Cancer

[0130] Twenty-one biomarkers were selected for their correlation with circulating colorectal cancer cells and immune responses associated with precancerous and cancerous states. They encompass a diverse genetic signature associated with circulating tumor cells (CTCs), including epithelial cell-specific genes, genes associated with proliferation and immortalization, and genes important for the epithelial-to-mesenchymal transition (EMT). Furthermore, genes involved in adenocarcinoma sequences and genes expressed during colorectal polyp formation, such as those associated with the innate immune response (IIR) and neutrophil extracellular traps (NETs), are included (Table 5).

[0131] Order Biomarker Gene Name Marker Characteristics 1CCR1C-C Motif Chemokine Receptor1IIR2CD177CD177 MoleculeIIR3DLG5Discs Large MAGUK Scaffold Protein 5CTC4EPCAMEpithelial Cell Adhesion MoleculeCTC5ERBB2Erb-B2 Receptor Tyrosine Kinase 2CTC6FCGR1AFc Gamma Receptor laCTC7FOXA2Forkhead Box A2CTC8GKGlycerol KinaseIIR9KRT19Keratin 19CTC10KRT73Keratin 73IIR1118S rRNA18S rRNAInternal control12MKI67Marker Of Proliferation Ki-67CTC13MPOMyeloperoxidaseIIR14MUC1Mucin 1CTC15NME1NME / NM23 Nucleoside Diphosphate Kinase 1CTC16NPTNNeuroplastinCTC17NQO2N-Ribosyldihydronicotinamide: Quinone Reductase 2IIR18SH2D1BSH2 Domain Containing 1BIIR19SNAI2Snail Family Transcriptional Repressor 2CTC20TERTTelomerase Reverse TranscriptaseCTC21TUG1Taurine Up-Regulated 1CTC22VIMVimentinCTC23GAPDHGlyceraldehyde-3-Phosphate DehydrogenaseReference Gene

[0132] Table 5 shows the final 21 biomarker genes selected by the present invention, as well as the internal control (18S rRNA) and reference gene (GAPDH).

[0133] Table 6 shows the primer and probe mix base sequences for PCR for the biomarker genes and internal control (18S rRNA) and reference gene (GAPDH) of Table 5 above.

[0134] 표적유전자서열번호염기서열(5‘ -> 3’)1CCR1Forward1CGG GAT GGA AAC TCC AAA CAProbe2FAM-ACA GAG GAC TAT GAC ACG AC-MGBReverse3GCA TCC CCA TAG TCA AAC TCT GT2CD177Forward4GAC GCT GCT GCT CCT AGA TGTProbe5FAM-ACT CAC ATC AAC CCT G-MGBReverse6CAG TGC TGC AGC CTT TTG TC3DLG5Forward7ACA GGC TGA ATC CTG ACT ATG AGAProbe8FAM-CTG AAG ATC CAG TGC GTG-MGBReverse9GCT CTG CAG GTC CGA CAT G4EPCAMForward10GTG TAC TTC AGT TGG TGC ACA AAA TProbe11FAM-TTG CTC AAA GCT GGC TGC-MGBReverse12CAT TTC TGC CTT CAT CAC CAA ACA T5ERBB2Forward13TGG GCA TCT GCC TGA CAT CProbe14FAM-ACG GTG CAG CTG GT-MGBReverse15AGC CAT AGG GCA TAA GCT GTG T6FCGR1AForward16GCC CTG AGT TGG AGC TTC AAProbe17FAM-TGC TTG GCC TCC AGT TA-MGBReverse18AAC TCC TGT CTG GTT TCA TGT CC7FOXA2Forward19GGA TGA ACG GCA TGA ACA CGT AProbe20FAM-ATG AGC ATG TCG GCG GC-MGBReverse21GCC CGC GCT CAT GTT G8GKForward22AGA AGG AGT CGG CGT ATG GAProbe23FAM-TCT CGA ACC CGA GGA T-MGBReverse24CCA TCG TGA CGG CAG ACA9KRT19Forward25GTC TTG AGA TTG AGC TGC AGT CAProbe26FAM-CTGAGC ATG AAA GCT G-MGBReverse27TTT CTG CCA GTG TGT CTT CCA A10KRT73Forward28TGC GCG AGT ACC AAG AGC TTProbe29FAM-TGA GCG TGA AGC TGT-MGBReverse30AGG TGG CGA TCT CAA TAT CCA1118S rRNAForward31GCA GCC GCG GTA ATT CCProbe32FAM-CTC CAA TAG CGT ATA TTA A-MGBReverse33ACG AGC TTT TTA ACT GCA GCA A12MKI67Forward34GTC CTT TGG TGG GCA CCT AProbe35FAM-CTG AAC TAT TTG ATG AAA ACT-MGBReverse36CCC TTT TGA GAG GCG TAT TAG GA13MPOForward37GCC CGG AGC AGG ACA AAT AProbe38FAM-CGC ACC ATC ACC G-MGBReverse39GAT GTG CAA CAA CAG ACG CAG14MUC1Forward40TCT TTC CAG CCC GGG ATA CProbe41FAM-CCA TCC TAT GAG CGA GTA C-MGBReverse42GCC CAT GGG TGT GGT AGG T15NME1Forward43GAC CTG AAG GAC CGT CCA TTCProbe44FAM-TTG CCG GCC TGG TG-MGBReverse45CGG CCC TGA GTG CAT GTA T16NPTNForward46CCT GTC ACC CTG CAG TGT AAProbe47FAM-CTC ACC TCC AGC TCT C-MGBReverse48CCC ATT CTT TGT CCA GTA GCT GTA T17NQO2Forward49GCC ACA GAC AAA GAT ATC ACT GGT AProbe50FAM-TCT TTC TAA TCC TGA GGT TTT-MGBReverse51CGT GGG TTT CCA CTC CAT AATT18SH2D1BForward52GGA CGT CTG ACC AAG CAA GACProbe53FAM-TGA GAC CTT GCT GCT C-MGBReverse54TGC CAT CCA CCC CTT CCT19SNAI2Forward55TGG GCG CCC TGA AGA TGProbe56FAM-ATA TTC GGA CCC ACA CAT-MGBReverse57CCG CAG ATC TTG CAA ACA CAA20TERTForward58CGT CCA GAC TCC GCT TCA TCProbe59FAM-CTG CGG CCG ATT GTG-MGBReverse60CCC ACG ACG TAG TCC ATG TT21TUG1Forward61GGA GTC CCC TTA CCT AAC AGC ATProbe62FAM-CAA GGC TGC ACC AGA T-MGBReverse63TGA TGT ATC AAG AGA AGC CTT TTC TG22VIMForward64CCT TGA ACG CAA AGT GGA ATC TTProbe65FAM-CAC GAA GAG GAA ATC CAG-MGBReverse66GGA CAT GCT GTT CCT GAA TCT GA23GAPDHForward67GGA AGC TTG TCA TCA ATG GAA ATC CProbe68FAM-CAG GAG CGA GAT CCC T-MGBReverse69CAT CGC CCC ACT TGA TTT TGG

[0135] Table 6 shows the primer and probe mix base sequences for PCR for the biomarker genes and internal control (18S rRNA) and reference gene (GAPDH) of Table 5 above.

[0136] (4) Diagnostic model performance evaluation

[0137] 1) Performance of ONCOCHECK™Colorectal Cancer

[0138] Using the developed IgxAI model, 413 test clinical samples were tested, and the sensitivity for detecting colorectal cancer was 93% (45 / 48), the sensitivity for detecting advanced adenoma was 91% (87 / 95), and the specificity was 91% (147 / 161).

[0139] Colonoscopy and Histopathology Test ResultsCRCAANANegative FindingsTotal Diagnostic Kit of the Present Invention (ONCOCHECK™Colorectal Cancer)Positive(+)45879714243Negative(-)3812147170Total4895109161413

[0140] Table 7 is a table showing the test results using the diagnostic kit of the present invention.

[0141] In particular, it was confirmed to have a sensitivity of 96% for early-stage colon cancer (stages 1-2) and a specificity of 93% for asymptomatic and symptomatic patients with negative colonoscopy findings (Table 8).

[0142] Early colorectal cancer (stages 1-2) with negative colonoscopy findings. Diagnostic kit for this invention (ONCOCHECK™ Colorectal Cancer). Positive (+) 269 Negative (-) 1118 Total 27127

[0143] Table 8 shows the results of confirming a sensitivity of 96% for early-stage colon cancer (stages 1-2) and a specificity of 93% for asymptomatic and symptomatic patients with negative colonoscopy findings.

[0144] From the table, it can be seen that 96% (26 / 27) had early colorectal cancer (stages 1-2) and 93% (118 / 127) were asymptomatic or symptomatic with negative colonoscopy findings.

Claims

1. A substance capable of measuring the relative expression level of CCR1, CD177, DLG5, EPCAM, ERBB2, FCGR1A, FOXA2, GK, KRT19, KRT73, 18S rRNA, MKI67, MPO, MUC1, NME1, NPTN, NQO2, SH2D1B, SNAI2, TERT, TUG1, VIM and GAPDH genes or proteins encoded by the genes, Here, the relative expression level is 2 -ΔCq It is indicated as 2 -ΔCq can be obtained by the following relationship: [Relationship] 2 -ΔCq = 2-(target gene Cq - GAPDH gene Cq) Here, the target genes are CCR1, CD177, DLG5, EPCAM, ERBB2, FCGR1A, FOXA2, GK, KRT19, KRT73, MKI67, MPO, MUC1, NME1, NPTN, NQO2, SH2D1B, SNAI2, TERT, TUG1, VIM, 18S rRNA is an internal control, GAPDH is a reference gene, and the Cq value is a composition for diagnosing colon cancer, characterized in that a certain fluorescence value is set as a threshold and the cycle number at the time of reaching this threshold is indicated.

2. A composition for diagnosing colon cancer in the first paragraph, wherein the substance capable of measuring the relative expression level of the gene or the relative expression level of the protein encoded by the gene is a primer and probe set or an antibody.

3. A composition for diagnosing colon cancer, characterized in that in the second paragraph, the primer and probe sequences are described in SEQ ID NOs: 1 to 69.

4. A kit for diagnosing colon cancer comprising a composition according to any one of claims 1 to 3.

5. A method for providing information for predicting the risk of developing colon cancer, characterized in that the risk of developing colon cancer is predicted to be high when the relative expression of the CD177, FOXA2, SNAI2, TERT, DLG5, FCGR1A, ERBB2, FCGR1A, FOXA2, GK, KRT19, KRT73, MKI67, MPO, MUC1, NME1, NPTN, NQO2, SH2D1B, SNAI2, TERT, TUG1 and VIM genes or proteins encoded by the genes is measured for a genetic sample obtained from a subject, and the relative expression of the CD177, FOXA2, SNAI2, TERT, DLG5, FCGR1A, NPTN and NQO2 genes or proteins is increased or the relative expression of the MKI67, KRT73, EPCAM and NME1 genes or proteins is decreased.

Citation Information

Patent Citations

  • A method for sorting colorectal cancer and advanced neoplasia and use of the same

    KR102548873B1

  • A method for sorting colon polyp and colorectal cancer and use of the same

    KR102591596B1

  • Colon cancer diagnostic method and means

    WO2014041185A2