A gene combination biomarker for colorectal cancer diagnosis and its application
By detecting gene combination markers such as TIPARP in blood cell RNA, combining high-throughput sequencing and PCR methods, a machine learning diagnostic model was established, which solved the problem of low diagnostic sensitivity of early colorectal cancer and achieved efficient and low-cost early colorectal cancer diagnosis.
Patent Information
- Application Number
- CN202410846063.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-27
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-06-27
AI Technical Summary
The prior art has low sensitivity in the diagnosis of early stage colorectal cancer, especially the detection sensitivity of stage 0 and stage I tumors is less than 50%. Common methods are susceptible to bacterial contamination and sample storage time, and cannot meet clinical needs.
Combined markers of TIPARP, PLVAP, GZMB, B3GNT7, RPL13, CCR1, PDK2, IDI1, S1PR5, SH2D2A, SH3TC1 and PNP genes were used to detect the RNA expression level of blood cells by high-throughput sequencing or PCR, and a diagnostic model was established in combination with machine learning algorithms.
It has achieved a high sensitivity and specific diagnosis of early-stage colorectal cancer, which is low in cost, only a small number of peripheral blood samples, wide applicability, and high population compliance, which significantly improves the diagnostic efficiency of early-stage colorectal cancer.
Smart Images

Figure CN118581222B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of biomedicine, and particularly relates to a gene combination biomarker for colorectal cancer diagnosis and its application. Background Art
[0002] In China, the incidence and mortality of colorectal cancer rank among the top five in the tumor spectrum, and the incidence shows an increasing trend year by year. The occurrence and development of colorectal cancer usually require a slow process of more than ten years. Detecting and treating colorectal cancer in the early stage can significantly increase the survival chance of patients, reduce the treatment cost, and relieve the pain and burden of patients. The commonly used tumor markers and fecal occult blood test have low sensitivity in diagnosing colorectal cancer and cannot meet the clinical needs. Detecting circulating free nucleic acids (such as cfDNA, microRNA, etc.) in the blood of patients by means of high-throughput sequencing, polymerase chain reaction, etc. is the most widely studied early diagnosis method for colorectal cancer. However, due to the small size of early colorectal cancer and the small number of tumor cells, the copy number of circulating nucleic acids released is extremely low, resulting in low detection sensitivity of these methods for early tumors (especially stage 0 and stage I tumors), even less than 50%. In addition, another widely studied fecal ribonucleic acid methylation method is also easily interfered by factors such as bacterial contamination and specimen storage time. Therefore, there is an urgent need to develop new tumor biomarkers to improve the diagnostic efficiency of early colorectal cancer.
[0003] Colorectal cancer is a systemic disease that not only affects the tumor microenvironment but also affects the immune system and other distal tissues and organs. The systemic effects induced by colorectal cancer can produce detectable tumor-related signals in other tissues and organs such as the blood system, which can be used as biomarkers for tumor diagnosis. For example, colorectal cancer affects various cells in the blood, causing changes in the ribonucleic acid (RNA) expression in white blood cells or platelets. Currently, RNA derived from white blood cells or platelets has been used to detect various cancers, and multiple studies have shown that the nucleic acids in these blood cells are tumor liquid biopsy materials with great application prospects. Summary of the Invention
[0004] One object of the present invention is to provide a gene combination biomarker for colorectal cancer diagnosis, and the gene combination biomarker is composed of the following genes: TIPARP, PLVAP, GZMB, B3GNT7, RPL13, CCR1, PDK2, IDI1, S1PR5, SH2D2A, SH3TC1, and PNP.
[0005] Another object of the present invention is to provide a product for colorectal cancer diagnosis, and the product contains primers for detecting the expression levels of the genes TIPARP, PLVAP, GZMB, B3GNT7, RPL13, CCR1, PDK2, IDI1, S1PR5, SH2D2A, SH3TC1, and PNP.
[0006] Preferably, the product contains primers for detecting the expression levels of TIPARP, PLVAP, GZMB, B3GNT7, RPL13, CCR1, PDK2, IDI1, S1PR5, SH2D2A, SH3TC1, and PNP genes, and their specific sequences are as follows:
[0007] The primers for detecting the expression level of TIPARP are shown in SEQ ID NO.1-2;
[0008] The primers for detecting the expression level of PLVAP are shown in SEQ ID NO.3-4;
[0009] The primers for detecting the expression level of GZMB are shown in SEQ ID NO.5-6;
[0010] The primers for detecting the expression level of B3GNT7 are shown in SEQ ID NO.7-8;
[0011] The primers for detecting the expression level of RPL13 are shown in SEQ ID NO.9-10;
[0012] The primers for detecting the expression level of CCR1 are shown in SEQ ID NO.11-12;
[0013] The primers for detecting the expression level of PDK2 are shown in SEQ ID NO.13-14;
[0014] The primers for detecting the expression level of IDI1 are shown in SEQ ID NO.15-16;
[0015] The primers for detecting the expression level of S1PR5 are shown in SEQ ID NO.17-18;
[0016] The primers for detecting the expression level of SH2D2A are shown in SEQ ID NO.19-20;
[0017] The primers for detecting the expression level of SH3TC1 are shown in SEQ ID NO.21-22;
[0018] The primers for detecting the expression level of PNP are shown in SEQ ID NO.23-24.
[0019] More preferably, the product is a kit.
[0020] The third object of the present invention is to provide the use of a reagent for detecting the expression levels of TIPARP, PLVAP, GZMB, B3GNT7, RPL13, CCR1, PDK2, IDI1, S1PR5, SH2D2A, SH3TC1 and PNP genes in the preparation of a colorectal cancer diagnostic product.
[0021] Preferably, the product includes a kit, a gene chip or a high-throughput sequencing system.
[0022] More preferably, the kit contains primers for detecting the expression levels of TIPARP, PLVAP, GZMB, B3GNT7, RPL13, CCR1, PDK2, IDI1, S1PR5, SH2D2A, SH3TC1 and PNP genes.
[0023] More preferably, the primer sequences for detecting the expression levels of TIPARP, PLVAP, GZMB, B3GNT7, RPL13, CCR1, PDK2, IDI1, S1PR5, SH2D2A, SH3TC1 and PNP genes are as follows:
[0024] The primers for detecting the expression level of TIPARP are shown as SEQ ID NO.1-2;
[0025] The primers for detecting the expression level of PLVAP are shown as SEQ ID NO.3-4;
[0026] The primers for detecting the expression level of GZMB are shown as SEQ ID NO.5-6;
[0027] The primers for detecting the expression level of B3GNT7 are shown as SEQ ID NO.7-8;
[0028] The primers for detecting the expression level of RPL13 are shown as SEQ ID NO.9-10;
[0029] The primers for detecting the expression level of CCR1 are shown as SEQ ID NO.11-12;
[0030] The primers for detecting the expression level of PDK2 are shown as SEQ ID NO.13-14;
[0031] The primers for detecting the expression level of IDI1 are shown as SEQ ID NO.15-16;
[0032] The primers for detecting the expression level of S1PR5 are shown as SEQ ID NO.17-18;
[0033] The primers for detecting the expression level of SH2D2A are shown as SEQ ID NO.19-20;
[0034] The primers for detecting the expression level of SH3TC1 are shown in SEQ ID NO.21-22;
[0035] The primers for detecting the expression level of PNP are shown in SEQ ID NO.23-24.
[0036] More preferably, the method for using the kit is as follows:
[0037] 1) Obtain a sample from a subject;
[0038] 2) Determine the expression level of the gene in the sample by primer amplification.
[0039] The present invention also provides a method for diagnosing tumors using the aforementioned gene combination markers, including the following steps:
[0040] Step 1: Collect a blood sample from a subject;
[0041] Step 2: Measure the expression level of a set of gene combinations in the sample, and the gene combination is TIPARP, PLVAP, GZMB, B3GNT7, RPL13, CCR1, PDK2, IDI1, S1PR5, SH2D2A, SH3TC1, PNP genes;
[0042] Step 3: a) Calculate the relative expression level of each gene using a gene expression level normalization algorithm, or b) Compare with the expression levels of internal reference genes GAPDH and ACTB to calculate the relative expression level of each gene;
[0043] Step 4: Calculate a diagnostic score for the relative expression level of each gene through an algorithm and compare it with a pre-defined critical value. If the diagnostic score is higher than the critical value, the subject has colorectal cancer; if the diagnostic score is lower than the critical value, the subject does not have colorectal cancer.
[0044] Experimental results show that the present invention detects the expression level of a gene combination in blood cell RNA for colorectal cancer detection, with high sensitivity and specificity. The present invention is a non-invasive liquid biopsy method, with a low blood collection volume, and only a small amount of peripheral blood sample is required to complete the detection, and the population compliance is high. The present invention has important application value.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] (1) The detected gene combination of the present invention is obtained by analyzing high-throughput sequencing data of actual tumor cases and healthy human blood cells from Union Hospital, Tongji Medical College, Huazhong University of Science and Technology through a specific machine learning algorithm, and the data has high reliability and credibility, and can accurately diagnose colorectal cancer.
[0047] (2) The gene combination of the present invention has a high diagnostic sensitivity for early colorectal cancer.
[0048] (3) Compared with methods such as circulating tumor DNA (cfDNA), the present technology has low cost and only requires ordinary transcriptome high-throughput sequencing or PCR methods, which can significantly reduce costs and has wide universality. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a comparison chart of the expression levels of TIPARP, PLVAP, GZMB, B3GNT7, RPL13, CCR1, PDK2, IDI1, S1PR5, SH2D2A, SH3TC1, and PNP genes in colorectal cancer patients and healthy people. The P value of the statistical analysis is calculated by two-tailed T test or two-tailed Mann-Whitney test. A P value less than 0.05 is considered to have statistical significance.
[0050] Figure 2 It is the predicted values of tumor patients and healthy populations obtained by the support vector machine algorithm in Example 3.
[0051] Figure 3 It is the ROC curve graph of the 12-gene combination marker composed of TIPARP, PLVAP, GZMB, B3GNT7, RPL13, CCR1, PDK2, IDI1, S1PR5, SH2D2A, SH3TC1, and PNP for diagnosing colorectal cancer patients at each stage; the left graph is for all stages, the middle graph is for early stage (stage 0 and stage I); the right graph is for middle and late stage (stage II-IV). DETAILED DESCRIPTION OF THE INVENTION
[0052] Example 1
[0053] Selection of gene combination markers for colorectal cancer screening and diagnosis
[0054] 1. Samples and data
[0055] The inventors mainly used the high-throughput transcriptome sequencing database of peripheral blood blood cells of the Cancer Institute of Union Hospital, Tongji Medical College, Huazhong University of Science and Technology for screening to find gene combination markers that can be used for colorectal cancer diagnosis.
[0056] 2. Data standardization
[0057] Under the Linux system environment of the workstation, the transcriptome sequencing data was aligned to the human reference genome GRCh37 / hg19 using the alignment software STAR, and the number of reads aligned to each gene was calculated using the quantMode-GeneCounts command. Subsequently, using the "DESeq2" toolkit in R language, the number of reads aligned to each gene was normalized using the "vst" command to obtain the normalized gene expression matrix.
[0058] 3. Calculate the contribution degree of each gene to differentiating colorectal cancer patients from healthy people
[0059] In R language, using the "e1071" package and the recursive feature elimination algorithm, calculate the contribution degree of each gene to differentiating colorectal cancer patients from healthy people. The 12 genes in the gene combination marker described in this patent are the top 12 genes, and their contribution degrees are shown in Table 1. Among them, the contribution degree ranking is determined according to the score given by the feature elimination algorithm. The lower the score, the higher the ranking.
[0060] 4. Differential analysis
[0061] The differential expression of TIPARP, PLVAP, GZMB, B3GNT7, RPL13, CCR1, PDK2, IDI1, S1PR5, SH2D2A, SH3TC1, and PNP in colorectal cancer patients and healthy people is shown in Figure 1 , among which TIPARP, PLVAP, GZMB, B3GNT7, RPL13, CCR1, PDK2, S1PR5, SH2D2A, SH3TC1, and PNP have statistical differences.
[0062] Table 1 Ranking of the contribution degree of each gene to differentiating tumor patients from healthy people
[0063] Rank Gene Name Score 1 TIPARP 4.3 2 PLVAP 9.1 3 GZMB 33.4 4 B3GNT7 35.6 5 RPL13 39.5 6 CCR1 50.3 7 PDK2 53.2 8 IDI1 55.9 9 S1PR5 57.1 10 SH2D2A 70.3 11 SH3TC1 71.7 12 PNP 79.6
[0064] TIPARP is numbered ENSG00000163659 in the Ensembl database (version GRCh37), PLVAP is numbered ENSG00000130300 in the Ensembl database (version GRCh37), GZMB is numbered ENSG00000100453 in the Ensembl database (version GRCh37), B3GNT7 is numbered ENSG00000156966 in the Ensembl database (version GRCh37), RPL13 is numbered ENSG00000167526 in the Ensembl database (version GRCh37), CCR1 is numbered ENSG00000163823 in the Ensembl database (version GRCh37), PDK2 is numbered ENSG00000005882 in the Ensembl database (version GRCh37), IDI1 is numbered ENSG00000067064 in the Ensembl database (version GRCh37), S1PR5 is numbered ENSG00000180739 in the Ensembl database (version GRCh37), SH2D2A is numbered ENSG00000027869 in the Ensembl database (version GRCh37), SH3TC1 is numbered ENSG00000125089 in the Ensembl database (version GRCh37), and PNP is numbered ENSG00000198805 in the Ensembl database (version GRCh37).
[0065] The gene information of TIPARP, PLVAP, GZMB, B3GNT7, RPL13, CCR1, PDK2, IDI1, S1PR5, SH2D2A, SH3TC1, and PNP involved in the present invention is shown in Table 2.
[0066] Table 2 Gene Information
[0067]
[0068] As can be seen from Table 1, TIPARP, B3GNT7, RPL13, PDK2, SH3TC1, and PNP are located on the positive strand of the chromosome, while PLVAP, GZMB, CCR1, IDI1, S1PR5, and SH2D2A are located on the negative strand of the chromosome.
[0069] Example 2
[0070] According to the gene combination markers screened in Example 1, based on the sequence information of the genes in the NCBI database, the following primers were designed for the detection of gene expression levels.
[0071] Primers for Detecting Gene Expression Levels in Table 3
[0072]
[0073]
[0074] Among them, "Gene Name - F" and "Gene Name - R" respectively represent the upstream primer and downstream primer for detecting the messenger RNA of this gene. GAPDH and ACTB are internal reference genes.
[0075] Example 3
[0076] Diagnostic Performance of a Gene - Combination Biomarker Composed of TIPARP, PLVAP, GZMB, B3GNT7, RPL13, CCR1, PDK2, IDI1, S1PR5, SH2D2A, SH3TC1, and PNP for Colorectal Cancer
[0077] 1. Samples and Data
[0078] Use a vacuum blood collection tube to collect 1 - 5 ml of blood specimens from colorectal cancer patients and healthy controls. The blood samples can be whole blood anticoagulated with EDTAK2, sodium citrate, or heparin; the healthy controls are healthy individuals determined not to have tumors. In this example, whole blood samples anticoagulated with EDTAK2 are selected, RNA is extracted by conventional methods, and then a transcriptome sequencing library is constructed by conventional methods to obtain the expression level data of each gene in the samples.
[0079] 2. Obtaining Normalized Expression Levels
[0080] Use the method in Example 1. After aligning and normalizing the sequencing data, obtain the normalized expression level matrix of each gene.
[0081] 3. Establishing a Colorectal Cancer Diagnostic Model
[0082] Randomly select 50% of all samples (888 cases) as the training set for establishing a colorectal cancer diagnostic model. Using the normalized expression levels of TIPARP, PLVAP, GZMB, B3GNT7, RPL13, CCR1, PDK2, IDI1, S1PR5, SH2D2A, SH3TC1, and PNP as features, use the "e1071" software package in R language and adopt the support vector machine (SVM) algorithm to establish a binary classification model for distinguishing colorectal cancer patients from healthy controls. When establishing the model, use the "tune.svm" command in the "e1071" package to select the best model parameters.
[0083] 4. ROC Curve Analysis
[0084] Use the remaining 50% of all samples as the test set (444 cases, including 220 colorectal cancer patients and 224 healthy controls). The performance of the diagnostic model for diagnosing colorectal cancer was evaluated using the area under the receiver operating characteristic (ROC) curve (AUC), sensitivity, and specificity. The prediction of the test set samples by the colorectal cancer diagnostic model and the acquisition of the predicted values were carried out using the "predict" command in the "e1071" package in R software. As Figure 2 shown, there were significant differences in the predicted values between colorectal cancer patients and healthy controls. Subsequently, the ROC curve of the subjects was plotted using the "ROCR" package in R software, and then the AUC value, sensitivity, and specificity were obtained to evaluate the diagnostic performance of the gene combination markers for colorectal cancer. The performance of the gene combination markers in diagnosing all 220 cases of colorectal cancer (36 cases in the early stage (0-I stage), 159 cases in the middle and advanced stage (II-IV stage), and 25 cases with unclear staging) can be seen in Table 4 and Figure 3 .
[0085] Table 4 Analysis of the diagnostic performance of gene combination markers for colorectal cancer (including some patients without pathological staging in "all")
[0086] AUC Sensitivity Specificity All 0.917 79.82% 91.15% Stage 0 - I 0.936 86.11% 86.61% Stages II - IV 0.923 81.76% 91.96%
[0087] In summary, the gene combination markers based on the present invention can effectively diagnose colorectal cancer, especially early colorectal cancer.
[0088] The above-described embodiments are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A gene combination marker for colorectal cancer diagnosis, characterized in that, The gene combination biomarker consists of the following genes: TIPARP, PLVAP, GZMB, B3GNT7, RPL13, CCR1, PDK2, IDI1, S1PR5, SH2D2A, SH3TC1, and PNP.
2. A product for colorectal cancer diagnosis, characterized in that, The product contains primers for detecting the expression levels of the genes TIPARP, PLVAP, GZMB, B3GNT7, RPL13, CCR1, PDK2, IDI1, S1PR5, SH2D2A, SH3TC1, and PNP.
3. The product according to claim 2, wherein, The product contains primers for detecting the expression levels of the genes TIPARP, PLVAP, GZMB, B3GNT7, RPL13, CCR1, PDK2, IDI1, S1PR5, SH2D2A, SH3TC1, and PNP, and the specific sequences are as follows: The primers for detecting the expression level of TIPARP are shown as SEQ ID NO.1-2; The primers for detecting the expression level of PLVAP are shown as SEQ ID NO.3-4; The primers for detecting the expression level of GZMB are shown as SEQ ID NO.5-6; The primers for detecting the expression level of B3GNT7 are shown as SEQ ID NO.7-8; The primers for detecting the expression level of RPL13 are shown as SEQ ID NO.9-10; The primers for detecting the expression level of CCR1 are shown as SEQ ID NO.11-12; The primers for detecting the expression level of PDK2 are shown as SEQ ID NO.13-14; The primers for detecting the expression level of IDI1 are shown as SEQ ID NO.15-16; The primers for detecting the expression level of S1PR5 are shown as SEQ ID NO.17-18; The primers for detecting the expression level of SH2D2A are shown as SEQ ID NO.19-20; The primers for detecting the expression level of SH3TC1 are shown as SEQ ID NO.21-22; The primers for detecting the expression level of PNP are shown as SEQ ID NO.23-24.
4. The product according to claim 2 or 3, characterized in that, The product is a kit.
5. Use of a reagent for detecting the expression levels of TIPARP, PLVAP, GZMB, B3GNT7, RPL13, CCR1, PDK2, IDI1, S1PR5, SH2D2A, SH3TC1 and PNP genes in the preparation of a colorectal cancer diagnostic product, characterized in that, The reagent contains primers for detecting the expression levels of the genes TIPARP, PLVAP, GZMB, B3GNT7, RPL13, CCR1, PDK2, IDI1, S1PR5, SH2D2A, SH3TC1, and PNP.
6. The application according to claim 5, wherein The product is a kit.
7. The application according to claim 6, wherein The ones for detecting TIPARP, PLVAP, GZMB, The primer sequences for detecting the expression levels of B3GNT7, RPL13, CCR1, PDK2, IDI1, S1PR5, SH2D2A, SH3TC1, and PNP are as follows: The primers for detecting the expression level of TIPARP are shown as SEQ ID NO.1-2; The primers for detecting the expression level of PLVAP are shown as SEQ ID NO.3-4; The primers for detecting the expression level of GZMB are shown as SEQ ID NO.5-6; The primers for detecting the expression level of B3GNT7 are shown as SEQ ID NO.7-8; The primers for detecting the expression level of RPL13 are shown in SEQ ID NO. 9-10; The primers for detecting the expression level of CCR1 are shown in SEQ ID NO. 11-12; The primers for detecting the expression level of PDK2 are shown in SEQ ID NO. 13-14; The primers for detecting the expression level of IDI1 are shown in SEQ ID NO. 15-16; The primers for detecting the expression level of S1PR5 are shown in SEQ ID NO. 17-18; The primers for detecting the expression level of SH2D2A are shown in SEQ ID NO. 19-20; The primers for detecting the expression level of SH3TC1 are shown in SEQ ID NO. 21-22; The primers for detecting the expression level of PNP are shown in SEQ ID NO. 23-24.
8. The application according to claim 6 or claim 7, characterized in that, The usage method of the kit is as follows: 1) Obtain a sample from the subject; 2) Determine the expression level of the gene in the sample by primer amplification.
Citation Information
Patent Citations
Molecular marker for diagnosing colorectal cancer
CN118166095A