Application of gene detection reagent in preparation of colorectal cancer screening kit and colorectal cancer screening kit, system or computer readable storage medium
By screening for gene combinations that continuously change in typical cancerous sequences of colorectal cancer and combining them with machine learning models, a highly efficient early screening method for colorectal cancer has been achieved. This solves the problems of insufficient sensitivity and specificity in existing technologies and provides a non-invasive and highly accurate diagnostic method.
Patent Information
- Application Number
- CN202511356610.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2026-01-30
Smart Images

Figure CN121428094A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biological detection technology, specifically relating to the use of reagents for detecting genes in the preparation of colorectal cancer screening kits, as well as kits, systems, or computer-readable storage media for screening colorectal cancer. Background Technology
[0002] Colorectal cancer (CRC) is one of the most common malignant tumors worldwide, ranking among the top in both incidence and mortality, posing a serious threat to human health. The typical development of CRC follows a "normal mucosa-adenoma-cancer" (NAC) sequence, accounting for approximately 85%-90% of all cases. This multi-stage carcinogenesis process usually takes more than 10 years. If detection and intervention can be achieved at the adenoma or even early cancer stage, the 5-year survival rate of patients can be significantly improved. Therefore, early diagnosis and universal screening of CRC are widely recognized as key measures to reduce its mortality rate.
[0003] Currently, early diagnosis of colorectal cancer (CRC) mainly relies on several methods: colonoscopy, stool testing, serological marker testing, and the rapidly developing liquid biopsy technique. Colonoscopy is considered the gold standard, but its invasiveness and poor patient compliance make it difficult to widely implement in large populations. Fecal occult blood test (FOBT), fecal immunochemical test (FIT), and fecal DNA testing have certain non-invasive advantages, but their sensitivity and specificity are limited, especially in the adenoma stage where the detection rate is low. Serological markers such as CEA and CA19-9 are widely used clinically, but they often only show significant increases in the middle and late stages of CRC, lacking early screening value.
[0004] Against this backdrop, liquid biopsy has become an important direction in molecular diagnostic research for CRC. Liquid biopsy utilizes circulating tumor DNA (ctDNA), circulating RNA (mRNA, miRNA, lncRNA), exosomes, and methylation patterns in blood samples to non-invasively obtain information related to tumor development and progression. Some results have already entered clinical applications; for example, commercial products based on SEPT9 gene methylation detection have shown value in early screening, and some miRNAs (such as miR-21 and miR-92a) have also been reported as potential early diagnostic biomarkers.
[0005] While liquid biopsy is non-invasive and promising, current research largely focuses on single indicators (such as ctDNA hotspot mutations or individual miRNAs), lacking stable multi-gene combinations capable of traversing the typical cancerous lesion sequence of CRC (Normal→Adenoma→Cancer). Furthermore, while existing multi-gene prediction models have shown good performance in some cohorts, they often lack external validation across populations and datasets, resulting in insufficient accuracy and generalizability. In addition, most models' diagnostic AUC remains between 0.75 and 0.85, failing to achieve the high accuracy required for clinical application. Finally, existing liquid biopsy biomarker research largely lacks integration with the immune microenvironment and molecular mechanisms, resulting in a lack of biological interpretation and limiting its clinical acceptance.
[0006] Therefore, how to systematically screen out stable biomarkers with continuous changing trends in CRC cancer sequences and combine them with advanced algorithms such as machine learning to establish high-precision prediction models has become a key area for breakthroughs in the field of early CRC screening and diagnosis. Summary of the Invention
[0007] The purpose of this invention is to provide the use of reagents for detecting genes in the preparation of colorectal cancer screening kits, as well as kits, systems, or computer-readable storage media for screening colorectal cancer.
[0008] This invention provides the use of reagents for detecting gene expression levels in the preparation of colorectal cancer screening kits, wherein the genes are one or more combinations of HECW2, WARS1, SLC16A3, SECTM1, IFITM3, ADAMTSL4, FCGR1A, F2RL1, OPLAH, SERPINA1, and FCGR1CP.
[0009] Furthermore, the genes are HECW2, WARS1, SLC16A3, SECTM1, IFITM3, ADAMTSL4, FCGR1A, F2RL1, OPLAH, SERPINA1, and FCGR1CP.
[0010] Furthermore, the reagent is used to detect gene expression levels in human feces, blood, or tissues.
[0011] Preferably, the reagent for detecting gene expression levels is a qPCR detection reagent, a bulk-RNA detection reagent, or a reagent used in single-cell sequencing methods.
[0012] The present invention also provides a colorectal cancer screening kit, which includes reagents for detecting gene expression levels; the gene is one or more of HECW2, WARS1, SLC16A3, SECTM1, IFITM3, ADAMTSL4, FCGR1A, F2RL1, OPLAH, SERPINA1 and FCGR1CP.
[0013] Furthermore, the genes are HECW2, WARS1, SLC16A3, SECTM1, IFITM3, ADAMTSL4, FCGR1A, F2RL1, OPLAH, SERPINA1, and FCGR1CP.
[0014] Furthermore, the reagent is used to detect gene expression levels in human feces, blood, or tissues.
[0015] Preferably, the reagent for detecting gene expression levels is a qPCR detection reagent, a bulk-RNA detection reagent, or a reagent used in single-cell sequencing methods.
[0016] The present invention also provides a colorectal cancer screening system, the system comprising the following parts: Input module: Used to input gene expression levels, wherein the gene is one or more of the following: HECW2, WARS1, SLC16A3, SECTM1, IFITM3, ADAMTSL4, FCGR1A, F2RL1, OPLAH, SERPINA1, and FCGR1CP. Prediction module: used to input the gene expression level as a feature into a machine learning model to obtain colorectal cancer screening results; Output module: Used to output colorectal cancer screening results.
[0017] Furthermore, the genes are HECW2, WARS1, SLC16A3, SECTM1, IFITM3, ADAMTSL4, FCGR1A, F2RL1, OPLAH, SERPINA1, and FCGR1CP; And / or, the machine learning model is a random forest model, vector machine model, logistic regression model, LASSO regression model, neural network model, or XGBoost model.
[0018] Furthermore, the gene expression level refers to the gene expression level in human feces, blood, or tissues.
[0019] The present invention also provides a computer-readable storage medium having a computer program stored thereon for implementing the aforementioned system.
[0020] The present invention has achieved the following beneficial effects: (1) It realizes non-invasive early diagnosis The 11 gene combinations and diagnostic models proposed in this invention are applicable to various detection methods, including fecal samples, blood samples (PBMCs, plasma, exosomes, etc.), and tissue samples. They offer the advantage of being non-invasive in fecal and blood tests, and demonstrate excellent diagnostic performance in tissue tests (AUC reaching 0.92 in an independent validation dataset), showing broad clinical application value.
[0021] (2) Screening out stable gene combinations that continuously change Through systematic transcriptomic analysis of multi-stage samples from normal tissues, adenomas, and colorectal cancer, this invention, for the first time, identified 11 genes (HECW2, WARS1, SLC16A3, SECTM1, IFITM3, ADAMTSL4, FCGR1A, F2RL1, OPLAH, SERPINA1, and FCGR1CP) that exhibit persistent monotonic changes in their NAC sequences. These genes demonstrate cross-stage consistency and good diagnostic capabilities both individually and in combination. Their diagnostic capabilities are further enhanced when used in combination.
[0022] (3) Constructing a high-performance random forest model Using the aforementioned genes as feature variables, the random forest model constructed in this invention exhibits high accuracy in multiple independent datasets and validations. The best result achieved an AUC value of 0.9293, demonstrating significantly better sensitivity and specificity than traditional serum biomarkers (AUC generally below 0.7) and some previously reported polygenic models (0.75–0.85). This model not only demonstrates superior diagnostic performance but also exhibits stability and universality in cross-cohort validation.
[0023] In summary, this invention is the first to discover 11 genes (HECW2, WARS1, SLC16A3, SECTM1, IFITM3, ADAMTSL4, FCGR1A, F2RL1, OPLAH, SERPINA1, and FCGR1CP) that exhibit continuous monotonic changes in their NAC sequences. One or more of these genes, in combination, can be used as biomarkers for colorectal cancer screening, thus aiding in the early diagnosis of colorectal cancer. This invention demonstrates excellent accuracy in screening for colorectal cancer, significantly improving upon existing techniques. Therefore, this invention has significant potential for clinical application in the diagnosis of colorectal cancer.
[0024] The key to this invention lies in identifying 11 genes (HECW2, WARS1, SLC16A3, SECTM1, IFITM3, ADAMTSL4, FCGR1A, F2RL1, OPLAH, SERPINA1, and FCGR1CP) that exhibit sustained monotonic changes in their NAC sequences and are significantly associated with colorectal cancer. Therefore, colorectal cancer can be screened by detecting the expression levels of these 11 genes in the subject's stool, blood, or tissue. As for the specific methods for detecting the expression of these genes, various methods disclosed in the prior art can be employed.
[0025] Obviously, based on the above description of the present invention, and according to common technical knowledge and conventional methods in the field, various other modifications, substitutions or alterations can be made without departing from the basic technical concept of the present invention.
[0026] The following detailed embodiments further illustrate the above-described content of the present invention. However, this should not be construed as limiting the scope of the present invention to the following embodiments. All technologies implemented based on the above-described content of the present invention fall within the scope of the present invention. Attached Figure Description
[0027] Figure 1 The graph shows the expression levels of 11 genes in the normal group, colorectal adenoma group, and colorectal cancer group.
[0028] Figure 2 This is a graph showing the diagnostic accuracy of a single-gene random forest model for normal individuals, patients with colorectal adenomas, and patients with colorectal cancer.
[0029] Figure 3 This is a graph showing the diagnostic accuracy of a multi-gene combination random forest model for normal individuals, patients with colorectal adenomas, and patients with colorectal cancer. Detailed Implementation
[0030] It should be noted that the raw materials and equipment used in the embodiments are all known products, obtained by purchasing commercially available products. Algorithms for data acquisition, transmission, storage, and processing steps not specifically described in the embodiments, as well as hardware structures and circuit connections not specifically described, can all be implemented using publicly available information in the prior art.
[0031] Example 1: Screening for diagnostic genes (1) Sample collection and preparation Blood samples were collected from volunteers in the general population, patients with colorectal adenomas, and patients with colorectal cancer. After centrifugation (10 min, 25°C, 3000 rpm), blood samples separated into plasma and blood cell layers. The liquid at the interface between the two layers was aspirated and added to a 15 mL centrifuge tube containing 4 mL of RPMI-1640 medium, bringing the medium volume to 8 mL to prepare a cell suspension. The cell suspension was slowly added to a centrifuge tube containing 4 mL of Ficoll lymphocyte separation medium, and then centrifuged (25 min, 25°C, 1500 rpm). After centrifugation, the suspension separated into four layers. Personal biomass cells (PBMCs) were aspirated and RPMI-1640 medium was added, bringing the medium volume to 10–12 mL. After mixing, the mixture was centrifuged (10 min, 25°C, 1500 rpm). The supernatant was discarded, and 10 mL of medium was added, followed by another centrifugation (10 min, 25°C, 1500 rpm). PBMCs were collected for subsequent analysis, and the supernatant was discarded.
[0032] (2) RNA sequencing and data preprocessing RNA integrity of PBMC samples was assessed using the RNA Nano 6000 assay kit on a Bioanalyzer 2100 system (Agilent Technologies, Inc.). RNA samples were prepared using total RNA as input. Briefly, mRNA was purified from total RNA using magnetic beads with poly-T oligonucleotides. Fragmentation was performed at high temperature using divalent cations in First Stand Synthesis Reaction Buffer (5X). First-strand cDNA was synthesized using random hexamer primers and m-MuLV reverse transcriptase (RNase H-). Second-strand cDNA was subsequently synthesized using DNA polymerase I and RNase H. Remaining overhangs were converted to blunt ends by exonuclease / polymerase activity. After adenylation at the 3' end of the DNA fragments, hairpin adapters were ligated to prepare for hybridization. Library fragments were purified using the AMPure XP system (Beckman Coulter, Inc.) to select cDNA fragments of approximately 370 to 420 base pairs in length. PCR was then performed using Phusion high-fidelity DNA polymerase, universal PCR primers, and index (X) primers. Finally, the PCR products were purified (using the AMPureXP system), and library quality was assessed on an Agilent 2100 Bioanalyzer system.
[0033] (3) Differential expression analysis and screening of genes with persistent changes Differential expression analysis was performed on the two groups (normal group vs. adenoma group, and adenoma group vs. cancer group) using the DESeq2 R software package. DESeq2 provides a statistical method based on a negative binomial distribution model to determine differential expression in digital gene expression data. The obtained p-values were adjusted using the Benjamini and Hochberg method to control for false positives. Genes with adjusted p-values ≤0.05 and |Log2FC|>1 determined by DESeq2 were considered to have differential expression for subsequent analysis.
[0034] First, gene sets belonging to the normal group, colorectal adenoma group, and colorectal cancer group are constructed, as shown in formula (1): (1) This represents the set of expression values of gene g across all samples at different stages.
[0035] Calculate the average value of each gene in each set using formulas (2)-(4): (2) : The average expression level of gene g in the normal group; |Normal|: Number of samples in the normal group; : Expression level of gene g in the i-th normal sample; : Sum the expression levels of gene g in all normal samples.
[0036] (3) : The average expression level of gene g in the adenoma group; |Adenoma|: Number of samples in the adenoma group; : Expression level of gene g in the j-th adenoma sample; : Sum the expression levels of gene g in all adenoma samples.
[0037] (4) : The average expression level of gene g in the cancer group; |Cancer|: Sample size in the cancer group; : Expression level of gene g in the k-th cancer sample; : Sum the expression levels of gene g across all cancer samples.
[0038] The significance of each gene between the normal group and the colorectal adenoma group, and between the colorectal adenoma group and the colorectal cancer group, was calculated using the nonparametric rank-sum test, as shown in the formulas (5)-(6) below: (5) The significance of the difference in gene g between the normal group and the adenoma group (p-value); WilcoxP: The p-value is obtained from the Wilcoxon rank-sum test; The expression set of gene g in all samples of the normal group; : The expression set of gene g in all samples of the adenoma group.
[0039] (6) The significance of the difference in gene g between the adenoma group and the normal group (p-value). WilcoxP: The p-value is obtained from the Wilcoxon rank-sum test; : The expression set of gene g in all samples of the adenoma group; : The expression set of gene g in all samples of the cancer group.
[0040] When a gene conforms to formula (7) or (8), it can be considered a gene that is continuously increasing or continuously decreasing: (7) (8) Mean expression of gene g in the normal group; Mean expression of gene g in the adenoma group; Mean expression of gene g in the cancer group; The significance of the difference in gene g between the normal group and the adenoma group (p-value); The significance of the difference in gene g between the adenoma group and the normal group (p-value).
[0041] The final screening identified 11 persistently altered genes: HECW2, WARS1, SLC16A3, SECTM1, IFITM3, ADAMTSL4, FCGR1A, F2RL1, OPLAH, SERPINA1, and FCGR1CP. The expression levels of these 11 genes in the normal group, colorectal adenoma group, and colorectal cancer group are shown below. Figure 1 As shown.
[0042] Example 2: Validation of the diagnostic capability of a single-gene model in three types of samples A random forest model for colorectal cancer screening was constructed and validated using the dataset GSE20916 (samples from intestinal tissue) from the NCBI-GEO (Gene Expression Omnibus) public database. Using selected genes as feature variables and individual genes as input, a random forest algorithm was employed to construct the colorectal cancer screening model for early diagnosis. The number of estimators was set to 100, and entropy was used as the criterion. The model underwent 5-fold cross-validation to evaluate its performance, and the mean accuracy was calculated and ROC curves were plotted. This was used to assess the diagnostic potential of these genes in disease identification.
[0043] The results of the 5-fold cross-validation experiment are as follows Figure 2 As shown: the average AUC value of the single-gene OPLAA distinguishing the normal group, colorectal adenoma group, and colorectal cancer group was 0.5138; the average AUC value of the single-gene IFITM3 was 0.7; the average AUC value of the single-gene HECW2 was 0.6897; the average AUC value of the single-gene FCGR1A was 0.5707; the average AUC value of the single-gene F2RL1 was 0.6069; the average AUC value of the single-gene WARS1 was 0.6483; the average AUC value of the single-gene SLC16A3 was 0.5862; the average AUC value of the single-gene SERPINA1 was 0.7466; the average AUC value of the single-gene SECTM1 was 0.6379; the average AUC value of the single-gene ADAMTSL4 was 0.6143; and the average AUC value of FCGR1CP was 0.7643.
[0044] It is evident that the AUC values of the random forest colorectal cancer screening models trained with single genes are all below 0.8 in distinguishing colorectal cancer.
[0045] Example 3: Validation of the diagnostic capability of the multi-gene model in three types of samples A random forest model for colorectal cancer screening was constructed and validated using the dataset GSE20916 (samples from intestinal tissue) from the NCBI-GEO (Gene Expression Omnibus) public database. Using selected genes as feature variables and 11 genes (HECW2, WARS1, SLC16A3, SECTM1, IFITM3, ADAMTSL4, FCGR1A, F2RL1, OPLAH, SERPINA1, FCGR1CP) as input, a random forest algorithm was used to construct the colorectal cancer screening model for early diagnosis. The number of estimators was set to 100, and entropy was used as the criterion. The model underwent 5-fold cross-validation to evaluate its performance, and the mean accuracy was calculated and ROC curves were plotted. This was used to evaluate the diagnostic potential of these gene combinations in disease identification.
[0046] The results of the 5-fold cross-validation experiment are as follows Figure 3 As shown, the average AUC value of the combination of 11 genes (HECW2, WARS1, SLC16A3, SECTM1, IFITM3, ADAMTSL4, FCGR1A, F2RL1, OPLAH, SERPINA1, FCGR1CP) distinguishing between the normal group, colorectal adenoma group, and colorectal cancer group was 0.9293, significantly higher than the results of each single gene. This indicates that the accuracy of screening for colorectal cancer using the above 11 gene combinations is significantly better than that of each single gene, demonstrating excellent early diagnosis results for colorectal cancer.
[0047] The inventors' research also found that the expression levels of the above 11 genes (HECW2, WARS1, SLC16A3, SECTM1, IFITM3, ADAMTSL4, FCGR1A, F2RL1, OPLAH, SERPINA1, FCGR1CP) in fecal and blood samples also have excellent early diagnostic effects for colorectal cancer.
[0048] In summary, this invention is the first to discover 11 genes (HECW2, WARS1, SLC16A3, SECTM1, IFITM3, ADAMTSL4, FCGR1A, F2RL1, OPLAH, SERPINA1, and FCGR1CP) that exhibit continuous monotonic changes in their NAC sequences. One or more of these genes, in combination, can be used as biomarkers for colorectal cancer screening, thus aiding in the early diagnosis of colorectal cancer. This invention demonstrates excellent accuracy in screening for colorectal cancer, significantly improving upon existing techniques. Therefore, this invention has significant potential for clinical application in the diagnosis of colorectal cancer.
Claims
1. Use of a reagent for detecting the level of gene expression in the manufacture of a kit for the screening of colorectal cancer, characterized in that: The gene is one or more of a combination of HECW2, WARS1, SLC16A3, SECTM1, IFITM3, ADAMTSL4, FCGR1A, F2RL1, OPLAH, SERPINA1 and FCGR1CP.
2. Use according to claim 1, characterized in that: The gene is one or more of a combination of HECW2, WARS1, SLC16A3, SECTM1, IFITM3, ADAMTSL4, FCGR1A, F2RL1, OPLAH, SERPINA1 and FCGR1CP.
3. Use according to claim 1 or 2, characterized in that: The reagent is a reagent for detecting the gene expression level in human feces, blood or tissue.
4. A colorectal cancer screening kit characterized by: It comprises a reagent for detecting the gene expression level; the gene is one or more of a combination of HECW2, WARS1, SLC16A3, SECTM1, IFITM3, ADAMTSL4, FCGR1A, F2RL1, OPLAH, SERPINA1 and FCGR1CP.
5. The kit of claim 4, wherein: The gene is one or more of a combination of HECW2, WARS1, SLC16A3, SECTM1, IFITM3, ADAMTSL4, FCGR1A, F2RL1, OPLAH, SERPINA1 and FCGR1CP.
6. The kit of claim 4 or 5, wherein: The reagent is a reagent for detecting the gene expression level in human feces, blood or tissue.
7. A colorectal cancer screening system characterized by: The system comprises the following parts: An input module for inputting the gene expression level, the gene being one or more of a combination of HECW2, WARS1, SLC16A3, SECTM1, IFITM3, ADAMTSL4, FCGR1A, F2RL1, OPLAH, SERPINA1 and FCGR1CP; A prediction module for inputting the gene expression level as a feature into a machine learning model to obtain a colorectal cancer screening result; An output module for outputting the colorectal cancer screening result.
8. The system of claim 7, wherein: The gene is one or more of a combination of HECW2, WARS1, SLC16A3, SECTM1, IFITM3, ADAMTSL4, FCGR1A, F2RL1, OPLAH, SERPINA1 and FCGR1CP. And / or, the machine learning model is a random forest model, a vector machine model, a logistic regression model, a LASSO regression model, a neural network model or an XGBoost model.
9. The system of claim 7 or 8, characterized in that: The gene expression level is the gene expression level in human feces, blood or tissue.
10. A computer-readable storage medium, characterized in that: A computer program for implementing the system of any one of claims 7-9 is stored thereon. A computer program for implementing the system of any one of claims 7-9 is stored thereon.