Colorectal cancer marker microorganism beta-glucuronidase and application thereof

By screening and combining β-glucuronidase genes from different microorganisms, a biomarker combination was constructed for early screening of colorectal cancer. This solved the problems of high invasiveness and low sensitivity of existing methods, and enabled the provision of highly accurate, non-invasive early diagnosis and personalized treatment plans.

CN121495906APending Publication Date: 2026-02-10UNIV OF MACAU
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511558878.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-29
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing colorectal cancer screening methods are highly invasive and have low sensitivity, failing to effectively cover the population that should be screened. Furthermore, the lack of colonoscopy resources and physicians results in low early diagnosis rates, failing to meet the needs of early screening.

Method used

By employing a combination of β-glucuronidase genes derived from different microorganisms, and using metagenomic research strategies to screen for microbial β-glucuronidase genes with diagnostic potential, a biomarker combination was constructed for early screening of colorectal cancer. Combined with data analysis and model building, a non-invasive and comfortable diagnostic method was provided.

Benefits of technology

It achieves highly accurate and non-invasive early screening for colorectal cancer, improves the diagnostic rate, provides a basis for individualized treatment plans, and can be used as a potential target for CRC treatment, with potential for drug development and clinical application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention belongs to the technical field of biology, and discloses colorectal cancer marking microorganism beta-glucuronidase and application thereof. The invention provides a marker combination of beta-glucuronidase genes derived from different microorganisms. The marker combination is analyzed in a CRC (cyclic redundancy check) classification model and an adenoma classification model, and results show that the marker combination has good diagnostic ability. Beta-glucuronidase gene combinations based on different microorganisms can be used for detecting CRC subjects, screening and diagnosis of CRC in different stages are assisted, and an important basis is provided for individualized treatment strategy selection.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of biotechnology, and particularly relates to a colorectal cancer marker microorganism β-glucuronidase and application thereof. BACKGROUND

[0002] Colorectal cancer (CRC) is the third most common and deadly cancer in the world, and shows a trend of youth. The existing treatment methods of CRC, including surgery, radiotherapy and chemotherapy, etc., will have the problems of postoperative recurrence and metastasis. Therefore, early screening is very important for reducing the incidence rate, improving the cure rate and reducing the cost and economic burden of surgical treatment for patients. Common colorectal cancer screening methods include fecal occult blood test, digital rectal examination and colonoscopy, etc. The detection sensitivity of fecal occult blood test is low, and the rejection rate of methods such as digital rectal examination is high, which cannot completely and effectively cover the people who should be screened; colonoscopy is an invasive examination, not only has a low compliance rate, but also the resources of colonoscopy and colonoscopy physicians are lacking in China, which cannot be completely covered. In order to improve this situation, it is necessary to develop more comfortable and non-invasive CRC early screening methods. In recent years, some non-invasive diagnostic techniques have emerged, including circulating tumor DNA detection (such as gene mutation and methylation detection), non-coding RNA detection (such as blood miRNA, lncRNA and circRNA detection) and circulating tumor cell detection, etc. These non-invasive diagnostic methods are simple, non-invasive or minimally invasive and have potential high accuracy, which is expected to improve the accuracy and acceptance rate of CRC early diagnosis, thereby reducing the mortality rate and improving the survival rate and quality of life of patients.

[0003] The development process of CRC is affected by multiple factors such as genetics, physiology and environment. Among the environmental factors, lifestyle, especially dietary intake, is considered to play an important role in the development of CRC. Western-style diet rich in red meat and processed meat is usually associated with an increased risk of CRC, while Mediterranean diet rich in fruits and vegetables is reported to be beneficial for preventing cancer. Intestinal microbiota plays an important role in the utilization, absorption and metabolism of nutrients, and has a key role in the occurrence and development of CRC. In recent years, intestinal microbiota in feces has shown its potential as a non-invasive detection method as a biomarker, which can be used as an important tool for early screening of CRC.

[0004] There is a class of important β-glucuronidases (GUSs) in human intestinal microbes, which coordinate with host uridine diphosphate glucuronosyltransferases (UGTs) to participate in the in vivo disposal of a large number of endogenous compounds (such as bilirubin, steroids, neurotransmitters, serotonin, bile acids, etc.) and exogenous components (such as drugs, carcinogens from food or environmental sources, etc.) with important biological functions. In the body, through the catalytic conjugation reaction of host UGTs, the generation of glucuronides is the main metabolic pathway of many endogenous and exogenous components (including drugs). This metabolic transformation can increase the polarity of the product and promote the excretion of foreign substances, which is an important "detoxification" way in the body. However, the glucuronide compounds excreted into the intestine through bile can be hydrolyzed by β-glucuronidases in the intestinal flora to generate free aglycone. The aglycone can be reabsorbed by the intestine and returned to the liver via the portal vein, forming a liver-intestine cycle.

[0005] In 2001, scholars found that the fecal β-glucuronidase activity of CRC patients was about 1.7 times higher than that of healthy people. In addition, food-derived polycyclic aromatic hydrocarbons (PAHs) and heterocyclic aromatic amines (HCAs) are risk factors for CRC. PAHs and HCAs can extend the exposure and / or circulating levels of toxic substances in the intestine through the host UGTs-intestinal bacterial GUSs axis, thereby causing colorectal cell DNA damage and promoting the occurrence and development of CRC. In addition, it has been found that the concentrations of primary and secondary bile acids in the feces of high-risk groups for CRC are significantly higher than those in low-risk groups for CRC. Bile acids can also undergo glucuronide conjugation under the action of liver UGTs, and then be excreted into the intestine through bile, extend their exposure and / or circulating levels in the intestine under the action of intestinal bacterial GUSs, and further promote the increase in the concentration of secondary bile acids (especially deoxycholic acid) with carcinogenic activity in the intestinal lumen. Therefore, based on the important significance of early non-invasive screening for CRC and the important role of intestinal bacterial β-glucuronidase, screening of microbial β-glucuronidase genes with the ability to serve as a biomarker for early screening of CRC is urgent to be solved.

[0006] In view of this, the present application is provided. SUMMARY

[0007] The first aspect of the present application aims to provide a marker combination.

[0008] The second aspect of the present application aims to provide a use of a substance for detecting the marker combination of the first aspect of the present application in the preparation of a diagnosis of a disease.

[0009] The third aspect of the present application aims to provide a method for constructing a model for disease diagnosis.

[0010] The fourth aspect of the present application aims to provide a model running module for disease diagnosis.

[0011] The fifth aspect of the present application aims to provide a system for disease diagnosis.

[0012] The sixth aspect of the present application aims to provide an electronic device.

[0013] To achieve the above-mentioned object, the technical solution adopted by the present application is: The first aspect of the present application provides a marker combination comprising β-glucuronidase genes derived from different microorganisms, including Bacteroides ovatus Bacteroides ovatus, Bacteroides nordii Bacteroides norimbergensis, Oscillospiraceae bacterium Oscillospiraceae bacterium, Faecalibacterium prausnitzii Parabacteroides chongii Faecalibacterium prausnitzii, Hungatella hathewayi Parabacteroides distasonis, Echinicola strongylocentroti Bacteroides cellulosilyticus Hongita hongii, Parabacteroides goldsteinii Echinoclostridium ramosum, Enterocloster clostridioformis Bacteroides helcogenes Bacteroides cellulosolvens, Bifidobacterium bifidum Parabacteroides gordinii, Bacteroides faecium Intestinibacter rectus, Stutzerimonas stutzeri Spirosoma spiroforme, Ruminococcus torques Bifidobacterium bifidum, Lachnospira eligens Bacteroides caccae, Bacteroides thetaiotaomicron Pseudomonas stutzeri, Dorea longicatena Ruminococcus torques, Cellulosimicrobium cellulans Lachnospira pectinoschiza, Bacteroides ovatus Bacteroides thetaiotaomicron, Bacteroides nordii Dorea longicatena, Oscillospiraceae bacterium Fibrobacter succinogenes.

[0014] In some embodiments of the present application, the β-glucuronidase gene derived from Faecalibacterium prausnitzii Bacteroides ovatus includes at least one of GUS3, GUS4, and is sequentially recorded as Bac_ovatus.GUS3 or Bac_ovatus.GUS4.

[0015] In some embodiments of the present application, the β-glucuronidase gene derived from Parabacteroides chongiiThe β-glucuronidase genes of Bacteroides nodosum include at least one of GUS1, GUS2, GUS3, GUS4, GUS5, and GUS6, which are respectively designated as Bac_nordii.GUS1, Bac_nordii.GUS2, Bac_nordii.GUS3, Bac_nordii.GUS4, Bac_nordii.GUS5, and Bac_nordii.GUS6.

[0016] In some embodiments of the present invention, the source Hungatella hathewayi The β-glucuronidase gene of (Osc_bacterium.GUS6 and Osc_bacterium.GUS7) includes at least one of GUS6 and GUS7.

[0017] In some embodiments of the present invention, the source Echinicola strongylocentroti The β-glucuronidase gene of *Clostridium prausnitzii* includes at least one of GUS3, GUS7, and GUS10, which are respectively denoted as Fae_prausnitzii.GUS3, Fae_prausnitzii.GUS7, and Fae_prausnitzii.GUS10.

[0018] In some embodiments of the present invention, the source Bacteroides cellulosilyticus The β-glucuronidase gene of *Parchongii* includes at least one of GUS1 and GUS7, referred to as Par_chongii.GUS1 and Par_chongii.GUS7, respectively.

[0019] In some embodiments of the present invention, the β-glucuronidase gene derived from Hungatella hathewayi includes GUS1, denoted as Hun_hathewayi.GUS1.

[0020] In some embodiments of the present invention, the source Parabacteroides goldsteinii The β-glucuronidase gene of *Echinococcus strongylocentroti* includes GUS, denoted as Ech_strongylocentroti.GUS.

[0021] In some embodiments of the present invention, the source Enterocloster clostridioformisThe β-glucuronidase genes of *Bacteroides cellulosilyticus* include at least one of GUS1, GUS2, GUS3, GUS5, GUS15, and GUS18, referred to as Bac_cellulosilyticus.GUS1, Bac_cellulosilyticus.GUS2, Bac_cellulosilyticus.GUS3, Bac_cellulosilyticus.GUS5, Bac_cellulosilyticus.GUS15, and Bac_cellulosilyticus.GUS18, respectively.

[0022] In some embodiments of the present invention, the source Bacteroides helcogenes The β-glucuronidase genes of *Parostenobacter gondii* include at least one of GUS2, GUS3, and GUS6, referred to as Par_goldsteinii.GUS2, Par_goldsteinii.GUS3, and Par_goldsteinii.GUS6, respectively.

[0023] In some embodiments of the present invention, the source Bifidobacterium bifidum The β-glucuronidase gene of *Enterobacter* includes GUS2, denoted as Ent_clostridioformis.GUS2.

[0024] In some embodiments of the present invention, the source Bacteroides faecium The β-glucuronidase gene of *Bacteroides helicogenes* includes GUS2, denoted as Bac_helcogenes.GUS2.

[0025] In some embodiments of the present invention, the source Stutzerimonas stutzeri The β-glucuronidase gene of Bifidobacterium bifidum includes at least one of GUS1 and GUS2, referred to as Bif_bifidum.GUS1 and Bif_bifidum.GUS2, respectively.

[0026] In some embodiments of the present invention, the source Ruminococcus torques The β-glucuronidase gene of Bac_faecium includes GUS7, denoted as Bac_faecium.GUS7.

[0027] In some embodiments of the present invention, the source Lachnospira eligens The β-glucuronidase gene of *Pseudomonas stutzeri* includes GUS, denoted as Stu_stutzeri.GUS.

[0028] In some embodiments of the present invention, the source Bacteroides thetaiotaomicronThe β-glucuronidase gene of *Ruminococcus tortifolius* includes GUS, denoted as Rum_torques.GUS.

[0029] In some embodiments of the present invention, the source Dorea longicatena The β-glucuronidase gene of (selected spirulina) includes GUS2, denoted as Lac_eligens.GUS2.

[0030] In some embodiments of the present invention, the source Cellulosimicrobium cellulans The β-glucuronidase gene of *Bacteroides polymorpha* includes GUS4, denoted as Bac_thetaiotaomicron.GUS4.

[0031] In some embodiments of the present invention, the source Bacteroides ovatus The β-glucuronidase gene of (Dor_longicatena) includes at least one of GUS1 and GUS2, referred to as Dor_longicatena.GUS1 and Dor_longicatena.GUS2, respectively.

[0032] In some embodiments of the present invention, the source Bacteroides nordii The β-glucuronidase gene of (fibrotic microbacteria) includes GUS, denoted as Cel_cellulans.GUS.

[0033] In some embodiments of the present invention, the source Oscillospiraceae bacterium The nucleic acid molecules of the β-glucuronidase genes GUS3 and GUS4 of *Bacteroides ovalis* are, respectively, the nucleic acid molecules encoding the β-glucuronidase shown in SEQ ID NO:1 and SEQ ID NO:2.

[0034] In some embodiments of the present invention, the source Faecalibacterium prausnitzii The nucleic acid molecules of the β-glucuronidase genes GUS1, GUS2, GUS3, GUS4, GUS5 and GUS6 of Bacteroides noriformis are, respectively, the nucleic acid molecules encoding the β-glucuronidase shown in SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:18, SEQ ID NO:19 and SEQ ID NO:20.

[0035] In some embodiments of the present invention, the source Parabacteroides chongii The nucleic acid molecules of the β-glucuronidase genes GUS6 and GUS7 of (Oscillatoriaceae bacteria) are the nucleic acid molecules encoding the β-glucuronidase shown in SEQ ID NO:15 and SEQ ID NO:16, respectively.

[0036] In some embodiments of the present invention, the source Hungatella hathewayi The nucleic acid molecules of the β-glucuronidase genes GUS3, GUS7, and GUS10 of *Clostridium praosporum* are, respectively, the nucleic acid molecules encoding the β-glucuronidases shown in SEQ ID NO:17, SEQ ID NO:26, and SEQ ID NO:31.

[0037] In some embodiments of the present invention, the source Echinicola strongylocentroti The nucleic acid molecules of the β-glucuronidase genes GUS1 and GUS7 of *Pseudomonas chinensis* are the nucleic acid molecules encoding the β-glucuronidase shown in SEQ ID NO:22 and SEQ ID NO:29, respectively.

[0038] In some embodiments of the present invention, the source Bacteroides cellulosilyticus Parabacteroides goldsteinii Enterocloster clostridioformis Bacteroides helcogenes Bifidobacterium bifidum Bacteroides faecium Stutzerimonas stutzeri Ruminococcus torques Lachnospira eligens Bacteroides thetaiotaomicron Dorea longicatena Cellulosimicrobium cellulans Bacteroides ovatus Bacteroides nordii Oscillospiraceae bacterium Faecalibacterium prausnitzii Parabacteroides chongii Hungatella hathewayi Echinicola strongylocentroti Bacteroides cellulosilyticus Parabacteroides goldsteinii Enterocloster clostridioformis Bacteroides helcogenes Bifidobacterium bifidum Bacteroides faecium Stutzerimonas stutzeri Ruminococcus torques Lachnospira eligens Bacteroides thetaiotaomicron Dorea longicatena Cellulosimicrobium cellulans Bacteroides ovatus Bacteroides nordii Oscillospiraceae bacterium Faecalibacterium prausnitzii Parabacteroides chongii Hungatella hathewayi Echinicola strongylocentroti Bacteroides cellulosilyticus Parabacteroides goldsteinii Enterocloster clostridioformis Bacteroides hel The nucleic acid molecule of the β-glucuronidase gene GUS1 of *Hunnerella haematobium* is the nucleic acid molecule encoding the β-glucuronidase shown in SEQ ID NO:23.

[0039] In some embodiments of the present invention, the source Echinicola strongylocentroti The nucleic acid molecule of the β-glucuronidase gene GUS of *Echinococcus spp.* is the nucleic acid molecule encoding the β-glucuronidase shown in SEQ ID NO:25.

[0040] In some embodiments of the present invention, the source Bacteroides cellulosilyticus The nucleic acid molecules of the β-glucuronidase genes GUS1, GUS2, GUS3, GUS5, GUS15 and GUS18 of *Bacteroides celluloseis* are, respectively, the nucleic acid molecules encoding the β-glucuronidases shown in SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:30 and SEQ ID NO:36.

[0041] In some embodiments of the present invention, the source Parabacteroides goldsteinii The nucleic acid molecules of the β-glucuronidase genes GUS2, GUS3 and GUS6 of *Pseudomonas gossyflores* are, respectively, the nucleic acid molecules encoding the β-glucuronidases shown in SEQ ID NO:13, SEQ ID NO:14 and SEQ ID NO:27.

[0042] In some embodiments of the present invention, the source Enterocloster clostridioformis The nucleic acid molecule of the β-glucuronidase gene GUS2 of Enterobacter fusiformis is the nucleic acid molecule encoding the β-glucuronidase shown in SEQ ID NO:28.

[0043] In some embodiments of the present invention, the source Bacteroides helcogenes The nucleic acid molecule of the β-glucuronidase gene GUS2 of *Bacteroides spirochetes* is the nucleic acid molecule encoding the β-glucuronidase shown in SEQ ID NO:38.

[0044] In some embodiments of the present invention, the source Bifidobacterium bifidum The nucleic acid molecules of the β-glucuronidase genes GUS1 and GUS2 of Bifidobacterium bifidum are the nucleic acid molecules encoding the β-glucuronidase shown in SEQ ID NO:12 and SEQ ID NO:21, respectively.

[0045] In some embodiments of the present invention, the source Bacteroides faecium The nucleic acid molecule of the β-glucuronidase gene GUS7 of *Bacteroides coccidioides* is the nucleic acid molecule encoding the β-glucuronidase shown in SEQ ID NO:33.

[0046] In some embodiments of the present invention, the source Stutzerimonas stutzeri The nucleic acid molecule of the β-glucuronidase gene GUS of *Pseudomonas stearothermia* is the nucleic acid molecule encoding the β-glucuronidase shown in SEQ ID NO:34.

[0047] In some embodiments of the present invention, the source Ruminococcus torques The nucleic acid molecule of the β-glucuronidase gene GUS of *Ruminococcus twitchis* is the nucleic acid molecule encoding the β-glucuronidase shown in SEQ ID NO:11.

[0048] In some embodiments of the present invention, the source Lachnospira eligens The nucleic acid molecule of the β-glucuronidase gene GUS2 of (selected *Helicobacter pylori*) is the nucleic acid molecule encoding the β-glucuronidase shown in SEQ ID NO:32.

[0049] In some embodiments of the present invention, the source Bacteroides thetaiotaomicron The nucleic acid molecule of the β-glucuronidase gene GUS4 of Bacteroides polymorpha is the nucleic acid molecule encoding the β-glucuronidase shown in SEQ ID NO:10.

[0050] In some embodiments of the present invention, the source Dorea longicatena The nucleic acid molecules of the β-glucuronidase genes GUS1 and GUS2 of *Dorhizium anisopliae* are the nucleic acid molecules encoding the β-glucuronidase shown in SEQ ID NO:24 and SEQ ID NO:35, respectively.

[0051] In some embodiments of the present invention, the source Cellulosimicrobium cellulans The nucleic acid molecule of the β-glucuronidase gene GUS of (fibrotic microbacteria) is the nucleic acid molecule encoding the β-glucuronidase shown in SEQ ID NO:37.

[0052] In some embodiments of the invention, the biomarker combination comprises Bac_ovatus.GUS3, Bac_nordii.GUS3, Osc_bacterium.GUS6, Osc_bacterium.GUS7, Fae_prausnitzii.GUS3, Bac_nordii.GUS5, Bac_nordii.GUS6, Par_chongii.GUS1, Hun_hathewayi.GUS1, Ech_strongylocentroti.GUS, Bac_nordii.GUS2, Bac_nordii.GUS4, and Par_goldsteinii.GUS6. This biomarker combination can be used to diagnose adenomas.

[0053] In some embodiments of the present invention, the biomarker combination comprises Bac_ovatus.GUS3, Bac_ovatus.GUS4, Bac_nordii.GUS1, Bac_nordii.GUS2, Bac_cellulosilyticus.GUS1, Bac_cellulosilyticus.GUS3, Bac_cellulosilyticus.GUS5, Bac_nordii.GUS4, Bac_nordii.GUS5, Par_goldsteinii.GUS6, Ent_clostridioformis.GUS2, Bac_cellulosilyticus.GUS15, and Bac_helcogenes.GUS2. This biomarker combination can be used to diagnose stage S0 colorectal cancer.

[0054] In some embodiments of the present invention, the biomarker combination consists of Bac_ovatus.GUS3, Bac_ovatus.GUS4, Bac_nordii.GUS1, Bac_nordii.GUS2, Bif_bifidum.GUS1, Bac_nordii.GUS4, Bac_nordii.GUS5, Bac_nordii.GUS6, Bif_bifidum.GUS2, Par_chongii.GUS7, Bac_faecium.GUS7, and Stu_stutzeri.GUS. This biomarker combination can be used to diagnose stage I / II colorectal cancer.

[0055] In some embodiments of the invention, the combination of markers comprises Rum_torques.GUS, Lac_eligens.GUS2, Bac_nordii.GUS2, Bac_cellulosilyticus.GUS1, Bac_cellulosilyticus.GUS2, Bac_cellulosilyticus.GUS3, Bac_cellulosilyticus.GUS5, Bac_thetaiotaomicron.GUS4, and Par_goldsteinii. The biomarker combination consists of GUS2, Par_goldsteinii.GUS3, Bac_nordii.GUS4, Bac_nordii.GUS5, Dor_longicatena.GUS1, Fae_prausnitzii.GUS7, Bac_cellulosilyticus.GUS15, Fae_prausnitzii.GUS10, Dor_longicatena.GUS2, Bac_cellulosilyticus.GUS18, and Cel_cellulans.GUS. This biomarker combination can be used to diagnose stage SIII / IV colorectal cancer.

[0056] In some embodiments of the invention, the combination of markers comprises Bac_ovatus.GUS3, Bac_nordii.GUS5, Bac_nordii.GUS6, Bac_ovatus.GUS4, Bac_nordii.GUS1, Bac_nordii.GUS2, Bac_cellulosilyticus.GUS1, Bac_cellulosilyticus.GUS3, Bac_cellulosilyticus.GUS5, Bac_nordii.GUS4, Par_goldsteinii.GUS6, Ent_clostridioformis.GUS2, Bac_cellulosilyticus.GUS15, Bac_helcogenes.GUS2, Bif_bifidum.GUS1, and Bi The combination of f_bifidum.GUS2, Par_chongii.GUS7, Bac_faecium.GUS7, Stu_stutzeri.GUS, Rum_torques.GUS, Lac_eligens.GUS2, Bac_cellulosilyticus.GUS2, Bac_thetaiotaomicron.GUS4, Par_goldsteinii.GUS2, Par_goldsteinii.GUS3, Dor_longicatena.GUS1, Fae_prausnitzii.GUS7, Fae_prausnitzii.GUS10, Dor_longicatena.GUS2, Bac_cellulosilyticus.GUS18, and Cel_cellulans.GUS is used to diagnose colorectal cancer.

[0057] This invention employs a metagenomic research strategy to mine β-glucuronidases in gut microbiota and, through data analysis, identifies microbial β-glucuronidase genes as biomarkers for colorectal cancer. It also provides an application for diagnosing colorectal cancer risk in subjects based on microbial β-glucuronidase genes. These microbial β-glucuronidase biomarkers, as a type of gene biomarker or gut microbiota biomarker, are of significant value for early CRC screening, offering advantages such as being non-invasive, comfortable, convenient, safe, and highly acceptable. Furthermore, the biomarker combination of this invention can also serve as a potential target for CRC treatment, possessing potential for drug development and clinical application. It also plays an important role in determining individualized treatment plans, prognostic assessment, and efficacy prediction.

[0058] A second aspect of the invention provides the use of a substance for detecting a combination of markers of the first aspect of the invention in the preparation of a diagnostic disease, said disease including colorectal cancer and / or adenoma.

[0059] In some embodiments of the present invention, the diagnosis includes early diagnosis of colorectal cancer, staging diagnosis of colorectal cancer, and diagnosis of adenoma.

[0060] In some embodiments of the invention, the substance includes a substance that detects the combination of markers at the gene or protein level.

[0061] In some embodiments of the present invention, the substance used to detect the combination of biomarkers at the protein level is selected from substances of one or more detection methods of the group consisting of: chemiluminescence, immunofluorescence, protein chip, proteometry, immunohistochemistry, patch tracing based on labeling technology, Western blotting, and enzyme-linked immunosorbent assay (ELISA).

[0062] In some embodiments of the present invention, the substance used to detect the combination of biomarkers at the gene level is selected from substances of one or more detection methods from the group consisting of: DNA sequencing, RNA sequencing, RNA-in situ hybridization, digital PCR, and quantitative real-time PCR.

[0063] In some embodiments of the present invention, the product includes reagents, kits, test strips, systems, or chips.

[0064] In some embodiments of the present invention, the test sample of the product is selected from the tissue, blood, plasma, and feces of the test subject.

[0065] A third aspect of the present invention provides a method for constructing a model for disease diagnosis, comprising the following steps: Models are constructed using a combination of biomarkers from the first aspect of the invention; the diseases include colorectal cancer and / or adenomas.

[0066] In some embodiments of the present invention, the diagnosis includes early diagnosis of colorectal cancer, staging diagnosis of colorectal cancer, diagnosis of adenoma, and diagnosis of colorectal cancer.

[0067] In some embodiments of the present invention, the construction method includes at least one of random forest, log regression, linear discriminant analysis, support vector machine, Bayesian network, KthNearestNeighbour, and decision tree; preferably random forest.

[0068] In some embodiments of the present invention, the construction method includes at least one of receiver operating characteristic (ROC) curve analysis, correlation analysis, survival analysis, regression analysis, stratified analysis, continuous variable analysis, categorical variable analysis, and propensity score matching.

[0069] Those skilled in the art can use different construction methods to combine the above-mentioned markers to obtain diagnostic models with different proportions of each marker, and achieve the same or similar diagnostic results. These diagnostic models obtained through conventional changes in samples and algorithms all fall within the protection scope of this invention.

[0070] A fourth aspect of the present invention provides a model running module for disease diagnosis, the model running module including a computing component; The computing component performs calculations on the data of the marker combination of the first aspect of the present invention; The diseases mentioned include colorectal cancer and / or adenoma.

[0071] In some embodiments of the present invention, the diagnosis includes early diagnosis of colorectal cancer, staging diagnosis of colorectal cancer, diagnosis of adenoma, and diagnosis of colorectal cancer.

[0072] A fifth aspect of the present invention provides a system for disease diagnosis, said system comprising at least one of 1) to 3): 1) Data input module, also known as detection result input device, can specifically be one or more of the following: mouse, keyboard, touch screen display, one or more buttons, one or more switches, one or more triggers, etc. 2) The model running module, also known as the diagnostic result display device, can specifically be one or more of the following: liquid crystal display (LCD), light-emitting diode (LED) display, plasma display, projection display, touch screen display, etc. 3) Results output module: The results output module can send the results of distinguishing whether the subject is in a high- or low-risk group to an information communication terminal device that can be viewed by the patient or medical staff. The model running module includes the model running module for disease diagnosis according to the fourth aspect of the present invention.

[0073] A sixth aspect of the present invention provides an electronic device including a storage device, a processor, and a computer program stored in the storage device and executable on the processor, wherein the computer program stored in the storage device and executable on the processor includes the system of the sixth aspect of the present invention.

[0074] The present invention also provides a storage medium storing processor-executable instructions, which, when executed by a processor, are used to perform the construction method as described in the third aspect of the present invention.

[0075] The beneficial effects of this invention are: This invention provides a combination of biomarkers derived from β-glucuronidase genes of different microorganisms. Analysis using this biomarker combination in CRC classification models and adenoma classification models showed that the combination exhibited good diagnostic capabilities. This suggests that combinations of β-glucuronidase genes based on different microorganisms can be used to detect CRC subjects, assisting in the screening and diagnosis of CRC at different stages, and providing important evidence for the selection of individualized treatment strategies. Attached Figure Description

[0076] The present invention will be further described below with reference to the accompanying drawings and embodiments, wherein: Figure 1 Species classification tree diagram, intergroup abundance distribution heatmap, upregulation and detection rate of 38 β-glucuronidase genes that showed significant differences between the adenoma group and the CRC case group and the control group.

[0077] Figure 2 ROC curves and AUC values ​​for diagnosing colorectal cancer (A) and adenoma (B) using 38 differentially expressed microbial β-glucuronidase genes based on a discovery cohort.

[0078] Figure 3 The ROC curves and AUC values ​​for the CRC model and adenoma model in the validation queue are shown. Detailed Implementation

[0079] The following will describe the concept and technical effects of the present invention clearly and completely with reference to embodiments, so as to fully understand the purpose, features and effects of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are all within the scope of protection of the present invention.

[0080] Unless otherwise specified in the examples, the procedures should be performed under standard conditions or conditions recommended by the manufacturer. Reagents or instruments whose manufacturers are not specified are all commercially available products.

[0081] The features and performance of the present invention will be further described in detail below with reference to embodiments.

[0082] Example 1: Discovery of β-glucuronidase gene in colorectal cancer marker microorganisms using metagenomic research strategy. 1. Discovery of sample collection in the queue Fecal metagenomic sequencing data of 576 samples (DRA006684 and DRA008156) were downloaded from the DDBJ DRA database. These 576 samples (the discovery cohort) were divided into five groups: a healthy control group (normal and with a few polyps) containing 251 samples; an adenoma group (multiple polypoid adenomas with low-grade dysplasia) containing 67 samples; the S0 group (polypoid adenomas with high-grade dysplasia) containing 73 samples; the SI / II group (stage I and II colorectal cancer) containing 111 samples; and the SIII / IV group (stage III and IV colorectal cancer) containing 74 samples. The study samples for this discovery cohort were obtained from the National Cancer Center Hospital in Tokyo, Japan, and the subjects were participants who underwent a full colonoscopy. Subjects with hereditary or suspected hereditary diseases (such as familial adenomatous polyposis, hereditary nonpolyposis colorectal cancer, high microsatellite instability), inflammatory bowel disease, a history of abdominal surgery, or insufficient stool samples to collect data were excluded from the study.

[0083] Starting from 576 samples, 5 samples with sequencing data volume less than 1G were excluded (including 4 healthy control samples and 1 adenoma group sample), leaving 571 samples for analysis. Basic clinical information on age, sex, and BMI of the discovery cohort is shown in Tables 1 and 2.

[0084] Table 1. Basic clinical information of the discovery cohort

[0085] Note: The distribution values ​​of age and BMI in each group represent mean ± SD. Gender (F:M) represents the female:male sample size.

[0086] Table 2 shows the results of intergroup differences in basic clinical information of the cohort.

[0087] Note: The p-value was obtained by using the rank-sum test for age and BMI and the chi-square test for gender.

[0088] 2. Based on 571 samples, a microbial gene set was constructed. Specifically, for each sample, starting from the downloaded sequencing data (DRA006684 and DRA008156), genome assembly was performed using SOAP denovo (Version 2.04, parameters: -d 1 -M 3 -u -F -V 1 -K 57). Subsequently, the sequencing data was aligned to the generated contigs using SOAP2 (Version 2.21, parameters: -r 2 -m 200 -x 400). Sequencing data from 571 samples that did not align were then mixed and assembled using MEGAHIT (Version 1.2.9, parameters: --k-list 57 --min-contig-len 500). Based on the contigs generated from individual sample assembly and mixed assembly, gene prediction was performed using MetaGeneMark (prokaryotic GeneMark.hmm Version 3.38, parameters: -a -d -fG). The sequences of all predicted genes were deredundant using CD-HIT (Version 4.8.1, parameters: -c0.95 -G 0 -aS 0.9 -g 1 -n 5 -d 0) to form a microbial gene set containing 10,947,761 genes.

[0089] 3. Obtain species annotations of genes in the microbial gene pool and their abundance information in each sample. Gene species annotation was performed using Kraken2 software (Version 2.0.7-bet, default parameters) and the microNT database constructed by the inventors. In constructing the microNT database, the NCBI genebank database was first downloaded, and then sequences from bacteria, archaea, fungi, and viruses were extracted based on their species classification information to form the microNT database.

[0090] For each sample, the number of sequencing sequences aligned to genes in each microbial gene set was calculated using BWA software (Version 0.7.17-r1188, default parameters). The sequence was then normalized using gene length and further normalized to the number of millions of sequencing sequences for each individual sample, which was taken as the abundance of the gene in that sample.

[0091] 4. Identify β-glucuronidase genes from microbial gene pools.

[0092] First, on September 25, 2023, the inventors searched the NCBI database using "beta glucuronidase" as the keyword, limiting the search to "bacteria" species, and obtained the initial public sequence of microbial β-glucuronidase. Further, sequences containing the keywords "candidate" or "putative," sequences labeled as other enzyme types, and sequences with incomplete information were filtered, ultimately retaining 114 sequences as reference sequences for identifying the β-glucuronidase gene.

[0093] During the identification process, starting from a microbial gene set containing 10,947,761 genes, the following three steps were used to progressively filter and screen for β-glucuronidase genes: 1) The microbial gene set was compared with 114 reference sequences using blastp (Version 2.12.0+, default parameters), retaining sequences with similarity greater than 25% and E values ​​lower than 0.05; 2) Hmmsearch (Version 3.3.2, default parameters) was used to retain sequences with E values ​​lower than 0.05 and containing three conserved domains of the GH2 family (PF02836, PF02837, and PF00703); 3) Sequences containing seven key and specific motifs of microbial β-glucuronidase (N and Y motifs, NxKG motif, and catalytic E motif) were retained. After these three progressive filtering steps, a total of 550 microbial β-glucuronidase genes were obtained.

[0094] 5. Identification of microbial β-glucuronidase genes that can be used as biomarkers for colorectal cancer. The inventors used a two-sided Wilcoxon rank-sum test and FDR correction method to compare and analyze the abundance profiles of 550 microbial β-glucuronidase genes in various groups. Using this method, based on adjusted... P With a screening condition of <0.05, microbial β-glucuronidase genes that can be used as biomarkers for colorectal cancer were obtained.

[0095] Through screening, the inventors obtained a total of 38 significantly different microbial β-glucuronidase genes (see Table 3 and...). Figure 1 Among them, the number of differentially expressed β-glucuronidase genes in the MP group, S0 group, SI / II group, and SIII / IV group compared with the healthy control group were 13, 17, 14, and 19, respectively. These differentially expressed β-glucuronidase genes were enriched in the case groups or the control group, suggesting that they may play a role in the etiology and pathogenesis of CRC, and also implying that they could serve as potential biomarkers for early CRC.

[0096] The amino acid sequences of the above 38 microbial β-glucuronidases are shown in Table 4.

[0097] Table 3. 38 β-glucuronidase genes that showed significant differences between the CRC case group and the control group.

[0098]

[0099] Note that NA in Table 3 represents P value>0.05.

[0100] Table 4. Amino acid sequences of the 38 screened β-glucuronidases

[0101]

[0102]

[0103]

[0104]

[0105]

[0106]

[0107]

[0108]

[0109] Example 2: Evaluation of the predictive power of microbial β-glucuronidase gene markers in the early diagnosis of CRC The inventors used 13, 12, and 19 differentially expressed microbial β-glucuronidase genes from groups S0, SI / II, and SIII / IV compared to the healthy control group, respectively. Based on Cox regression survival analysis, they measured the risk coefficients of these differentially expressed microbial β-glucuronidase genes. The results are detailed in Table 5. When analyzing group S0, the sample survival time was uniformly set to 1, the survival events in group S0 were set to 1, and the survival events in the healthy control group were set to 0. The risk coefficient exp(coef), or HR, of the differentially expressed microbial β-glucuronidase genes was obtained using the survival package in R software. The same strategy was used to analyze and obtain the risk coefficients for groups SI / II and SIII / IV as for group S0.

[0110] The results are shown in Table 5. The overall analysis results show that the likelihood test p-values ​​for groups S0, SI / II, and SIII / IV are 0.001062, 4.449e-05, and 0.0001471, respectively.

[0111] Table 5. Risk coefficient analysis results of differentially expressed β-glucuronidase genes based on Cox.

[0112]

[0113] Furthermore, the inventors constructed a random forest model using 38 differentially expressed microbial β-glucuronidase genes to distinguish between CRC case groups and healthy controls, as well as adenoma case groups and healthy controls, and employed ROC (receiver over time) analysis. The area under the curve (AUC) of the microbial β-glucuronidase gene (a larger AUC indicates higher diagnostic ability) was used to evaluate the ability of microbial β-glucuronidase genes as biomarkers for early diagnosis of CRC (i.e., adenoma vs. healthy) and for CRC diagnosis (i.e., CRC vs. healthy). In constructing and evaluating a random forest model to distinguish between CRC cases and healthy controls, 258 samples (DRA006684 and DRA008156) from the CRC group and 251 samples from the healthy control group were first collected. Using the sample function in R software with no replacement, 80% of the samples were randomly selected as the training set and 20% as the test set. Subsequently, based on the training set samples, 38 significantly different microbial β-glucuronidase genes were included as feature variables, and a random forest model was constructed using the randomForest package. For the constructed model, ROC curves were plotted on both the training and test sets using the ROSE package, and the AUC was calculated using the caret package. The same strategy was used to construct the model to distinguish between adenoma cases and healthy controls as for the CRC group model.

[0114] The model distinguishing between CRC cases and healthy controls was evaluated on samples. The AUC on the training set (80% of samples) was 0.855, and the AUC on the test set (20% of samples) was 0.863. The model distinguishing between adenoma cases and healthy controls was evaluated on samples. The AUC on the training set (80% of samples) was 0.838, and the AUC on the test set (20% of samples) was 0.816. Figure 2 (A to B).

[0115] Example 3: Predicting the risk of CRC in subjects based on 38 β-glucuronidase gene markers A validation cohort of 218 subjects was constructed from samples collected from Shenzhen People's Hospital. These 218 samples were divided into three groups: a healthy control group (60 samples), an adenoma case group (8 samples), and a CRC case group (150 samples). Basic clinical information regarding age, sex, and BMI of the validation cohort is summarized in Tables 6 and 7.

[0116] Table 6 Basic Clinical Information of the Validation Cohort

[0117] Note: The distribution values ​​of age and BMI in each group represent mean ± SD. Gender (F:M) represents the female:male sample size.

[0118] Table 7 Results of intergroup differences in basic clinical information of the validation cohort.

[0119] Note: The p-value was obtained by using the rank-sum test for age and BMI and the chi-square test for gender.

[0120] Fecal samples were collected from the subjects, and DNA was extracted from the samples for metagenomic sequencing. After quality control and host removal, the raw sequencing sequences were aligned to a microbial gene set containing 10,947,761 genes to obtain the abundance information of these genes in the subjects. Optionally, based on the metagenomic sequencing data, reassembly and gene prediction could be performed, and 38 β-glucuronidase gene sequences could be merged. Through redundancy removal and extraction of representative sequences, a new microbial gene set and abundance information could be obtained.

[0121] Abundance information of 38 β-glucuronidase gene markers screened in Example 1 was extracted. Based on the abundance information of β-glucuronidase gene markers and the aforementioned constructed random forest model, the risk of CRC in the subjects was predicted.

[0122] The results showed that, when evaluating samples from the model distinguishing between CRC case groups and healthy controls, the subject AUC was 0.844. When evaluating samples from the model distinguishing between adenoma case groups and healthy controls, the subject AUC was 0.807. Figure 3 ).

[0123] In summary, the microbial β-glucuronidase gene markers obtained by this invention showed good diagnostic capabilities in both the CRC classification model and the adenoma classification model. This indicates that the microbial β-glucuronidase gene based on these differences can be used to detect CRC subjects, assist in the early screening and diagnosis of CRC, and provide another powerful method for CRC diagnosis.

[0124] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made within the scope of knowledge possessed by those skilled in the art without departing from the spirit of the present invention. Furthermore, the embodiments of the present invention and the features thereof can be combined with each other unless otherwise specified.

Claims

1. A combination of biomarkers comprising β-glucuronidase genes derived from various microorganisms, said microorganisms including Bacteroides ovatus (Ovobacteroides) Bacteroides nordii (Bacteroides norsi) Oscillospiraceae bacteria (Bacteria of the family Oscillatoria) Faecalibacterium prausnitzii (Clostridium plasminogen lysate) Parabacteroides chongii (Pseudomonas chrysogenum) Hungatella hathawayi (Hund's bacillus) Echinicola strongylocentrotii (Echinococcus spp.) Bacteroides cellulolyticus (Bacteroides cellulose-degrading) Parabacteroides goldsteinii (Pseudomonas gondii) Enteroclostrum clostridial (Enterobacter fusiformis) Bacteroides helcogenes (Bacteroides spirochetes) Bifidobacterium bifidum (Bifidobacterium bifidum) Bacteroides faecium (Bacteroides foetida) Stutzerimonas stutzeri (Pseudomonas stearothermia) Ruminococcus torques (Ruminococcus twitchis) Lachnospira eligens (Selected *Trichophyton spp.*) Bacteroides thetaiotaomicron (Bacteroides polymorpha) Long-chained dorea (Long-chain Doraemon) Cellulosimicrobium cellulans At least two of the (fibrinolytic microbacteria).

2. The marker combination according to claim 1, characterized in that, The source Bacteroides ovatus The β-glucuronidase gene of (Bacteroides ovalis) includes at least one of GUS3 and GUS4; The source Bacteroides nordii The β-glucuronidase genes of Bacteroides noriensis include at least one of GUS1, GUS2, GUS3, GUS4, GUS5, and GUS6. The source Oscillospiraceae bacteria The β-glucuronidase gene of (Oscillatoriaceae bacteria) includes at least one of GUS6 and GUS7; The source Faecalibacterium prausnitzii The β-glucuronidase gene of (Clostridium praosporum) includes at least one of GUS3, GUS7, and GUS10; The source Parabacteroides chongii The β-glucuronidase gene of *Pseudomonas chinensis* includes at least one of GUS1 and GUS7. The source Hungatella hathawayi The β-glucuronidase gene of *Hunnerella haematobium* includes GUS1; The source Echinicola strongylocentrotii The β-glucuronidase gene of *Echinococcus spp.* includes GUS; The source Bacteroides cellulolyticus The β-glucuronidase gene of *Bacteroides cellulosicus* includes at least one of GUS1, GUS2, GUS3, GUS5, GUS15, and GUS18. The source Parabacteroides goldsteinii The β-glucuronidase gene of *Pseudomonas gossypii* includes at least one of GUS2, GUS3, and GUS6; The source Enterocloster clostridioformis The β-glucuronidase gene in *Enterobacter* includes GUS2; The source Bacteroides helcogenes The β-glucuronidase gene of *Bacteroides spirulinae* includes GUS2; The source Bifidobacterium bifidum The β-glucuronidase gene of Bifidobacterium bifidum includes at least one of GUS1 and GUS2; The source Bacteroides faecium The β-glucuronidase gene of (Bacteroides foetida) includes GUS7; The source Stutzerimonas stutzeri The β-glucuronidase gene of *Pseudomonas stearothermia* includes GUS; The source Ruminococcus torques The β-glucuronidase gene of *Ruminococcus tortifolius* includes GUS; The source Lachnospira eligens The β-glucuronidase gene of (selected spirulina) includes GUS2; The source Bacteroides thetaiotaomicron The β-glucuronidase gene of *Bacteroides multiforme* includes GUS4; The source Long-chained dorea The β-glucuronidase gene of (Long Chain Doraemon) includes at least one of GUS1 and GUS2; The source Cellulosimicrobium cellulans The β-glucuronidase gene of (fibrotic microbacteria) includes GUS.

3. a combination of logos according to 2, characterized by The source Bacteroides ovatus The nucleic acid molecules of the β-glucuronidase genes GUS3 and GUS4 of *Bacteroides ovalis* are, respectively, the nucleic acid molecules encoding the β-glucuronidases shown in SEQ ID NO:1 and SEQ ID NO:2; and / or, The source Bacteroides nordii The nucleic acid molecules of the β-glucuronidase genes GUS1, GUS2, GUS3, GUS4, GUS5, and GUS6 of *Bacteroides nigra* are, respectively, the nucleic acid molecules encoding the β-glucuronidases shown in SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:18, SEQ ID NO:19, and SEQ ID NO:20; and / or, The source Oscillospiraceae bacteria The nucleic acid molecules of the β-glucuronidase genes GUS6 and GUS7 of (bacteria of the family Oscillatoria) are, respectively, the nucleic acid molecules encoding the β-glucuronidases shown in SEQ ID NO:15 and SEQ ID NO:16; and / or, The source Faecalibacterium prausnitzii The nucleic acid molecules of the β-glucuronidase genes GUS3, GUS7, and GUS10 of *Clostridium plasmidonum* are, respectively, nucleic acid molecules encoding the β-glucuronidases shown in SEQ ID NO:17, SEQ ID NO:26, and SEQ ID NO:31; and / or, The source Parabacteroides chongii The nucleic acid molecules of the β-glucuronidase genes GUS1 and GUS7 of *Pseudomonas chinensis* are, respectively, the nucleic acid molecules encoding the β-glucuronidases shown in SEQ ID NO:22 and SEQ ID NO:29; and / or, The source Hungatella hathewayi The nucleic acid molecule of the β-glucuronidase gene GUS1 of *Hunnerella harzianum* is the nucleic acid molecule encoding the β-glucuronidase shown in SEQ ID NO:23; and / or, The source Echinicola strongylocentroti The nucleic acid molecule of the β-glucuronidase gene GUS of *Echinococcus spp.* is the nucleic acid molecule encoding the β-glucuronidase shown in SEQ ID NO:25; and / or, The source Bacteroides cellulosilyticus The nucleic acid molecules of the β-glucuronidase genes GUS1, GUS2, GUS3, GUS5, GUS15, and GUS18 of *Bacteroides cellulosicum* are, respectively, the nucleic acid molecules encoding the β-glucuronidases shown in SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:30, and SEQ ID NO:36; and / or, The source Parabacteroides goldsteinii The nucleic acid molecules of the β-glucuronidase genes GUS2, GUS3, and GUS6 of *Pseudomonas gondii* are, respectively, nucleic acid molecules encoding the β-glucuronidases shown in SEQ ID NO:13, SEQ ID NO:14, and SEQ ID NO:27; and / or, The source Enterocloster clostridioformis The nucleic acid molecule of the β-glucuronidase gene GUS2 of *Enterobacter* is the nucleic acid molecule encoding the β-glucuronidase shown in SEQ ID NO:28; and / or, The source Bacteroides helcogenes The nucleic acid molecule of the β-glucuronidase gene GUS2 of *Bacteroides spirochetes* is the nucleic acid molecule encoding the β-glucuronidase shown in SEQ ID NO:38; and / or, The source Bifidobacterium bifidum The nucleic acid molecules of the β-glucuronidase genes GUS1 and GUS2 of *Bifidobacterium bifidum* are, respectively, the nucleic acid molecules encoding the β-glucuronidases shown in SEQ ID NO:12 and SEQ ID NO:21; and / or, The source Bacteroides faecium The nucleic acid molecule of the β-glucuronidase gene GUS7 of *Bacteroides coccidioides* is the nucleic acid molecule encoding the β-glucuronidase shown in SEQ ID NO:33; and / or, The source Stutzerimonas stutzeri The nucleic acid molecule of the β-glucuronidase gene GUS of *Pseudomonas stearothermia* is the nucleic acid molecule encoding the β-glucuronidase shown in SEQ ID NO:34; and / or, The source Ruminococcus torques The nucleic acid molecule of the β-glucuronidase gene GUS of *Ruminococcus tortifolia* is the nucleic acid molecule encoding the β-glucuronidase shown in SEQ ID NO:11; and / or, The source Lachnospira eligens The nucleic acid molecule of the β-glucuronidase gene GUS2 of *Syngonium tumefaciens* is the nucleic acid molecule encoding the β-glucuronidase shown in SEQ ID NO:32; and / or, The source Bacteroides thetaiotaomicron The nucleic acid molecule of the β-glucuronidase gene GUS4 of *Bacteroides polymorpha* is the nucleic acid molecule encoding the β-glucuronidase shown in SEQ ID NO:10; and / or, The source Dorea longicatena The nucleic acid molecules of the β-glucuronidase genes GUS1 and GUS2 of *Dorhizium anisopliae* are, respectively, the nucleic acid molecules encoding the β-glucuronidases shown in SEQ ID NO:24 and SEQ ID NO:35; and / or, The source Cellulosimicrobium cellulans The nucleic acid molecule of the β-glucuronidase gene GUS of (fibrotic microbacteria) is the nucleic acid molecule encoding the β-glucuronidase shown in SEQ ID NO:

37.

4. The use of the combination of markers according to any one of claims 1 to 3 in the preparation of diagnostic drugs for diseases including colorectal cancer and / or adenoma.

5. The application according to claim 4, characterized in that, The substance includes substances that detect the combination of the markers at the gene or protein level; Preferably, the substance used to detect the combination of biomarkers at the protein level is selected from substances of one or more detection methods from the group consisting of: chemiluminescence, immunofluorescence, protein chip, proteometry, immunohistochemistry, patch tracing based on labeling technology, Western blotting, and enzyme-linked immunosorbent assay (ELISA). Preferably, the substance used to detect the combination of biomarkers at the gene level is selected from substances selected by one or more detection methods from the group consisting of: DNA sequencing, RNA sequencing, RNA-in situ hybridization, digital PCR, and quantitative real-time PCR. Preferably, the product includes reagents, kits, test strips, systems, or chips.

6. The application according to claim 4 or 5, characterized in that, The diagnosis includes at least one of early diagnosis of colorectal cancer, staging diagnosis of colorectal cancer, and diagnosis of adenoma.

7. A method for constructing a model for disease diagnosis, comprising the following steps: Model construction was performed using the combination of biomarkers according to any one of claims 1 to 3; the disease includes colorectal cancer and / or adenoma; Preferably, the construction method includes at least one of random forest, log regression, linear discriminant analysis, support vector machine, Bayesian network, KthNearestNeighbour, and decision tree.

8. A model execution module for disease diagnosis, characterized in that: The model execution module includes a computing component; The computing component calculates the data of the marker combination according to any one of claims 1 to 3; The diseases mentioned include colorectal cancer and / or adenoma.

9. A system for disease diagnosis, characterized in that: The system includes at least one of 1) to 3): 1) Data input module; 2) Model execution module; 3) Result output module; The model running module includes the model running module for disease diagnosis as described in claim 8.

10. An electronic device comprising a storage device, a processor, and a computer program stored in the storage device and executable on the processor, characterized in that: The computer program stored on the memory and capable of running on the processor includes the system of claim 9.