Tissue resident bacterium marker for identifying normal colorectal tissue, polyp and intestinal cancer and application of tissue resident bacterium marker
By using large-sample sequencing analysis and machine learning models, and utilizing tissue-resident bacterial markers such as Pseudomonas, the accuracy problem of differentiating normal colorectal tissue, polyps, and colorectal cancer has been solved, enabling early detection and treatment guidance and improving the accuracy of identification.
Patent Information
- Application Number
- CN202510916593.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-11-14
AI Technical Summary
Current technologies are insufficient to accurately distinguish between normal colorectal tissue, polyps, and colorectal cancer at the microbial level, leading to a high possibility of missed diagnoses during colonoscopy biopsies and a lack of effective early detection and treatment methods.
Through large-sample sequencing analysis, tissue-resident bacterial markers such as Pseudomonas, Pigmentiphaga, Fusobacterium, Parabacteroides, Bacteroides, Collinsella, and Parvimonas were identified and utilized. Combined with 16S sequencing, whole-genome sequencing, and other methods, machine learning models were constructed for identification.
It significantly improved the accuracy of differentiating normal colorectal tissue, polyps, and colorectal cancer by 12.5% and 15.3%, respectively, solving the problem of sample heterogeneity and providing a new approach for early detection and treatment.
Smart Images

Figure CN120945033A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedical technology, specifically relating to tissue resident bacterial markers for identifying normal colorectal tissue, polyps and colorectal cancer and their applications. Background Technology
[0002] The diagnosis and treatment of colorectal cancer remain challenging. Early detection and treatment of polyps and colorectal cancer are particularly crucial. However, due to tissue heterogeneity, current colonoscopy biopsy methods still have a certain possibility of missed diagnoses, and the search for new biomarkers for the early detection of polyps and colorectal cancer remains a long and arduous task.
[0003] Alterations in the human gut microbiome are increasingly being used as biomarkers for disease pre-screening and diagnosis, as well as targets for disease treatment and intervention. Growing evidence suggests the importance of the tumor microbiome [Garrett WS. Cancer and the microbiota. Science. 2015 Apr 3; 348(6230):80-6]. Despite differences in microbiota between normal and colorectal cancer tissues, large-scale studies analyzing the characteristics of tissue-resident bacteria to aid clinical diagnosis and treatment are still lacking [Curry KD. It takes guts to learn: machine learning techniques for disease detection from the gut microbiome. Emerg Top Life Sci. 2021 Dec 21; 5(6):815-82], and corresponding identification methods are also lacking.
[0004] Therefore, at the microbial level, differentiating normal intestinal tissue, polyps, and colorectal cancer may be used to guide clinical diagnosis and treatment, providing a new approach for the early detection and treatment of polyps or colorectal cancer. Summary of the Invention
[0005] The purpose of this invention is to provide tissue resident bacterial markers for differentiating normal colorectal tissue, polyps, and colorectal cancer through large-sample sequencing analysis. These markers can reflect the biological specificity of patients from a host-microbe perspective, and can more accurately identify normal colorectal tissue, polyps, and colorectal cancer. They are expected to guide clinical diagnosis and treatment, enabling early detection and effective treatment, and potentially significantly improving patient prognosis.
[0006] The technical solution adopted in this invention is:
[0007] In the first aspect, a combination of markers is provided, including Pseudomonas, Pigmentiphaga, Fusobacterium, Parabacteroides, Bacteroides, and Collinsella.
[0008] The second aspect provides a combination of markers, including Pseudomonas, Pigmentiphaga, Fusobacterium, and Parvimonas.
[0009] Thirdly, the invention provides the application of the combination of markers described in the first or second aspect of the invention or the reagent for detecting the combination of markers described in the first or second aspect of the invention in the preparation of a product, wherein the product has the function of at least one of the following:
[0010] (a) Differentiate between normal intestinal tissue and colorectal polyps.
[0011] (b) Differentiate between colorectal polyps and colorectal cancer.
[0012] The reagents include those for detecting relative abundance using one or more of the following methods: 16S sequencing, whole-genome sequencing, quantitative polymerase chain reaction, PCR-pyrosequencing, fluorescence in situ hybridization, microarray, and PCR-ELISA. The reagents include primers, probes, antisense oligonucleotides, aptamers, or antibodies. The sequences of the primers are as follows:
[0013]
[0014] A fourth aspect of the present invention provides a product comprising the reagents used in the applications described in the third aspect of the present invention. The product includes reagents, kits, test strips, or chips.
[0015] A fifth aspect of the present invention provides a method for constructing a model for differentiating normal intestinal tissue from colorectal polyps, and / or for differentiating colorectal polyps from colorectal cancer, the method comprising the step of identifying differentially expressed substances in samples between patients and healthy controls, said differentially expressed substances being a combination of biomarkers described in the first aspect of the present invention, and / or a combination of biomarkers described in the second aspect.
[0016] A sixth aspect of the present invention provides a detection system for distinguishing between normal intestinal tissue and colorectal polyps, and / or for distinguishing between colorectal polyps and colorectal cancer, comprising:
[0017] (a) Detection module: used to receive the relative abundance of markers in the marker combination described in the first and / or second aspects of the present invention as measured in the sample, and output the relative abundance data of each marker to the analysis output module;
[0018] (b) Analysis Output Module: Based on the trained model used to distinguish between normal intestinal tissue and colorectal polyps, and / or to distinguish between colorectal polyps and colorectal cancer, the module calculates a risk score for each sample to classify it. The input data is typically in CSV format, containing relative abundance information for each bacterial community.
[0019] According to a specific embodiment of the present invention, the criteria for classifying and judging samples are as follows: The model outputs a probability (a value between 0 and 1). Based on a first-aspect combined threshold suggestion: when the score is greater than 0.5, it is predicted as "polyp". Based on a second-aspect combined threshold suggestion: when the probability is greater than 0.5, it is predicted as "colorectal cancer".
[0020] A seventh aspect of the present invention provides a computing device comprising: at least one processing unit and at least one memory, the memory being coupled to the processing unit and storing instructions for execution by the processing unit, wherein, when executed, the device is capable of analyzing at least one of the following tissue conditions: normal, polyp, and colorectal cancer tissue, comprising the following steps:
[0021] (a) Receive the relative abundance of markers in the marker combination described in the first and / or second aspects of the present invention as measured in the sample, and output the relative abundance data of each marker for subsequent analysis output;
[0022] (b) Analysis Output: Based on the trained model used to distinguish between normal intestinal tissue and colorectal polyps, and / or to distinguish between colorectal polyps and colorectal cancer, a risk score is calculated for each sample to classify and distinguish them. The beneficial effects of this invention are:
[0023] This invention is the first to discover a set of tissue-resident bacterial markers for differentiating normal colorectal tissue, polyps, and colorectal cancer through large-sample sequencing analysis. These markers include *Pseudomonas*, *Pigmentiphaga*, *Fusobacterium*, *Parabacteroides*, *Bacteroides*, *Collinsella*, and *Parvimonas*. The relative abundance of these seven tissue-resident bacteria forms a marker that reflects the biological specificity of patients. A combined model of six bacteria (*Pseudomonas*, *Pigmentiphaga*, *Fusobacterium*, *Parabacteroides*, *Bacteroides*, and *Collinsella*) can more accurately distinguish between normal colorectal tissue and polyps, improving accuracy by 12.5% compared to existing non-invasive screening methods. A combined model of four bacteria (*Pseudomonas*, *Pigmentiphaga*, *Fusobacterium*, and *Parvimonas*) can more accurately distinguish between polyps and colorectal cancer, improving accuracy by 15.3% compared to existing non-invasive screening methods. Furthermore, this invention solves the problem of sample heterogeneity in biopsy-based diagnosis, and can be effectively and hopefully used to guide clinical diagnosis and treatment, providing a new approach for the early detection and treatment of polyps and colorectal cancer. Attached Figure Description
[0024] Figure 1 The AUC values of a random forest machine learning model based on tissue-resident bacteria show that normal tissue and polyp tissue can be accurately distinguished using Pseudomonas, Pigmentiphaga, Fusobacterium, Parabacteroides, Bacteroides, and Collinsella. Figure 1 A) Pseudomonas, Pigmentiphaga, Fusobacterium, and Parvimonas can accurately distinguish between polyp tissue and colorectal cancer tissue. Figure 1 B).
[0025] Figure 2 Training cohorts based on tissue-resident bacteria ( Figure 2 A) and verification queue ( Figure 2 B) ROC curve of machine learning random forest model. Detailed Implementation
[0026] The following will describe the concept and technical effects of the present invention clearly and completely with reference to embodiments, so as to fully understand the purpose, features and effects of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are all within the scope of protection of the present invention.
[0027] 1. In this study, we collected 710 colorectal cancer tissue samples, 227 normal tissue samples, and 85 polyp tissue samples, and extracted DNA from the tissues using the QIAamp PowerFecal(pro) DNAkit (QIAGEN#51804) in conjunction with a tissue disruptor.
[0028] 2. To ensure the success rate of subsequent sequencing, the 16S fragment (bacterial content) in the sample DNA must first be quantitatively detected by qPCR. Samples with negative qPCR results should be discarded, and the remaining samples should be enriched and amplified with the 16S V4 fragment, as follows:
[0029] a. Synthetic biotin-labeled 16S V4 fragment primers (containing sequencing adapter sequences),
[0030]
[0031] b. Perform PCR amplification for 25-30 cycles using the kit required by the sequencer (the number of cycles needs to be determined based on the bacterial content in the tissue);
[0032] c. Enrichment and precipitation of biotin-tagged amplified fragments using Dynabeads MyOne Streptavidin C1 Beads (Thermo 65002).
[0033] 3. The enriched and precipitated fragments (without eluting the magnetic beads) are then used to construct sequencing libraries.
[0034] 4. Identifying biomarkers for tissue-resident bacteria:
[0035] Effect size analysis (LEfSe) was performed using the microeco package (v0.11.0) in R software (Zhang et al., 2023). This analysis employed a three-stage approach to determine the phylogenetic differences between the tumor microbiome of normal tissues and colorectal cancer (CRC):
[0036] 1) The Kruskal-Wallis test (Kruskal & Wallis, 1952) examined the difference in abundance at the genus level between groups (α = 0.25).
[0037] 2) Pairwise Wilcoxon rank-sum test (Wilcoxon, 1945) confirmed subgroup consistency;
[0038] 3) Linear discriminant analysis (LDA) quantified the effect size (log10 LDA score > 2.0). Low abundance features were filtered out before analysis (prevalence threshold: 5% of the sample by filter_thres = 0.05). The top 20 taxa with the highest LDA scores were selected for visualization, and their relative abundance distribution was labeled. Multiple test corrections were performed using the Benjamini-Hochberg procedure (Benjamini & Hochberg, 1995) during feature selection. Microbial characterization was performed to determine genus-level abundance patterns associated with clinical phenotypes. The top 30 genera with the highest abundance across all samples were selected based on cumulative relative abundance. The group-specific mean of log10(x+1) transformed abundance for each group was calculated.
[0039] 5. Modeling process based on random forest of tissue-resident bacteria
[0040] Before machine learning modeling, the microbiome data was preprocessed using the metagenomeSeq package (v1.44.0) (Paulson et al., 2013). First, samples with low prevalence (<5%) were filtered out by removing genera with high occurrence frequency (prevalence threshold = 0.05) to reduce noise caused by sparsity. Then, the retained microbial features were normalized using cumulative scaling (CSS) normalization through a three-stage processing. The final normalized matrix was used as input to the downstream machine learning pipeline. Initially, the dataset was randomly split into a training set (80%) and a test set (20%). For each classification problem, models were built using the caret (Kuhn, 2008) and randomForest (Liaw & Wiener, 2002) R packages. The training set underwent three iterations of 5-fold cross-validation to optimize model performance, and the resulting models were then validated on the test set. The importance of features was evaluated using an average reduction in precision metric derived from random forest analysis. This ranking was then combined with pairwise correlations between microbial occurrence and taxa. We iteratively add candidate classes, build models, and evaluate their performance on training and testing datasets, ultimately selecting a limited number of key microbial genera. In binary classification scenarios, we calculate the area under the receiver operating characteristic (AUC) curve based on the cross-validation results of the training and independent test sets, and visualize the ROC curve.
[0041] 6. The efficacy of tissue-resident bacteria in classifying normal tissues, polyps, and colorectal cancer tissues.
[0042] The classification efficacy of the aforementioned tissue resident bacterial markers was validated in 182 normal tissue samples, 68 polyp samples, and 568 colorectal cancer samples in the training cohort, and in 45 normal tissue samples, 17 polyp samples, and 142 colorectal cancer samples in the validation cohort. Application models for differentiating between normal tissue, polyp tissue, and polyp tissue and colorectal cancer tissue were constructed as follows: Figure 1 As shown in A and 1B. The training and validation queues of the model input the corresponding microbial abundance of the above samples into the model. In the training queue, the accuracy in distinguishing between normal tissue and polyp tissue was 0.911, and the accuracy in distinguishing between polyps and colorectal cancer was 0.982. Figure 2 A), the accuracy in distinguishing between normal tissue and polyp tissue in the validation cohort was 0.954, and the accuracy in distinguishing between polyps and colorectal cancer was 1 ( ). Figure 2 B). The results showed that the accuracy of the above-mentioned model of resident bacteria in distinguishing normal tissues and polyps was 12.5% higher than that of the model based solely on gut bacteria in the existing literature [Xiang J. Alterations of Gut Microbiome in Patients with Colorectal Advanced Adenoma by Metagenomic Analyses. Turk J Gastroenterol. 2024 Nov 1; 35(11):859-868.], and the accuracy of distinguishing polyps and colorectal cancer tissues was 15.3% higher than that of the model based solely on gut bacteria in the existing literature [Yang Y. Dysbiosis of human gutmicrobiome in young-onset colorectal cancer. Nat Commun. 2021 Nov 19; 12(1):6757.].
[0043] The above detailed embodiments have provided a comprehensive description of the present invention. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention. Furthermore, unless otherwise specified, the embodiments of the present invention and the features thereof can be combined with each other.
Claims
1. A combination of biomarkers comprising Pseudomonas, Pigmentiphaga, Fusobacterium, Parabacteroides, Bacteroides, and Collinsella.
2. A combination of biomarkers comprising Pseudomonas, Pigmentiphaga, Fusobacterium, and Parvimonas.
3. The use of the biomarker combination or reagent for detecting the biomarker combination of claim 1 or 2 in the preparation of a product, wherein the product functions to include at least one of the following: (a) Differentiating normal intestinal tissue from colorectal polyps; (b) Differentiate between colorectal polyps and colorectal cancer.
4. The application according to claim 3, characterized in that, The reagent is used to detect the relative abundance of the biomarker combination described in claim 1 or 2.
5. The application according to claim 4, characterized in that, The reagents include those for detecting relative abundance using one or more of the following methods: 16S sequencing, whole genome sequencing, quantitative polymerase chain reaction, PCR-pyrosequencing, fluorescence in situ hybridization, microarray, and PCR-ELISA.
6. The application according to claim 5, characterized in that, The reagent includes one or more of primers, probes, antisense oligonucleotides, aptamers, or antibodies; the sequence of the primer is:
7. A product comprising the reagent used in any one of claims 2 to 5.
8. The product according to claim 6, characterized in that, The products mentioned are reagents, reagent kits, test strips, or chips.
9. A method for constructing a model for differentiating normal intestinal tissue from colorectal polyps, and / or for differentiating colorectal polyps from colorectal cancer, characterized in that, The method includes the step of identifying differentially expressed substances in samples between patients and healthy controls, said differentially expressed substances being the biomarker combination of claim 1 and / or the biomarker combination of claim 2.
10. A detection system for differentiating normal intestinal tissue from colorectal polyps, and / or for differentiating colorectal polyps from colorectal cancer, characterized in that, include (a) Detection module: used to receive the relative abundance of markers in the marker combination as described in claim 1 and / or claim 2 as measured in the sample, and output the relative abundance data of each marker to the analysis output module; (b) Analysis output module: Calculates the risk of samples based on the trained model used to distinguish between normal intestinal tissue and colorectal polyps, and / or to distinguish between colorectal polyps and colorectal cancer, in order to classify and distinguish the samples.
Citation Information
Patent Citations
Microbial marker of colorectal cancer and application of marker
CN109943636A
Intestinal cancer biomarker composition and application thereof
CN110512015A
Microbial marker related to colorectal cancer and application thereof
CN112410449A
Microbial markers for predicting risk of colorectal cancer and application thereof
CN112609015A
Colorectal cancer biomarker for overweight population, kit and early screening model
CN117025798A