Identification of immune-related microorganisms in colorectal cancer based on multi-omics signatures
By screening intestinal microorganisms with immune regulatory functions in colorectal cancer through methods such as LASSO, SVM-RFE and WGCNA, combined with CIBERSORTx analysis, the problem of difficulty in identifying intestinal microorganisms in colorectal cancer in existing technologies was solved, and efficient intestinal microbial diagnosis and the formulation of personalized treatment plans were achieved.
Patent Information
- Application Number
- CN202211313242.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-25
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2042-10-25
AI Technical Summary
Existing technologies make it difficult to efficiently identify intestinal microorganisms with immune regulatory functions in colorectal cancer through traditional statistical analysis methods, and are unable to systematically analyze complex intestinal microbial data to identify key intestinal flora, which affects the immunotherapy and diagnosis of colorectal cancer.
The LASSO regression algorithm and SVM-RFE method were used to screen key intestinal microorganisms, combined with the WGCNA method for module division and immune correlation analysis. The CIBERSORTx algorithm was used to evaluate the level of immune cell infiltration, and the mechanism of action of intestinal microorganisms was predicted by biological function to construct a colorectal cancer diagnostic model.
Accurately identifying intestinal microorganisms with immune regulatory functions in colorectal cancer provides personalized precision medication plans, improving the accuracy of colorectal cancer diagnosis and the ability of systems biology analysis.
Smart Images

Figure CN115472214B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for identifying immune-related microorganisms in colorectal cancer based on multi-omics characteristics, and belongs to the technical field of biomedicine. Background Art
[0002] Colorectal cancer is one of the most common malignancies. According to cancer statistics in the United States, over 100,000 new cases of colorectal cancer are diagnosed annually, ranking third among all malignant tumors in both morbidity and mortality. In recent years, with the rapid economic growth of my country, changes in dietary habits and the environment have led to a significant increase in the incidence and mortality of colorectal cancer, with a trend toward younger age. Risk factors for colorectal cancer include genetic and environmental factors, such as age, gender, obesity, chronic inflammation, lack of physical activity, and diet. The pathogenesis of colorectal cancer, whether congenital or acquired, is multifactorial and multistep. Increasing evidence indicates that chronic inflammation, host genetic susceptibility, and environmental factors are associated with the progression of colorectal cancer. Among these environmental factors, the role of the gut microbiome in cancer biology is gaining increasing attention. The gut microbiome is a complex microbial community composed of over 1,000 bacterial species. The intestinal bacterial community is primarily concentrated in the colorectum and is primarily composed of Bacteroides and Firmicutes, with other common species including Actinobacteria, Proteobacteria, and Verrucomicrobia. Studies have shown that the intestinal microbiome can influence the onset and progression of colorectal cancer by regulating systemic and local immune function. Therefore, studying the intestinal microbiome and its role in tumor immunoregulation will help identify immunotherapy targets for colorectal cancer. Currently, some studies have compared the fecal microbiome composition of colorectal cancer patients with that of healthy controls through bioinformatics analysis, identifying some tumor-associated bacterial species. However, only a few intestinal microbes have been shown to be involved in regulating the immune microenvironment, and the exact mechanisms of immunoregulatory effects are currently unclear. Therefore, identifying and studying intestinal microbes with key biological functions in colorectal cancer and the possible molecular mechanisms are of great significance for a deeper understanding of this disease and may also provide new insights into the immunotherapy and evaluation of colorectal cancer.
[0003] 16S rRNA amplicons sequencing and shotgun metagenomic sequencing are both important methods for studying the gut microbiome. With the advancement of second- and third-generation sequencing technologies, gut microbiome analysis has generated a vast amount of data. Traditional statistical analysis methods face a significant challenge in interpreting this complex data: how to convert this vast amount of omics data into biologically meaningful information. First, traditional statistical analysis methods struggle to identify microorganisms with small fold changes in abundance, leading to biased analysis results. Second, traditional analysis methods are more suited to pairwise comparisons of samples and are unable to systematically integrate high-throughput data from multiple sources. In addition to traditional unsupervised and supervised methods, a recently emerging method, weighted gene co-expression network analysis (WGCNA), has also been applied to metagenomic data analysis. WGCNA, an algorithm for mining module information from high-throughput data, overcomes these two limitations. WGCNA was originally used for transcriptome data analysis. In this method, a module is defined as a group of genes with similar expression profiles. If certain genes consistently show similar expression changes in a physiological process or across different tissues, it is reasonable to assume that these genes are functionally related and can be defined as a module. This seems somewhat similar to the results obtained through cluster analysis, but the difference is that the clustering criteria used in WGCNA are biologically meaningful, rather than conventional clustering methods, resulting in higher confidence in the results obtained.
[0004] Currently, studies have used methods such as correlation analysis to identify gut microbes that regulate the immune microenvironment. However, no studies have used WGCNA to identify key gut microbes that potentially regulate the immune microenvironment in colorectal cancer. Therefore, we propose a research method to identify key gut microbes with immunomodulatory effects in colorectal cancer, laying the foundation for understanding the mechanisms of action of the gut microbiome, discovering key gut microbes, and developing personalized precision medicine regimens. Summary of the Invention
[0005] The purpose of the present invention is to provide a bioinformatics analysis method and model for identifying intestinal microorganisms with immune regulatory functions in colorectal cancer based on multi-omics characteristics in response to the shortcomings of the existing technology. This method can screen out key intestinal microorganisms with immune regulatory functions in colorectal cancer, and the model can predict the immune regulatory functions and fine molecular mechanisms of action of intestinal microorganisms, laying the foundation for analyzing the mechanism of action of intestinal flora, discovering key intestinal bacteria, and formulating personalized precision medication plans.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is to provide a method for identifying immune-related microorganisms in colorectal cancer based on multi-omics features, including analysis of the diversity and composition of the intestinal flora in colorectal cancer patients and healthy controls; identifying intestinal microorganisms significantly associated with colorectal cancer patients through the LASSO regression algorithm; selecting characteristic intestinal microorganisms of colorectal cancer by using the SVM-RFE method; obtaining overlapping intestinal microorganisms between the two algorithms, and incorporating the overlapping intestinal microorganisms into ROC and DCA analysis verification.
[0007] Preferably, the method includes screening key intestinal microorganisms related to immune regulation; the screening adopts a screening method for differential abundance intestinal microorganisms, a WGCNA method and / or an intestinal microbial interaction network analysis method.
[0008] Preferably, the WGCNA method includes performing cluster analysis on samples from colorectal cancer patients to eliminate outlier samples; using intestinal microbial abundance data to calculate the correlation coefficients between intestinal microorganisms and intestinal microorganisms, and constructing a correlation matrix of intestinal microbial co-expression; by selecting a soft threshold parameter β, the correlation matrix is converted into an adjacency matrix, and then the adjacency matrix is converted into a topological overlap matrix, and the intestinal microorganisms are divided into modules, and intestinal microorganisms with the same function are divided into one module; the sample phenotypic data and the intestinal microbial module eigenvalues are associated to construct an intestinal microbial module-trait correlation heat map, and significant microbial modules and core (hubs) microorganisms that are highly correlated with the sample phenotypic traits are obtained.
[0009] Preferably, the method includes predictive analysis of the biological functions of immune-related intestinal microorganisms, analysis of immune-related intestinal microorganisms and immune microenvironment, and prediction of the mechanism of action of immune-related intestinal microorganisms.
[0010] Preferably, the immune-related intestinal microbial function prediction analysis includes using intestinal microbial abundance data and gene expression profile data to calculate the correlation coefficient between intestinal microorganisms and genes, and constructing a correlation matrix between intestinal microorganisms and genes; performing GO and KEGG analysis on the genes in the matrix to predict the biological functions of intestinal microorganisms.
[0011] Preferably, the immune-related microorganism and immune microenvironment analysis includes deconvolution analysis using the CIBERSORTx algorithm; using the transcriptome expression profile of cancer tissue, quantifying the relative infiltration levels of 25 immune cells in the tumor and surrounding tissues of each patient by LM22 markers and 1,000 sampling tests; using the Spearman algorithm to perform co-expression analysis of the intestinal microbial abundance matrix and the immune cell abundance matrix to screen out a certain number of intestinal microorganisms that are strongly correlated with immune cell infiltration.
[0012] Preferably, the prediction of the mechanism of action of immune-related intestinal microorganisms includes obtaining immune regulatory genes; using intestinal microbial abundance data and expression spectrum data of immune regulatory genes to calculate the correlation coefficient between intestinal microorganisms and genes, and constructing a correlation matrix between intestinal microorganisms and immune regulatory genes; based on the correlation matrix between intestinal microorganisms and immune regulatory genes, constructing an intestinal microorganism-immune regulatory gene correlation heat map to obtain the mechanism of action of immune-related intestinal microorganisms to exert biological effects.
[0013] The present invention also provides an intestinal flora-based colorectal cancer diagnostic model for use in preparing a reagent for detecting colorectal cancer. The model detects colorectal cancer by detecting 28 overlapping characteristic intestinal microorganisms; the overlapping characteristic intestinal microorganisms are Lactobacillus, Lachnospira, Ruminiclostridium, Collinsella, Lachnoclostridium, Paludibacter, Sutterella, Parabacteroides, Bacillus, Ruminococcaceae_UCG, Holdemanella, Lactonifactor, Allisonella, Anaerofustis, Coriobacterium, Fusobacterium, Neisseria, Lachnoanaerobaculum, Selenomonas, Shuttleworthia, Coprobacillus, Intestinimonas, Ruminococcus, Eisenbergiella, Eubacterium, Catenibacterium, Mogibacterium, and Coprococcus.
[0014] Compared with the prior art, the present invention has the following beneficial effects:
[0015] 1. To screen key gut microbial diagnostic markers, we used a combination of the least absolute shrinkage and selection operator (LASSO) algorithm and the support vector machine-recursive feature elimination (SVM-RFE) algorithm. Receiver operating characteristic curve (ROC) and decision curve analysis (DCA) were then performed on the key gut microbes obtained by intersecting the results of the two methods to identify the optimal gut microbial model for diagnosing colorectal cancer.
[0016] 2. The present invention uses weighted gene co-expression network analysis (WGCNA) to analyze gut microbiome data, identify significantly different gut microbial modules and core (hubs) microorganisms; determine the association between gut microbial modules and specific phenotypes, and ultimately achieve the purpose of finding target microorganisms that regulate the tumor immune microenvironment. Through systems biology methods, the present invention can efficiently and accurately identify gut microorganisms that have a regulatory effect on the immune microenvironment of colorectal cancer tumors, and explain their molecular mechanism of action, laying the foundation for analyzing the mechanism of action of the gut microbiome, discovering key gut bacteria, and formulating personalized precision medication plans. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 To analyze the diversity and composition of intestinal microbiota in healthy controls (H) and patients with colorectal cancer (CRC).
[0018] Figure 2 To screen and validate the intestinal microbiota as diagnostic biomarkers.
[0019] Figure 3 To perform cluster analysis of intestinal flora based on WGCNA algorithm.
[0020] Figure 4 Analysis of interactions between intestinal microbiota modules.
[0021] Figure 5 The gut microbial network in each module identified by the WGCNA method.
[0022] Figure 6Figure 2 is a heat map showing the correlation between gut microbial modules and clinicopathological characteristics.
[0023] Figure 7 Heat map showing the correlation between immune cell infiltration levels and gut microbiota in CRC patients based on CIBERSORTx and MCP algorithms.
[0024] Figure 8 Heat map showing the correlation between immune cell infiltration levels and gut microbiota in CRC patients based on TIMER and xCell algorithms.
[0025] Figure 9 Functional annotation and pathway analysis of gut microbiome-related genes.
[0026] Figure 10 Correlation network analysis between immune regulatory genes and intestinal microorganisms. DETAILED DESCRIPTION
[0027] In order to make the present invention more clearly understood, preferred embodiments are described in detail below with reference to the accompanying drawings:
[0028] like Figure 1-10 As shown, the present invention provides a bioinformatics analysis method and model for identifying intestinal microorganisms with immune regulatory functions in colorectal cancer based on multi-omics features, wherein the method comprises the following steps:
[0029] 1) Data preprocessing step: Obtain two single-omics data, the sample's microbiome data and the gene transcriptome data, and perform data preprocessing;
[0030] 2) Cluster analysis step: integration and cluster analysis of the two single-omics data of microbiome data and gene transcriptome data;
[0031] 3) Establishment of diagnostic marker models based on intestinal microorganisms:
[0032] 1. Diagnostic marker screening:
[0033] LASSO, also known as the Lasso algorithm, can adjust the regression coefficients of independent variables, even compressing the coefficients of variables with minimal impact to zero, thereby reducing data overfitting and screening for relatively important variables. LASSO regression analysis was performed using the "glmnet" package in the R language to reduce data overfitting. The minimum lambda value was used as the optimal reference value to screen for intestinal microbes that are more critical for colorectal cancer diagnosis.
[0034] SVM-RFE is a generalized prediction model that performs binary classification of data in a supervised learning manner. It can filter relevant features and remove relatively unimportant feature variables, thereby obtaining higher classification performance. Through feature selection, redundant and noisy data are filtered, computing time is reduced, and classification accuracy is improved. In the present invention, RBF with nonlinear projection is used for model fitting. First, gamma and cost (cost) fitting models with different value ranges are used, and the parameters corresponding to the model with the smallest error in the 10-fold cross validation results in the training set are selected. Gamma is 0.001, and the cost parameter cost is 10 as the final prediction model parameters. Based on this parameter combination, the final prediction model is fitted to all samples.
[0035] 2. Evaluation of diagnostic marker models:
[0036] The most commonly used indicator for evaluating diagnostic marker models is the ROC curve. The area under the curve (AUC), also known as the C statistic, can directly evaluate the model's ability to diagnose colorectal cancer. The ROC curve is mainly used for the evaluation of binary classification models. It is based on the true category and predicted probability of the sample, with the false positive rate (FP) and true positive rate (TP) as the vertical coordinates to calculate the AUC value. Since ROC is not sensitive to whether the sample category is balanced, the ROC evaluation classification can be selected for the imbalanced samples between the research groups of the present invention, and the classifier model is trained by optimizing the AUC value.
[0037] Decision curve analysis (DCA) is a method for evaluating the net benefit of diagnostic methods or prediction models. It focuses on clinical utility and is a good supplement to ROC curve analysis.
[0038] 4) Screening of immune-related intestinal microorganisms:
[0039] Using one or more of the following methods: screening of differentially abundant intestinal microorganisms, WGCNA method, and intestinal microbial interaction network analysis;
[0040] The differential abundance intestinal microorganism screening method includes using the "edgeR" software package of R language to screen for differential abundance intestinal microorganisms, and setting the standard for differential intestinal microorganism screening to |log2(fold change)|≥1.0, false discovery rate (FDR)<0.05, where fold change is the difference in intestinal microorganism abundance between the colorectal cancer patient group and the healthy control group, and FDR is obtained by correcting the significant P value (Pvalue).
[0041] The WGCNA method includes the following steps:
[0042] a. Data cleaning: First, cluster analysis is performed on colorectal cancer patient samples to remove outlier samples;
[0043] b. Constructing a gut microbial interaction network: Using gut microbial abundance data, we calculated the correlation coefficients between gut microbes and constructed a gut microbial co-expression correlation matrix. To establish a scale-free network among gut microbes, we converted the correlation matrix into an adjacency matrix by selecting a soft threshold parameter β (weighting coefficient). The soft threshold parameter is generally required to achieve an R² > 0.8 or a soft threshold that minimizes the plateau phase, i.e., the correlation coefficient between the logarithm of the number of connected nodes (log(k))) and the logarithm of the probability of a node appearing (log(p(k)))) is greater than 0.8. The adjacency matrix is converted into a topological overlap matrix (TOM), and gut microbes are divided into modules. Microbes with the same function are grouped into one module.
[0044] c. Screening of trait-related modules: The sample phenotypic data were correlated with the characteristic values of the gut microbial module. The Spearman correlation coefficient and P value between the characteristic values of the gut microbial module and the sample phenotypic trait data were calculated by Spearman analysis. A gut microbial module-trait correlation heat map was constructed to obtain significant microbial modules and hub microorganisms that were highly correlated with the sample phenotypic traits.
[0045] 5) Functional prediction analysis of immune-related intestinal microorganisms.
[0046] Gene function enrichment analysis and immune microenvironment analysis include any one of GO and KEGG analysis based on differentially expressed genes, GSEA analysis, and immune microenvironment analysis.
[0047] a. Using intestinal microbial abundance data and tumor tissue gene expression profile data, calculate the correlation coefficient between intestinal microbes and genes, and construct a correlation matrix between intestinal microbes and genes; perform GO enrichment analysis and KEGG signaling pathway enrichment analysis on the genes in the matrix to predict biological processes involved in intestinal microbes. The GO database includes gene functions including cellular components, biological processes, and molecular functions, and the KEGG database includes a database established based on signaling pathway-related data. The GO and KEGG analysis based on intestinal microbe-related genes includes the following steps:
[0048] b. Use the "org.Hs.eg.db" package in the R language in Bioconductor to convert gene names;
[0049] c. Convert the gene name to the unified gene number used in the Entrez gene database;
[0050] d. Metascape (https: / / metascape.org / ) and DAVID (https: / / david.ncifcrf.gov / ) were then used to perform GO enrichment analysis and KEGG signaling pathway enrichment analysis, respectively, to predict the biological processes involved in intestinal microorganisms.
[0051] 6) Analysis of immune-related microorganisms and immune microenvironment
[0052] a. Gene expression data acquisition and data preprocessing: Colorectal cancer-related data were obtained from the GEO database (https: / / www.ncbi.nlm.nih.gov / gds). Data preprocessing included gene expression distribution analysis and principal component analysis. Principal component analysis was used to identify outliers in the processed microarray data.
[0053] b. Acquisition of immune cell infiltration level data. Since GEO data does not provide ready-made immune cell infiltration level data, in order to obtain various immune cells in the sample, such as macrophages, CD8 + T cells, CD4 + The proportion of T cells and other cells was determined using the CIBERSORTx algorithm (https: / / cibersortx.stanford.edu / ). CIBERSORTx is an analytical tool developed by the Alizadeh and Newman labs that estimates gene expression profiles and uses gene expression data to estimate the abundance of member cell types in mixed cell populations. First, a standardized mRNA dataset was uploaded to CIBERSORTx. Using the web analysis tool, 1000 permutations were selected to obtain the relative infiltration levels of 25 immune cell types in each sample.
[0054] c. The Spearman algorithm was used to perform co-expression analysis on the intestinal microbial abundance matrix and the immune cell abundance matrix, and a certain number of intestinal microorganisms strongly correlated with immune cell infiltration were extracted using the correlation coefficient |r| > 0.7 and p < 0.05 as the screening criteria.
[0055] 7) Prediction of immune-related gut microbial mechanisms:
[0056] a. Acquisition of immune regulatory gene data.
[0057] b. Using intestinal microbial abundance data and immune regulatory gene expression profile data, the Spearman algorithm was used to obtain the correlation coefficient between intestinal microbes and genes, and a correlation matrix between intestinal microbes and immune regulatory genes was constructed;
[0058] c. Based on the correlation matrix between intestinal microorganisms and immune regulatory genes, a heat map of the correlation between intestinal microorganisms and immune regulatory genes was constructed to confirm the mechanism by which immune-related intestinal microorganisms exert biological effects.
[0059] Example
[0060] 1. Analysis of intestinal microbiota diversity and composition in colorectal cancer patients and healthy controls:
[0061] 1.1 Patients and Data Sources:
[0062] The PubMed database (https: / / pubmed.ncbi.nlm.nih.gov) was searched for 16S rRNA sequencing data of colorectal cancer patients and healthy controls, and four studies were finally included: the Yamada et al. cohort (Nat Med 2019, 25:968-976.), the Blekhman et al. cohort (PLoS Genet 2018, 14:e1007376.), the Chen et al. cohort (mBio 2020, 18; 11(1):e03186-19.), and the Ma et al. cohort (Theranostics 2019, 9:4101-4114.). These studies contained publicly available fecal metabolomics data, high-throughput fecal microbiome sequencing, and sample phenotypic data. To explore the impact of intestinal microorganisms on gene expression in colorectal cancer, the original chip arrays and intestinal microbial profiles (Esteban Gil A et al.al) were downloaded from ColPortal (https: / / colportal.imib.es), which included 16S rRNA sequencing data of 88 stool samples and chip-based gene expression data of 58 tumor samples (SciData 2019, 6:255).
[0063] Data Processing
[0064] The operational taxonomic unit (OTU) clustering method was used by UPARSE software (v7.0.1001, www.drive5.com / uparse), and the optimized sequences of all samples were clustered according to 97% similarity to maximize the identification of known and unknown OUT sequences. UCHIME (http: / / drive5.com / uchime / ) was used to identify and remove chimeric sequences. RDP classifier (https: / / sourceforge.net / projects / rdp-classifier / ) was used to annotate each sequence with species classification, and the community composition of each sample was statistically analyzed at different classification levels of kingdom, phylum, class, order, family, genus, and species. The data of each sample were homogenized, and on this basis, the α and β diversity of each sample were analyzed using QIIME software (http: / / qiime.org / ).
[0065] Key results:
[0066] In order to compare the differences in the structure and composition of the intestinal microbiota between colorectal cancer patients and healthy controls, we first analyzed the species diversity differences between the two groups using α and β diversity. Figure 1 As shown in Figure 2, compared with healthy controls, the α diversity of intestinal flora in patients with colorectal cancer is reduced to a certain extent, as shown by a significant decrease in the Shannon index and the Inv Simpson index (e.g. Figure 1 Figures A and B in the middle. To examine the global differences in the composition of the intestinal flora between the two groups of samples, the Bray-Curtis dissimilarity coefficient and the Jaccard algorithm were used to calculate the distance between each sample, and the PCoA analysis was performed on the samples of each group for β diversity. The results showed that, similar to the results of the α diversity analysis, the intestinal flora of colorectal cancer patients clustered into a community that was significantly distinguishable from that of healthy people, indicating that the intestinal flora composition of the two groups was significantly different (e.g. Figure 1 In summary, the results of both α- and β-diversity analyses indicated that the diversity of the intestinal microbiota was significantly reduced in patients with colorectal cancer.
[0067] Specific differences in intestinal flora between colorectal cancer patients and healthy people. First, at the phylum level, the intestinal flora of both groups were mainly composed of Firmicutes, Bacteroidetes, Proteobacteria, Actinobacteria and Fusobacteria. Among them, Bacteroidetes had the highest relative abundance, accounting for 50.32% and 60.73% in the two groups respectively (e.g. Figure 1In addition, the relative abundances of Actinobacteria and Proteobacteria were higher in the CRC group, while the relative abundances of Firmicutes and Fusobacteria were lower (e.g. Figure 1 Figure E). At the genus level, 59 gut microbes with significantly different abundances were identified between colorectal cancer patients and healthy controls. Among them, 34 bacterial genera were reduced in abundance in colorectal cancer patients, while 23 bacterial genera were more abundant in colorectal cancer patients than in healthy controls (e.g. Figure 1 Figure F). It is worth noting that the abundance of beneficial or pathogenic bacteria was also increased in healthy controls and colorectal cancer patients, respectively. For example, some well-known probiotics, such as Lachnobacterium, Bifidobacterium, Lachnoclostridium, and Ruminococcus, which are the main producers of immunomodulatory short-chain fatty acids and are closely related to the healthy intestinal ecosystem of the body, were significantly reduced in colorectal cancer patients; while pathogenic bacteria, including Fusobacterium, were significantly increased in colorectal cancer patients (such as Figure 1 (Figure F in the middle).
[0068] 2. Gut flora as diagnostic biomarkers:
[0069] Using the colorectal cancer patients and healthy controls from Example 1, the LASSO regression algorithm was performed using the "glmnet" package in R software to identify gut microbes significantly associated with colorectal cancer. The SVM-RFE method was then used to select characteristic gut microbes for colorectal cancer patients. Overlapping gut microbes between the two algorithms were included in subsequent receiver operating characteristic (ROC) and direct current analysis (DCA) analyses for validation.
[0070] The LASSO algorithm uses the L1 criterion to build the model. The regression coefficient changes under the L1 criterion are as follows: Figure 2 As shown in Figure A, each line in the figure represents a regression coefficient, and the direction of each curve represents the change in the regression coefficient path, indicating that most coefficients tend to 0, which can effectively screen variables and avoid data overfitting. After 1000 times of 10-fold cross-validation, the optimal parameters were selected, and 32 microorganisms were obtained as relevant biomarkers for colorectal cancer diagnosis (such as Figure 2 49 feature subsets (such as Figure 2Finally, 28 overlapping characteristic gut microorganisms between the two algorithms were selected: Lactobacillus, Lachnospira, Ruminiclostridium, Collinsella, Lachnoclostridium, Paludibacter, Sutterella, Parabacteroides, Bacillus, Ruminococcaceae_UCG, Holdemanella, Lactonifactor, Allisonella, Anaerofustis, Coriobacterium, Fusobacterium, Neisseria, Lachnoanaerobaculum, Selenomonas, Shuttleworthia, Coprobacillus, Intestinimonas, Ruminococcus, Eisenbergiella, Eubacterium, Catenibacterium, Mogibacterium, Coprococcus (e.g., Figure 2 The AUC of the diagnostic model is 0.836 (Figure C). Figure 2 DCA analysis showed that the diagnostic model showed good net clinical benefits at almost all threshold probabilities (e.g. Figure 2 These characteristic microorganisms have good clinical applicability in distinguishing colorectal cancer patients from healthy controls.
[0071] 3. WGCNA analysis based on intestinal microorganisms:
[0072] The "WGCNA" package in R software was used to construct a weighted gene co-expression network for high-throughput intestinal microbial sequencing data. After inputting the standardized microbial abundance matrix, the pickSoftThreashold in the "WGCNA" package was used to calculate the weight value, the power value was selected as 8.0, and the blockwiseModules was used to construct a scale-free network with the parameters set as default. The softConnectivity function was used to calculate the connectivity of the genes, and the intestinal microorganisms with the top 10 connectivity within the module were screened out as the core intestinal microorganisms of the module. Finally, the weight values between different intestinal microorganisms in the module were calculated based on the topological overlap matrix to obtain the intestinal microbial similarity association network, and the hierarchical clustering of intestinal microbial dissimilarity was performed based on the topological overlap, and a hierarchical clustering tree was established. The dynamic shear tree algorithm was used to cluster intestinal microorganisms with high correlation together, and finally 24 intestinal microbial modules with biological significance were found (such as Figure 3). Further, the interactions between modules were analyzed and correlation heat maps were drawn. The results suggested that the 24 intestinal microbial modules were independent of each other (such as Figure 4 ).
[0073] The identification of core (hubs) microorganisms is based on the significant modules determined by WGCNA analysis. The significant microbial module data are further input into Visant and Cytoscape software for visualization. After sorting the connectivity of microorganisms, 8 intestinal microorganisms with high connectivity are screened out, namely hub microorganisms (such as Figure 5 ).
[0074] Next, we correlated the microbial modules with clinicopathological characteristics, including age, sex, tumor site, AJCC stage, and MSI status. Figure 6 The bacterial genera in the gray module were positively correlated with tumor status. Tumor site was particularly important, as six modules were found to be significantly correlated with tumor site, indicating that tumor site is the main factor leading to changes in intestinal immunity. The green module and royalblue module were significantly correlated with MSI status, indicating that the intestinal microorganisms in these two modules may affect the tumor mutation status. Notably, taxonomic analysis showed that the royalblue module was rich in Bifidobacterium, and its abundance could predict the effect of tumor immunotherapy. In addition, three modules (magenta, darkgreen and turquoise modules) were significantly correlated with AJCC stage. In summary, these results revealed intestinal microbial modules that are closely related to the clinical pathological characteristics of colorectal cancer patients, especially the immune microenvironment status, at the level of systems biology.
[0075] 4. Identification of microorganisms with immunomodulatory effects:
[0076] Using the gene expression profiles in the Esteban Gil A et.al cohort, the CIBERSORTx and MCP-counter algorithms were used to assess the immune cell infiltration level of each tumor sample, and then Spearman correlation rank analysis was used to identify intestinal microorganisms that were correlated with the immune cell infiltration level. The results showed that many bacteria associated with immune cell infiltration were immune-related (such as Figure 7 ). In addition, these bacteria were also more abundant in colorectal cancer samples compared to healthy controls (e.g. Figure 1 For example, among the Firmicutes genera that were found to be increased in abundance in the colorectal cancer group, 13 bacterial genera were significantly associated with T cell infiltration. In particular, the beneficial microorganisms Lactobacillus and Bifidobacterium were found to be associated with CD8 +T cell infiltration was significantly positively correlated with the expression of T cells, suggesting their beneficial role in antitumor and immunomodulatory activities. In addition, the infiltration level of immune cells in each sample was assessed using two other independent algorithms: TIMER and xCell, and Spearman correlation rank analysis was performed with the intestinal microbiota (e.g. Figure 8 The results suggest that the immune-related microorganisms identified by the TIMER and xCell algorithms have significant overlap with those identified by the CIBERSORTx algorithm. For example, Lactobacillus and CD8 + T cells, CD4 + The infiltration level of T cells was significantly positively correlated with that of CD8 + T cells, CD4 + T cells play a key role in the body's anti-tumor immune response.
[0077] 5. Functional prediction analysis of immune-related intestinal microorganisms:
[0078] The discovery of clinically relevant microorganisms, especially those related to immune regulation, has prompted us to study the relationship between intestinal microorganisms and tumor progression, as well as the potential impact of intestinal microbiota on the host, from the perspective of host gene expression and biological processes. By obtaining a dataset from the Esteban Gil A.et.al cohort, which contains parallel paired intestinal microbial and gene sequencing data of 60 CRC patient samples. Using Spearman correlation rank analysis, 50 significant unique gene-intestinal microbial associations were identified, with a correlation strength of more than 0.90. It is worth noting that in the TCGA cohort, many genes that were significantly associated with intestinal microorganisms also had significant differences in expression levels between tumor tissues and normal adjacent cancer tissues. Functional enrichment analysis showed that these genes were mainly concentrated in immune signaling pathways, such as T cell proliferation, defense response to virus by host, chemotaxis, and positive regulation of interleukin-12 production, which are key components of the tumor immune microenvironment ( Figure 9 Figures A and B in the middle).
[0079] exist Figure 9Figure C shows the top 30 most correlated gene-gut microbial associations, and found significant positive correlations between Lactobacillus and CXCL9, Bifidobacterium and CXCR6, and Faecalibacteriu and FFAR2. CXCR6 is an important marker for T cells, and FFAR2 is a member of the G protein-coupled receptor family. Both play an important role in immune regulation and metabolic homeostasis. In addition, in the TCGA-COAD and READ cohorts, these 30 genes were significantly correlated with Stromal, Immune, ESTIMATE score, and TumorPurity, confirming that intestinal microorganisms may affect the occurrence and development of tumors by shaping the state of the tumor immune microenvironment ( Figure 9 Figures D and E in the middle).
[0080] 6. Prediction of immune-related gut microbial mechanisms of action:
[0081] In order to gain a deeper understanding of the molecular mechanism by which intestinal microbiota regulates immune responses, the relationship between intestinal microbes and immune regulatory molecules was studied. It was found that intestinal microbes have significant correlations with a large number of immune regulatory molecules (such as Figure 10 For example, three types of CD8 + The key chemokines (CCL5, CXCL9, and CXCL10) for T cell infiltration into the tumor immune microenvironment are significantly associated with most intestinal microorganisms. Other chemokines, including CCL1, CCL3, CXCL1, CXCL13, CXCL14, and CXCL2, as well as chemokine receptors, including CCR2, CCR3, CCR4, CCR5, CXCR1, CXCR3, and CXCR6, are also significantly associated with a variety of intestinal microorganisms. The literature reports that these chemokines and their receptors promote the infiltration of effector immune cells (including CD8 + In addition, intestinal microbes were significantly correlated with most immune stimulatory factors and MHC molecules, suggesting that the intestinal microbiota may shape the tumor microenvironment by affecting antigen processing and presentation. Given the key role of immunomodulatory molecules in the immune system and their strong association with intestinal microbes, these results are sufficient to demonstrate the overall immunomodulatory effect of intestinal microbes on the tumor immune microenvironment.
[0082] Figure 1Figure 5: Diversity and composition of the gut microbiota between healthy controls (H) and patients with colorectal cancer (CRC). Figures A and B use the Shannon index (A) and the Inv Simpson index (B) to assess the alpha diversity of the gut microbiota between the H and CRC groups. Figures C and D use PCoA analysis based on the Bray-Curtis dissimilarity coefficient (C) and the Jaccard distance (D) to assess the beta diversity of the gut microbiota between the H and CRC groups. Figure E shows the composition ratio of the gut microbiota between the H and CRC groups at the phylum level. Figure F shows a Manhattan plot of the genus-level differential analysis. Genus-level differential microbial analysis was performed using the edgeR package in the R language. A threshold of |log2(fold change)| ≥ 1.0 and a false discovery rate (FDR) < 0.05 were used to identify gut microbes with significantly different abundances.
[0083] Figure 2 Figure 1. Screening and validation of the gut microbiome as a diagnostic biomarker. Figure A shows the dynamic process of the LASSO screening variables. Figure B shows the SVM-RFE algorithm. Figure C shows a Venn diagram of the characteristic microorganisms identified by the LASSO or SVM-RFE algorithm, which yields a diagnostic biomarker model for the gut microbiome. Figure D shows receiver operating characteristic (ROC) analysis of the diagnostic biomarker model for the gut microbiome. Figure E shows a direct current analysis (DCA) analysis of the diagnostic biomarker model for the gut microbiome. The X-axis represents the percentage of threshold probability, and the Y-axis represents the net clinical benefit, calculated by adding the true positive value and subtracting the false positive value.
[0084] Figure 3 This image shows a cluster analysis of the gut microbiome using the WGCNA algorithm. Based on relative abundance at the genus level, gut microbes were grouped into 24 modules. Different colors represent different modules, and gray modules represent gut microbes that do not belong to any known modules. Height represents the shear height, Cluster Dendrogram represents the cluster tree, and Module represents the gut microbial module. The selected cluster tree branches correspond to the corresponding gut microbial module.
[0085] Figure 4 Analysis of interactions between intestinal microbiota modules.
[0086] Figure 5 The gut microbial network in each module identified for WGCNA. Different colors represent different modules.
[0087] Figure 6 This is a heatmap of the correlation between gut microbial modules and clinical pathological characteristics. Darker colors indicate stronger correlations; red indicates positive correlations, and green indicates negative correlations.
[0088] Figure 7 Heat map showing the correlation between immune cell infiltration levels and gut microbiota in colorectal cancer patients based on CIBERSORTx and MCP algorithms.
[0089] Figure 8 Heat map showing the correlation between immune cell infiltration levels and gut microbiota in colorectal cancer patients based on TIMER and xCell algorithms.
[0090] Figure 9 Functional annotation and pathway analysis of intestinal microbial-related genes. Figure A is the GO functional enrichment network of intestinal microbial-related genes using Metascape (https: / / metascape.org / ). Each node corresponds to a GO biological function. Figure B is the KEGG signaling pathway analysis of intestinal microbial-related genes using DAVID (https: / / david.ncifcrf.gov / ). Figure C shows the top 30 gene-intestinal microbial associations with the strongest correlation. Figures D and E are Sankey diagrams of the correlation between tumor immune status, genes, and intestinal microbes in the TCGA-COAD (Figure D) and READ (Figure E) cohorts. Each rectangle represents a feature, and the connection coefficient of each feature is visualized according to the size of the rectangle.
[0091] Figure 10 Network analysis of the correlation between immune regulatory molecules and intestinal microorganisms.
[0092] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form or substance. It should be pointed out that ordinary technicians in this technical field can make several improvements and supplements without departing from the present invention, and these improvements and supplements should also be regarded as the scope of protection of the present invention. Any equivalent changes, modifications and evolutions made by technicians familiar with this profession without departing from the spirit and scope of the present invention by using the technical content disclosed above are all equivalent embodiments of the present invention; at the same time, any equivalent changes, modifications and evolutions made to the above embodiments based on the essential technology of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A method for identifying immune-related intestinal microorganisms in colorectal cancer based on multi-omics features, characterized in that: This includes analyzing the diversity and composition of intestinal flora in patients with colorectal cancer and healthy controls; identifying intestinal microorganisms significantly associated with colorectal cancer patients using the LASSO regression algorithm; selecting characteristic intestinal microorganisms of colorectal cancer using the SVM-RFE method; obtaining overlapping intestinal microorganisms between the two algorithms and incorporating the overlapping intestinal microorganisms into ROC and DCA analysis verification; screening key intestinal microorganisms related to immune regulation using differential abundance intestinal microorganism screening methods, WGCNA methods, and / or intestinal microbial interaction network analysis methods; predicting and analyzing the biological functions of immune-related intestinal microorganisms, analyzing the relationship between immune-related intestinal microorganisms and the immune microenvironment, and predicting the mechanism of action of immune-related intestinal microorganisms; The WGCNA method includes performing cluster analysis on samples from colorectal cancer patients to eliminate outlier samples; using intestinal microbial abundance data to calculate the correlation coefficients between intestinal microorganisms and construct a correlation matrix of intestinal microbial co-expression; by selecting a soft threshold parameter β, the correlation matrix is converted into an adjacency matrix, and then the adjacency matrix is converted into a topological overlap matrix to divide the intestinal microorganisms into modules, and intestinal microorganisms with the same function are divided into one module; sample phenotypic data and intestinal microbial module eigenvalues are correlated to construct a intestinal microbial module-trait correlation heat map, and significant intestinal microbial modules and hub intestinal microorganisms that are highly correlated with the sample phenotypic traits are obtained.
2. The method for identifying immune-related intestinal microorganisms in colorectal cancer based on multi-omics features according to claim 1, characterized in that: The immune-related intestinal microbial function prediction analysis includes using intestinal microbial abundance data and gene expression profile data to calculate the correlation coefficient between intestinal microorganisms and genes, and constructing a correlation matrix between intestinal microorganisms and genes; performing GO and KEGG analysis on the genes in the matrix to predict the biological functions of intestinal microorganisms.
3. The method for identifying immune-related intestinal microorganisms in colorectal cancer based on multi-omics features according to claim 1, characterized in that: The immune-related microorganism and immune microenvironment analysis includes deconvolution analysis using the CIBERSORTx algorithm; using the transcriptome expression profile of cancer tissue, quantifying the relative infiltration levels of 25 immune cells in each patient's tumor and surrounding tissues through LM22 markers and 1,000 sampling tests; and using the Spearman algorithm to perform co-expression analysis of the intestinal microbial abundance matrix and the immune cell abundance matrix to screen out a certain number of intestinal microorganisms that are strongly correlated with immune cell infiltration.
4. The method for identifying immune-related intestinal microorganisms in colorectal cancer based on multi-omics features according to claim 1, characterized in that: The prediction of the mechanism of action of immune-related intestinal microorganisms includes obtaining immune regulatory genes; using intestinal microorganism abundance data and expression spectrum data of immune regulatory genes to calculate the correlation coefficient between intestinal microorganisms and genes, and constructing a correlation matrix between intestinal microorganisms and immune regulatory genes; based on the correlation matrix between intestinal microorganisms and immune regulatory genes, constructing an intestinal microorganism-immune regulatory gene correlation heat map to obtain the mechanism of action of immune-related intestinal microorganisms to exert biological effects.
5. Application of a colorectal cancer diagnostic model based on intestinal flora in the preparation of a reagent for detecting colorectal cancer, characterized in that: The model detects colorectal cancer by detecting 28 overlapping characteristic intestinal microorganisms; the overlapping characteristic intestinal microorganisms are Lactobacillus, Lachnospira, Ruminiclostridium, Collinsella, Lachnoclostridium, Paludibacter, Sutterella, Parabacteroides, Bacillus, Ruminococcaceae_UCG, Holdemanella, Lactonifactor, Allisonella, Anaerofustis, Coriobacterium, Fusobacterium, Neisseria, Lachnoanaerobaculum, Selenomonas, Shuttleworthia, Coprobacillus, Intestinimonas, Ruminococcus, Eisenbergiella, Eubacterium, Catenibacterium, Mogibacterium, and Coprococcus; the 28 overlapping characteristic intestinal microorganisms are obtained by the method of any one of claims 1 to 4.
Citation Information
Patent Citations
New method of intestinal flora multi-genomics in micro-ecological traditional Chinese medicine screening
CN111575393A
Intelligent colorectal cancer early warning system based on intestinal microbial diversity detection
CN114496270A