A biomarker screening method and related applications
By establishing a strain genome and metabolite gene cluster sequence library, and using metagenomic sequencing data comparison to screen out strains and metabolites with significant differences, the problem of time-consuming and labor-intensive strain screening in existing technologies is solved, and rapid and accurate strain screening and drug development are achieved, which are applied to disease diagnosis and prognosis risk prediction.
Patent Information
- Application Number
- CN202210770641.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-30
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2042-06-30
AI Technical Summary
Existing technologies for screening strains that are beneficial for immunotherapy have the following problems: the strain class is too large, the screening process is time-consuming and labor-intensive, and it is impossible to screen the strain's metabolite gene clusters through blood, urine or fecal metabolites, resulting in a long and costly drug development process.
By establishing a representative strain genome sequence library and a metabolite gene cluster sequence library, and using metagenomic sequencing data comparison, we can screen out strains and metabolites with significant differences, directly predict the strain metabolites, and form biomarkers.
It has achieved rapid and accurate screening of strains beneficial for immunotherapy, shortened drug development time, reduced R&D costs, and has been applied to disease diagnosis and prognosis risk prediction.
Smart Images

Figure CN114974432B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biological detection technology, and in particular to a method for screening biomarkers and related applications thereof. Background Art
[0002] Metagenomics, also known as metagenomics, primarily involves high-throughput sequencing of the genomic DNA of all microorganisms in a given environment, followed by analysis of strain classification and functional annotation information for the microorganisms in the sample, and calculation of the relative abundance information of the strains and functions. Metagenomics eliminates the need for isolation and culturing, directly extracting DNA from environmental samples for sequencing, thus avoiding the difficulty of studying uncultivable microorganisms. Furthermore, metagenomic sequencing can not only analyze the structural composition of microbial communities, but also deeply explore the genetic and functional information of the communities. In recent years, with the discovery of immune checkpoints, tumor immunotherapy has rapidly developed, and many studies have shown that immunotherapy has shown excellent results in the treatment of solid tumors such as melanoma and non-small cell lung cancer. However, this therapy only benefits a small number of patients, so developing methods to improve the efficacy of immunotherapy is of great significance.
[0003] Studies have shown that the gut microbiome can directly influence the effectiveness of cancer immunotherapy. The presence, composition, and diversity of the gut microbiome directly affect the patient's tumor treatment outcome, and the efficiency of the immune response can be improved through fecal microbiota transplantation or intervention with beneficial bacteria. For example, Gajewski's team analyzed the fecal microbiome composition of 42 patients with metastatic melanoma and found that the composition of the patients' gut microbiota was significantly correlated with the efficacy of PD-1 inhibitor immunotherapy. The gut microbiota of patients who responded to treatment showed high abundance of Bifidobacterium longum, Collinsella aerofaciens, and Enterococcus faecium. Transplanting the gut microbiota of patients who responded to treatment into germ-free mice improved tumor control, T cell responses, and the efficacy of PD-L1 inhibitor treatment. (Science. 2018, 359:104-108). Professor Guido Kroemer and Professor Laurence Zitvogel's team found through sampling analysis of patients with lung cancer and kidney cancer that those patients who could not benefit from immunotherapy lacked a bacterium called Akkermansia muciniphila in their bodies. They then demonstrated the benefits of this bacterium using mouse experiments. First, they used fecal transplantation to transplant the bacterial communities of "immunotherapy-responsive" and "immunotherapy-unresponsive" patients into antibiotic-treated mice (which themselves did not respond to immunotherapy). As the researchers expected, the former restored their response to immunotherapy, while the latter still did not respond to immunotherapy. Even more interesting is that if the latter is given oral administration of Akkermansia muciniphila, the efficacy of immunotherapy can be restored. Therefore, it is very important to screen for strains that can improve the effectiveness of immunotherapy.
[0004] In recent years, it has been discovered that the metabolites of intestinal flora can directly cross the intestinal barrier to act on the host, affecting host physiology and activating host immunity. These substances include short-chain fatty acids, indolepropionic acid, tryptamine and secondary bile acids, etc. Among them, short-chain fatty acids and secondary bile acids are important research directions in the field of intestinal flora metabolites and tumor immunity. Short-chain fatty acids (SCFAs) are one of the most characteristic microbial metabolic strains known to affect host immunity. It can affect the production of cytokines, the function of macrophages and dendritic cells, and the conversion of B cell categories. For example, on July 1, 2021, a research team from the Philipps University of Marburg, Germany, published an article in the journal Nature Communications entitled: Microbial short-chain fatty acids modulate CD8 +This study is the first to experimentally demonstrate that two microbial metabolites, valerate and butyrate, can enhance CD8 T cell responses and improve adoptive immunotherapy for cancer. + The production of effector cytokines in T cells, thereby enhancing the anti-tumor activity of immune cells.
[0005] Although many studies have shown that intestinal flora can directly affect the effectiveness of cancer immunotherapy, many problems are encountered in the actual drug development process. For example, the microbial strains are too large and how to select strains that are beneficial to immunotherapy? If too many strains are selected, the subsequent animal experiments and clinical trials will be very long, which will increase a lot of R&D costs. Therefore, how to accurately and effectively screen strains is an urgent problem that needs to be solved to accelerate the development of microbial drugs.
[0006] Usually, the abundance of each strain is calculated using metagenomic sequencing data, and strains that are beneficial to immunotherapy are screened based on the strain abundance combined with statistical methods. In addition, the metabolite gene cluster of each strain can also be detected, and the detection value of the strain's metabolite gene cluster can be used to infer whether the strain has the ability to produce metabolites that are beneficial to immunotherapy, and further screen strains that are beneficial to immunotherapy. In the prior art, the detection of the metabolites of the strain requires first of all a special detection instrument; secondly, it takes a certain amount of time to culture the strain, and the detection also costs a certain amount of money, which is time-consuming and labor-intensive, hindering the screening process of strains that are beneficial to immunotherapy.
[0007] Although existing technologies include nuclear magnetic resonance spectrometers for measuring relevant metabolites in blood, urine, or feces, they cannot obtain the metabolic product gene clusters of the strains, and cannot use the metabolites of blood, urine, and feces as the metabolites of the strains to screen microbial drugs for treatment.
[0008] In view of this, the present invention is proposed. Summary of the Invention
[0009] The purpose of the present invention is to provide a biomarker screening method and related applications.
[0010] The present invention is achieved in that:
[0011] In a first aspect, an embodiment of the present invention provides a method for screening a biomarker, comprising the following steps: S1: establishing a representative strain genome sequence library: clustering the sequences in the obtained genome sequence library according to a set threshold to obtain strain clusters at different strain levels or species levels; screening and obtaining representative strain sequences of each strain cluster to establish a representative strain genome sequence library; S2: establishing a metabolite gene cluster sequence library: annotating the sequences of the obtained genome sequence library and predicting the metabolite gene clusters of each strain or each bacterial species, performing similarity clustering on the metabolite gene clusters to obtain gene cluster families, and merging the gene cluster families to obtain a metabolite gene cluster sequence library; the order of operations of steps S1 and S2 can be interchanged; S3: obtaining metagenomic sequencing data of the sample, comparing the metagenomic sequencing data with the representative strain genome sequence library described in step S1 and the metabolite gene cluster sequence library described in step S2, respectively, to obtain the relative abundance of each strain or each bacterial species and the relative abundance of metabolites; S4: screening for significantly different strains and significantly different metabolites; and S5: selecting candidate strains or candidate bacterial species (strains) having significantly different strains or bacterial species and their metabolites as biomarkers.
[0012] In a second aspect, an embodiment of the present invention provides a biomarker screening device, which includes: a representative strain genome sequence library construction module, which is used to establish a representative strain genome sequence library: clustering the sequences in the obtained genome sequence library according to a set threshold to obtain strain clusters at different strain levels or species levels; screening to obtain representative strain sequences of each strain cluster to establish a representative strain genome sequence library; a metabolite gene cluster sequence library construction module, which is used to establish a metabolite gene cluster sequence library: annotating the sequences of the obtained genome sequence library and predicting the metabolite gene clusters of each strain or each bacterial species, and performing relative annotating on the metabolite gene clusters. Similarity clustering is used to obtain gene cluster families, and the gene cluster families are merged to obtain the metabolite gene cluster sequence library; a relative abundance calculation module is used to obtain the metagenomic sequencing data of the sample, and the metagenomic sequencing data is compared with the representative strain genome sequence library described in step S1 and the metabolite gene cluster sequence library described in step S2 to obtain the relative abundance of each strain or each bacterial species and the relative abundance of metabolites; a screening module is used to screen significantly different strains and significantly different metabolites; a marker generation module is used to use candidate strains or candidate bacterial species (strains) with significantly different strains or bacterial species and their metabolites as biomarkers.
[0013] In a third aspect, an embodiment of the present invention provides a method for training a prediction model, comprising: obtaining the detection results of biomarkers in training samples and corresponding annotation results; wherein the biomarkers are obtained by screening by the screening method as described in the aforementioned embodiment or by the biomarker screening device of the aforementioned embodiment, and the annotation results are labels representing the disease risk and / or prognosis efficacy of the sample; the detection results of the biomarkers are input into a pre-constructed prediction model to obtain prediction results; the pre-constructed prediction model is a classifier that can judge the disease risk and / or prognosis efficacy of the sample based on the detection results of the biomarkers; and the parameters of the prediction model are updated based on the label results and the prediction results.
[0014] In a fourth aspect, embodiments of the present invention provide a prediction device for predicting disease risk and / or prognosis and efficacy, comprising: an acquisition module for acquiring detection results of biomarkers in a sample to be tested, wherein the biomarkers are obtained by screening using the screening method described in the preceding embodiments or by the biomarker screening device described in the preceding embodiments; a prediction model for inputting the obtained detection results into a prediction model trained using the training method described in the preceding embodiments to obtain a prediction result;
[0015] In a fifth aspect, an embodiment of the present invention provides an electronic device, comprising: a processor and a memory; the memory is used to store a program, and when the program is executed by the processor, the processor implements: the biomarker screening method as described in the aforementioned embodiment; or, implements the prediction model training method of the aforementioned embodiment: the biomarker is obtained by the screening method as described in the aforementioned embodiment and by utilizing the screening, and the obtained biomarker detection results are input into the prediction model trained by the training method as described in the aforementioned embodiment to obtain a prediction result.
[0016] In a sixth aspect, an embodiment of the present invention provides a computer-readable medium, and when the computer program is executed by a processor, it implements: the biomarker screening method as described in the aforementioned embodiment; or, implements the training method as described in the aforementioned embodiment: obtaining the detection results of biomarkers in the sample to be tested, the biomarkers are obtained by screening the screening method as described in the aforementioned embodiment or by screening the biomarker screening device of the aforementioned embodiment; inputting the obtained detection results into the prediction model trained by the training method as described in the aforementioned embodiment to obtain the prediction results.
[0017] In a seventh aspect, embodiments of the present invention provide the use of reagents for detecting biomarkers in the preparation of products for treating diseases or products for predicting disease risk and / or prognosis of efficacy, wherein the biomarkers are obtained by screening by the screening method described in the preceding embodiments or by the biomarker screening device described in the preceding embodiments.
[0018] In an eighth aspect, an embodiment of the present invention provides a product for predicting the risk of colorectal cancer or the prognosis and efficacy, which includes a reagent for detecting a biomarker as described in the above embodiment.
[0019] The present invention has the following beneficial effects:
[0020] The present invention provides a new method for screening biomarkers. It does not require the detection of the metabolites of the strains by measuring instruments. Instead, the metabolites of each strain are predicted directly by functional prediction of the strain genome sequence, and a metabolite sequence library is formed. Subsequently, the metagenomic sequencing data is used to compare with the strain sequence library and the metabolite sequence library respectively, and the abundance of each bacterial species and the abundance of each metabolite are calculated. Then, based on the strain abundance and metabolite abundance and the correspondence between metabolites and strains, the strains are screened based on statistical methods. Even without actually measuring the metabolites, strains can be screened based on both strains and metabolites. This method has the following advantages:
[0021] (1) Simple and rapid diagnosis, drug screening, and prediction of diagnosis / treatment effects anytime and anywhere;
[0022] (2) Improve the accuracy and effectiveness of screening; screening strains (strains) based on both strains (strains) and metabolites can greatly reduce the number of strains screened compared to screening strains (strains) based only on strain (strain) abundance. Moreover, because metabolites play a major role in immunotherapy, screening using metabolites can make the screening results more accurate.
[0023] (3) Shorten drug development time and reduce R&D costs; the more accurate the strains (bacteria) screened, the more time will be shortened for subsequent animal experiments and clinical trials, and the R&D costs will also be greatly reduced.
[0024] (4) Biomarkers can be used to diagnose diseases or predict prognostic risks. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0026] Figure 1 Schematic diagram for calculating the abundance of bacterial species;
[0027] Figure 2 This is a flow chart of the technical solution of Example 1;
[0028] Figure 3 This is a flow chart of the technical solution of Example 2;
[0029] Figure 4 This is a flow chart of the technical solution of Example 3;
[0030] Figure 5 The optimal number of model strains for predicting whether a disease occurs in Example 4;
[0031] Figure 6 is the importance of the strain in Example 4 in the optimal model;
[0032] Figure 7 The evaluation results of the random forest model training set in Example 4 are as follows;
[0033] Figure 8 This is the evaluation result of the random forest model test set of Example 4.
[0034] Figure 9 This is a flow chart of the biomarker screening technology of this application. DETAILED DESCRIPTION
[0035] To make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely below. Where specific conditions are not specified in the embodiments, conventional conditions or conditions recommended by the manufacturer are used. Where the manufacturer of the reagents or instruments is not specified, all are conventional products that can be purchased commercially.
[0036] An embodiment of the present invention provides a method for screening biomarkers, which comprises the following steps:
[0037] S1: Establishing a representative strain genome sequence library: Clustering the sequences in the obtained genome sequence library according to a set threshold to obtain strain clusters at different strain levels and / or species levels; Screening and obtaining representative strain sequences of each strain cluster to establish a representative strain genome sequence library;
[0038] S2: Establishing a metabolite gene cluster sequence library: Gene annotation is performed on the sequences of the obtained genomic sequence library, and the metabolite gene clusters of each strain and / or each bacterial species are predicted. Gene cluster families are obtained by clustering the metabolite gene clusters based on similarity, and the gene cluster families are merged to obtain a metabolite gene cluster sequence library. The order of the operations of step S1 and step S2 can be interchanged.
[0039] S3: Obtain metagenomic sequencing data of the sample, and compare the metagenomic sequencing data with the representative strain genome sequence library described in step S1 and the metabolite gene cluster sequence library described in step S2 to obtain the relative abundance of each strain and / or each bacterial species and the relative abundance of metabolites;
[0040] S4: Screening of significantly different strains and significantly different metabolites;
[0041] S5: Candidate strains or candidate strains (strains) with significant differences in strains or strains and their metabolites are used as biomarkers.
[0042] In some embodiments, in step S1, the set threshold is ≥95%.
[0043] In some embodiments, the set threshold is ≥99%.
[0044] In some embodiments, in the S1 step, strain clusters at different strain levels are obtained according to a set threshold value ≥99%, and strain clusters at different species levels are obtained according to a set threshold value ≥95%, and representative strains of the strain clusters are obtained.
[0045] In some embodiments, the step of screening and obtaining a representative strain sequence of each strain cluster to establish a representative strain genome sequence library includes: for each strain in each strain cluster, selecting the gene sequence with the longest gene sequence length as the representative strain sequence of the same strain cluster.
[0046] In some embodiments, in the steps S1 and / or S2, the obtained genome sequence library includes a genome database and / or a strain database.
[0047] In some embodiments, the genome database comprises at least one of a UHGG database and a human gut microbial genome sequence database.
[0048] In some embodiments, in step S2, after gene annotation, the annotated genes are analyzed to predict the metabolite gene clusters of each strain and / or each bacterial species.
[0049] In some embodiments, in step S2, the step of obtaining a gene cluster family by performing similarity clustering on the metabolite gene clusters includes any one of (a) to (c):
[0050] (a) extracting protein sequences of metabolite gene clusters, performing redundancy filtering on the predicted metabolite gene clusters based on sequence similarity between the protein sequences to obtain a non-redundant gene cluster set, and clustering the metabolite gene clusters in the non-redundant gene cluster set based on a set similarity threshold to obtain the gene cluster family;
[0051] (b) merging gene clusters of the same metabolite into a gene cluster set, selecting a representative gene cluster from the gene cluster set; clustering the representative gene clusters according to a set similarity threshold to obtain the gene cluster family;
[0052] (c) extracting the protein sequences of the metabolite gene clusters, performing redundancy filtering on the predicted metabolite gene clusters based on the sequence similarity between the protein sequences to obtain a non-redundant gene cluster set, calculating the distance between each pair of gene clusters in each non-redundant gene cluster set, and selecting the gene cluster with the smallest distance value as the representative gene cluster of this gene cluster set for merging to obtain the metabolite gene cluster sequence library.
[0053] In some embodiments, in item (b), the step of selecting a representative gene cluster of the gene cluster set includes: calculating the distance between each pair of gene clusters in each gene cluster set, and selecting the gene cluster with the smallest distance value as the representative gene cluster of this gene cluster set, for merging to obtain the metabolite gene cluster sequence library.
[0054] In some embodiments, the similarity threshold is ≥ 0.3.
[0055] In some embodiments, the tool for performing said gene annotation comprises prokka.
[0056] In some embodiments, the tools used to predict the metabolite gene clusters of each strain or each bacterial species include gutsmash and / or antismash.
[0057] In some embodiments, in step S3, after obtaining the metagenomic sequencing data, the screening method further includes filtering the metagenomic sequencing data (raw reads) to remove low-quality sequences and host sequences to obtain clean reads. For details, see Figure 1 .
[0058] In some embodiments, in the S4 step, the step of screening for significant differences and / or strains and significantly different metabolites includes: for the data of each cohort, the strains and / or strains with significant differences in their types or their abundances are regarded as significantly different strains, and the metabolites with significant differences in their types and abundances are regarded as significantly different metabolites.
[0059] In step S4, after calculating and obtaining the significantly different strains and / or bacterial species, and the significantly different metabolites, the corresponding relationship between the significantly different metabolites and the significantly different strains can be obtained based on the gene clusters of the metabolites.
[0060] In some embodiments, in the S5 step, when there are data from multiple cohorts, the step further includes calculating the heterogeneity of strains or bacterial species between different cohorts, and retaining candidate strains or bacterial species with less heterogeneity as biomarkers.
[0061] In some embodiments, the criteria for low heterogeneity include: I2<40% and P>0.1.
[0062] The samples in step S3 can be used to collect relevant cohort data based on different pre-defined purposes. The cohort data can be sourced from the collected samples and existing databases. The different pre-defined purposes include drug screening, disease risk and / or prognosis prediction, etc. The disease can include existing tumors, including but not limited to at least one of esophageal cancer, breast cancer, lung cancer, gastric cancer, liver cancer, pancreatic cancer, gallbladder cancer, small intestine cancer, colon cancer, rectal cancer, kidney cancer, cervical cancer, ovarian cancer, peritoneal cancer, tongue cancer, lip cancer, laryngeal cancer, and thyroid cancer.
[0063] Specifically, the cohort data can be collected based on the preset use of the biomarker, and each cohort can include negative samples and positive samples. Multiple cohorts can be sample data collected for different preset uses or sample data collected for the same preset use. When the preset use of the biomarker is disease diagnosis, the cohort data can include a diseased group and a non-disease group, and in some embodiments, can also include sample groups with different disease processes; when the preset use of the biomarker is efficacy prediction, the cohort data can include a good efficacy group and a poor efficacy group; when the preset use of the biomarker is drug screening, different groupings can include a post-treatment effective group and a post-treatment ineffective group, and in some embodiments, can also include a good efficacy group and a poor efficacy group.
[0064] In some embodiments, the biomarkers obtained through screening are used as indicators for constructing a prediction model, and a cross-validation method is used to screen out strains or bacterial species that can be effectively predicted as the final biomarkers. It is understandable that the method for constructing a prediction model can be the same as the training method for the prediction model described in any of the following embodiments, and no further details will be given. "Bacterial species that can be effectively predicted" can be understood as the construction indicators (bacterial species or bacterial species combination) used to construct the prediction model when the prediction accuracy reaches 70% to 90% or more. Specifically, the prediction accuracy can be reflected by the ROC curve. If the AUC reaches 0.7 to 0.9 or more, it is considered that the prediction model can effectively predict and can be applied in clinical practice.
[0065] Alternatively, the biomarker screening technology flow chart of this application can be referred to Figure 9 .
[0066] An embodiment of the present invention further provides a biomarker screening device, comprising:
[0067] The representative strain genome sequence library construction module is used to establish a representative strain genome sequence library: the sequences in the obtained genome sequence library are clustered according to the set threshold to obtain strain clusters at different strain levels or species levels; the representative strain sequence of each strain cluster is screened to establish a representative strain genome sequence library;
[0068] The metabolite gene cluster sequence library construction module is used to establish the metabolite gene cluster sequence library: the sequences of the obtained genomic sequence library are annotated and the metabolite gene clusters of each strain or each bacterial species are predicted. The metabolite gene clusters are clustered by similarity to obtain gene cluster families, and the gene cluster families are merged to obtain the metabolite gene cluster sequence library;
[0069] A relative abundance calculation module is used to obtain metagenomic sequencing data of the sample, and compare the metagenomic sequencing data with the representative strain genome sequence library described in step S1 and the metabolite gene cluster sequence library described in step S2 to obtain the relative abundance of each strain or each bacterial species and the relative abundance of metabolites;
[0070] Screening module, used to screen significantly different strains and significantly different metabolites;
[0071] The marker generation module is used to use candidate strains or candidate bacterial species (strains) with significant differences in strains or bacterial species and their metabolites as biomarkers.
[0072] It can be understood that each of the executed steps may correspond to any of the aforementioned embodiments and will not be described in detail.
[0073] An embodiment of the present invention further provides a method for training a prediction model, comprising: obtaining biomarker detection results and corresponding annotation results in a training sample; wherein the biomarker is obtained by screening according to any of the screening methods described in any of the preceding embodiments or by the biomarker screening device described in any of the preceding embodiments, and the annotation result is a label representing the disease risk and / or prognostic efficacy of the sample;
[0074] Inputting the test results of the biomarkers into a pre-built prediction model to obtain a prediction result; the pre-built prediction model is a classifier that can determine the disease risk and / or prognosis of the sample based on the test results of the biomarkers;
[0075] Parameters of the prediction model are updated based on the marking result and the prediction result.
[0076] In some embodiments, when the prediction model is used to predict the risk of colorectal cancer and / or the prognostic efficacy, the biomarkers include at least four of SGB6629, SGB6012, MGYG-HGUT-04562, MGYG-HGUT-01613, SGB6006, SGB6017, SGB5997, MGYG-HGUT-04629, MGYG-HGUT-01464, MGYG-HGUT-01459, MGYG-HGUT-01347, MGYG-HGUT-01326, SGB6013 and MNH0000.
[0077] In some embodiments, the prediction model can be used to predict disease risk and / or prognosis of treatment efficacy.
[0078] In some embodiments, the classifier includes any one of LR, SVM, KNN, RF, GNB, DT, GBDT and AdaBoost, preferably RF.
[0079] In some embodiments, the detection result of the biomarker includes the detection result of the type and / or abundance of the marker.
[0080] In some embodiments, the number of training samples can be selected based on actual conditions, and the sample size can be ≥10 to 100. The training samples include both positive and negative samples. When used to predict disease risk, they can include a non-diseased group and groups with different disease risks. When used to predict therapeutic efficacy, they can include a non-diseased group and groups with different therapeutic efficacy.
[0081] In some embodiments, the label may specifically be a character or a string of characters.
[0082] The present invention also provides a prediction device for predicting disease risk and / or prognosis of treatment efficacy, comprising:
[0083] an acquisition module, configured to acquire a detection result of a biomarker in a sample to be tested, wherein the biomarker is obtained by screening using the screening method described in any of the preceding embodiments or by screening using the biomarker screening device described in any of the preceding embodiments;
[0084] The prediction model is used to input the obtained detection results into the prediction model trained by the training method described in any of the above embodiments to obtain the prediction results.
[0085] In some embodiments, when the prediction model is used to predict the risk of colorectal cancer and / or the prognosis and efficacy, the biomarkers include at least four of SGB6629, SGB6012, MGYG-HGUT-04562, MGYG-HGUT-01613, SGB6006, SGB6017, SGB5997, MGYG-HGUT-04629, MGYG-HGUT-01464, MGYG-HGUT-01459, MGYG-HGUT-01347, MGYG-HGUT-01326, and SGB6013.
[0086] Optionally, the above modules can be stored in a memory in the form of software or firmware or fixed in the operating system (OS) of the electronic device provided in this application, and can be executed by a processor in the electronic device. At the same time, the data, program code, etc. required to execute the above modules can be stored in the memory.
[0087] An embodiment of the present invention further provides an electronic device, comprising: a processor and a memory; the memory is used to store a program, and when the program is executed by the processor, the processor implements: the biomarker screening method according to any of the above embodiments;
[0088] Alternatively, a method for training a prediction model as described in any of the foregoing embodiments is implemented: the biomarker is obtained by the screening method as described in any of the foregoing embodiments and by utilizing the screening, and the obtained biomarker detection results are input into the prediction model trained by the training method as described in any of the foregoing embodiments to obtain a prediction result.
[0089] It can be understood that the method for predicting disease risk and / or prognosis of treatment efficacy corresponds to the steps performed by the prediction device described in any of the aforementioned embodiments.
[0090] The memory may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0091] A processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including a central processing unit (CPU) or a network processor (NP). It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0092] In actual applications, the electronic device can be a server, a cloud platform, a mobile phone, a tablet computer, a laptop computer, an ultra-mobile personal computer (UMPC), a handheld computer, a netbook, a personal digital assistant (PDA), a wearable electronic device, a virtual reality device, etc. Therefore, the embodiments of the present application do not limit the type of electronic device.
[0093] In some embodiments, when the prediction model is used to predict the risk of colorectal cancer and / or the prognosis and efficacy, the biomarkers include at least four of SGB6629, SGB6012, MGYG-HGUT-04562, MGYG-HGUT-01613, SGB6006, SGB6017, SGB5997, MGYG-HGUT-04629, MGYG-HGUT-01464, MGYG-HGUT-01459, MGYG-HGUT-01347, MGYG-HGUT-01326, and SGB6013.
[0094] An embodiment of the present invention provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements: the biomarker screening method as described in any of the aforementioned embodiments; or, implements the training method as described in any of the aforementioned embodiments: obtaining the detection results of biomarkers in a sample to be tested, wherein the biomarkers are obtained by screening the screening method as described in any of the aforementioned embodiments or by screening the biomarker screening device as described in any of the aforementioned embodiments; and inputting the obtained detection results into a prediction model trained by the training method as described in any of the aforementioned embodiments to obtain a prediction result.
[0095] Computer-readable media include: USB flash drives, mobile hard drives, read-only memories, random access memories, magnetic disks, or optical disks, among other media that can store program codes.
[0096] Embodiments of the present invention also provide the use of a reagent for detecting a biomarker in the preparation of a product for treating a disease or a product for predicting the risk of a disease and / or the prognosis of the therapeutic effect, wherein the biomarker is obtained by screening using the screening method described in any of the preceding embodiments or by screening using the biomarker screening device described in any of the preceding embodiments.
[0097] The embodiments of the present invention also provide the use of a reagent for detecting a biomarker in the preparation of a product for treating a disease or a product for predicting the risk of colorectal cancer and / or the prognosis and efficacy, wherein the biomarkers include at least four of SGB6629, SGB6012, MGYG-HGUT-04562, MGYG-HGUT-01613, SGB6006, SGB6017, SGB5997, MGYG-HGUT-04629, MGYG-HGUT-01464, MGYG-HGUT-01459, MGYG-HGUT-01347, MGYG-HGUT-01326, and SGB6013.
[0098] In some embodiments, the product for treating a disease may be a drug, reagent or kit for treating a disease, and the product for predicting the risk of a disease and / or the prognosis of the efficacy may be a related reagent, kit or prediction model.
[0099] In addition, an embodiment of the present invention further provides a product for predicting the risk of colorectal cancer and / or the prognosis efficacy, which includes a reagent for detecting a biomarker as described in any of the aforementioned embodiments.
[0100] The features and performance of the present invention are further described in detail below with reference to the embodiments.
[0101] Example 1 Metagenomic determination of bacterial communities and metabolite gene clusters for disease prediction
[0102] The obtained samples were divided into a diseased group vs. a non-diseased group, and biomarker screening and disease diagnosis for disease diagnosis were performed according to the following method.
[0103] 1. Establish a strain genome sequence library
[0104] The sequences in the UHGG database and the Muen self-tested strain database were clustered into different strain clusters. The strain sequence with the longest sequence length in each strain cluster was selected as the representative strain sequence of this strain cluster, and all representative strain sequences were merged into the strain genome sequence library finally used for analysis.
[0105] 2. Establish a metabolite gene cluster sequence library
[0106] (1) Use Prokka to annotate the gene functions of the strain genome sequence, perform gutsmash / antismash analysis on the annotated genes, and predict the metabolite gene clusters of each strain;
[0107] (2) extracting the protein sequences of gene clusters and performing redundancy filtering on the predicted gene clusters based on their mutual sequence similarity to obtain a non-redundant gene cluster set;
[0108] (3) Calculate the distance between each gene cluster in each gene cluster set, and select the gene cluster with the smallest distance value as the representative gene cluster of this gene cluster set;
[0109] (4) All representative gene clusters were newly clustered, and gene clusters with similarity higher than 0.3 were considered to be the same gene cluster family. All gene cluster families were merged to form a metabolite gene cluster sequence library for subsequent analysis;
[0110] (5) Based on the prediction results of the above steps, the metabolite gene cluster of each bacterial species (strain) can be obtained, that is, the correspondence between each bacterial species (strain) and the metabolite gene cluster produced by the bacterial species (strain); the metabolite gene cluster sequence library used for subsequent analysis of metabolites can also be obtained.
[0111] 3. Calculate strain abundance and metabolite abundance
[0112] (1) Obtain metagenomic sequencing data of relevant sample sets, i.e., raw reads; the relevant sample sets can be multiple cohorts collected from public databases;
[0113] (2) Remove low-quality sequences and host sequences from raw reads to obtain clean reads;
[0114] (3) Compare the clean reads with the strain genome sequence library to obtain candidate bacterial species (strains) and calculate the relative abundance of each candidate bacterial species (strain);
[0115] (4) Compare the clean reads with the metabolite sequence library, obtain the candidate metabolites in the clean reads, and calculate the relative abundance of each candidate metabolite;
[0116] 4. Calculate the significantly different strains and significantly different metabolites;
[0117] (1) Calculate the significance of the difference between the diseased group and the non-disease group for each candidate bacterial species (strain) using statistical methods;
[0118] (2) Using this statistical method to calculate the heterogeneity of each candidate bacterial species (strain) in different cohort data;
[0119] (3) Calculate the difference significance of each candidate metabolite between the diseased group and the non-disease group using statistical methods;
[0120] (4) Using the correspondence between candidate bacterial species (strains) and metabolites, screen out candidate bacterial species (strains) with significant differences in species and metabolites of these strains;
[0121] (5) Based on the heterogeneity of bacterial species (strains) in different cohorts, strains with high heterogeneity in different cohorts were removed, and only bacterial species (strains) with low heterogeneity were retained;
[0122] 5. Build models using machine learning algorithms
[0123] Based on the bacterial species (strains) screened out in step 4, all samples were merged, 70% of the samples were randomly selected as the training set, and the rest were used as the test set. A model was established for the training set data, and the cross-validation method was used to screen out bacterial species (strains) that were relatively important to the model as the final markers. Based on these bacterial species (strains), a model was established to predict the test set data, and the AUC was calculated to evaluate whether the screened markers could be used as markers for disease prediction. When there are new samples, this model can be used to predict diseases for the new samples.
[0124] Please refer to the flowchart for details. Figure 2 .
[0125] Example 2 Metagenomic determination of bacterial flora and metabolite gene clusters for predicting therapeutic efficacy
[0126] The obtained samples were divided into an effective group vs. an ineffective group, and biomarker screening and efficacy prediction for efficacy prediction were performed according to the following method.
[0127] 1. Establish a strain genome sequence library
[0128] The sequences in the UHGG database and the Muen self-tested strain database were clustered into different strain clusters. The strain sequence with the longest sequence length in each strain cluster was selected as the representative strain sequence of this strain cluster, and all representative strain sequences were merged into the strain genome sequence library finally used for analysis.
[0129] 2. Establish a metabolite gene cluster sequence library
[0130] (1) Use Prokka to annotate the gene functions of the strain genome sequence, perform gutsmash / antismash analysis on the annotated genes, and predict the metabolite gene clusters of each strain;
[0131] (2) extracting the protein sequences of gene clusters and performing redundancy filtering on the predicted gene clusters based on their mutual sequence similarity to obtain a non-redundant gene cluster set;
[0132] (3) Calculate the distance between each gene cluster in each gene cluster set, and select the gene cluster with the smallest distance value as the representative gene cluster of this gene cluster set;
[0133] (4) All representative gene clusters were newly clustered, and gene clusters with similarity higher than 0.3 were considered to be the same gene cluster family. All gene cluster families were merged to form a metabolite gene cluster sequence library for subsequent analysis;
[0134] (5) Based on the prediction results of the above steps, the metabolite gene cluster of each strain can be obtained, that is, the correspondence between each bacterial species (strain) and the metabolite gene cluster produced by the bacterial species (strain); the metabolite gene cluster sequence library used for subsequent metabolite analysis can also be obtained.
[0135] 3. Calculate the abundance of bacterial species (strains) and metabolites
[0136] (1) Obtain metagenomic sequencing data of relevant sample sets, i.e., raw reads; the relevant sample sets can be multiple cohorts collected from public databases;
[0137] (2) Remove low-quality sequences and host sequences from raw reads to obtain clean reads;
[0138] (3) Compare the clean reads with the strain genome sequence library to obtain candidate bacterial species (strains) and calculate the relative abundance of each candidate bacterial species (strain);
[0139] (4) Compare the clean reads with the metabolite sequence library, obtain the candidate metabolites in the clean reads, and calculate the relative abundance of each candidate metabolite;
[0140] 4. Calculation of significantly different bacterial species (strains) and significantly different metabolites
[0141] (1) Calculate the significance of the difference between the effective group and the ineffective group for each candidate bacterial species (strain) using statistical methods;
[0142] (2) Using this statistical method to calculate the heterogeneity of each candidate bacterial species (strain) in different cohort data;
[0143] (3) Calculate the difference significance of each candidate metabolite between the effective group and the ineffective group using statistical methods;
[0144] (4) Using the correspondence between candidate bacterial species (strains) and metabolites, screen out candidate bacterial species (strains) with significant differences in species and metabolites of these strains;
[0145] (5) Based on the heterogeneity of bacterial species (strains) in different cohorts, strains with high heterogeneity in different cohorts were removed, and only bacterial species (strains) with low heterogeneity were retained;
[0146] 5. Build models using machine learning algorithms
[0147] Based on the bacterial species (strains) screened out in step 4, all samples were merged, 70% of the samples were randomly selected as the training set, and the rest were used as the test set. A model was established for the training set data, and the cross-validation method was used to screen out bacterial species (strains) that were relatively important to the model as the final markers. Based on these strains, a model was established to predict the test set data, and the AUC was calculated to evaluate whether the screened markers could be used as markers for efficacy prediction; when there are new samples, this model can be used to predict the efficacy of the new samples.
[0148] Please refer to the flowchart for details. Figure 3 .
[0149] Example 3 Metagenomic determination of bacterial populations and metabolite gene clusters for drug screening
[0150] The obtained samples were divided into a treatment-effective group and a treatment-ineffective group, and screened for biomarkers for drug screening according to the following method.
[0151] 1. Establish a strain genome sequence library
[0152] The sequences in the UHGG database and the Muen self-tested strain database were clustered into different strain clusters. The strain sequence with the longest sequence length in each strain cluster was selected as the representative strain sequence of this strain cluster, and all representative strain sequences were merged into the strain genome sequence library finally used for analysis.
[0153] 2. Establish a metabolite gene cluster sequence library
[0154] (1) Use Prokka to annotate the gene functions of the strain genome sequence, perform gutsmash / antismash analysis on the annotated genes, and predict the metabolite gene clusters of each strain;
[0155] (2) extracting the protein sequences of gene clusters and performing redundancy filtering on the predicted gene clusters based on their mutual sequence similarity to obtain a non-redundant gene cluster set;
[0156] (3) Calculate the distance between each gene cluster in each gene cluster set, and select the gene cluster with the smallest distance value as the representative gene cluster of this gene cluster set;
[0157] (4) All representative gene clusters were newly clustered, and gene clusters with similarity higher than 0.3 were considered to be the same gene cluster family. All gene cluster families were merged to form a metabolite gene cluster sequence library for subsequent analysis;
[0158] (5) Based on the prediction results of the above steps, the metabolite gene cluster of each strain can be obtained, that is, the correspondence between each strain and the metabolite gene cluster produced by the strain; the metabolite gene cluster sequence library used for subsequent metabolite analysis can also be obtained.
[0159] 3. Calculate strain abundance and metabolite abundance
[0160] (1) Obtain metagenomic sequencing data, i.e., raw reads, of a relevant sample set; the relevant sample set is the cohort data that is found and downloaded and whether the disease treatment is effective;
[0161] (2) Remove low-quality sequences and host sequences from raw reads to obtain clean reads;
[0162] (3) Compare the clean reads with the strain genome sequence library to obtain candidate bacterial species (strains) and calculate the relative abundance of each candidate bacterial species (strain);
[0163] (4) Compare the clean reads with the metabolite sequence library, obtain the candidate metabolites in the clean reads, and calculate the relative abundance of each candidate metabolite;
[0164] 4. Calculation of significantly different strains and significantly different metabolites
[0165] (1) Calculate the significance of the difference between the effective treatment group and the ineffective treatment group for each candidate bacterial species (strain) using statistical methods;
[0166] (2) Using this statistical method to calculate the heterogeneity of each candidate bacterial species (strain) in different cohort data;
[0167] (3) Calculate the difference significance of each candidate metabolite between the effective treatment group and the ineffective treatment group using statistical methods;
[0168] (4) Using the correspondence between candidate bacterial species (strains) and metabolites, screen out candidate bacterial species (strains) with significant differences in species and metabolites of these strains;
[0169] (5) Based on the heterogeneity of bacterial species in different cohorts, the bacterial species (strains) with high heterogeneity in different cohorts were removed, and only the bacterial species (strains) with low heterogeneity were retained;
[0170] 5. Build models using machine learning algorithms
[0171] Based on the bacterial species (strains) screened out in step 4, all samples were merged, 70% of the samples were randomly selected as the training set, and the rest were used as the test set. A model was established for the training set data, and the cross-validation method was used to screen out bacterial species (strains) that were relatively important to the model as the final markers. Based on these strains (species), a model was established to predict the test set data, and the AUC was calculated to evaluate whether the screened markers could be used for drug screening. When new samples are available, this model can be used to screen drugs for the new samples.
[0172] Please refer to the flowchart for details. Figure 4 .
[0173] Example 4
[0174] Steps 1 and 2 are the same as in Examples 1 to 3.
[0175] 3. Calculate strain abundance and metabolite abundance
[0176] 3.1 Sample Collection
[0177] Metagenomic sequencing data of three cohorts totaling 341 samples were collected and downloaded from public databases. The cohort numbers were PRJEB12449, PRJEB10878, and ERP008279, respectively. The specific grouping information of the samples is shown in Table 1.
[0178] Table 1 Sample grouping information
[0179]
[0180] 3.2 Removal of low-quality sequences and host sequences
[0181] The sequencing data of each sample was filtered to remove adapter contamination sequences, low-quality sequences, and host genome sequences to obtain high-quality sequences.
[0182] 3.3 Calculation of strain abundance
[0183] The high-quality sequences obtained in the above steps were aligned with the representative genome sequence library of the bacterial species (strains). The number of reads aligned to each bacterial species (strain) was calculated based on the alignment results. Combined with the length of the representative genome sequence of the bacterial species (strain), the relative abundance of the bacterial species (strains) was further calculated. Table 2 shows some of the results.
[0184] Table 2 Partial results
[0185]
[0186]
[0187] Note: The horizontal columns in the table are sample numbers, and the vertical columns are gene sequences of representative strains.
[0188] 3.4 Calculation of metabolite abundance
[0189] The high-quality sequences obtained in the above steps were aligned with the metabolite sequence library. Based on the alignment results, the number of reads aligned to each metabolite was calculated. Combined with the metabolite sequence length, the relative abundance of the metabolites was further calculated. Table 3 shows some of the results.
[0190] Table 3 Partial results
[0191]
[0192] 4. Screening of significantly different bacterial species (strains) and significantly different metabolites
[0193] 4.1 Calculation of significance of differences among bacterial species (strains)
[0194] For each cohort data set, based on the abundance and grouping information of bacterial species (strains), the Wilcoxon Rank-Sum test was used to calculate the significance of differences between each strain in different groups (p-value < 0.05 was considered significant; p-value > = 0.05 was considered insignificant). Table 4 shows some of the results.
[0195] Table 4 Partial results
[0196] Species ID qvalue pvalue CRC_mean control_mean enriched GCA_003466705 1.15E-02 5.93E-04 1.44E-03 3.60E-03 control GCA_003466705 1.15E-02 5.93E-04 1.44E-03 3.60E-03 control GCA_003466705 1.15E-02 5.93E-04 1.44E-03 3.60E-03 control GCA_003466705 1.15E-02 5.93E-04 1.44E-03 3.60E-03 control GCA_003482185 3.42E-02 3.88E-03 9.42E-04 1.58E-03 control GCA_003482185 3.42E-02 3.88E-03 9.42E-04 1.58E-03 control GCA_003482185 3.42E-02 3.88E-03 9.42E-04 1.58E-03 control GCA_003482185 3.42E-02 3.88E-03 9.42E-04 1.58E-03 control GCA_003482185 3.42E-02 3.88E-03 9.42E-04 1.58E-03 control GCA_003482185 3.42E-02 3.88E-03 9.42E-04 1.58E-03 control GCA_003482185 3.42E-02 3.88E-03 9.42E-04 1.58E-03 control GCA_003482185 3.42E-02 3.88E-03 9.42E-04 1.58E-03 control GCA_003482185 3.42E-02 3.88E-03 9.42E-04 1.58E-03 control GCA_003482185 3.42E-02 3.88E-03 9.42E-04 1.58E-03 control GCA_014287335 4.84E-02 6.85E-03 1.24E-04 1.61E-05 CRC GCA_014287475 2.47E-02 2.13E-03 2.29E-05 4.80E-06 CRC GCA_014287475 2.47E-02 2.13E-03 2.29E-05 4.80E-06 CRC GCA_014287475 2.47E-02 2.13E-03 2.29E-05 4.80E-06 CRC GCA_014287475 2.47E-02 2.13E-03 2.29E-05 4.80E-06 CRC GCA_014287475 2.47E-02 2.13E-03 2.29E-05 4.80E-06 CRC
[0197] 4.2 Calculation of the significance of metabolite differences
[0198] For each cohort data, based on metabolite abundance and grouping information, the Wilcoxon Rank-Sum test was used to calculate the significance of differences between different groups for each metabolite (p-value < 0.05 was considered significant; p-value > = 0.05 was considered insignificant). Table 5 shows the results.
[0199] Table 5 Partial results
[0200]
[0201]
[0202]
[0203] 4.3 Preliminary screening of different bacterial species (strains)
[0204] For each cohort, the results of significant differences in strains, metabolites, and the corresponding relationships between strains and metabolites (Table 5) were combined to preliminarily screen out differential strains. The screening rule was that both strains and metabolites must be significantly different. According to the screening rule, a total of 792 differential strains were screened out in the PRJEB10878 cohort, 361 differential strains were screened out in the PRJEB12449 cohort, and 1239 differential strains were screened out in the ERP008279 cohort. Table 6 shows some of the results.
[0205] Table 5 Correspondence between bacterial species (strains) and metabolites
[0206]
[0207]
[0208] Table 6 Partial results of different bacterial species (strains)
[0209]
[0210]
[0211] The above steps were used to screen out the differential strains in the three cohorts and summarize the results. The summary rule was: as long as it was screened out in one cohort, it was considered a differential strain. According to the screening rules, a total of 2060 differential strains were screened out.
[0212] 4.4 Further screening of differential strains using strain heterogeneity
[0213] The heterogeneity of strains in the three cohorts was calculated, and strains were further screened based on the magnitude of their heterogeneity. The screening rule was to remove strains with high heterogeneity and retain only those with low heterogeneity as differential strains. Ultimately, 1,320 differential strains were identified. The heterogeneity classification rule was: if I<40% and P>0.1, heterogeneity was considered low; otherwise, heterogeneity was considered high. Table 7 shows some of the results.
[0214] Table 7 Partial results of strain heterogeneity
[0215]
[0216]
[0217] 5. Use machine learning algorithms to build models to select candidate bacterial species (strains) for disease treatment
[0218] 5.1 Training and test set samples
[0219] 341 samples (172 Case group samples; 169 Control group samples) were randomly divided into a training set (172 Case group samples; 169 Control group samples) and a test set (172 Case group samples; 169 Control group samples) according to a 7:3 ratio.
[0220] 5.2 Screening strain markers using training set data
[0221] Based on the abundance of bacterial species (strains) in the training set, a random forest classifier was used for five cross-validations. The average error of the five cross-validations was calculated. The minimum error in the average error plus the standard deviation was used as the critical value. All bacterial species (strain) marker sets with an average error less than the critical value were listed, and the marker set with the least number of bacterial species (strains) was selected as the optimal set, namely the bacterial species (strain) marker set. The model also outputs the importance of each bacterial species (strain). The higher the importance, the more important the species (strain) is in predicting whether the disease occurs.
[0222] Through the model, 14 bacterial species (strains) that can well predict whether a disease occurs were selected, and the importance of each bacterial species (strain) in the optimal model was output. The results are as follows. The number of bacterial species (strains) in the optimal model for predicting whether a disease occurs is shown in Figure 5 The importance of each bacterial species (strain) in the optimal model is shown in Figure 6 .
[0223] 5.3 Using the Model to Predict the Training and Test Sets: A model was established based on the abundance of strain markers. The model was used to predict the probability of CRC in the training set samples and a receiver operating characteristic (ROC) curve was plotted. The model was used to predict the probability of CRC in the test set samples and a ROC curve was plotted.
[0224] The evaluation results of the random forest model training set are shown in Figure 7 The evaluation results of the random forest model test set are shown in Figure 8 .
[0225] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A method for screening microbial markers, characterized in that: It includes the following steps: S1: Establishing a representative strain genome sequence library: Clustering the sequences in the obtained genome sequence library according to a set threshold to obtain strain clusters at different strain levels and / or species levels; Screening and obtaining representative strain sequences of each strain cluster to establish a representative strain genome sequence library; S2: Establishing a metabolite gene cluster sequence library: performing gene annotation on the sequences of the obtained genomic sequence library and predicting the metabolite gene clusters of each strain and / or each bacterial species, performing similarity clustering on the metabolite gene clusters to obtain gene cluster families, and merging the gene cluster families to obtain a metabolite gene cluster sequence library; the order of the operations of step S1 and step S2 can be interchanged or steps S1 and S2 can be performed simultaneously; S3: Obtain metagenomic sequencing data of the sample, and compare the metagenomic sequencing data with the representative strain genome sequence library described in step S1 and the metabolite gene cluster sequence library described in step S2 to obtain the relative abundance of each strain and / or each bacterial species and the relative abundance of metabolites; S4: Screening for significantly different strains and / or bacterial species and significantly different metabolites; S5: Candidate strains or candidate strains with significant differences in strains or strains and their metabolites are used as microbial markers.
2. The method for screening microbial markers according to claim 1, characterized in that The microbial markers obtained by screening are used as indicators for constructing predictive models, and cross-validation methods are used to screen out strains or bacterial species that can be used for drug screening or effectively predict the disease risk and / or prognosis of samples as the final microbial markers.
3. The method for screening microbial markers according to claim 1 or 2, characterized in that: In step S1, the threshold is set to be ≥95%.
4. The method for screening microbial markers according to claim 3, characterized in that: The set threshold is ≥99%.
5. The method for screening microbial markers according to claim 1 or 2, characterized in that: In the step S1, strain clusters at different strain levels are obtained according to a set threshold value ≥99%, and strain clusters at different species levels are obtained according to a set threshold value ≥95%, and representative strains of the strain clusters are obtained.
6. The method for screening microbial markers according to claim 1 or 2, characterized in that: The step of screening and obtaining the representative strain sequence of each strain cluster to establish a representative strain genome sequence library includes: for the strains in each strain cluster, selecting the gene sequence with the longest gene sequence length as the representative strain sequence of the same strain cluster.
7. The method for screening microbial markers according to claim 1 or 2, characterized in that: In the steps S1 and / or S2, the obtained genome sequence library includes a genome database and / or a strain database.
8. The method for screening microbial markers according to claim 7, characterized in that: The genome database includes a human intestinal microbial genome sequence database.
9. The method for screening microbial markers according to claim 7, wherein: The genome database includes the UHGG database.
10. The method for screening microbial markers according to claim 1 or 2, characterized in that: In step S2, after gene annotation, the annotated genes are analyzed to predict the metabolite gene clusters of each strain and / or each bacterial species.
11. The method for screening microbial markers according to claim 1 or 2, characterized in that: In step S2, the step of obtaining a gene cluster family by performing similarity clustering on the metabolite gene clusters includes any one of (a) to (c): (a) extracting protein sequences of metabolite gene clusters, performing redundancy filtering on the predicted metabolite gene clusters based on sequence similarity between the protein sequences to obtain a non-redundant gene cluster set, and clustering the metabolite gene clusters in the non-redundant gene cluster set based on a set similarity threshold to obtain the gene cluster family; (b) merging gene clusters of the same metabolite into a gene cluster set, and selecting a representative gene cluster of the gene cluster set; Clustering the representative gene clusters according to a set similarity threshold to obtain the gene cluster family; (c) extracting the protein sequences of the metabolite gene clusters, performing redundancy filtering on the predicted metabolite gene clusters based on the sequence similarity between the protein sequences to obtain a non-redundant gene cluster set, calculating the distance between each pair of gene clusters in each non-redundant gene cluster set, and selecting the gene cluster with the smallest distance value as the representative gene cluster of this gene cluster set for merging to obtain the metabolite gene cluster sequence library.
12. The method for screening microbial markers according to claim 11, characterized in that: In item (b), the step of selecting a representative gene cluster of the gene cluster set includes: calculating the distance between each pair of gene clusters in each gene cluster set, selecting the gene cluster with the smallest distance value as the representative gene cluster of this gene cluster set, and merging it to obtain the metabolite gene cluster sequence library.
13. The method for screening microbial markers according to claim 11, characterized in that: The similarity threshold is ≥0.
3.
14. The method for screening microbial markers according to claim 1 or 2, characterized in that: Tools for performing such gene annotation include prokka.
15. The method for screening microbial markers according to claim 1 or 2, characterized in that: Tools for predicting metabolite gene clusters for each strain or species include gutsmash and / or antismash.
16. The method for screening microbial markers according to claim 1 or 2, characterized in that: In the S4 step, the step of screening for significantly different strains and significantly different metabolites includes: for the data of each cohort, the strains and / or bacterial species with significantly different types or their abundances are regarded as significantly different strains, and the metabolites with significantly different types and their abundances are regarded as significantly different metabolites.
17. The method for screening microbial markers according to claim 1 or 2, characterized in that: In the step S5, when there are data from multiple cohorts, the step also includes calculating the heterogeneity of strains and / or bacterial species between different cohorts, and retaining candidate strains or bacterial species with less heterogeneity as microbial markers.
18. The method for screening microbial markers according to claim 16, wherein: The criteria for small heterogeneity include: I2<40% and P>0.
1.
19. A screening device for microbial markers, characterized in that: It includes: The representative strain genome sequence library construction module is used to establish a representative strain genome sequence library: the sequences in the obtained genome sequence library are clustered according to the set threshold to obtain strain clusters at different strain levels or species levels; the representative strain sequence of each strain cluster is screened to establish a representative strain genome sequence library; The metabolite gene cluster sequence library construction module is used to establish the metabolite gene cluster sequence library: the sequences of the obtained genomic sequence library are annotated and the metabolite gene clusters of each strain or each bacterial species are predicted. The metabolite gene clusters are clustered by similarity to obtain gene cluster families, and the gene cluster families are merged to obtain the metabolite gene cluster sequence library; A relative abundance calculation module is used to obtain metagenomic sequencing data of the sample, and compare the metagenomic sequencing data with the representative strain genome sequence library described in step S1 and the metabolite gene cluster sequence library described in step S2 to obtain the relative abundance of each strain or each bacterial species and the relative abundance of metabolites; Screening module, used to screen significantly different strains and significantly different metabolites; The marker generation module is used to use candidate strains or candidate bacterial species with significant differences in strains or bacterial species and their metabolites as microbial markers.
20. The screening device according to claim 19, characterized in that The microbial markers obtained by screening are used as indicators for constructing a prediction model, and the cross-validation method is used to screen out the strains or strains that can effectively predict as the final microbial markers.
21. A method for training a prediction model, characterized in that: It includes: Obtaining detection results and corresponding annotation results of microbial markers in training samples; wherein the microbial markers are obtained by screening using the screening method according to any one of claims 1 to 18 or by the screening device for microbial markers according to claim 19 or 20, and the annotation results are labels representing the disease risk and / or prognostic efficacy of the samples; Inputting the detection results of the microbial markers into a pre-constructed prediction model to obtain a prediction result; the pre-constructed prediction model is a classifier that can determine the disease risk and / or prognosis of the sample based on the detection results of the microbial markers; Parameters of the prediction model are updated based on the marking result and the prediction result.
22. The training method according to claim 21, characterized in that The classifier includes any one of LR, SVM, KNN, RF, GNB, DT, GBDT and AdaBoost.
23. The training method according to claim 22, characterized in that: The classifier is RF.
24. The training method according to any one of claims 21 to 23, characterized in that: The detection result of the microbial marker includes the detection result of the type and / or abundance of the marker.
25. A prediction device for predicting disease risk and / or prognosis of treatment efficacy, characterized in that: It includes: an acquisition module, configured to acquire a detection result of a microbial marker in a sample to be tested, wherein the microbial marker is obtained by screening by the screening method according to any one of claims 1 to 18 or by screening by the microbial marker screening device according to claim 19 or 20; The prediction model is used to input the obtained detection results into the prediction model trained by the training method according to any one of claims 21 to 24 to obtain the prediction results.
26. The prediction device according to claim 25, characterized in that When the prediction model is used to predict the risk of colorectal cancer and / or the prognosis and efficacy, the microbial markers include at least four of SGB6629, SGB6012, MGYG-HGUT-04562, MGYG-HGUT-01613, SGB6006, SGB6017, SGB5997, MGYG-HGUT-04629, MGYG-HGUT-01464, MGYG-HGUT-01459, MGYG-HGUT-01347, MGYG-HGUT-01326 and SGB6013.
27. An electronic device, characterized in that: The electronic device comprises: a processor and a memory; the memory is used to store a program, and when the program is executed by the processor, the processor implements: the microbial marker screening method according to any one of claims 1 to 18; Or, a training method for the prediction model described in any one of claims 21 to 24 is implemented: the microbial marker is obtained by the screening method described in any one of claims 1 to 18 and by using the screening, and the obtained microbial marker detection results are input into the prediction model trained by the training method described in any one of claims 21 to 24 to obtain a prediction result.
28. The electronic device according to claim 27, wherein: When the prediction model is used to predict the risk of colorectal cancer and / or the prognosis and efficacy, the microbial markers include at least four of SGB6629, SGB6012, MGYG-HGUT-04562, MGYG-HGUT-01613, SGB6006, SGB6017, SGB5997, MGYG-HGUT-04629, MGYG-HGUT-01464, MGYG-HGUT-01459, MGYG-HGUT-01347, MGYG-HGUT-01326 and SGB6013.
29. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements: the microbial marker screening method according to any one of claims 1 to 18; or, the training method according to any one of claims 21 to 24: obtaining the detection results of the microbial markers in the sample to be tested, the microbial markers being obtained by screening the screening method according to any one of claims 1 to 18 or by screening the microbial marker screening device according to claim 19 or 20; inputting the obtained detection results into the prediction model trained by the training method according to any one of claims 21 to 24 to obtain the prediction results.
30. The computer-readable medium of claim 29, wherein: When the prediction model is used to predict the risk of colorectal cancer and / or the prognosis and efficacy, the microbial markers include at least four of SGB6629, SGB6012, MGYG-HGUT-04562, MGYG-HGUT-01613, SGB6006, SGB6017, SGB5997, MGYG-HGUT-04629, MGYG-HGUT-01464, MGYG-HGUT-01459, MGYG-HGUT-01347, MGYG-HGUT-01326 and SGB6013.
31. Use of a reagent for detecting microbial markers in the preparation of a product for treating a disease or a product for predicting the risk of a disease and / or the prognosis of the therapeutic effect, characterized in that: The microbial marker is obtained by screening by the screening method according to any one of claims 1 to 18, or by screening by the microbial marker screening device according to claim 19 or 20; When the disease is colorectal cancer, the microbial markers include at least four of: SGB6629, SGB6012, MGYG-HGUT-04562, MGYG-HGUT-01613, SGB6006, SGB6017, SGB5997, MGYG-HGUT-04629, MGYG-HGUT-01464, MGYG-HGUT-01459, MGYG-HGUT-01347, MGYG-HGUT-01326 and SGB6013.
32. A product for predicting the risk of colorectal cancer and / or prognosis, characterized in that: It comprises the reagent for detecting a microbial marker as claimed in claim 31.
Citation Information
Patent Citations
Microbial marker of colorectal cancer and application of marker
CN109943636A