Gall-stone disease risk prediction method and system based on intestinal flora

By constructing a gut microbiota prediction model based on a random forest model and selecting key gut microbiota species as input characteristics, the problem of insufficient accuracy of the existing gallstone risk prediction model is solved, and higher prediction accuracy and distinction are achieved, laying the foundation for microbial targeted treatment and bacterial intervention of gallstones.

CN120221089APending Publication Date: 2025-06-27SICHUAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510330567.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing gallstone risk prediction models mainly rely on clinical biochemical indicators and demographic information, and lack prediction methods to utilize intestinal flora, resulting in insufficient accuracy and distinction of the model.

Method used

By constructing a gut microbiota prediction model based on a random forest model, several gut microbiota species (such as Christensenellaceae_R-7_group, etc.) were selected as input feature species, cross-verification and area verification under the curve were carried out, and key bacterial species were screened to improve prediction accuracy.

Benefits of technology

It significantly improves the prediction accuracy and distinction of gallstone disease risk, expands the understanding of the pathogenic microbial mechanism of gallstones, and provides a basis for subsequent targeted microbial therapy and bacterial flora intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120221089A_ABST
    Figure CN120221089A_ABST
Patent Text Reader

Abstract

The invention discloses a gall-stone disease risk prediction method and system based on intestinal flora. The method comprises the following steps: constructing a prediction model; selecting a plurality of intestinal bacterial species as input feature species of the prediction model; according to the present invention, the selected intestinal bacterial species comprise at least one selected from the group consisting of ChostenosenaceaeR-7 group, norkfEubacteria coprostanooligopeptide group, Ruminococcus acidilacticusgroup, UCG-002, and the like, and the selected intestinal bacterial species comprise at least one selected from the group consisting of ChostenosenaceaeR-7 group, norkfEubacteria coprostanooligopeptide group, Ruminococcus acidilacticusgroup, UCG-001, UCG-002, UCG-002, and the like; and inputting the feature value corresponding to the input feature type of the person to be tested into the constructed prediction model to predict the gall-stone disease risk of the person to be tested. 38 kinds of key bacteria related to the gallstone risk are screened out through the model, cognition on a gallstone pathogenic microorganism mechanism is expanded, and a foundation is laid for research of subsequent microorganism targeted therapy, flora intervention and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of gallstone disease risk prediction, and particularly relates to a method and system for predicting gallstone disease risk based on gut microbiota. Background Art

[0002] Gallstones are common digestive system diseases. Symptoms such as abdominal pain and vomiting caused by gallstones will lead to a decline in the quality of life of patients. When the symptoms are severe, patients must undergo cholecystectomy for treatment, which brings a great economic burden to patients and society. Clinical screening of gallstone disease risk can timely take preventive intervention / treatment for the screened high-risk population to prevent the occurrence and development of gallstones.

[0003] Previous gallstone risk prediction models mostly focused on clinical biochemical indicators and demographic information, and used methods such as logistic regression models, cox regression models, and machine learning to predict the risk of gallstones. The inventor also explored the relationship between oral microbiota and the risk of gallstone disease, and made a patent application CN2024118305972, a method and system for predicting gallstone disease risk based on oral microbiota, using oral microbiota as a prediction index to predict the risk of gallstone disease.

[0004] However, the reason for using oral microbiota as an indicator is that it is convenient for sampling, which is more conducive to the popularization of the method. However, previous studies have shown that compared with oral microbiota, the compositional differences of gut microbiota are more closely related to the occurrence and development of gallstones. Incorporating gut microbiota into the gallstone risk prediction model is beneficial to further improve the accuracy and discrimination of the model. Although some researchers have tried to use gut microbiota to predict the risk of gallstone disease in recent years, the machine learning models they used cannot reflect the genera of bacteria that play an important role in the prediction process, and their clinical application value is limited. Summary of the Invention

[0005] The purpose of the present invention is to provide a method and system for predicting gallstone disease risk based on gut microbiota, which can partially solve or alleviate the above deficiencies in the prior art, and can select the genera of gut bacteria that play an important role in the prediction result as the prediction index, thereby improving the accuracy of the prediction.

[0006] To solve the above-mentioned technical problems, the present invention specifically adopts the following technical solutions: In the first aspect of the present invention, there is provided a method for predicting gallstone disease risk based on gut microbiota, including: Constructing a prediction model; Select several intestinal bacterial genus species as the input feature species of the prediction model; the selected intestinal bacterial genus species are at least one of Christensenellaceae_R-7_group, norank_f__Eubacterium_coprostanoligenes_group, Ruminococcaceae_gnavus_group, UCG-002, Lactococcus, Gemella, Lachnospiraceae_NK4A136_group, Lachnoodoristidium, Faecalibales, UCG-005, Finegoldia, Bifidobacterium, Anaerofuncis, norank_f__Christensenellaceae, Butyricimonas, Defluviitaleaceae_UCG-011, Lachniflactor, Rothia, Veillonella, Murdoshella, Eubacterium_saphenum_group, Megamonas, Fusobacerium, Euzaklella, Dilma, Abiotrophia, norank_f__Pepesiellaceae, Propionibacterium, Raoultibacter, Family_XIII_AD3011_group, Alistipes, UBA1819, Eubacterium_hallii_group, Parabacteroides, Lachnospiraceae_UCG_010, Peptoniphilus, Bradytrophobacter, DTU089; Input the eigenvalue corresponding to the input feature species of the person to be tested into the constructed prediction model to predict the risk of gallstone disease of the person to be tested.

[0007] As an improvement, the method for selecting intestinal bacterial genus species as the input feature species of the prediction model includes: Collect samples, where the samples include the intestinal bacterial genus species and their abundances of the gallstone patient group and the healthy control group; Compare the samples of the gallstone patient group with the samples of the healthy control group to determine the correlation between the intestinal flora and the risk of gallstone disease; Use the samples to train a random forest model as the prediction model and output the importance scores of the intestinal bacterial genus species; The prediction results of the prediction model are verified using the cross-validation method and the area under the curve (AUC) verification method, and the error rate and the area under the curve value of the prediction model are obtained respectively. The intestinal genus species are sorted from largest to smallest according to the importance score, and several intestinal genus species with a ranking higher than the threshold are selected as the input feature species of the prediction model based on the error rate and the area under the curve value.

[0008] As an improvement, the random forest model is: ; where I(.) is the indicator function, which has a value of 1 when the condition in the parentheses holds, and 0 otherwise; hm(x) is the decision tree function with serial number m in the random forest rule; k is the predicted value; H(x) is the random forest model function; and n is the number of decision trees.

[0009] As an improvement, the steps of the cross-validation method include: The collected samples are divided into several sub-datasets of equal size. One of the sub-datasets is used as the test set, and the remaining sub-datasets are used as the training set to perform cross-validation on the random forest model, and the single-validation error rate is calculated; this step is looped until all sub-datasets have been used as the test set. The error rate of the random forest model is calculated using the single-validation error rate.

[0010] As an improvement, the formula: ; is used to calculate the single-validation error rate; where ErrorRate i is the single-validation error rate of the i-th cross-validation; m i is the number of samples with prediction errors; is the test set 's sample size; The formula: ; is used to calculate the error rate of the random forest model; where ErrorRate cv is the error rate; ErrorRate i is the single-validation error rate of the i-th cross-validation; and j is the number of sub-datasets.

[0011] As an improvement, the area under the curve (AUC) verification method uses the formula: ; to calculate the area under the curve value; where AUC is the area under the curve value; TPR is the true positive rate, and FPR is the false positive rate.

[0012] As an improvement, when collecting samples, saliva samples of the gallstone patient group and the healthy control group are collected, and the types and abundances of intestinal bacteria in the samples are obtained through 16S rRNA gene sequencing.

[0013] As an improvement, the method for determining the correlation between the intestinal flora and the risk of suffering from gallstones includes: Perform inter-group difference tests on the samples of the gallstone patient group and the healthy control group using the Ace index, Chao1 index, Shannon index, and Simpson index based on the OTU level; Compare the results obtained by several distance algorithms according to the contribution degree to the sample differences; Perform difference tests on the intestinal flora at the genus level of the samples of the gallstone patient group and the healthy control group.

[0014] As an improvement, after selecting the types of input features, the accuracy of the prediction model is verified by using the collected samples through receiver operating characteristic analysis.

[0015] The present invention also provides a gallstone disease risk prediction system based on intestinal flora, including: A model construction module for constructing a prediction model; A feature type selection module for selecting several gut bacterial genera as the input feature types of the prediction model; the selected gut bacterial genera are at least one of Christensenellaceae_R-7_group, norank_f__Eubacterium_coprostanoligenes_group, Ruminococcaceae_gnavus_group, UCG-002, Lactococcus, Gemella, Lachnospiraceae_NK4A136_group, Lachnoodoristidium, Faecalibales, UCG-005, Finegoldia, Bifidobacterium, Anaerofuncis, norank_f__Christensenellaceae, Butyricimonas, Defluviitaleaceae_UCG-011, Lachniflactor, Rothia, Veillonella, Murdoshella, Eubacterium_saphenum_group, Megamonas, Fusobacerium, Euzaklella, Dilma, Abiotrophia, norank_f__Pepesiellaceae, Propionibacterium, Raoultibacter, Family_XIII_AD3011_group, Alistipes, UBA1819, Eubacterium_hallii_group, Parabacteroides, Lachnospiraceae_UCG_010, Peptoniphilus, Bradytrophobacter, DTU089; A prediction module for inputting the feature values corresponding to the input feature types of the person to be tested into the constructed prediction model to predict the risk of gallstone disease in the person to be tested.

[0016] Beneficial effects: The present invention breaks the traditional mode of gallstone prediction relying on clinical indicators and demographic information, incorporates gut microbiota into the prediction system, explores the pathogenesis of gallstones from the microbial level, and provides a new dimension for disease prediction.

[0017] Thirty-eight key bacterial genera related to gallstone risk (such as Christensenellaceae_R-7_group, etc.) are screened out by the model, expanding the understanding of the pathogenic microbial mechanism of gallstones and laying a foundation for subsequent research on microbial targeted therapy, microbiota intervention, etc.

[0018] The present invention constructs a model using the random forest algorithm. Through 10-fold cross-validation (error rate 0.287) and ROC analysis (AUC = 0.73) verification, the model's discrimination ability for the risk of gallstones is significantly better than the random level, and it has practical diagnostic reference value. The random forest algorithm can quantify the importance of each genus (such as sorting by Mean Decrease Accuracy), clarify the contribution of key genera to the model, facilitate understanding the association between the gut microbiota and gallstones, and promote the transformation of the model to clinical applications.

[0019] The present invention covers the α-diversity (Ace, Chao indices), β-diversity (PCoA analysis) of the sample microbiota, and the analysis of genus-level differences, analyzing the microbiota characteristics of the gallstone group and the control group from multiple perspectives, providing comprehensive data support for model construction. Multiple verification methods such as 10-fold cross-validation, AUC evaluation, and ROC curve analysis are used to ensure the stability and reliability of the model in different scenarios and improve the credibility of the results. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts do not necessarily draw according to the actual scale. Obviously, the following-described drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 (a) is a comparison chart of Ace index results.

[0022] Figure 1 (b) is a comparison chart of Chao1 index results.

[0023] Figure 1 (c) is a comparison chart of Shannon index results Figure 1 (d) is a comparison chart of Simpson index results.

[0024] Figure 2 It is a PCoA analysis chart based on the OTU level.

[0025] Figure 3 It is a chart of species with differences between groups at the genus level.

[0026] Figure 4 It is a schematic diagram of the cross-validation results of the random forest prediction model for important features.

[0027] Figure 5 Schematic diagram of the verification result of the AUC of the random forest prediction model with important features

[0028] Figure 6 Schematic diagram of the top 38 genera in importance ranking in the random forest analysis (after removing the genera with "unclassified" in their names)

[0029] Figure 7 Schematic diagram of the ROC analysis of the intestinal flora species obtained based on the random forest analysis Specific implementation manners

[0030] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0031] In this article, suffixes such as "module", "component" or "unit" used to represent elements are only for the convenience of describing the present invention, and have no specific meaning in themselves. Therefore, "module", "component" or "unit" can be used interchangeably.

[0032] In this article, the orientation or positional relationship indicated by terms such as "upper", "lower", "inner", "outer", "front", "rear", "one end", "the other end", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be construed as a limitation of the present invention. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0033] In this article, unless otherwise clearly specified and limited, terms such as "installed", "provided with", "connected", etc. shall be understood in a broad sense. For example, "connected" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, a direct connection, or an indirect connection through an intermediate medium, and can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0034] In this article, "and / or" includes any and all combinations of one or more of the listed related items.

[0035] In this article, "a plurality of" means two or more, that is, it includes two, three, four, five, etc.

[0036] Example 1: This example provides a method for predicting the risk of gallstones based on gut microbiota, and its specific steps include: S1 Collect samples, where the samples include the gut bacterial genera and their abundances in the gallstone patient group and the healthy control group.

[0037] In this example, 34 gallstone patients and 80 healthy control subjects were included. When collecting samples, fecal samples of the gallstone patient group and the healthy control group were collected, and the species and abundances of various oral bacterial genera in the saliva samples were obtained through 16S rRNA gene sequencing.

[0038] The so-called bacterial genus abundance refers to the content or quantity of a certain bacterial genus in a specific environment. Bacterial genus abundance reflects the distribution of different bacterial genera in a specific environment or sample, and is an important parameter in microbial ecology research.

[0039] S2 Compare the samples of the gallstone patient group with those of the healthy control group to determine the correlation between gut microbiota and the risk of gallstones.

[0040] Although previous studies have shown a close correlation between gut microbiota and the risk of gallstone disease, in order to make the whole study more rigorous, the present invention first determined that there is a statistical significance between gut microbiota and the risk of gallstones, providing corresponding theoretical support for subsequent prediction work. In this step, the α-diversity and β-diversity of gut microbiota between the two groups were compared, and at the same time, species difference analysis was carried out to prove whether there is a correlation between gut microbiota changes and gallstone disease.

[0041] α-diversity is mainly used to describe the diversity of species within a single sample or community, and measures the number and evenness of different species within the sample. β-diversity is used to describe the differences in species composition between different samples or communities, and measures the changes in species composition between samples.

[0042] In this example, the method for determining the correlation between oral microbiota and the risk of gallstones includes: S21 Perform between-group difference tests on the samples of the gallstone patient group and the healthy control group using the Ace index, Chao1 index, Shannon index, and Simpson index based on the OTU level.

[0043] OTU (Operational Taxonomic Unit) refers to the "operational taxonomic unit", which is a classification unit used to represent microbial groups with relatively high similarity in microbiome research. OTUs are usually based on sequence data. By calculating the similarity between sequences, sequences with similarity higher than a certain threshold (such as 97%) are clustered together to form an OTU. Each OTU represents a small group composed of similar sequences in the microbial community.

[0044] Ace index: The full name is "Abundance-based Coverage Estimator", which is used to estimate the number of species in a sample, especially suitable for low-abundance species.

[0045] Chao1 index: It is used to estimate the species richness in a sample, with special attention to the discovery of low-frequency species.

[0046] Shannon index: It measures the diversity of the microbial community, taking into account the species richness and evenness.

[0047] Simpson index: It is also a method to measure diversity, with special attention to the influence of dominant species. The lower the value, the higher the diversity.

[0048] Figure 1 (a) is the Wilcoxon rank-sum test of the Ace index, which evaluates the species richness of two groups of samples (Ace index at the OTU level). The Ace index of the gallstone group (blue) is significantly lower than that of the control group (red). The bar chart shows obvious differences in the data distribution of the two groups. The statistical test p = 0.000008, indicating that the control group has higher species richness.

[0049] Figure 1 (b) is the Wilcoxon rank-sum test of the Chao index, which also reflects the species richness (Chao index at the OTU level). The Chao index of the gallstone group (blue) is significantly lower than that of the control group (red), p = 0.002979, further verifying that the species richness of the control group is better than that of the gallstone group.

[0050] Figure 1 (c) is the Wilcoxon rank-sum test of the Shannon index. It measures species diversity (combining species richness and evenness). The Shannon index of the gallstone group (blue) is significantly lower than that of the control group (red), p = 0.0187, indicating that the microbial community diversity of the control group is more abundant.

[0051] Figure 1(d) is the Wilcoxon rank sum test of the Simpson index. It evaluates species dominance (focusing on the proportion of dominant species). There is no significant difference in the Simpson index between the gallstone group (blue) and the control group (red), p=0.3172, indicating that there is no statistically significant difference in the distribution of dominant species between the two groups.

[0052] The analysis of Ace, Chao, and Shannon indices showed that the richness and diversity of intestinal flora in the gallstone group were significantly lower than those in the control group, while the Simpson index did not reflect the difference between the groups, suggesting that gallstones may be associated with reduced richness and diversity of intestinal flora.

[0053] S22 compares the results obtained by several distance algorithms according to their contribution to sample differences.

[0054] This step uses PCoA analysis (principal co-ordinates analysis) based on the unweighted-unifrac distance algorithm to perform Beta diversity analysis to study the similarities or differences in the sample community composition. The Adonis test results showed that P < 0.05, and the changes in the bacterial community structure between the two groups were significant.

[0055] The results are as follows Figure 2 As shown in the figure, the distribution of the gallstone group (blue dots) and the control group (red dots) in the figure does not completely overlap, and the ellipse range has a separation trend, indicating that there are obvious differences in the intestinal flora community structure of the two groups of samples. Although some samples overlap, the overall trend supports the conclusion that "the flora structure of the gallstone group and the control group is different", further proving the association between intestinal flora and gallstones.

[0056] The p-value is an important parameter used in statistics to judge the results of hypothesis testing. It indicates the probability of observing a more extreme result than the current sample result when the hypothesis is true. In this embodiment, the p-value is used to measure whether there is a significant difference in the overall community structure between the gallstone group and the healthy control group.

[0057] S23 performed a difference test on the oral flora at the genus level between the gallstone patient group samples and the healthy control group.

[0058] Taxa with missing taxonomic information in the database were excluded, e.g. Figure 3As shown, there were 53 genera with significant differences between the two groups (p < 0.05). Compared with the healthy control group, the relative abundances of 15 genera such as Bifidobacterium and Lachnoclostridium were higher in the gallstone case group, while the relative abundances of 38 genera such as Ruminococcus, Christensenellaceae_R-7_group, and Alistipes were lower in the gallstone group.

[0059] Based on the results of the above three-group comparison, it can be seen that there is a considerable correlation between the gut microbiota and the incidence of gallstones. Therefore, the oral microbiota is used as an indicator to predict the risk of gallstone disease.

[0060] S3 Select several oral genus species as the input feature species of the prediction model.

[0061] It can be foreseen that there are numerous genus species in the intestine, but their importance for the prediction model varies. The error rates and AUC values of the models constructed by inputting different genus species and numbers are different. Therefore, the purpose of this step is to screen out the oral genus species that can best predict the incidence of gallstones as the input feature species of the prediction target. In this embodiment, the prediction model is a random forest model, and the random forest model can use the collected samples to predict the risk of gallstone disease and output the importance scores of oral genus species. The mathematical expression of the random forest model is: ; where I(.) is an indicator function, which has a value of 1 when the condition in the parentheses holds, and 0 otherwise; hm(x) is the decision tree function with serial number m in the random forest rule; k is the predicted value; H(x) is the random forest model function; and n is the number of decision trees.

[0062] The random forest model is an ensemble learning method that improves the prediction accuracy and stability by constructing multiple decision trees and combining their prediction results. The random forest algorithm combines the simplicity of decision trees and the powerful ability of ensemble learning, and is widely used in tasks such as classification and regression. Its construction process is as follows: 1) Random sampling: The random forest randomly selects multiple subsets from the training dataset by sampling with replacement (Bootstrap), and these subsets are used to construct different decision trees.

[0063] 2) Random feature selection: When constructing each decision tree, the random forest algorithm also randomly selects a part of the features (instead of all features) to divide the nodes, which further increases the diversity and generalization ability of the model.

[0064] 3) Construct multiple trees: A random forest consists of multiple decision trees, and each tree is independently constructed on different data subsets and feature subsets.

[0065] 4) Predict the result: For classification tasks, the random forest determines the final classification result through majority voting; for regression tasks, the random forest obtains the final regression result by averaging the predicted values of each tree.

[0066] The final prediction result of the random forest model for a sample is 1 or 0, where 1 indicates the risk of having gallstones and 0 indicates no risk of having gallstones. Another reason for the present invention to select the random forest model as the prediction model is that the random forest model provides multiple methods to calculate and output feature importance scores, which can help users understand which features contribute the most to the prediction result of the model.

[0067] Based on the prediction result of the random forest model, the specific steps for screening oral genus species include: S31 Use the cross-validation method and the area under the curve validation method to verify the prediction result of the prediction model, and obtain the error rate and the area under the curve value of the prediction model respectively.

[0068] More specifically, the specific steps of the cross-validation method are as follows: 1) Divide the collected samples into several equal-sized sub-datasets; for example, the ten-fold cross-validation used in this embodiment divides the samples into ten equal-sized sub-datasets. The dataset D contains N = 112 samples, and the dataset D is divided into 10 subsets D1, D2,..., D10 of approximately equal size, and the number of samples in each N subset is , that is, N = 11 (rounded down).

[0069] 2) Use a certain sub-dataset as the test set and the remaining sub-datasets as the training set to perform cross-validation on the random forest model, and calculate the single-validation error rate; loop this step until all sub-datasets have been used as the test set. For example, for the i-th cross-validation (i = 1, 2,..., 10), the training set consists of all subsets except i, that is , and the test set .

[0070] 3) Calculate the error rate of the random forest model using the single-validation error rate.

[0071] In the i-th cross-validation of the model, the number of mispredicted samples is , where It is an indicator function. When the condition in the parentheses holds, its value is 1 (suffering from gallstones); otherwise, its value is 0 (not suffering from gallstones). Then the error rate of the i-th cross-validation is , is the test set and the number of samples in it. The average error rate after ten-fold cross-validation is , where j = 10. Among them, ErrorRate i is the single-validation error rate of the i-th cross-validation; m i is the number of samples with prediction errors; is the test set and the number of samples in it; j is the number of sub-datasets. In this embodiment, the results after cross-validation are as Figure 4 shown.

[0072] Area Under Curve (AUC) verification method The Area Under Curve (AUC) verification method is a commonly used model performance evaluation method, mainly used to evaluate the classification ability of classification models, especially in binary classification problems. AUC is the area under the Receiver Operating Characteristic Curve (ROC curve), and the ROC curve is obtained by plotting the relationship curve between the True Positive Rate (TPR, also known as sensitivity) and the False Positive Rate (FPR, also known as 1 - specificity). Specifically, its formula is: ; Calculate the area value under the curve; among them, AUC is the area value under the curve; TPR is the true positive rate, and FPR is the false positive rate.

[0073] In this embodiment, the results after AUC verification are as Figure 5 shown.

[0074] S32 sorts the intestinal genus species from largest to smallest according to the importance score, and selects several intestinal genus species with rankings higher than the threshold based on the error rate and the area value under the curve as the input feature species of the prediction model.

[0075] Combining the results after cross-validation and ACU verification, it can be seen that using the top 39 intestinal genus species ranked by importance for prediction not only has a low error rate but also a high AUC value. In this embodiment, after removing the genus with "Unclassfied" (meaning that the corresponding classification information cannot be found in the database) in the name, finally 38 genus are selected as the input feature species of the model, and the results are as Figure 7 shown.

[0076] Of course, according to the specific situation of model deployment and computing power, at least one of the above 38 genera can be selected as the input feature type. More specifically, the 38 genera are arranged in descending order of importance as follows: Christensenellaceae_R-7_group norank_f__Eubacterium_coprostanoligenes_group Ruminococcaceae_gnavus_group UCG-002 Lactococcus Gemella Lachnospiraceae_NK4A136_group Lachnoodoris tidium Faecalibales UCG-005 Finegoldia Bifidobacterium Anaerofuncis norank_f__Christensenellaceae Butyricimonas Defluviitaleaceae_UCG-011 Lachniflactor Rothia Veillonella Murdoshella Eubacterium_saphenum_group Megamonas Fusobacerium Euzaklella Dilma Abiotrophia norank_f__Pepesiellaceae Propionibacterium Raoultibacter Family_XIII_AD3011_group Alistipes UBA1819 Eubacterium_hallii_group Parabacteroides Lachnospiraceae_UCG_010 Peptoniphilus Bradytrophobacter DTU089 S4. The accuracy of the prediction model is verified by using the collected samples through ROC (Receiver Operating Characteristic) analysis.

[0077] After determining the input features, in order to verify the performance of the model, the accuracy and discrimination of the model can also be verified again by using ROC analysis.

[0078] The so-called discrimination refers to the ability of the model to distinguish different categories of samples (such as the gallstone patient group and the healthy group), that is, the performance of judging whether the model can effectively distinguish the target categories (diseased / non-diseased).

[0079] Such as Figure 7As shown, the area under the curve (AUC) was 0.73, and the 95% confidence interval (95% CI) was 0.63 - 0.83. Generally, 0.7 < AUC < 0.8 indicates moderate predictive efficacy, suggesting that the model has certain diagnostic value for the risk of gallstones and can better distinguish between the gallstone group and the control group.

[0080] S5 inputs the feature values corresponding to the types of input features of the person to be tested into the random forest model to predict the risk of gallstones in the person to be tested.

[0081] After the prediction model is constructed, fecal samples of the person to be tested are collected and 16S rRNA gene sequencing is performed to obtain the types and abundances of various intestinal bacterial genera in the saliva samples. Selecting the abundances of one or several of the 38 types of oral bacterial genera determined in step S3 and inputting them into the prediction model can obtain the prediction result.

[0082] Example 2: This example also provides a gallstone disease risk prediction system based on the intestinal flora, including: A model construction module for constructing a prediction model; A feature type selection module for selecting several intestinal bacterial genera as the input feature types of the prediction model; the selected intestinal bacterial genera are at least one of Christensenellaceae_R-7_group, norank_f__Eubacterium_coprostanoligenes_group, Ruminococcaceae_gnavus_group, UCG-002, Lactococcus, Gemella, Lachnospiraceae_NK4A136_group, Lachnoodoristidium, Faecalibales, UCG-005, Finegoldia, Bifidobacterium, Anaerofuncis, norank_f__Christensenellaceae, Butyricimonas, Defluviitaleaceae_UCG-011, Lachniflactor, Rothia, Veillonella, Murdoshella, Eubacterium_saphenum_group, Megamonas, Fusobacerium, Euzaklella, Dilma, Abiotrophia, norank_f__Pepesiellaceae, Propionibacterium, Raoultibacter, Family_XIII_AD3011_group, Alistipes, UBA1819, Eubacterium_hallii_group, Parabacteroides, Lachnospiraceae_UCG_010, Peptoniphilus, Bradytrophobacter, DTU089; A prediction module for inputting the feature values corresponding to the input feature types of the person to be tested into the constructed prediction model to predict the risk of gallstone disease of the person to be tested.

[0083] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the element.

[0084] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a computer terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0085] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims. All of these are within the protection scope of the present invention.

Claims

1. A method for predicting the risk of gallstones based on intestinal flora, characterized in that include: Build predictive models; Select several intestinal bacteria species as input feature species of the prediction model; The selected enterobacteria species are Christensenellaceae_R-7_group, norank_f__Eubacterium_coprostanoligenes_group, Ruminococcaceae_gnavus_group, UCG-002, Lactococcus, Gemella, Lachnospiraceae_NK4A136_group, Lachnoodoristidium, Faecalibales, UCG-005, Finegoldia , Bifidobacterium, Anaerofuncis, norank_f__Christensenellaceae, Butyricimonas, Defluviitaleaceae_UCG-011, Lachniflactor, Rothia, Veillonella, Murdoshella, Eubacterium_saphenum_group, Megamonas, Fusobacerium, Euzaklella, Dilma, Abiotrophia, norank_f__At least one of Pepesiellaceae, Propionibacterium, Raoultibacter, Family_XIII_AD3011_group, Alistipes, UBA1819, Eubacterium_hallii_group, Parabacteroides, Lachnospiraceae_UCG_010, Peptoniphilus, Bradytrophobacter, DTU089; The characteristic values ​​corresponding to the input characteristic types of the subject are input into the constructed prediction model to predict the risk of gallstones in the subject.

2. A method for predicting the risk of gallstones based on intestinal flora according to claim 1, characterized in that Methods for selecting intestinal bacteria species as input feature species for prediction models include: Collecting samples, wherein the samples include intestinal bacteria species and their abundance in a gallstone patient group and a healthy control group; Compare samples from the gallstone patient group with those from the healthy control group to determine the association between intestinal flora and the risk of gallstones; The samples were used to train a random forest model as a prediction model, and the importance scores of intestinal bacteria species were output; The prediction results of the prediction model are verified by using the cross-validation method and the area under the curve validation method, and the error rate and area under the curve value of the prediction model are obtained respectively; The intestinal bacteria species were ranked from large to small according to their importance scores, and several intestinal bacteria species ranked higher than the threshold were selected as input feature species of the prediction model based on the error rate and the area under the curve value.

3. A method for predicting the risk of gallstones based on intestinal flora according to claim 2, characterized in that The random forest model is: ; Among them, I(.) is the indicator function, and its value is 1 when the condition in the brackets is met, otherwise it is 0; hm(x) is the decision tree function with sequence number m in the random forest rule; k is the predicted value; H(x) is the random forest model function; and n is the number of decision trees.

4. A method for predicting the risk of gallstones based on intestinal flora according to claim 2, characterized in that The steps of the cross validation method include: Divide the collected samples into several sub-datasets of equal size; Use a sub-dataset as the test set and the remaining sub-datasets as the training set to cross-validate the random forest model and calculate the single validation error rate; repeat this step until all sub-datasets are used as test sets; The error rate of the random forest model is calculated using the single-shot validation error rate.

5. A method for predicting the risk of gallstones based on intestinal flora according to claim 4, characterized in that Using the formula: ; Calculate the single verification error rate; where ErrorRate i is the single validation error rate of the i-th cross validation; m i is the number of samples with prediction errors; For the test set The number of samples; Using the formula: ; Calculate the error rate of the random forest model; where ErrorRate cv ErrorRate i is the single validation error rate of the i-th cross validation; j is the number of sub-datasets.

6. A method for predicting the risk of gallstones based on intestinal flora according to claim 2, characterized in that The area under the curve validation method uses the formula: ; Calculate the area under the curve; AUC is the area under the curve; TPR is the true positive rate, and FPR is the false positive rate.

7. The method for predicting the risk of gallstones based on intestinal flora according to claim 2, characterized in that: When collecting samples, saliva samples were collected from the gallstone patient group and the healthy control group, and the species and abundance of intestinal bacteria in the samples were obtained by 16S rRNA gene sequencing.

8. The method for predicting the risk of gallstones based on intestinal flora according to claim 1, characterized in that Methods to determine the association between gut microbiota and gallstone risk include: The Ace index, Chao1 index, Shannon index, and Simpson index based on OTU levels were used to test the differences between the gallstone patient group and the healthy control group. Compare the results of several distance algorithms based on their contribution to sample differences; The differences in intestinal flora at the genus level were tested between the gallstone patient group samples and the healthy control group.

9. The method for predicting the risk of gallstones based on intestinal flora according to claim 1, characterized in that: After selecting the types of input features, the accuracy of the prediction model was verified using the collected samples through receiver operating characteristic analysis.

10. A gallstone risk prediction system based on intestinal flora, characterized in that include: Model building module, used to build prediction models; A feature type selection module is used to select several intestinal bacteria species as input feature types of the prediction model; The selected enterobacteria species are Christensenellaceae_R-7_group, norank_f__Eubacterium_coprostanoligenes_group, Ruminococcaceae_gnavus_group, UCG-002, Lactococcus, Gemella, Lachnospiraceae_NK4A136_group, Lachnoodoristidium, Faecalibales, UCG-005, Finegoldia , Bifidobacterium, Anaerofuncis, norank_f__Christensenellaceae, Butyricimonas, Defluviitaleaceae_UCG-011, Lachniflactor, Rothia, Veillonella, Murdoshella, Eubacterium_saphenum_group, Megamonas, Fusobacerium, Euzaklella, Dilma, Abiotrophia, norank_f__At least one of Pepesiellaceae, Propionibacterium, Raoultibacter, Family_XIII_AD3011_group, Alistipes, UBA1819, Eubacterium_hallii_group, Parabacteroides, Lachnospiraceae_UCG_010, Peptoniphilus, Bradytrophobacter, DTU089; The prediction module is used to input the feature value corresponding to the input feature type of the test subject into the constructed prediction model to predict the risk of gallstones in the test subject.