Method, device and equipment for evaluating and screening characteristics of intestinal microorganisms and medium
By constructing a screening method for intestinal microbial characteristics evaluation and screening methods, using multiple mainstream tree algorithms and interpretable algorithms, the problem of lack of targetedness and effectiveness of intestinal microbial characteristics in existing studies was solved, and a more accurate and effective evaluation of microbial characteristics was achieved.
Patent Information
- Application Number
- CN202411861954.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-16
AI Technical Summary
Existing research only relies on simple intestinal microbial sequencing data, and attempts to find valuable features by directly observing basic information such as the abundance and species of microorganisms, resulting in the discovered microbial characteristics that may lack targeting and effectiveness.
A method for evaluating and screening of intestinal microbial characteristics is provided. By obtaining the fecal sample data set and clinical data set of multiple users, sequencing and processing of multiple mainstream tree algorithms, the target prediction model is constructed, and the interpretable algorithm is used to evaluate and screen the model to obtain the target intestinal microbial characteristics set.
This method can accurately quantify the abundance of different microorganisms in the intestine, improve the fit and generalization ability of the model, simplify the complexity of the model, enhance the interpretability and practicality of the model, and the excavated microbial characteristics are targeted and effective.
Smart Images

Figure CN120015124A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of feature screening, and in particular to a method, device, equipment and medium for evaluating and screening intestinal microorganism features. Background Art
[0002] In the field of modern medicine and biological research, intestinal microbiota plays a critical role in human health and disease progression. Numerous studies have shown that intestinal microbes are closely related to a variety of diseases, such as gastrointestinal cancer, inflammatory bowel disease, obesity, diabetes, cardiovascular disease, and certain neurological diseases. The composition and function of intestinal microbes are extremely complex, containing thousands of microbial species that interact with each other and influence the host's physiological functions.
[0003] Currently, there are many methods for analyzing intestinal microbial characteristics, but existing research only relies on simple intestinal microbial sequencing data, and attempts to find valuable features by directly observing basic information such as the abundance and species of microorganisms. As a result, the discovered microbial characteristics may lack specificity and effectiveness. Summary of the invention
[0004] In view of this, the present invention provides a method, device, equipment and medium for evaluating and screening intestinal microbial characteristics to solve the problem that existing research only relies on simple intestinal microbial sequencing data and attempts to find valuable characteristics by directly observing basic information such as the abundance and type of microorganisms, resulting in the lack of specificity and effectiveness of the mined microbial characteristics.
[0005] In a first aspect, the present invention provides a method for evaluating and screening intestinal microbial characteristics, the method comprising:
[0006] A first target stool sample data set and a first clinical data set of multiple first users are obtained; the first target stool sample data set is sequenced to obtain an intestinal microbial abundance data set; based on the first clinical data set and the intestinal microbial abundance data set, a target prediction model is constructed after being processed by a plurality of preset mainstream tree algorithms; an interpretable algorithm is used to perform feature evaluation and screening on the target prediction model to obtain a target intestinal microbial feature set.
[0007] The intestinal microbial feature evaluation and screening method provided by the present invention can obtain an intestinal microbial abundance data set by sequencing the first target fecal sample data set of multiple first users, and can more accurately quantify the actual abundance of different microorganisms in the intestine. A plurality of preset mainstream tree algorithms are used to construct a target prediction model, which combines the advantages of different algorithms, improves the model's fitting ability and generalization ability for complex data relationships, and allows the model to function more stably when facing data from different individuals and different health conditions. Furthermore, by using an interpretable algorithm to evaluate and screen the features of the target prediction model, those intestinal microbial features that may not be highly relevant or redundant to the prediction results can be discarded, thereby focusing on the truly critical, representative and effective target intestinal microbial feature set, which simplifies the complexity of the model and enhances the interpretability and practicality of the model. Therefore, by implementing the present invention, the excavated microbial features are targeted and effective.
[0008] In an optional implementation, obtaining first target stool sample datasets and first clinical datasets of multiple first users includes:
[0009] Acquire initial stool sample data sets and first clinical data sets of a plurality of first users; and determine a first target stool sample data set based on the initial stool sample data sets and the first clinical data set.
[0010] The intestinal microbial characteristics evaluation and screening method provided by the present invention can screen and determine the first clinical data set based on the initial stool sample data set and the first clinical data set, and can eliminate data that does not meet the preset requirements. Furthermore, the factors of both microbial samples and clinical conditions are fully considered, thereby ensuring that the stool sample data finally screened out is highly correlated with the microbial characteristics to be evaluated and screened.
[0011] In an optional embodiment, determining a first target stool sample dataset based on the initial stool sample dataset and the first clinical dataset includes:
[0012] Based on the initial stool sample data set and the first clinical data set, multiple second users are determined from multiple first users according to preset requirements; based on the initial stool sample data set, first target stool sample data sets of the multiple second users are acquired.
[0013] The intestinal microbial characteristic evaluation and screening method provided by the present invention can accurately screen out sample sources that are valuable for subsequent analysis from multiple first users, focus on fecal sample data corresponding to user groups that are more representative and more in line with research conditions, further optimize the accuracy of input data, and provide support for the subsequent accurate construction of prediction models and the mining of reliable intestinal microbial characteristics.
[0014] In an optional embodiment, based on the first clinical data set and the intestinal microbial abundance data set, a target prediction model is constructed after being processed by a preset mainstream tree algorithm, including:
[0015] Based on the first clinical data set and the intestinal microbial abundance data set, an initial intestinal microbial feature set is obtained through a preset feature screening method; based on the initial intestinal microbial feature set, a target prediction model is constructed after processing with a preset mainstream tree algorithm.
[0016] The intestinal microbial feature evaluation and screening method provided by the present invention obtains an initial intestinal microbial feature set from an intestinal microbial abundance data set through a preset feature screening method, which can remove some irrelevant microbial feature information that does not contribute much to prediction, effectively reducing data dimensions and noise interference. Furthermore, by constructing a target prediction model through the initial intestinal microbial feature set, the efficiency and accuracy of model training are improved, the prediction ability and interpretability of the model are enhanced, and the model can better make accurate predictions based on valuable features.
[0017] In an optional embodiment, based on the initial intestinal microbial feature set, a target prediction model is constructed after being processed by a plurality of preset mainstream tree algorithms, including:
[0018] Based on the initial intestinal microbial feature set, multiple machine learning models are constructed after being processed by a variety of preset mainstream tree algorithms; multiple machine learning models are evaluated and the initial prediction model is determined; the initial prediction model is adjusted and evaluated to obtain the target prediction model.
[0019] The intestinal microbial characteristic evaluation and screening method provided by the present invention constructs multiple machine learning models through a variety of preset mainstream tree algorithms, then evaluates these models to determine the initial prediction model, and further performs parameter adjustment and evaluation processing to obtain the target prediction model. After multiple rounds of screening and optimization processes, the advantages of different algorithms can be fully utilized to find the model parameter settings that best suit the data characteristics and prediction targets, which significantly improves the accuracy, stability and generalization ability of the model, thereby ensuring that the final target prediction model can give reliable prediction results in practical applications.
[0020] In an optional embodiment, the method further includes:
[0021] A second target stool sample data set and a second clinical data set of a plurality of second users are obtained; and a target prediction model is verified based on the second target stool sample data set and the second clinical data set.
[0022] The intestinal microbial characteristic evaluation and screening method provided by the present invention can verify the generalization ability of the target prediction model by testing it on independent sample data that is different from that used to build the model, and determine whether it can maintain stable and accurate evaluation and screening effects in different sample groups.
[0023] In a second aspect, the present invention provides a device for evaluating and screening intestinal microbial characteristics, the device comprising:
[0024] The first acquisition module is used to obtain the first target stool sample data set and the first clinical data set of multiple first users; the processing module is used to sequence the first target stool sample data set to obtain the intestinal microorganism abundance data set; the construction module is used to construct a target prediction model based on the first clinical data set and the intestinal microorganism abundance data set after processing by multiple preset mainstream tree algorithms; the feature evaluation and screening module is used to use an interpretable algorithm to perform feature evaluation and screening on the target prediction model to obtain a target intestinal microorganism feature set.
[0025] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the intestinal microbial characteristic evaluation and screening method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.
[0026] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the intestinal microbial characteristic evaluation and screening method of the first aspect or any corresponding embodiment thereof.
[0027] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions for causing a computer to execute the intestinal microbial characteristic assessment and screening method of the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0029] Figure 1 is a schematic diagram of the process of the intestinal microbial characteristic evaluation and screening method according to an embodiment of the present invention;
[0030] Figure 2is a schematic flow chart of another intestinal microbial characteristic assessment and screening method according to an embodiment of the present invention;
[0031] Figure 3 is a flow chart of a method for constructing a prediction model of immunotherapy response degree according to an embodiment of the present invention;
[0032] Figure 4 is a schematic diagram of a receiver operating characteristic curve and an area under the curve at the species level according to an immunotherapy response prediction model according to an embodiment of the present invention;
[0033] Figure 5 is a species-level species contribution graph of the immunotherapy response prediction model according to an embodiment of the present invention;
[0034] Figure 6 is a structural block diagram of a device for evaluating and screening intestinal microbial characteristics according to an embodiment of the present invention;
[0035] Figure 7 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0036] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0037] The embodiment of the present invention provides a method for evaluating and screening intestinal microbial characteristics, which achieves the effect of mining more targeted and effective microbial characteristics by constructing a target prediction model and using an interpretable algorithm to perform characteristic evaluation and screening on the target prediction model.
[0038] According to an embodiment of the present invention, an embodiment of a method for evaluating and screening intestinal microbial characteristics is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in an order different from that shown here.
[0039] In this embodiment, a method for evaluating and screening intestinal microbial characteristics is provided, which can be used in electronic devices such as computers, mobile phones, tablet computers, etc. Figure 1 is a flow chart of a method for evaluating and screening intestinal microbial characteristics according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:
[0040] Step S101 : acquiring a first target stool sample dataset and a first clinical dataset of a plurality of first users.
[0041] Specifically, the intestinal microbiome plays an extremely critical role in human health and disease progression. Therefore, it is necessary to clarify which individuals are the research subjects for the intestinal microbial characteristics assessment and screening. That is, the corresponding multiple first users are usually selected based on specific disease types, age ranges, genders and other factors. For example, if the focus is on screening out intestinal microbial characteristics related to the response to immunotherapy in patients with gastrointestinal cancer, then the "first user" can be a group of patients with gastrointestinal cancer who meet certain inclusion criteria (such as being in a specific cancer stage, not having received certain specific previous treatments, etc.). It can also be for other diseases or health conditions related research, and people with corresponding characteristics are selected as the first users.
[0042] Furthermore, the first target stool sample data set is used to reflect the composition characteristics of the microbial community in the intestine of each first user, and may include information on the microbial community in the stool of multiple first users.
[0043] Furthermore, the first clinical data set may be a summary of the comprehensive clinical information of multiple first users, and may describe the health status, disease conditions, and related background factors of multiple first users from multiple perspectives, and may include basic demographic information (such as age, gender, etc.), disease-related information (disease type, stage, treatment history, etc.), and other health-related indicators (BMI, comorbidities, etc.). For example, if the focus is on screening out intestinal microbial features related to immunotherapy response in patients with gastrointestinal cancer, then the first clinical data set may include the immunotherapy response results of each first user, i.e., each gastrointestinal cancer patient, six months later.
[0044] Step S102: sequencing the first target stool sample dataset to obtain an intestinal microbial abundance dataset.
[0045] Specifically, the microbial DNA in the first target stool sample data set can be extracted, and then the extracted microbial DNA sample can be transferred to an adapted sequencing instrument for a next generation sequencing (NGS) operation.
[0046] Furthermore, the obtained raw sequencing data can be imported into the MetaPhIAn2 analysis tool, through which the composition of the microbial community can be analyzed based on its built-in microbial reference genome database and specific algorithms.
[0047] Among them, MetaPhIAn2 can compare each sequencing read with the reference genome, identify the type of microorganism it comes from, and then count the occurrence of each microorganism in the sample and the relative abundance information. After analyzing and processing all sequencing reads, it can output relevant results on the composition of the microbial community and finally generate an expression spectrum matrix at the level of multiple microbial species.
[0048] Furthermore, considering that low-abundance and rare microorganisms will have an adverse effect on the construction of subsequent prediction models, the expression spectrum matrix generated at the level of multiple microbial species can be filtered and screened.
[0049] For example, the threshold for low occurrence rate is set to less than 10%, that is, if the frequency of a certain type of microorganism in all samples is lower than this ratio, its occurrence rate is considered to be low; at the same time, the threshold for low relative abundance is set to less than 0.01%, that is, if the relative content of a certain type of microorganism in each sample is lower than this value, its relative abundance is judged to be low.
[0050] Furthermore, based on the above two threshold conditions, the microorganisms in the expression spectrum matrix can be screened one by one, and the microorganisms that meet the low occurrence rate or low relative abundance conditions can be eliminated, and finally the expression spectrum matrix at the level of multiple microbial species after screening, that is, the intestinal microbial abundance dataset, can be obtained.
[0051] Step S103, based on the first clinical data set and the intestinal microbial abundance data set, a target prediction model is constructed after being processed by a plurality of preset mainstream tree algorithms.
[0052] Among them, the preset mainstream tree algorithms can be decision tree algorithm, random forest algorithm, XGBoost algorithm and lightGBM algorithm, etc.
[0053] Specifically, based on the first clinical data set and the intestinal microbial abundance data set, model training is performed respectively through a variety of preset mainstream tree algorithms and the trained optimal model is selected as the target prediction model.
[0054] Step S104, using an interpretable algorithm to perform feature evaluation and screening on the target prediction model to obtain a target intestinal microbial feature set.
[0055] Among them, the interpretable algorithm can be the SHAP (SHapley Additive exPlanations) value algorithm, the LIME (Local Interpretable Model-agnostic Explanations) algorithm, etc.
[0056] Specifically, taking the SHAP value algorithm as an example, the SHAP value of each microbial feature in the model can be calculated by the SHAP value algorithm, and the importance of the microbial feature can be measured by the SHAP value. For example, if the average SHAP value of a microorganism is greater than 0, it is considered to be an important microorganism.
[0057] Therefore, the SHAP tool can be called to evaluate the important features of the target prediction model, and the selected features can be used as potential microbial markers to form the corresponding target intestinal microbial feature set.
[0058] The SHAP value can be calculated using the following equation (1):
[0059]
[0060] Where: represents the SHAP value of feature j; N represents the set of all features; S represents any subset of N except feature j; f x (S) represents the predicted value of the model when considering the feature set S; |S| represents the number of features in the set S; |N| represents the total number of all features; f x (S∪{j})-f x (S) represents the marginal contribution of adding feature j to the model prediction.
[0061] The intestinal microbial feature evaluation and screening method provided in this embodiment can obtain an intestinal microbial abundance data set by sequencing the first target fecal sample data set of multiple first users, and can more accurately quantify the actual abundance of different microorganisms in the intestine. A variety of preset mainstream tree algorithms are used to construct a target prediction model, which combines the advantages of different algorithms, improves the model's fitting ability and generalization ability for complex data relationships, and allows the model to function more stably when facing data from different individuals and different health conditions. Further, using an interpretable algorithm to evaluate and screen the target prediction model can discard those intestinal microbial features that may not be highly relevant or redundant to the prediction results, thereby focusing on the truly critical, representative and effective target intestinal microbial feature set, which simplifies the complexity of the model and enhances the interpretability and practicality of the model. Therefore, by implementing the present invention, the excavated microbial features are targeted and effective.
[0062] In this embodiment, a method for evaluating and screening intestinal microbial characteristics is provided, which can be used in electronic devices such as computers, mobile phones, tablet computers, etc. Figure 2 is a flow chart of a method for evaluating and screening intestinal microbial characteristics according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:
[0063] Step S201 : acquiring a first target stool sample dataset and a first clinical dataset of a plurality of first users.
[0064] Specifically, the above step S201 includes:
[0065] Step S2011, obtaining initial stool sample data sets and first clinical data sets of multiple first users.
[0066] The initial stool sample data set is unprocessed stool sample data.
[0067] For a specific description, please refer to the description in the above step S101, which will not be repeated here.
[0068] Step S2012: determining a first target stool sample dataset based on the initial stool sample dataset and the first clinical dataset.
[0069] Specifically, the first clinical data set may be combined to screen out a corresponding first target stool sample data set from the first target stool sample data set.
[0070] In some optional implementations, the above step S2012 includes:
[0071] Step a1: based on the initial stool sample data set and the first clinical data set, determine a plurality of second users from a plurality of first users according to preset requirements.
[0072] Step a2: acquiring first target stool sample data sets of multiple second users based on the initial stool sample data set.
[0073] Specifically, the initial stool sample data set and the first clinical data set may be combined to screen out a plurality of second users who meet preset requirements from a plurality of first users.
[0074] Furthermore, first target stool sample data sets belonging to a plurality of second users may be screened out from the initial stool sample data set.
[0075] Step S202: Sequencing the first target stool sample dataset to obtain an intestinal microbial abundance dataset. Figure 1 Step S102 of the illustrated embodiment will not be described in detail here.
[0076] Step S203, based on the first clinical data set and the intestinal microbial abundance data set, a target prediction model is constructed after being processed by a plurality of preset mainstream tree algorithms.
[0077] Specifically, the above step S203 includes:
[0078] Step S2031, based on the first clinical data set and the intestinal microbial abundance data set, an initial intestinal microbial feature set is obtained through a preset feature screening method.
[0079] Specifically, the first clinical dataset and the intestinal microbial abundance dataset are preprocessed. For example, clinical information such as cancer type is converted into binary features, and species with low occurrence and low abundance are filtered out from the intestinal microbial data.
[0080] Furthermore, the Boruta algorithm can be used to perform feature extraction operations and obtain the corresponding initial intestinal microbial feature set. The Boruta algorithm can judge the value of features and perform feature extraction by comparing the importance of the recombined features with the importance of the original features.
[0081] First, the preprocessed first clinical data set and the intestinal microbial abundance data set can be provided as input data to the Boruta algorithm, and the first clinical data set and the intestinal microbial abundance data set are integrated to ensure that the clinical data and intestinal microbial data corresponding to each sample are accurately matched.
[0082] Furthermore, the Boruta algorithm can evaluate the importance of microbial features by performing operations such as multiple random permutations and recombinations of data, and comparing the importance of the recombined features with the importance of the original features.
[0083] Furthermore, according to the rules set by the Boruta algorithm, if the importance of an original microbial feature is greater than the importance of the recombined feature, then the feature is considered to be an important and valuable feature; conversely, if the importance of an original feature is always less than the importance of its recombined feature, then the feature may be judged to be a relatively unimportant or redundant feature.
[0084] Furthermore, through the above process, the final multiple microbial features can be screened out and a corresponding initial intestinal microbial feature set can be formed.
[0085] Step S2032, based on the initial intestinal microbial feature set, after being processed by a preset mainstream tree algorithm, a target prediction model is constructed.
[0086] Specifically, based on the initial intestinal microbial characteristics, model training is performed using a variety of preset mainstream tree algorithms and the trained optimal model is selected as the target prediction model.
[0087] In some optional implementations, the above step S2032 includes:
[0088] Step b1, based on the initial intestinal microbial feature set, multiple machine learning models are constructed after being processed by a variety of preset mainstream tree algorithms.
[0089] Step b2, evaluate multiple machine learning models and determine the initial prediction model.
[0090] Step b3: adjust parameters and evaluate the initial prediction model to obtain the target prediction model.
[0091] Specifically, the obtained initial intestinal microbial feature set can be divided into a training data set and a test data set at 80% and 20%, respectively.
[0092] Furthermore, you can choose to apply a variety of preset mainstream tree algorithms for model training and build multiple machine learning models of corresponding types.
[0093] Furthermore, different evaluation indicators can be used as specific indicators for testing and evaluating the effect of machine learning models. For example, if the final screening is for intestinal microbial characteristics related to the response to immunotherapy in patients with gastrointestinal cancer, the subject (second user) operating characteristic curve (ROC) and the area under the curve (AUC) can be used as specific indicators for testing and evaluating the effect of machine learning models. Among them, the ROC curve is a curve drawn with the false positive rate (False Positive Rate) as the horizontal axis and the true positive rate (True Positive Rate) as the vertical axis, which can intuitively show the classification performance of the model under different thresholds; and AUC is the area under the ROC curve, and its value range is between 0 and 1. The larger the value, the better the classification performance of the model, that is, the better the model performance.
[0094] Furthermore, various machine learning models (such as the random forest, XGBoost, and LightGBM models built earlier) can be used to predict the test data set, and then the ROC curve and AUC value corresponding to each model can be calculated based on the prediction results. For example, after calculation, the AUC of the LightGBM algorithm is equal to 0.85, the AUC of the random forest algorithm is 0.67, and the AUC of the XGBoost algorithm is 0.79.
[0095] Furthermore, the AUC values of different models are compared, and the model with the largest AUC value is selected as the initial prediction model. In the above example, since the model constructed by the LightGBM algorithm has the largest AUC value, the prediction model constructed by the LightGBM algorithm is selected as the initial prediction model, with the aim of using its relatively better performance to better predict the patient's immunotherapy response results.
[0096] The lightGBM algorithm includes a calculation function, which is the loss function and gradient boosting formula of LightGBM, as shown in the following equation (2):
[0097]
[0098] Where: L(y,F(x)) represents the loss function; N represents the total number of samples; y i represents the actual label of the i-th sample, usually 0 or 1; p i It indicates the probability that the model predicts that the i-th sample is a positive class (label 1).
[0099] After each iteration, the model is continuously updated, and the model update formula is shown in the following relationship (3):
[0100] F t (x) = F t-1 (x)+η*h t (x) (3)
[0101] Where: F t-1 (x) represents the cumulative prediction of the previous t-1 steps; h t (x) represents the decision tree at step t; η represents the learning rate.
[0102] Furthermore, after determining the prediction model constructed by the LightGBM algorithm as the initial prediction model, the grid search strategy can be used to optimize the parameters of the LightGBM algorithm. Among them, grid search is an exhaustive search parameter adjustment method, which can give a set of value ranges of parameters to be adjusted, and try all possible combinations of these parameters in turn, train the model on the training data set, and then evaluate the model performance on the test data set (or validation data set, if set), and finally find the parameter combination that optimizes the model performance.
[0103] Furthermore, according to the grid search strategy, different value ranges of the parameters of the LightGBM algorithm (such as learning rate, number of leaf nodes, tree depth, etc.) are set, and then the algorithm is allowed to traverse these parameter combinations, and multiple LightGBM models are trained using different parameter combinations on the training data set. The test data set is used to evaluate each trained model, and the corresponding evaluation indicators such as the AUC value are calculated.
[0104] Furthermore, with the default parameters, the AUC of the LightGBM prediction model in the test data set is equal to 0.85. After adjusting the parameters and optimizing the parameters, the model after the parameter adjustment is evaluated again using the training data set and the test data set. It is found that the AUC of the model in the training data set and the test data set reached 0.90 and 0.88 respectively. Comparing the model performance before and after the parameter adjustment, the model after the parameter adjustment has better performance on both the training set and the test set, and the AUC value is relatively higher, indicating that its classification accuracy and generalization ability have been improved.
[0105] Furthermore, the adjusted LightGBM model can be determined as the target prediction model based on the evaluation results of the adjusted model on the training dataset and the test dataset.
[0106] Step S204: Use an interpretable algorithm to evaluate and screen the target prediction model to obtain a target intestinal microbial feature set. Figure 1 Step S104 of the illustrated embodiment will not be described in detail here.
[0107] In an optional implementation, after the above step S203, the method further includes: acquiring a second target stool sample data set and a second clinical data set of multiple second users; and verifying the target prediction model based on the second target stool sample data set and the second clinical data set.
[0108] The second user is an individual completely independent of the first user.
[0109] Specifically, the second target stool sample data set and the second clinical data set may be preprocessed, and the preprocessed second target stool sample data set and the second clinical data set may be used to verify the target prediction model.
[0110] For example, if the target prediction model is used to screen out intestinal microbial characteristics associated with immunotherapy response in patients with gastrointestinal cancer, the microbial relative abundance data of the patient (second user) before treatment in the pre-processed second target stool sample data set and the second clinical data set can be extracted as input data for the target prediction model.
[0111] Furthermore, the target prediction model performs calculations based on the input data and inputs corresponding prediction data, and uses the prediction data as parameters of the target prediction model.
[0112] Furthermore, the area under the operating characteristic curve (AUC) can be selected as the key indicator for evaluating the generalization ability (universality) of the model.
[0113] Furthermore, the prediction results output by the target prediction model and the actual known patient immunotherapy response (extracted from clinical data) are used to calculate the corresponding AUC value, and then the AUC value can be used to verify whether the target prediction model has a certain generalization ability, that is, whether the target prediction model can make relatively reasonable predictions on immunotherapy responses in new data sample groups.
[0114] Therefore, through the above process, it can be verified whether the target prediction model can reasonably evaluate and screen the intestinal microbial characteristics in the new data sample population.
[0115] The intestinal microbial feature evaluation and screening method provided in this embodiment obtains an initial intestinal microbial feature set from an intestinal microbial abundance data set through a preset feature screening method, which can remove some microbial feature information that is insignificant or does not contribute much to the prediction, and effectively reduces the data dimension and noise interference. Further, multiple machine learning models are constructed through a variety of preset mainstream tree algorithms, and then these models are evaluated to determine the initial prediction model, and further parameter adjustment and evaluation processing are performed to obtain the target prediction model. After multiple rounds of screening and optimization processes, the advantages of different algorithms can be fully utilized to find the model parameter settings that best suit the data characteristics and prediction targets, significantly improving the accuracy, stability and generalization ability of the model, thereby ensuring that the final target prediction model can give reliable prediction results in practical applications. Further, by testing on independent sample data that is different from the sample data used to build the model, the generalization ability of the target prediction model can be tested to determine whether it can maintain a stable and accurate evaluation and screening effect in different sample groups.
[0116] In one example, a method for constructing a prediction model for the degree of immunotherapy response is provided by screening out intestinal microbial characteristics associated with immunotherapy response in patients with gastrointestinal cancer. Figure 3 As shown, including:
[0117] S1: Collect stool samples and clinical data from patients with MSI-H (MicroSatellite Instability-High) gastrointestinal cancer, including the patient's immunotherapy response results six months later. Screen eligible patients based on stool sample data and clinical data.
[0118] S2: Perform next-generation metagenomic sequencing on stool samples of patients who pass the screening criteria to obtain their intestinal microbial relative abundance data for subsequent modeling and analysis.
[0119] S3: Perform data preprocessing on clinical data and gut microbial data, such as converting clinical information such as cancer type into binary features and filtering out species with low occurrence and abundance from gut microbial data.
[0120] S4: Select mainstream tree algorithms, such as decision tree, random forest, XGBoost and lightGBM, and build prediction models by adjusting parameters. Finally, select the best prediction model based on the performance of the prediction model in the training data set and the test data set.
[0121] S5: Use interpretable algorithms to evaluate the features of the best prediction models and screen out important potential microbial markers.
[0122] S6: An independent cohort of patients with MSI-H gastrointestinal cancer was recruited, and their intestinal flora abundance data and six-month immunotherapy response results were collected to form an external validation dataset, and the prediction model was evaluated on this dataset.
[0123] The method for constructing a prediction model for immunotherapy response provided in this example has the following beneficial effects:
[0124] (1) Provide a model that accurately predicts the six-month immunotherapy response of patients with MSI-H gastrointestinal cancer.
[0125] (2) The model has high expressiveness in both the test dataset and the external validation dataset, which means that the model can provide clinicians with additional references to determine whether MSI-H gastrointestinal cancer patients can undergo immunotherapy, and provide decision-making reference opinions for clinicians, which has good clinical application potential.
[0126] (3) Efficiency: For the same data set, the best prediction model can be quickly screened out from multiple tree algorithm models.
[0127] Furthermore, based on the method for constructing a predictive model for the degree of immunotherapy response provided in the above examples, three specific embodiments are provided for description.
[0128] Example 1: Construction of a prediction model.
[0129] Step 1. Data acquisition: Between 2018 and 2023, volunteers with gastrointestinal cancer were recruited in a hospital, with a total of 1,100 subjects recruited. The following screening indicators were used for selection:
[0130] (1) Whether the subject has used antibiotics in the past two months (self-reported by the subject);
[0131] (2) whether the subject received treatment with immune checkpoint inhibitors (PD1 / PD-L1 / CTLA-4 inhibitors) (through the treatment plan given by the doctor);
[0132] (3) whether the subject was MSI-H status (through clinical report before treatment);
[0133] (4) Whether the subject has gastrointestinal cancer (colorectal cancer or gastric cancer, as reported by clinical reports before treatment).
[0134] Finally, a total of 48 subjects were included in the model construction, and stool samples and clinical data of these gastrointestinal cancer patients were collected before receiving immunotherapy. The clinical data included the immunotherapy response results six months later.
[0135] Step 2, data filtering: Extract the microbial DNA from the fecal samples of the above subjects, then use a sequencing instrument for second-generation sequencing, and then analyze the original sequencing data with MetaPhIAn2 (a tool for analyzing the composition of microbial communities), and finally obtain a relative abundance matrix at the level of 528 types of microbial species (row names are microbial names, column names are sample names). Low-abundance and rare microorganisms will affect the construction of the prediction model, so microorganisms with low occurrence rate (less than 10% threshold) and low relative abundance (less than 0.01% threshold) are filtered out, and finally a relative abundance matrix of 188 types of microorganisms is obtained.
[0136] Step 3, data preparation: Feature extraction can screen out important features, which can reduce computational costs and improve efficiency when used for modeling. The Boruta algorithm was used to extract features from the 188 selected microorganisms based on the fact that the importance of the recombined features was greater than the original feature importance, and 13 important microbial features were retained, as shown in Table 1. In addition, the python machine learning package (sklearn) only accepts numerical classification labels, which requires the six-month immunotherapy response binary classification variable labels (responder (R); non-responder (NR)) to be encoded into 1 and 0.
[0137] Table 1. 13 important microorganisms extracted by feature extraction
[0138] Serial number Species classification English name Chinese name 1 Species Bacteroides_caccae Bacteroides faecalis 2 Species Paraprevotella_clara Parapus clara 3 Species Paraprevotella_unclassified Unclassified Parapus 4 Species Alistipes_finegoldii Finegoldii 5 Species Alistipes_onderdonkii onderdonkii 6 Species Streptococcus_parasanguinis Streptococcus parahaemolyticus 7 Species Lachnospiraceae_bacterium_3_1_46FAA bacterium_3_1_46FAA Lachnospirillum 8 Species Subdoligranulum_unclassified Unclassified rare micrococcus 9 Species Veillonella_atypica Atypical Veillonella 10 Species Veillonella_parvula Veillonella parvula 11 Species Veillonella_unclassified Unclassified Veillonella 12 Species Clostridiales_bacterium_1_7_47FAA bacterium_1_7_47FAA Clostridium 13 Species Lachnospiraceae_bacterium_2_1_58FAA bacterium_2_1_58FAA Lachnospirillum
[0139] Furthermore, model training is performed on the training data set, and model parameter adjustment and model performance evaluation are performed on the test data set. Therefore, the relative abundance data set of the 13 microbial characteristics obtained above is divided into a training data set and a test data set at 80% and 20%, respectively.
[0140] Step 4, optimize the training model: In order to obtain a better prediction model, the 80% training data set obtained above is subjected to various machine learning including but not limited to random forest, XGBoost and lightGMB algorithms to build a prediction model, and the trained prediction model also corresponds to the type of the above model.
[0141] The receiver operating characteristic curve (ROC) and the area under the curve (AUC) are used as specific indicators to test and evaluate the effect of machine learning models. The area under the curve (AUC) is an indicator for evaluating model performance. The larger its value, the better the model performance.
[0142] The prediction accuracy of each prediction model for microbial variables is different or limited. For example, the AUC of the lightGBM algorithm is 0.85, while the AUC of other algorithms such as the random forest algorithm is 0.67 and the AUC of the XGBoost algorithm is 0.79. Therefore, the prediction model constructed by the lightGBM algorithm is selected for classification based on the size of the AUC, with the aim of better predicting the patient's immunotherapy response outcome.
[0143] Among them, the lightGBM algorithm includes a calculation function, which is the loss function and gradient boosting formula of LightGBM, as shown in the above relationship (2).
[0144] After each iteration, the model is continuously updated, and the model update formula is shown in the above relationship (3).
[0145] Step 5: Model parameter adjustment and evaluation: The training model only evaluates the quality of the model with the default parameters. After finding that the model built by the lightGBM algorithm is optimal, the algorithm needs to be adjusted and evaluated on the training data set and the test data set. Here, the grid search strategy is used to optimize the parameters of the lightGBM algorithm.
[0146] The receiver operating characteristic curve (ROC) and area under the curve (AUC) were used as specific indicators to test and evaluate the model effect. Under the default parameters, the AUC of the lightGBM prediction model in the test data set was equal to 0.85. After adjusting and optimizing the parameters, the AUC of the model in the training data set and the test data set were 0.90 and 0.88 respectively. This model was the optimal prediction model, and finally a prediction model composed of 13 types of microorganisms was obtained.
[0147] Example 2: Validation of the prediction model.
[0148] Step 1: Construction of prediction model (repeat the content of Example 1):
[0149] Sample collection: A gastrointestinal cancer patient cohort was recruited in a hospital, and volunteers who met the requirements were screened. Finally, 48 subjects were included in the analysis. Clinical data and stool samples of these subjects were collected, and the stool samples were subjected to second-generation sequencing to obtain sequencing data.
[0150] Data processing: Obtain the clinical data and intestinal microbial data of the subjects, perform data preprocessing, and finally obtain data for predictive model construction;
[0151] Model construction: After feature extraction, 13 types of microorganisms were finally selected as input data to build a prediction model based on the lightGBM algorithm, and the optimal prediction model was obtained after optimizing the parameters of the algorithm;
[0152] Step 2: Verification of prediction model:
[0153] External data verification of the best algorithm: Also in Peking University Cancer Hospital, another group of independent MSI-H gastrointestinal cancer subjects completely different from the previous cohort were recruited. They also met the previous filtering criteria, and a total of 31 patients constituted the external data set. Their clinical data and intestinal microbial relative abundance data were subjected to the same data preprocessing as the model construction embodiment. Their data were used to evaluate the generalization ability of the immunotherapy response prediction model, that is, the versatility of the model. The relative abundance of microorganisms before treatment in the external data set was used as the input data of the immunotherapy response prediction model, and the output data of the corresponding optimal prediction model was used as the parameters of the immunotherapy response prediction model for MSI-H gastrointestinal cancer patients. Finally, the AUC of the prediction model for this data set was 0.79. Figure 4 shown.
[0154] Example 3: Screening of bacterial composition markers.
[0155] Step 1: Construction of prediction model (repeat the content of Example 1):
[0156] 1. Recruitment of subjects: Volunteers were recruited in a certain hospital, with a total of 1,100 subjects recruited. The following screening indicators were used for selection:
[0157] (11) Whether antibiotics have been used in the past two months;
[0158] (12) Whether the subject received treatment with immune checkpoint inhibitors (PD1 / PD-L1 / CTLA-4 inhibitors);
[0159] (13) Whether the subject is MSI-H status;
[0160] (14) Whether the subject has gastrointestinal cancer (colorectal cancer or gastric cancer).
[0161] A total of 48 subjects were included in the analysis. Fecal samples and clinical data were collected from these patients before treatment, including the results of immunotherapy response six months later.
[0162] 2. Data collection: After evaluating the above patients as having the MSI-H phenotype, stool samples of these patients before immunotherapy were extracted from the collected data and the response results (divided into two types: response and non-response) after 6 months of treatment were recorded.
[0163] 3. Obtaining relative abundance data of bacterial flora: Performing second-generation sequencing on the stool samples of the above-mentioned population to obtain microbial data of the intestinal environment, which is composed of 528 species-level microorganisms.
[0164] 4. Data preprocessing: Filter the above-mentioned species-level microbial data, discard species with low occurrence rate (less than 10%) and low abundance (less than 0.01%), and then select the remaining 188 types of microbial species-level relative abundance data as the input features of the model. At the same time, convert the cancer type information (gastric cancer or colorectal cancer, etc.) into binary features and add them to the model training, so that the model can learn the differences between different cancer types. In addition, the six-month immunotherapy response binary classification variables (responder (R); non-responder (NR)) need to be encoded into 1 and 0 respectively.
[0165] Step 2: Feature selection and verification of the prediction model (repeat the contents of Examples 1 and 2):
[0166] Feature selection: Feature selection was performed on the microbial data obtained above, and the Boruta algorithm was used to screen important features. Finally, 13 types of microorganisms were screened out, which served as markers for subsequent modeling.
[0167] Model screening: Select mainstream tree models (random forest, decision tree, XGBoost, lightGBM) algorithms, select the best model parameters through parameter adjustment and cross-validation to build a prediction model, and evaluate the classification efficiency of the model on the test set (the AUC of the training set and test set are 0.90 and 0.88 respectively), and finally select the prediction model of the lightGBM algorithm. Figure 5 As shown, the evaluation details of the training data set (train ROC curve) and the test data set (test ROC curve) by the receiver operating characteristic curve (ROC) and the area under the curve (AUC) are shown.
[0168] External data validation of the prediction model: According to the recruitment criteria, another independent cohort was re-collected, including 31 MSI-H gastrointestinal cancer subjects, who constituted the external data set. The relative abundance of the patients' microorganisms before treatment was used as the input data of the immunotherapy response prediction model, and the output data of the corresponding optimal prediction model was used as the parameters of the immunotherapy response prediction model for patients with MSI-H gastrointestinal cancer. Finally, the AUC of the prediction model for this dataset was 0.79. Figure 4 The external ROC curve is shown.
[0169] Step 3: Screening of bacterial combination markers:
[0170] Important feature evaluation: Evaluate the features of the above optimal prediction model. If the average SHAP value of the microorganism is greater than 0, it is considered to be an important microorganism. Call the SHAP tool to evaluate the important features of the prediction model of the lightGBM algorithm, and select the features as potential microbial markers (see Figure 5 ), and finally four important microbial characteristics were screened out, as shown in Table 2.
[0171] Table 2: Four important microbial features selected by the SHAP algorithm
[0172] Serial number Species classification English name Chinese name 1 Species Bacteroides_caccae Bacteroides faecalis 2 Species Veillonella_atypica Atypical Veillonella 3 Species Veillonella_parvula Veillonella parvula 4 Species Clostridiales_bacterium_1_7_47FAA bacterium_1_7_47FAA Clostridium
[0173] The mathematical formula for calculating the SHAP value is shown in the above equation (1).
[0174] In this embodiment, a device for evaluating and screening intestinal microorganism characteristics is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0175] This embodiment provides a device for evaluating and screening intestinal microbial characteristics, such as Figure 6 As shown, the device comprises:
[0176] The first acquisition module 601 is used to acquire a first target stool sample data set and a first clinical data set of a plurality of first users.
[0177] The processing module 602 is used to perform sequencing processing on the first target stool sample dataset to obtain an intestinal microbial abundance dataset.
[0178] The construction module 603 is used to construct a target prediction model based on the first clinical data set and the intestinal microbial abundance data set after being processed by multiple preset mainstream tree algorithms.
[0179] The feature evaluation and screening module 604 is used to perform feature evaluation and screening on the target prediction model using an interpretable algorithm to obtain a target intestinal microbial feature set.
[0180] In some optional implementations, the first acquisition module 601 includes:
[0181] The acquisition submodule is used to acquire initial stool sample data sets and first clinical data sets of multiple first users.
[0182] The determination submodule is used to determine a first target stool sample dataset based on the initial stool sample dataset and the first clinical dataset.
[0183] In some optional implementations, the determining submodule includes:
[0184] The determining unit is used to determine a plurality of second users from a plurality of first users according to preset requirements based on the initial stool sample data set and the first clinical data set.
[0185] The acquiring unit is configured to acquire first target stool sample data sets of a plurality of second users based on the initial stool sample data set.
[0186] In some optional implementations, the construction module 603 includes:
[0187] The screening submodule is used to obtain an initial intestinal microbial feature set based on the first clinical data set and the intestinal microbial abundance data set through a preset feature screening method.
[0188] The processing submodule is used to construct a target prediction model based on the initial intestinal microbial feature set after processing by a preset mainstream tree algorithm.
[0189] In some optional implementations, the processing submodule includes:
[0190] The evaluation and determination unit is used to evaluate multiple machine learning models and determine an initial prediction model.
[0191] The processing unit is used to adjust the parameters and evaluate the initial prediction model to obtain the target prediction model.
[0192] In some optional embodiments, the device further comprises:
[0193] The second acquisition module is used to acquire a second target stool sample data set and a second clinical data set of multiple second users.
[0194] A verification module is used to verify the target prediction model based on a second target stool sample data set and a second clinical data set.
[0195] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0196] The intestinal microbial characteristic evaluation and screening device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0197] The embodiment of the present invention also provides a computer device having the above Figure 6The screening device for evaluating gut microbial signature is shown.
[0198] See also Figure 7 , Figure 7 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 7 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 7 A processor 10 is taken as an example.
[0199] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0200] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.
[0201] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0202] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.
[0203] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0204] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.
[0205] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.
[0206] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A method for evaluating and screening intestinal microbial characteristics, characterized in that: The method comprises: Acquire a first target stool sample dataset and a first clinical dataset of a plurality of first users; Performing sequencing on the first target stool sample dataset to obtain an intestinal microbial abundance dataset; Based on the first clinical data set and the intestinal microbial abundance data set, a target prediction model is constructed after being processed by a plurality of preset mainstream tree algorithms; The target prediction model is characterized by evaluation and screening using an interpretable algorithm to obtain a target intestinal microbial feature set.
2. The method according to claim 1, characterized in that Acquiring a first target stool sample dataset and a first clinical dataset of a plurality of first users, including: Acquire initial stool sample datasets of a plurality of first users and the first clinical dataset; The first target stool sample dataset is determined based on the initial stool sample dataset and the first clinical dataset.
3. The method according to claim 2, characterized in that Determining the first target stool sample dataset based on the initial stool sample dataset and the first clinical dataset includes: Based on the initial stool sample data set and the first clinical data set, determining a plurality of second users from the plurality of first users according to preset requirements; Based on the initial stool sample dataset, the first target stool sample datasets of the plurality of second users are acquired.
4. The method according to claim 1, characterized in that: Based on the first clinical data set and the intestinal microbial abundance data set, a target prediction model is constructed after being processed by a preset mainstream tree algorithm, including: Based on the first clinical data set and the intestinal microbial abundance data set, an initial intestinal microbial feature set is obtained through a preset feature screening method; Based on the initial intestinal microbial feature set, the target prediction model is constructed after being processed by a preset mainstream tree algorithm.
5. The method according to claim 4, characterized in that Based on the initial intestinal microbial feature set, the target prediction model is constructed after being processed by a variety of preset mainstream tree algorithms, including: Based on the initial intestinal microbial feature set, multiple machine learning models are constructed after being processed by the multiple preset mainstream tree algorithms; Evaluating the multiple machine learning models and determining an initial prediction model; The initial prediction model is adjusted and evaluated to obtain the target prediction model.
6. The method according to claim 1, characterized in that The method further comprises: acquiring a second target stool sample dataset and a second clinical dataset of a plurality of second users; The target prediction model is verified based on the second target stool sample dataset and the second clinical dataset.
7. A device for evaluating and screening intestinal microbial characteristics, characterized in that: The device comprises: A first acquisition module, used to acquire a first target stool sample data set and a first clinical data set of a plurality of first users; A processing module, used for performing sequencing processing on the first target stool sample dataset to obtain an intestinal microbial abundance dataset; A construction module, for constructing a target prediction model based on the first clinical data set and the intestinal microbial abundance data set after being processed by a plurality of preset mainstream tree algorithms; The feature evaluation and screening module is used to perform feature evaluation and screening on the target prediction model using an interpretable algorithm to obtain a target intestinal microbial feature set.
8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the intestinal microbial characteristic evaluation and screening method according to any one of claims 1 to 6 by executing the computer instructions.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the intestinal microbial characteristic evaluation and screening method according to any one of claims 1 to 6.
10. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the intestinal microbial characteristic evaluation and screening method according to any one of claims 1 to 6.