Construction method and application of individual blood glucose improvement prediction model after dietary fiber intervention

By constructing an individual blood sugar improvement prediction model based on intestinal flora, the problem that dietary fiber intervention schemes in the prior art cannot target individual differences is solved, and an accurate prediction of individual blood sugar improvement effect is achieved, and a personalized treatment strategy is provided.

CN120072203AActive Publication Date: 2025-05-30SHANGHAI JIAOTONG UNIV +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510151253.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

The existing dietary fiber intervention scheme cannot target individual differences, resulting in different effects of blood sugar improvement in different individuals after receiving dietary fiber intervention, and lack of effective predictive models.

Method used

A prediction model for individual blood glucose improvement after dietary fiber intervention based on intestinal flora was constructed. By collecting metabolites before subjects, extracting DNA, high-throughput amplicon sequencing of 16S rRNA genes, analyzing the original abundance matrix of amplicon sequence variants, building a LightGBM classification model, and using a stratified 10-fold cross-validation method to create training sets and test sets, constructing microbiome-fiber scoring formulas and blood glucose benefit scores, and making predictions.

Benefits of technology

It successfully predicted the improvement of blood sugar in different individuals after receiving dietary fiber intervention. The verification cohort was verified, showing that the prediction model is highly applicable in different populations and can help individuals choose effective treatment strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072203A_ABST
    Figure CN120072203A_ABST
Patent Text Reader

Abstract

The invention discloses a method for constructing an individual blood glucose improvement prediction model after dietary fiber intervention, which is characterized by comprising the following steps: S1, collecting metabolites of a subject before intervention, extracting DNA (Deoxyribose Nucleic Acid), and analyzing to obtain an amplicon sequence variant; aSVs) of the original abundance matrix; s2, constructing a Light GBM (Light Graded-Boosting Mach) classification model, and carrying out the classification of the light GBM (Light Graded-Boosting Mach) according to the classification model of the light GBM (Light Graded-Boosting Mach); s3, creating a training set and a test set by adopting a layered 10-fold cross validation method; s4, constructing a microbiome-fiber scoring formula and a blood glucose benefit score; and S5, according to the corresponding relationship between the microbiome-fiber score and the blood glucose benefit score, performing prediction through the microbiome-fiber score. The model provided by the invention can predict whether a subject will benefit from dietary fiber intervention, and help the subject to select an effective treatment strategy before intervention is started.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of bioengineering and gut health, and particularly relates to the research on the response of gut microbiota to dietary fiber intervention, the evaluation and prediction of individual intervention results. Specifically, it relates to a method for constructing and applying a prediction model for improving individual blood glucose after dietary fiber intervention. Background Art

[0002] The incidence of diabetes has been increasing year by year and has become a major threat to national health. Gut microbiota and dietary fiber play important roles in the intervention and treatment of diabetic patients. Not only does the intake of dietary fiber regulate blood glucose levels by slowing down the digestion and absorption of carbohydrates, thus helping to control blood glucose levels, but also the production of short-chain fatty acids (SCFAs). Gut microbiota produce acetic acid and butyric acid by fermenting dietary fiber, and these short-chain fatty acids can affect the release of gut hormones (such as GLP-1 and PYY), thereby regulating blood glucose levels. Certain dietary fiber components can selectively promote the growth of beneficial microbiota, such as microbiota that produce short-chain fatty acids, thus improving the metabolic health of the host. Therefore, gut microbiota can affect the host's glucose metabolism through multiple pathways, and the imbalance of gut microbiota is closely related to the occurrence and development of diabetes.

[0003] Dietary fiber intervention is an intervention method targeting gut microbiota, which can improve human health by regulating gut microbiota. Specific members of gut microbiota can ferment and utilize dietary fiber to produce products beneficial to human health. More and more evidence shows that the impact of dietary intervention on health shows great individual differences, which is related to the personalized response of gut microbiota to diet. Since humans do not have the ability to digest and utilize dietary fiber, the efficacy of dietary fiber intervention depends more on the individuality of gut microbiota.

[0004] Many dietary fiber intervention studies have observed that different individuals benefit differently after receiving dietary fiber intervention. Not all subjects benefit from dietary fiber intervention. Specifically, current dietary fiber programs cannot be targeted. The dietary programs prescribed by doctors may be the same, but due to individual differences, the actual effects are different. Therefore, for the same dietary program, some patients have good improvement effects, while some patients have poor improvement effects.

[0005] At the present stage, a successful model for predicting individual benefits based on gut microbiota has not been established. If a prediction model for improving individual blood glucose after dietary fiber intervention can be realized, it will be very meaningful to accurately evaluate the effectiveness of individuals in advance to a certain extent. Summary of the Invention

[0006] In view of the above defects of the prior art, the present invention provides a method for constructing and applying an individual blood glucose improvement prediction model after dietary fiber intervention. The intestinal flora before dietary fiber intervention was used to successfully predict the improvement of blood glucose in different individuals after receiving dietary fiber intervention.

[0007] The specific technical solution adopted by the present invention is as follows:

[0008] In the first aspect of the present invention, there is provided a method for constructing an individual blood glucose improvement prediction model after dietary fiber intervention, comprising the following steps:

[0009] S1: Collect metabolites before intervention of the subjects, extract DNA, and analyze to obtain the original abundance matrix of Amplicon Sequence Variants (ASVs);

[0010] S2: Construct a LightGBM (Light Gradient-Boosting Machine) classification model;

[0011] S3: Create a training set and a test set by using a stratified 10-fold cross-validation method;

[0012] S4: Construct a microbiome-fiber score formula and a blood glucose benefit score;

[0013] S5: According to the corresponding relationship between the microbiome-fiber score and the blood glucose benefit score, predict based on the microbiome-fiber score.

[0014] In some specific embodiments, in S1, after extracting DNA, the V3-V4 region of the 16S rRNA gene is subjected to high-throughput amplicon sequencing using a sequencing platform, and data analysis is performed using software to obtain the original abundance matrix of amplicon sequence variants.

[0015] In some specific embodiments, in S2, based on the abundances of amplicon sequence variants with a prevalence rate of more than 20% among the subjects, a LightGBM (Light Gradient-Boosting Machine) classification model is constructed.

[0016] In some specific embodiments, in S3, the data set is initially divided into 10 mutually exclusive subsets of equal size. In each round, one subset is used as the test set, and the remaining subsets are used as the training set. The final result is obtained by averaging the results of all 10 subsets. The area under the receiver operating characteristic curve (ROC-AUC) is used as an index to quantify the performance of the classifier.

[0017] In some specific embodiments, for feature selection, first calculate the average feature importance of stratified 10-fold cross-validation, and then calculate the weighted average feature importance, where the weighting coefficient is the frequency of feature importance greater than 0. Then, based on the stepwise forward selection method of stratified 10-fold cross-validation, determine the optimal number of features according to the weighted average importance ranking.

[0018] After determining the optimal number of features, use the optimal features to construct a LightGBM classification model, and use SHAP (Shapley additive explanations) to calculate the contribution value (i.e., SHAP value) of each feature to the prediction result of each sample in the LightGBM classification model. SHAP is a method for explaining the prediction results of machine learning models. By calculating the contribution of each feature to the model's prediction results, it provides global and local explanations for the model.

[0019] The calculation of SHAP values is based on the following formula

[0020]

[0021] Φi represents the SHAP value of feature i. F represents the set of all features. S represents the subset of features that does not include feature i. XS represents the values of the input features in subset S. fS(XS) represents the predicted output of the model when given subset S.

[0022] In this study, the shapviz function in the R language package shapviz (version 0.9.5) was used to calculate the SHAP value of each feature in the LightGBM classification model.

[0023] In some specific embodiments, according to the calculated SHAP values of the ASVs above, construct the Microbiome-Fiber Score (MFS) formula as follows:

[0024]

[0025] Where MFS i represents the Microbiome-Fiber Score value of individual i, S ij is the MFS value of the j-th ASV in individual i, and n is the total number of ASVs; S ij is defined as: if the contribution value (i.e., SHAP value) of the j-th ASV in individual i is greater than 0, then S ij = 1; otherwise, S ij = 0.

[0026] Further, the contribution value of the j-th ASV in individual i is obtained using the SHAP (Shapley additive explanations) method. For each sample and each feature, the model generates a SHAP value; the contribution value is the SHAP value.

[0027] In some specific embodiments, the blood glucose benefit score is obtained by adding up the scores of blood glucose indicators, and the blood glucose indicators include fasting plasma glucose (FPG), postprandial two-hour blood glucose (PBG), and glycated hemoglobin (HbA1c).

[0028] Further, if the change value (% change) of the blood glucose indicator is less than zero, it is recorded as 1 point, otherwise it is recorded as 0 point; the change value (% change) of the blood glucose indicator = (post-intervention indicator - baseline indicator) ÷ baseline indicator × 100%.

[0029] In some specific embodiments, individuals with a blood glucose benefit score greater than or equal to 2 are high responders, and those with a score less than 2 are low responders.

[0030] In some specific embodiments, the MFS is divided into three ranges; the first range, including MFS values less than or equal to 17, only contains subjects with blood glucose benefit scores of 0 and 1, indicating that these subjects have limited or no blood glucose benefit after receiving dietary fiber intervention; the second range, including subjects with MFS values greater than or equal to 18 and less than or equal to 22, involves all blood glucose benefit scores, indicating that the blood glucose benefit of these subjects after receiving dietary fiber intervention is unclear; the third range, including subjects with MFS values greater than or equal to 23, contains subjects with blood glucose benefit scores of 2 and 3, indicating that these subjects can benefit from dietary fiber intervention.

[0031] In the second aspect of the present invention, there is provided an application of an individual blood glucose improvement prediction model after dietary fiber intervention, for predicting whether a subject will benefit from dietary fiber intervention.

[0032] The application of the individual blood glucose improvement prediction model after dietary fiber intervention includes the following steps:

[0033] S1: Collect fecal samples from the subject before the intervention, extract DNA, perform high-throughput amplicon sequencing on the V3-V4 region of the 16S rRNA gene using a sequencing platform, and use software for data analysis to obtain the original abundance matrix of amplicon sequence variants (ASVs);

[0034] S2: Adopt the LightGBM (Light Gradient-Boosting Machine) classification model described in claim 3;

[0035] S3: Construct the microbiome-fiber score formula according to the construction method described in claim 4, and make predictions based on the scores.

[0036] For the further application described above, in step S1, after extracting DNA, use a sequencing platform to perform high-throughput amplicon sequencing on the V3-V4 region of the 16S rRNA gene, and use software for data analysis to obtain the original abundance matrix of amplicon sequence variants.

[0037] In some specific embodiments, the number of amplicon sequence variants is 44.

[0038] Compared with the prior art, the present invention has the following beneficial effects:

[0039] It has been successfully predicted whether a subject will benefit from dietary fiber intervention through a machine learning model based on baseline gut microbiota characteristics, and it has been verified in an external validation cohort. This indicates that the prediction model has strong applicability in different populations. The prediction model is constructed based on gut microbiota and can, through a non-invasive method, help subjects select effective treatment strategies before the start of the intervention.

[0040] The following will further illustrate the concept and technical effects of the present invention in conjunction with the accompanying drawings to fully understand the purpose, features, and effects of the present invention. Description of the Drawings

[0041] Figure 1 It is a flowchart of the population cohort intervention. Among them, Figure 1 a is the flowchart of the discovery cohort intervention; Figure 1 b is the flowchart of the validation cohort intervention.

[0042] Figure 2 It is the classification of the improvement of the blood glucose of the subjects in the population cohort. Among them, Figure 2 a is the classification of the improvement of the blood glucose of the subjects in the discovery cohort; Figure 2 b is the classification of the improvement of the blood glucose of the subjects in the validation cohort.

[0043] Figure 3 It is an ASV diagram that can predict high or low responders. Among them, Figure 3 a is the average ROC-AUC diagram under different numbers of ASVs. Figure 3 b is the weighted average importance and abundance diagram of the ASVs selected by the LightGBM model in the discovery cohort; the clustering tree shows the associations between ASVs, which are determined by the Spearman correlation coefficient based on the average abundance of all subjects, and the heatmap shows the average abundance (z-score transformation) of each ASV in low responders and high responders. Figure 3c is the abundance plot of ASVs selected by the LightGBM model in the validation cohort; the clustering tree shows the associations between ASVs, which are determined by the Spearman correlation coefficient based on the average abundance of all subjects, and the heatmap shows the average abundance (z-score transformed) of each ASV in low and high responders.

[0044] Figure 4 are the model prediction result plots for the discovery cohort and the validation cohort. Among them, Figure 4 a is the receiver operating characteristic (ROC) curve plot for the discovery cohort based on 10-fold cross-validation; the mean area under the receiver operating characteristic curve (AUC) is shown as mean ± standard deviation (SD). Figure 4 b is the ordered logistic regression analysis plot of the correlation between MFS and the glycemic benefit score in the discovery cohort; the bar chart (mean ± standard error (SEM)) shows the difference in MFS between low and high responders in the GPD study. Figure 4 c is the receiver operating characteristic (ROC) curve plot for the validation cohort. Figure 4 d is the ordered logistic regression analysis plot of the correlation between MFS and the glycemic benefit score in the validation cohort; the bar chart shows the difference in MFS between low and high responders in the GLC study.

[0045] Figure 5 is the corresponding relationship plot between MFS and the glycemic benefit score. Detailed implementation mode

[0046] In order to make the technical means, creative features, achieved purposes and effects of the invention easy to understand, the present invention will be further described below in conjunction with specific illustrations. However, the present invention is not limited to the following implemented cases.

[0047] It should be noted that the ratios, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the limiting conditions for the implementation of the present invention. Therefore, they do not have technical substantial significance. Any improvement without creative labor should still fall within the scope covered by the technical content disclosed in the present invention without affecting the effects that the present invention can produce and the purposes that can be achieved.

[0048] Example 1 Model construction

[0049] (I) Recruitment and intervention experiment for the discovery cohort

[0050] The intervention flow chart for the discovery cohort is as Figure 1As shown in Figure a. We recruited 802 volunteers with pre-diabetes (as defined by the World Health Organization) in the eastern region of China and randomly assigned them to two groups: a control group (Group U, n = 393) and a dietary fiber supplement intervention group (Group W, n = 409). The control group maintained their normal lives; the dietary fiber supplement intervention group received an intervention of 45 grams of dietary fiber supplement per day for 6 months. Finally, 331 volunteers in Group W completed the 6-month intervention. A prediction model was constructed using the abundance data of the gut microbiota at baseline of the volunteers in Group W of the discovery cohort to predict the improvement of blood glucose in different individuals after the intervention. The above dietary fiber supplement is from Shanghai Jiuben Technology Co., Ltd.

[0051] (II) Classification of the degree of blood glucose improvement

[0052] According to the change value (% change) of the blood glucose index of the subjects in the discovery cohort after the intervention, they were divided into low-responders and high-responders. The blood glucose indices include fasting plasma glucose (FPG), 2-hour postprandial blood glucose (PBG), and glycated hemoglobin (HbA1c). If the change value (% change) of a blood glucose index is less than zero, it is recorded as 1 point, otherwise it is recorded as 0 point. Then, the scores of the three blood glucose indices are added together, and the total score obtained is the blood glucose benefit score of this individual. Individuals with a blood glucose benefit score greater than or equal to 2 are high-responders, and those less than 2 are low-responders. The calculation formula of % change is as follows: % change = (post-intervention index - baseline index) ÷ baseline index × 100%. The responder classification of the discovery cohort is as shown in Figure 2 Figure a.

[0053] (III) Construction of the prediction model

[0054] Feces of the volunteers at baseline were collected, DNA was extracted, and the V3-V4 region of the 16S rRNA gene of the feces was subjected to high-throughput amplicon sequencing using the Illumina Miseq sequencing platform, and Qiime 2 software was used for data analysis. First, the "Cutadapt" plugin was used to remove adapters and primers in the sequences, and the sequences were trimmed with DADA2 according to the sequencing quality of the bases. After filtering, denoising, dereplicating, and merging, the original abundance matrix of amplicon sequence variants (ASVs) was obtained. Since there are differences in the sequencing depth among different samples, the ASV abundance table was transformed by relative log expression (RLE), and the subsequent analysis was carried out using the transformed abundance information.

[0055] Based on the abundances of ASVs with a prevalence exceeding 20% in the discovery cohort W group, a LightGBM (Light Gradient-Boosting Machine) classification model was constructed to distinguish low-responders and high-responders. A stratified 10-fold cross-validation method was used to create the training set and the test set. The dataset was initially divided into 10 mutually exclusive subsets of equal size. In each round, one subset was used as the test set, and the remaining subsets were used as the training set. The final result was obtained by averaging the results of all 10 subsets. The area under the receiver operating characteristic curve (ROC-AUC) was used as a metric to quantify the performance of the classifier. The average ROC-AUC of the stratified 10-fold cross-validation was considered as the accuracy measure of the model.

[0056] For feature selection, first, the average feature importance of the stratified 10-fold cross-validation was calculated, and then the weighted average feature importance was calculated, where the weighting coefficient was the frequency of the feature importance being greater than 0. Subsequently, based on the stepwise forward selection method of the stratified 10-fold cross-validation, the optimal number of features was determined according to the weighted average importance ranking. The ROC-AUC for different numbers of ASVs and the number, abundance, and taxonomic status of the finally selected ASVs are shown in Figure 3 a, 3b.

[0057] The hyperparameters of the classification model were optimized through stratified 10-fold cross-validation grid search and manual adjustment. The hyperparameters "num_leaves", "learning_rate", "max_depth", and "n_estimators" represent the number of leaves of the decision tree, the iteration speed, the maximum depth of the decision tree, and the number of boosting iterations, respectively. According to the suggestions in the LightGBM official documentation (https: / / lightgbm.readthedocs.io / en / latest / index.html), the hyperparameter ranges were set as follows: learning_rate [0.01, 0.1], n_estimators [100, 1000], max_depth [3, 7], num_leaves [7, 127]. The finally determined model parameters were: learning_rate = 0.05, n_estimators = 115, max_depth = 7, num_leaves = 80.

[0058] The LightGBM classification model was constructed using the lightgbm package (version 3.3.5). Stratified 10-fold cross-validation was performed using the createMultiFolds function of the caret package (version 6.0-94). Feature importance was obtained using the lgb.importance function of the lightgbm package. Hyperparameter tuning was performed using the tidymodels package (version 1.2.0) and the bonsai package (version 0.2.1). The pROC package (version 1.18.5) was used for ROC curve analysis, and the DeLong method was used to calculate the area under the curve.

[0059] Based on the above method, we identified 44 ASVs ( Figure 3 b, the DNA sequences are shown in Table 1), and these 44 ASVs can be used to effectively distinguish low-responders and high-responders.

[0060] Table 1 List of 44 amplicon sequence variants

[0061]

[0062]

[0063]

[0064]

[0065]

[0066]

[0067]

[0068]

[0069]

[0070]

[0071]

[0072]

[0073]

[0074]

[0075]

[0076]

[0077]

[0078]

[0079] To analyze the contribution of ASVs to the classification model, we used the SHAP (Shapley additive explanations) method to visualize the contribution of ASVs. SHAP is a game-theory-based method that can be used to analyze the contribution of features to a classification model. For each prediction sample and each feature, the model generates a SHAP value, which represents the contribution value of each feature. We used the shapviz function of the shapviz package (version 0.9.2) to calculate the SHAP values of each ASV in each sample. To quantify the individual responsiveness to dietary fiber, we constructed a microbiome-fiber score (MFS) based on the SHAP values of these 44 ASVs in each subject. The construction formula of the microbiome-fiber score (MFS) is shown as follows:

[0080]

[0081] where MFS i represents the microbiome-fiber score value of individual i, S ij is the MFS value of the j-th ASV in individual i, and n is the total number of ASVs. S ij is defined as: if the SHAP value of the j-th ASV in individual i is greater than 0, then S ij = 1; otherwise, S ij = 0.

[0082] The results showed that the MFS of high responders was significantly higher than that of low responders, and there was a significant positive correlation between MFS and the glycemic benefit score ( Figure 4 b). This indicates that the MFS based on gut microbiome features (44 ASVs) can be used to predict and quantify the individual responsiveness of prediabetic subjects in the discovery cohort to dietary fiber intervention. The regression analysis between MFS and the glycemic benefit score was performed using the lrm function in the rms package (version 6.8-0).

[0083] We further divided MFS into three ranges according to the correspondence between MFS and the glycemic benefit score. The first range, including MFS values less than or equal to 17, only contains subjects with glycemic benefit scores of 0 and 1, indicating that these volunteers had limited or no glycemic benefit after receiving dietary fiber intervention ( Figure 5)。The second range, including subjects with MFS values greater than or equal to 18 and less than or equal to 22, involves all blood glucose benefit scores, indicating that the blood glucose benefits of these volunteers after dietary fiber intervention are unclear( Figure 5 )。The third range, including subjects with MFS values greater than or equal to 23, includes subjects with blood glucose benefit scores of 2 and 3, indicating that these volunteers can benefit from dietary fiber intervention( Figure 5 )。

[0084] Example 2 Model Application

[0085] (I) Recruitment and Intervention Experiment of Validation Cohort

[0086] The intervention flow chart of the validation cohort is as shown in Figure 1 Figure b. The predictive model constructed in the discovery cohort was validated in an independent population cohort. Twenty volunteers were recruited in Shanghai, China, including 10 overweight or obese subjects (3 of whom were pre-diabetic) and 10 type 2 diabetic subjects. All volunteers received dietary fiber supplement intervention. The dietary fiber supplement intervention lasted for 14 days. Each person consumed 18 grams of dietary fiber per day for the first 7 days and 36 grams per day for the next 7 days. Finally, 19 volunteers completed the intervention. In addition, volunteers (n = 2) who could not be accurately grouped due to missing blood glucose data were excluded. The dietary fiber supplement used in the validation cohort was the same as that in the discovery cohort.

[0087] (II) Classification of Degree of Blood Glucose Improvement

[0088] According to the change value (%change) of blood glucose indicators after the intervention of volunteers, they were divided into low-responders and high-responders. Blood glucose indicators included fasting plasma glucose (FPG), mean continuous glucose (GCM), and glycated hemoglobin (HbA1c). If the change value (%change) of a blood glucose indicator was less than zero, 1 point was recorded; otherwise, 0 point was recorded. Then, the scores of the three blood glucose indicators were added up, and the total score obtained was the blood glucose benefit score of the individual. Individuals with a blood glucose benefit score greater than or equal to 2 were high-responders, and those less than 2 were low-responders. In the validation cohort, due to the lack of detection of postprandial two-hour blood glucose (PBG), continuous glucose monitoring values (CGM) were used instead. The calculation formula of %change is as follows: %change = (post-intervention indicator - baseline indicator) ÷ baseline indicator × 100%. The responder classification of the validation cohort is as shown in Figure 2 Figure b.

[0089] (III) Model Prediction Operation

[0090] The steps of the model prediction operation are as follows:

[0091] 1. Extract the DNA of the feces in the validation cohort using the same method as in the discovery cohort, and perform high-throughput amplicon sequencing on the V3-V4 region of the 16S rRNA gene of the feces using the Illumina Miseq sequencing platform.

[0092] 2. Process the sequencing data using the same method as in the discovery cohort to obtain the abundance matrix of Amplicon Sequence Variants (ASVs).

[0093] 3. Search for the 44 ASVs found in the discovery cohort in the validation cohort, and use the abundance information of these 44 ASVs in the validation cohort for subsequent analysis.

[0094] 4. Use the abundance information of the 44 ASVs in the discovery cohort to train the constructed LightGBM (Light Gradient-Boosting Machine) classification model. It is implemented using the lgb.train function of the lightgbm package (version 3.3.5).

[0095] 5. Use the abundance information of the 44 ASVs in the validation cohort as the input file to the trained model, and use the abundance information of the 44 ASVs in the validation cohort to predict the low-responders and high-responders in the validation cohort. It is implemented using the predict function of the lightgbm package (version 3.3.5).

[0096] 6. Input the grouping information of the validation cohort predicted by the model and the actual grouping information of the validation cohort into the roc function of the pROC package (version 1.18.5) to calculate the ROC curve and the area under the ROC curve.

[0097] 7. In the validation cohort, use the shapviz function of the shapviz package (version 0.9.5) to calculate the SHAP value of each ASV in each sample.

[0098] 8. Using the same method as in the discovery cohort, calculate the Microbiome-Fiber Score (MFS) of each sample in the validation cohort using the SHAP values obtained in the previous step.

[0099] 9. Perform a regression analysis between MFS and the blood glucose benefit score using the lrm function in the rms package (version 6.8-0).

[0100] (III) Model Prediction Results and Validation Cohort Results

[0101] The effectiveness of the classification model was verified using the data of the validation cohort, and it was found that the constructed LightGBM classification model could also well distinguish high responders and low responders in the validation cohort ( Figure 4 c). The microbiome-fiber score (MFS) was constructed using the same method, and it was also found in the validation cohort that the MFS of high responders was significantly higher than that of low responders, and the MFS was significantly positively correlated with the blood glucose benefit score in the validation cohort ( Figure 4 d). Figure 5 It is a corresponding relationship diagram between the MFS and the blood glucose benefit score. The MFS was divided into three score ranges, reflecting the degree of benefit from dietary fiber intervention, which was consistent with the blood glucose benefit score.

[0102] It can be seen that this model can be used to predict the benefit situation of different individuals after receiving dietary fiber intervention.

[0103] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative labor. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention based on the concept of the present invention through logical analysis, reasoning or limited experiments on the basis of the prior art should be within the protection scope determined by the claims.

Claims

1. A method for constructing a prediction model for individual blood sugar improvement after dietary fiber intervention, characterized in that: The steps include: S1: Collect metabolites of subjects before intervention, extract DNA, and analyze to obtain the original abundance matrix of Amplicon Sequence Variants (ASVs); S2: Build a LightGBM (Light Gradient-Boosting Machine) classification model; S3: Create training and test sets using stratified 10-fold cross validation method; S4: Construction of microbiome-fiber scoring formula and glycemic benefit score; S5: Based on the correspondence between the microbiome-fiber score and the glycemic benefit score, prediction was made by the microbiome-fiber score.

2. The method for constructing a prediction model for individual blood sugar improvement after dietary fiber intervention according to claim 1, characterized in that: In the S1 step, after DNA is extracted, a high-throughput amplicon sequencing is performed on the V3-V4 segment of the 16S rRNA gene using a sequencing platform, and data analysis is performed using software to obtain an original abundance matrix of amplicon sequence variants.

3. The method for constructing a prediction model for improving individual blood sugar after dietary fiber intervention according to claim 1, characterized in that: In the S2 step, a LightGBM (Light Gradient-Boosting Machine) classification model is constructed based on the abundance of amplicon sequence variants with a common rate exceeding 20% ​​in the subjects.

4. The method for constructing a prediction model for individual blood sugar improvement after dietary fiber intervention according to claim 1, characterized in that: Based on the amplicon sequence variants identified in step S3, the microbiome-fiber score (MFS) formula was constructed as follows: Among them, MFS i represents the microbiome-fiber score of individual i, S ij is the MFS value of the jth ASV in individual i, n is the total number of ASVs; S ij Defined as: If the contribution value of the jth ASV in individual i is greater than 0, then S ij =1; otherwise, S ij =0.

5. The method for constructing a prediction model for individual blood sugar improvement after dietary fiber intervention according to claim 4, characterized in that: The contribution value of the j-th ASV in individual i is obtained using the SHAP (Shapley additive explanations) method. For each sample and each feature, the model generates a SHAP value; the contribution value is the SHAP value.

6. The method for constructing a prediction model for individual blood sugar improvement after dietary fiber intervention according to claim 1, characterized in that: The blood glucose benefit score in step S4 is obtained by adding blood glucose index scores, and the blood glucose indexes include fasting blood glucose (FPG), two-hour postprandial blood glucose (PBG), and glycosylated hemoglobin (HbA1c).

7. The method for constructing a prediction model for improving individual blood sugar after dietary fiber intervention according to claim 6, characterized in that: The blood glucose index score is 1 point if the change value (% change) of the blood glucose index is less than zero, otherwise it is 0 point; the change value (% change) of the blood glucose index = (post-intervention index-baseline index) ÷ baseline index × 100%.

8. The method for constructing a prediction model for improving individual blood sugar after dietary fiber intervention according to claim 6, characterized in that: Individuals with a glycemic benefit score greater than or equal to 2 were high responders, and those with a glycemic benefit score less than 2 were low responders.

9. The method for constructing a prediction model for improving individual blood sugar after dietary fiber intervention according to claim 5, characterized in that: The MFS was divided into three ranges; the first range, including MFS values ​​less than or equal to 17, only included subjects with glycemic benefit scores of 0 and 1, indicating that these subjects had limited or no glycemic benefit after dietary fiber intervention; the second range, including subjects with MFS values ​​greater than or equal to 18 and less than or equal to 22, involved all glycemic benefit scores, indicating that these subjects had unclear glycemic benefit after dietary fiber intervention; the third range, including subjects with MFS values ​​greater than or equal to 23, included subjects with glycemic benefit scores of 2 and 3, indicating that these subjects could benefit from dietary fiber intervention.

10. Application of a prediction model for individual blood sugar improvement after dietary fiber intervention, characterized in that: Predict whether subjects will benefit from dietary fiber intervention.

11. The use according to claim 10, characterized in that: The steps include: S1: Fecal samples were collected from subjects before intervention, DNA was extracted, and high-throughput amplicon sequencing of the V3-V4 segment of the 16S rRNA gene was performed using a sequencing platform. The original abundance matrix of amplicon sequence variants (ASVs) was obtained using software for data analysis. S2: Using the LightGBM (Light Gradient-Boosting Machine) classification model described in claim 3; S3: Construct a microbiome-fiber scoring formula according to the construction method of claim 4, and make predictions based on the scores.

12. The use according to claim 11, characterized in that: In the S1 step, after DNA is extracted, a high-throughput amplicon sequencing is performed on the V3-V4 segment of the 16S rRNA gene using a sequencing platform, and data analysis is performed using software to obtain an original abundance matrix of amplicon sequence variants.

13. The use according to claim 11, characterized in that: There are 44 amplicon sequence variants.

Citation Information

Patent Citations

  • Gestational diabetes biomarkers of intestinal bacteria in early pregnancy as well as screening and application of gestational diabetes biomarkers

    CN113174444A

  • Bayesian optimization RF and Light GBM disease prediction method

    CN115050477A

  • Prediction method suitable for postprandial blood sugar response of type 1 diabetes patient

    CN118136246A

  • Microbial community-scale metabolic modeling predicts personalized short-chain-fatty-acid production profiles in the human gut

    US20240290416A1