A system and method for predicting the state of an individual's microbiome and providing personalized recommendations for maintaining or improving the state of the microbiome.

JP7912540B2Active Publication Date: 2026-08-28SOCIETE DES PRODUITS NESTLE SA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023528513
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-08-03
Filing Date
2021-11-24
Publication Date
2026-08-28
Estimated Expiration
2041-11-24

Smart Images

  • Figure 0007912540000002
    Figure 0007912540000002
  • Figure 0007912540000003
    Figure 0007912540000003
  • Figure 0007912540000004
    Figure 0007912540000004
Patent Text Reader

Abstract

The present invention relates to systems and methods for predicting an individual's microbiome status and providing personalized recommendations for maintaining or improving the microbiome status. In some embodiments of the invention, individual microbiome features are clustered based on the individual's responses to a questionnaire. In some embodiments, the method is implemented by a computer system. In some embodiments of the invention, personalized recommendations and dietary advice are given to the individual to maintain or improve the individual's microbiome status.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to systems and methods for predicting the state of an individual's microbiome and providing personalized recommendations for maintaining or improving the state of the microbiome. In some embodiments of the present invention, individual microbiome features are clustered based on their answers to a questionnaire. In some embodiments, the method is implemented by a computer system. In some embodiments of the present invention, personalized recommendations, dietary advice and nutritional advice are provided to the individual to maintain or improve the state of the individual's microbiome. [Background Art]

[0002] The gut microbiota hosts trillions of microorganisms, mainly bacteria living in the gut, particularly in the colon. Changes in the composition and function of the gut microbiota are associated with many diseases and conditions such as irritable bowel syndrome, inflammatory bowel disease, allergies, diabetes, cancer, asthma, and obesity (Dogra SK et al., "Front Microbiol.", 2020).

[0003] The composition of the microbiota is influenced by a variety of exogenous factors including diet, geographical location, ethnicity, exercise / physical activity, antibiotic use and the use of other types of medication (Rothschild D et al., "Nature.", 2018). However, these exogenous factors do not reliably predict the state of an individual's microbiome at any point in the individual's lifetime, which depends on both the composition of the endogenous microbiota and these exogenous factors.

[0004] Since no two microbiomes are identical between individuals, there is a need for methods and systems that provide personalized recommendations for microbiome health. Successful solutions for maintaining or improving microbiome health require assessment of the state of the microbiome before any recommendations or advice can be provided.

[0005] The methods and systems of the present invention for predicting microbiome state in relation to providing dietary and nutritional recommendations for maintaining or improving a healthy microbiome differ from prior art in which microbiome state has been used to predict disease outcomes such as type 2 diabetes (Reitmeier S et al., "Cell Host Microbe." 2020; Wu H et al., "Cell Metab." 2020), postprandial glucose response (Zeevi D et al., "Cell." 2015), NAFLD cirrhosis (Oh TG et al., "Cell Metab." 2020), NAFLD fibrosis (Loomba R et al., "Cell Metab." 2017), or host variables (physiological characteristics, lifestyle characteristics, and dietary characteristics) (Vujkovic-Cvijin I et al., "Nature." 2020). In short, the direction from input to output is reversed in the present invention compared to previous studies.

[0006] Current assessments of an individual's gut microbiome state are based on biological sampling (either fecal or plasma sampling, or the use of advanced technologies such as next-generation sequencing and complex bioinformatics analysis). This is time-consuming and requires numerous processing steps, ideally including the cryopreservation of fecal samples at -80 degrees Celsius; processing of fecal samples for DNA extraction; sequencing of extracted DNA; and complex bioinformatics processing to detect and identify the presence and abundance of numerous microorganisms (Jovel J et al., "Front Microbiol.", 2016). In addition to the time and cost required to process biological samples, many individuals are reluctant to provide fecal samples to the laboratory for processing to assess the diversity of their microbiome (Wilmanski T et al., "Nat Biotechnol.", 2019), and even do not provide plasma samples for processing (Vandeputte D et al., "FEMS Microbiol Rev.", 2017).

[0007] The present invention advantageously provides non-invasive methods and systems for evaluating the state of the microbiome in an individual, without requiring biological samples such as fecal or plasma samples. Furthermore, the present invention provides user-friendly systems and methods for an individual to evaluate the state of their own microbiome. In some embodiments, the systems and methods of the present invention are useful in helping an individual modify their diet, nutrition, and lifestyle in accordance with the state of their own microbiome. [Prior art documents] [Non-patent literature]

[0008] [Non-Patent Document 1] Dogra SK et al., "Front Microbiol," 2020. [Non-Patent Document 2] Rothschild D et al., "Nature," 2018. [Non-Patent Document 3] Reitmeier S et al., "Cell Host Microbe," 2020. [Non-Patent Document 4] Wu H et al., "Cell Metab." 2020. [Non-Patent Document 5] Zeevi D et al., "Cell." 2015. [Non-Patent Document 6] Oh TG et al., "Cell Metab." 2020. [Non-Patent Document 7] Loomba R et al., "Cell Metab." 2017. [Non-Patent Document 8] Vujkovic-Cvijin I et al., "Nature," 2020. [Non-Patent Document 9] Jovel J et al., "Front Microbiol," 2016. [Non-Patent Document 10] Wilmanski T et al., "Nat Biotechnol," 2019. [Non-Patent Document 11] Vandeputte D et al., "FEMS Microbiol Rev.", 2017. [Overview of the project]

[0009] The method and system of the present invention advantageously implement an artificial intelligence-based machine learning method to assess the state of an individual's gut microbiome from a set of questionnaires.

[0010] One advantage of the present invention is that individuals do not need to provide biological samples to obtain an estimate of the state of their microbiome. Instead, this is done by using a predictive model based on data provided by the user regarding their responses to a set of questionnaires to identify predictive features.

[0011] In some embodiments, the present invention determines the state of an individual's microbiome in relation to its position within a larger population distribution. For example, low, high, or low, not low, or high, not high, in terms of having low or not low, high or not low, or being combined to determine low, medium, and high, and possibly cross-confirmed by another low vs. high rating, are defined in various ways based on the distributions found in large general populations such as the American Gut Project (AGP) (McDonald D et al., "mSystems", 2018) and the Microba Discovery Database (MDD), Microba, Australia.

[0012] In some embodiments of the present invention, the system and method of the present invention evaluate and extract features from a questionnaire and rank these features in order of importance for determining the state of the microbiome.

[0013] One advantage of some embodiments of the present invention is that, for assessing the state of a microbiome, questionnaire responses from individual users are evaluated to personalize recommendations and advice, so as to provide advice for maintaining or improving the microbiome state of an individual.

[0014] Another advantage of some embodiments of the present invention is that the assessment of an individual's microbiome state is performed in consideration of weighting for the importance of individual characteristics; therefore, any relevant recommendations and proposals for improving the microbiome state are personalized.

[0015] Various embodiments of the disclosed system display a customized dashboard or other suitable user interface to the user, based on the user's input to the questionnaire, the predicted microbiome state, and the personalized advice for maintaining or improving the microbiome state.

[0016] In some embodiments, the disclosed system may be linked to automatically collect required input data from an activity meter, or other wearable devices such as a smart watch or a fitness tracker.

[0017] In some embodiments, the disclosed system may be linked to automatically collect required input data from food intake records captured by the user in various forms such as a food diary or an application that logs eating and drinking records.

[0018] In various embodiments, the system of the present disclosure may operate in cooperation with a laboratory or other testing facility that generates actual data relating to an individual using the system of the present disclosure. For example, in one embodiment, the disclosed system allows a user to submit a biological analysis report indicating biomarkers in a biological sample of the individual. In such embodiments, such testing and research reports may enable the system to optionally improve recommendations.

[0019] In some embodiments, the systems and methods disclosed herein can also be used by dietitians and healthcare professionals other than individual users.

[0020] Further advantages of the present disclosure will become apparent from the following "Description of Embodiments" and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] [Figure 1] FIG. 1 is a block diagram of an exemplary computer-implemented system for assessing microbiome status according to an embodiment of the present disclosure. [Figure 2] FIG. 2 is a schematic diagram showing a microbiome recommendation system with individual components, their mutual interfaces, and relevant inputs and outputs from the component units. [Figure 3A-1] FIG. 3 is a diagram showing the ROC performance of low-model versus non-low-model defined for three diversity metrics generated together with quartile-based bin definitions, wherein (i) shows the ROC performance of training in cross-validation mode. [Figure 3A-2] FIG. 4 is a diagram showing the ROC performance of low-model versus non-low-model defined for three diversity metrics generated together with quartile-based bin definitions, wherein (ii) shows the ROC performance of a hold-out / test set. [Figure 3B-1] FIG. 5 is a diagram showing the ROC performance of low-model versus non-low-model defined for three diversity metrics generated together with mean and standard deviation-based bin definitions, wherein (i) shows the ROC performance of training in cross-validation mode. [Figure 3B-2] FIG. 6 is a diagram showing the ROC performance of low-model versus non-low-model defined for three diversity metrics generated together with mean and standard deviation-based bin definitions, wherein (ii) shows the ROC performance of a hold-out / test set. [Figure 4A-1] FIG. 7 is a diagram showing the ROC performance of high-model versus non-high-model defined for three diversity metrics generated together with quartile-based bin definitions, wherein (i) shows the ROC performance of training in cross-validation mode. [Figure 4A-2] (ii) The ROC performance of high-model versus non-high-model defined for three diversity measures created together with bin definitions based on quartiles, and the figure shows the ROC performance of the holdout / test set. [Figure 4B-1] This figure shows the ROC performance of high-model versus non-high-model defined for three diversity measures created together with bin definitions based on mean and standard deviation, and (i) the ROC performance of training in cross-validation mode. [Figure 4B-2] (ii) The figure shows the ROC performance of high-model versus non-high-model defined for three diversity measures created together with bin definitions based on mean and standard deviation, and the ROC performance of the holdout / test set. [Figure 5A-1] This figure shows the ROC performance of low-model versus high-model defined for three diversity measures created together with bin definitions based on quartiles, and (i) the ROC performance of training in cross-validation mode. [Figure 5A-2] (ii) The ROC performance of the holdout / test set, defined for three diversity measures created together with bin definitions based on quartiles. [Figure 5B-1] This figure shows the ROC performance of low-model versus high-model defined for three diversity measures created together with bin definitions based on mean and standard deviation, and (i) the ROC performance of training in cross-validation mode. [Figure 5B-2] (ii) The figure shows the ROC performance of the holdout / test set, defined for three diversity measures created together with bin definitions based on mean and standard deviation, where (ii) the ROC performance of the holdout / test set. [Figure 6A-1]This figure shows feature importance plots for low vs. non-low models across three diversity measures, created with bin definitions based on quartiles and input data containing these features with a response rate greater than 0.65. The top 30 features of this model are shown in these plots, (i) showing the average impact of each feature on the model output, sorted from high to low importance. The gray-to-black color gradient indicates the low to high value of that feature, and the vertical line at 0.00 defines the direction of the impact on the baseline class "low" (left side represents a negative impact on the model output, and right side represents a positive impact on the model output). [Figure 6A-2] This figure shows feature importance plots for low vs. non-low models across three diversity measures, created with bin definitions based on quartiles and input data with these features having a response rate greater than 0.65. The top 30 features of this model are shown in these plots, and (ii) the influence of features on model output is shown in more detail. The color gradient from gray to black indicates the value of the feature from low to high, and the vertical line at 0.00 defines the direction of the influence on the baseline class "low" (the left side represents a negative influence on model output, and the right side represents a positive influence on model output). [Figure 6B-1] This figure shows feature importance plots for low vs. non-low models across three diversity measures, created with bins defined based on mean and standard deviation, and input data with these features having a response rate greater than 0.85. The top 30 features of this model are shown in these plots, (i) showing the average impact of each feature on the model output, sorted from high to low importance. The gray-to-black color gradient indicates the low to high value of that feature, and the vertical line at 0.00 defines the direction of the impact on the criterion class "low" (left side represents a negative impact on the model output, and right side represents a positive impact on the model output). [Figure 6B-2]This figure shows feature importance plots for low vs. non-low models across three diversity measures, created with bin definitions based on mean and standard deviation, and input data with these features having a response rate greater than 0.85. The top 30 features of this model are shown in these plots, and (ii) the influence of features on model output in more detail. The color gradient from gray to black indicates the value of the feature from low to high, and the vertical line at 0.00 defines the direction of the influence on the criterion class "low" (the left side represents a negative influence on model output, and the right side represents a positive influence on model output). [Figure 7A-1] This figure shows feature importance plots for high versus low models against three diversity measures, created with bin definitions based on quartiles and input data with these features having a response rate greater than 0.65. The top 30 features of this model are shown in these plots, showing (i) the average impact of each feature on the model output sorted from high to low importance. The gray to black color gradient indicates the low to high value of that feature, and the vertical line at 0.00 defines the direction of the impact on the reference class "high" (left side represents a negative impact on the model output, and right side represents a positive impact on the model output). [Figure 7A-2] This figure shows feature importance plots for high versus low models against three diversity measures, created with bin definitions based on quartiles and input data with these features having a response rate greater than 0.65. The top 30 features of this model are shown in these plots, and (ii) the influence of features on model output in more detail. The color gradient from gray to black indicates the value of the feature from low to high, and the vertical line at 0.00 defines the direction of the influence on the reference class "high" (the left side represents a negative influence on model output, and the right side represents a positive influence on model output). [Figure 7B-1]This figure shows feature importance plots for low vs. non-low models across three diversity measures, created with bins defined based on mean and standard deviation, and input data with these features having a response rate greater than 0.85. The top 30 features of this model are shown in these plots, (i) showing the average impact of each feature on the model output, sorted from high to low importance. The gray-to-black color gradient indicates the low to high value of that feature, and the vertical line at 0.00 defines the direction of the impact on the criterion class "non-high" (left side represents a negative impact on the model output, and right side represents a positive impact on the model output). [Figure 7B-2] This figure shows feature importance plots for low vs. non-low models across three diversity measures, created with bins defined based on mean and standard deviation, and input data with these features having a response rate greater than 0.85. The top 30 features of this model are shown in these plots, and (ii) the influence of features on model output in more detail. The color gradient from gray to black indicates the value of the feature from low to high, and the vertical line at 0.00 defines the direction of the influence on the criterion class "non-high" (the left side represents a negative influence on model output, and the right side represents a positive influence on model output). [Figure 8A-1] This figure shows feature importance plots for low vs. high models across three diversity measures, created with bin definitions based on quartiles and input data containing these features with a response rate greater than 0.65. The top 30 features of this model are shown in these plots, (i) showing the average impact of each feature on the model output, sorted from high to low importance. The gray-to-black color gradient indicates the low to high value of that feature, and the vertical line at 0.00 defines the direction of the impact on the baseline class "low" (left side represents a negative impact on the model output, and right side represents a positive impact on the model output). [Figure 8A-2]This figure shows feature importance plots for low vs. high models across three diversity measures, created with bin definitions based on quartiles and input data containing these features with response rates greater than 0.65. The top 30 features of this model are shown in these plots, and (ii) the influence of features on model output is shown in more detail. The color gradient from gray to black indicates the value of the feature from low to high, and the vertical line at 0.00 defines the direction of the influence on the baseline class "low" (the left side represents a negative influence on model output, and the right side represents a positive influence on model output). [Figure 8B-1] This figure shows feature importance plots for low vs. high models across three diversity measures, created with bins defined based on mean and standard deviation, and input data containing these features with a response rate greater than 0.85. The top 30 features of this model are shown in these plots, (i) showing the average impact of each feature on the model output, sorted from high to low importance. The gray-to-black color gradient indicates the low to high value of that feature, and the vertical line at 0.00 defines the direction of the impact on the criterion class "low" (left side represents a negative impact on the model output, and right side represents a positive impact on the model output). [Figure 8B-2] This figure shows feature importance plots for low vs. high models across three diversity measures, created with bins defined based on mean and standard deviation, and input data with these features having a response rate greater than 0.85. The top 30 features of this model are shown in these plots, and (ii) the influence of features on model output is shown in more detail. The color gradient from gray to black indicates the value of the feature from low to high, and the vertical line at 0.00 defines the direction of the influence on the criterion class "low" (the left side represents a negative influence on model output, and the right side represents a positive influence on model output). [Figure 9]This figure shows the decomposition performance of the low-performing model versus the non-low-performing model. This low-performing versus non-low-performing model was defined by three diversity measures, created along with the mean and standard deviation, and bin definitions based on input data with these features having a response rate greater than 0.85. This figure shows the improvement in model performance (AUC values ​​of the ROC curves for the training (cross-validation) and holdout / test sets) as features considered important to the model were added one by one through SHAP analysis. [Figure 10] This figure shows the decomposition performance of high-performance models versus low-performance models. These high-performance versus low-performance models were defined by three diversity measures, along with the mean and standard deviation, and bin definitions based on input data with these features (response rate greater than 0.85). The figure shows the improvement in model performance (AUC values ​​of ROC curves for training (cross-validation) and holdout / test sets) as features considered important to the model were added one by one through SHAP analysis. [Figure 11] This figure shows the decomposition performance of the low model versus the high model. The low model versus the high model was defined by three diversity measures, created along with the mean and standard deviation, and bin definitions based on input data with these features (response rate greater than 0.85). This figure shows the improvement in model performance (AUC values ​​of the ROC curves for the training (cross-validation) and holdout / test sets) as features considered important to the model were added one by one through SHAP analysis. [Figure 12-1] This is a sample computer user interface for a questionnaire, illustrating an example of information for an individual user answering a set of questionnaires. [Figure 12-2] This is a sample computer user interface for a questionnaire, illustrating an example of information for an individual user answering a set of questionnaires. [Figure 12-3] This is a sample computer user interface for a questionnaire, illustrating an example of information for an individual user answering a set of questionnaires. [Figure 13]This figure shows an example of how the predicted microbiome state is displayed. [Figure 14] This figure shows the SHAP force plot for each individual. This plot indicates features that positively or negatively influence the prediction of the user's microbiome state, based on the user's responses. The net result of these factors is the final prediction of the individual's microbiome state. [Figure 15] This figure shows examples of personalized recommendations and advice for maintaining or improving the state of the microbiome. [Figure 16-1] This figure shows an example questionnaire necessary for the model. [Figure 16-2] This figure shows an example questionnaire necessary for the model. [Figure 16-3] This figure shows an example questionnaire necessary for the model. [Figure 16-4] This figure shows an example questionnaire necessary for the model. [Figure 16-5] This figure shows an example questionnaire necessary for the model. [Figure 16-6] This figure shows an example questionnaire necessary for the model. [Figure 16-7] This figure shows an example questionnaire necessary for the model. [Figure 16-8] This figure shows an example questionnaire necessary for the model. [Figure 16-9] This figure shows an example questionnaire necessary for the model. [Figure 16-10] This figure shows an example questionnaire necessary for the model. [Figure 16-11] This figure shows an example questionnaire necessary for the model. [Figure 16-12] This figure shows an example questionnaire necessary for the model. [Figure 17-1]This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. non-low models for three diversity measures created with input data having these features with response rates greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 17-2] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. non-low models for three diversity measures created with input data having these features with response rates greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 17-3] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. non-low models for three diversity measures created with input data having these features with response rates greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 17-4] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. non-low models for three diversity measures created with input data having these features with response rates greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 17-5]This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. non-low models for three diversity measures created with input data having these features with response rates greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 17-6] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. non-low models for three diversity measures created with input data having these features with response rates greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 17-7] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. non-low models for three diversity measures created with input data having these features with response rates greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 17-8] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. non-low models for three diversity measures created with input data having these features with response rates greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 17-9]This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. non-low models for three diversity measures created with input data having these features with response rates greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 17-10] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. non-low models for three diversity measures created with input data having these features with response rates greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 17-11] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. non-low models for three diversity measures created with input data having these features with response rates greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 17-12] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. non-low models for three diversity measures created with input data having these features with response rates greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 18-1]This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for high vs. non-high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "high," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "high" class. [Figure 18-2] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for high vs. non-high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "high," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "high" class. [Figure 18-3] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for high vs. non-high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "high," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "high" class. [Figure 18-4] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for high vs. non-high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "high," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "high" class. [Figure 18-5]This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for high vs. non-high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "high," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "high" class. [Figure 18-6] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for high vs. non-high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "high," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "high" class. [Figure 18-7] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for high vs. non-high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "high," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "high" class. [Figure 18-8] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for high vs. non-high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "high," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "high" class. [Figure 18-9]This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for high vs. non-high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "high," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "high" class. [Figure 18-10] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for high vs. non-high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "high," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "high" class. [Figure 18-11] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for high vs. non-high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "high," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "high" class. [Figure 18-12] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for high vs. non-high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "high," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "high" class. [Figure 19-1]This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 19-2] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 19-3] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 19-4] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 19-5]This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 19-6] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 19-7] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 19-8] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 19-9]This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 19-10] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 19-11] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 19-12] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 19-13]This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 19-14] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 19-15] This figure shows key exemplary features related to the definition of bins based on quartiles and SHAP dependency plots for low vs. high models for three diversity measures created with input data having these features with a response rate greater than 0.65. Note that the reference class here was "low," so the positive coefficient of the SHAP value for the corresponding x value of the feature indicates how much the model was influenced by this feature when predicting the "low" class. [Figure 20] This figure shows an overview of the classification results for the high-low model. The confusion matrix obtained for the best-performing model on the "high vs. low" classification task is shown, and this model used 20 metadata features that were a mixture of binary, categorical, and continuous. [Figure 21] This figure shows the relative importance of the 20 metadata features used in the high-low model. The best-performing model for the "high vs. low" classification task used 20 metadata features that were a mixture of binary, categorical, and continuous. Here, the relative weights of each feature in the model are shown, with the most important features listed from top to bottom. [Figure 22]This figure shows an overview of the classification results for low-to-non-low models. The confusion matrix obtained for the best-performing model on the "low vs. non-low" classification task is shown. [Figure 23] This figure shows the relative importance of metadata features used in the low-non-low model. The relative weights of the 20 features used by the low-non-low model are shown, with the most important features listed from top to bottom. Many of the 20 features (and their relative importance) overlapped with those determined to be optimal for the high-low model, such as physical activity, height, weight, and alcohol consumption. Non-overlapping features in the low-non-low model included stress, vegetable serves, cat or dog ownership, overseas travel, smoking, and bloating. [Figure 24] This figure shows the architecture of the ensemble modeling. It is a schema used for ensemble modeling (also known as stacked modeling) to predict three categories of diversity (low, medium, and high), and binary classification for "high vs. low" with a continuous output model using thresholds performed best. [Figure 25] This figure shows the results of ensemble modeling, specifically the performance results for "low-medium-high" modeling tasks using the ensemble model approach. [Figure 26] This figure shows an overview of the classification results of the ensemble model, specifically the confusion matrix obtained for the "low-medium-high" classification of diversity. [Figure 27] This figure shows a summarized set of key results, representing an integrated set of features that combine the main results obtained from the AGP and MDD datasets. [Modes for carrying out the invention]

[0022] definition The "intestinal microbiota" is the composition of microorganisms (including bacteria, archaea, and fungi) that inhabit the digestive tract.

[0023] The term "gut microbiota" can encompass both the "gut microbiota" itself and its "theatre of activity," which may include its structural elements (nucleic acids, proteins, lipids, polysaccharides), metabolites (signaling molecules, toxins, organic molecules, and inorganic molecules), and molecules produced by the coexisting host and structured by the surrounding environmental conditions (e.g., Berg, G. et al. (2020) "Microbiome" Vol. 8 (No. 1) pp. 1-22).

[0024] Therefore, in this invention, the term "intestinal microbiota" can be used interchangeably with the term "intestinal microbiota."

[0025] The "state of the microbiome" can be assessed by several different measurements, including determining the alpha diversity of bacteria found in the gut.

[0026] "Alpha diversity of bacteria found in the gut" summarizes the composition of an ecological community in terms of its "richness" (number of taxa), "uniformity" (distribution of group abundances), or both. In gut microbial ecology, analyzing alpha diversity from amplicon sequencing data is a common first approach to assess differences between environments. Generally, improving or maintaining alpha diversity of microbial species in the gut is an indicator of a healthy microbiome.

[0027] An "operational taxonomic unit" (OTU) is an operational definition used to classify groups of closely related individuals. The term "OTU" also refers to a cluster of organisms grouped by the DNA sequence similarity of a particular taxonomic marker gene (molecular OTU). OTUs are a practical substitute for "species" (microorganisms or metazoans) at different taxonomic levels when there is no conventional biological classification system available for macroscopic organisms. For several years, OTUs were the most commonly used unit of diversity, particularly when analyzing small subunit 16S (for prokaryotes) or 18S rRNA (for eukaryotes) marker gene sequence datasets.

[0028] "Faith Phylogenetic Diversity" (Faith PD) is the most commonly used phylogenetic index. Faith PD is a phylogenetic analog of taxone abundance and is expressed as the number of tree units found in a sample. A decrease in microbial PD in the human body may indicate reduced resilience associated with many human diseases.

[0029] The "Shannon Index" is a measure of diversity, not abundance. It measures the number of OTUs (abundance) in a sample, but scales them based on the homogeneity of the community. For example, if a control has more OTUs, but only a minority of those OTUs are dominant in the sample, it will be reported as having lower Shannon diversity than a community where the OTUs were evenly distributed.

[0030] In this invention, the "subject" may be a mammal, particularly a human. A human may be a male and / or female. For example, a human may be an adult, for example, an adult aged 18 or older. Furthermore, for example, an adult may be 30 or older, 40 or older, or 50 or older. According to one embodiment of the present invention, an adult may be 18 to 99 years old, preferably 20 to 70 years old. A mammal may be a pet animal, particularly a dog or a cat.

[0031] In some embodiments, the systems and methods of the present invention contribute to the assessment of the state of a microbiome in question by providing different methods for estimating whether the microbiome diversity is "low" or "not low", "high" or "not high", or "low", "medium", or "high" with respect to the microbiome distribution observed in a normal population, with respect to the α-diversity parameter of gut bacteria.

[0032] In some embodiments, the low α-variability group and the non-low α-variability group are defined as "low" if they fall below the first or lower quartile of the population OTU distribution, and as "non-low" if they remain within the distribution.

[0033] In some embodiments, the low alpha diversity group and the non-low alpha diversity group are defined as "low" if they fall below the first or lower quartile of the population FAITH PD distribution, and as "non-low" if they remain within the distribution.

[0034] In some embodiments, the low α-variability group and the non-low α-variability group are defined as "low" if they fall below the first or lower quartile of the population Shannon distribution, and as "non-low" if they remain within the distribution.

[0035] In some embodiments, the low α diversity group and the non-low α diversity group are defined as "low," which is below the first or lower quartile in these three distributions of OBSERVEDOTU, FAITHPD, and SHANNON, and "non-low," which is the remaining distribution of OBSERVEDOTU, FAITHPD, and SHANNON combined.

[0036] In some embodiments, the high α diversity group and the non-high α diversity group are defined as "high" being above the third quartile or upper quartile of the population Observedotu distribution, and "non-high" being the remainder of the distribution.

[0037] In some embodiments, the high α diversity group and the non-high α diversity group are defined as follows: "high" is defined as being above the third quartile or upper quartile of the population FAITHPD distribution, and "non-high" is defined as the remainder of the distribution.

[0038] In some embodiments, the high α diversity group and the non-high α diversity group are defined as "high" being above the third quartile or upper quartile of the population Shannon distribution, and "non-high" being the remainder of the distribution.

[0039] In some embodiments, the high α diversity group and the non-high α diversity group are defined as follows: "High" refers to the third quartile or upper quartile in these three distributions of OBSERVEDOTU, FAITHPD, and SHANNON, while "Non-High" refers to the three distributions of OBSERVEDOTU, FAITHPD, and SHANNON combined for the rest of the distribution.

[0040] In some embodiments, the low α diversity group, the non-low α diversity group, the high α diversity group, and the non-high α diversity group are defined as those in different population datasets having different numerical cutoffs, which would be obvious to those skilled in the art.

[0041] In some embodiments, the low-alpha diversity group and the high-alpha diversity group are defined as "low," which is lower than the first quartile or lower quartile of the OBSERVEDOTU distribution, and "high," which is higher than the third quartile or upper quartile of the OTU distribution.

[0042] In some embodiments, the low-alpha diversity group and the high-alpha diversity group are defined as "low," being lower than the first quartile or lower quartile of the FAITHPD distribution, and "high," being higher than the third quartile or upper quartile of the FAITHPD distribution.

[0043] In some embodiments, the low-alpha diversity group and the high-alpha diversity group are defined as "low," being lower than the first quartile or lower quartile of the Shannon distribution, and "high," being higher than the third quartile or upper quartile of the Shannon distribution.

[0044] In some embodiments, the low-α diversity group and the high-α diversity group may be defined by different numerical cutoffs, which will be apparent to those skilled in the art.

[0045] In some embodiments, the “low” α diversity group is defined as data points smaller than the mean minus the standard deviation in any of the diversity measures, and the “non-low” α diversity group is defined as the remaining data points.

[0046] In some embodiments, the “low” α diversity group is defined as data less than the first quartile or the lower quartile minus the interquartile range in any of the diversity measures, and the “non-low” α diversity group is defined as the remaining data.

[0047] In some embodiments, the “low” α diversity group is defined as data smaller than the first quartile or the lower quartile minus an interquartile range of 1.5 times, and the “non-low” α diversity group is defined as the remaining data.

[0048] In some embodiments, the “high” α diversity group is defined as data greater than the sum of the mean and standard deviations in one of the diversity measures, and the “non-low” α diversity group is defined as the remaining data.

[0049] In some embodiments, the “high” α diversity group is defined as data greater than the first quartile or the lower quartiles added to the interquartile range in any of the diversity measures, while the “non-high” α diversity group is defined as the remaining data.

[0050] In some embodiments, the “low” α diversity group is defined as data points smaller than the mean minus the standard deviation, and the “high” α diversity group is defined as data points larger than the mean plus the standard deviation.

[0051] In some embodiments, the “low” α diversity group is defined as data points smaller than the first quartile or lower quartile minus the interquartile range in either of the diversity measures, and the “high” α diversity group is defined as data points larger than the first lower quartile plus the interquartile range.

[0052] In some embodiments, the groups, individually or together, are evaluated against any of the α-variability scales: Low vs. Non-Low: A comparison is made between the ≤10%, ≤15%, ≤20%, ≤25%, or ≤30% of the data in the low category and the >10%, >15%, >20%, >25%, or >30% of the data in the remaining non-low category. For high vs. non-high cases: compare the data in the high category (≥90%, ≥85%, ≥80%, ≥75%, or ≥70%) with the remaining data in the non-high category (<90%, <85%, <80%, <75%, or <70%). In the case of low vs. high, it is defined in the following format: ≤10% at low vs. ≥90% at high, or ≤15% at low vs. ≥85% at high, or ≤20% at low vs. ≥80% at high, or ≤25% at low vs. ≥75% at high, or ≤30% at low vs. ≥70% at high.

[0053] In some embodiments, both measures of alpha diversity are used to classify microbiome profiles as having one of the following levels of diversity: low, -notlow, medium, not high, and high, which is appropriate for defining population stratification groups.

[0054] Here, SD = Shannon diversity, SR = species richness, and the base of the Shannon exponent is the natural logarithm.

[0055] In some embodiments, high and low bins are defined as low = SD + SR being in the lowest tertile of the sample (≤0.33) and high = SD + SR being in the highest tertile of the sample (≥0.66). Therefore, the threshold is for both measurements, i.e., smaller than the third tertile of both. The median is discarded.

[0056] In some embodiments, low bins and non-low bins are defined as low = SD + SR is in the lowest quartile of the sample (≤0.25), and non-low = SD + SR is in the upper third quartile of the sample (>0.25). Thus, both measurements ≤0.25 are low, and otherwise are not low.

[0057] In some embodiments, low, medium, and high bins are defined based on ternary values. Low = SD + SR (≤ 0.33) for the lowest ternary, medium = SD + SR (0.34 to 0.65) for the middle ternary, and high = SD + SR (≥ 0.66) for the upper ternary. Thus, the threshold for both measurements is smaller than both third ternary values. All others are assigned as medium.

[0058] Please understand that, although the above is a variation, such groups can be defined in many other ways, such as by somewhat different definitions, like median / mean + / -1 standard deviation, median / mean + / -1 / 2 standard deviation, or median / mean + / -1 / 2 interquartile range, or by a different percentage of data points that would be obvious to those skilled in data analysis.

[0059] The Receiver Operating Characteristic (ROC) curve is one of the best-developed statistical tools for describing performance in diagnostic tests measured on a continuous scale. The use of ROC is based on having two outcomes from prediction. Numerical indices were used to summarize the ROC curves. These summarization measures were also used to compare ROC curves.

[0060] The Area Under the ROC Curve (AUC) is the most widely used summary metric. A perfectly predictive model with an ideal ROC curve has an AUC of 1.0, while a random predictive model has an AUC of 0.5. An ROC curve AUC moving from 0.5 to 1.0 indicates improvement and better performance of the predictive model.

[0061] Many other measures of model performance include, for example, true positives (TP), false positives (FP), true negatives (TN), false negatives (FN), total predicted positives, total predicted negatives, total actual positives, total actual negatives, sensitivity / hit rate / recall / true positive rate (TPR), specificity / selectivity / true negative rate (TNR), prevalence, accuracy / positive predictive value (PPV), negative predictive value (NPV), miss rate / false negative rate (FNR), fallout / false positive rate (FPR), false detection rate (FDR), false dropout rate (FOR), prevalence threshold (PT), threat score (TS) / critical success index (CSI), accuracy (ACC), equilibrium accuracy (BA), random accuracy, total accuracy, F1 score, Matthews correlation coefficient (MCC), Fowlkes-Mallows index (FM), and informedness / bookmark understanding. Informedness (BM), markedness (MK) / delta P, positive likelihood ratio (LR+), negative likelihood ratio (LR-), diagnostic odds ratio (DOR), and κ can all be calculated on the confusion matrix.

[0062] AUC-ROC is the area under the curve, created by plotting the true positive rate against the false positive rate at various probabilities. AUC-PR is the area under the precise recall curve.

[0063] Various embodiments of the disclosed system satisfy a general goal when given a questionnaire for assessing the overall microbiome state of an individual. The assessment of the microbiome state depends on the individual's general characteristics (e.g., sex, age, weight, anthropometric measurements, physical activity level, and other health-related conditions such as IBS or diabetes), and subsequent recommendations or advice for maintaining or improving the microbiome health incorporate the individual's characteristics, such that they are also collected from the individual's responses to a set of questionnaires.

[0064] In various embodiments, the systems disclosed herein calculate from various responses obtained from an individual's input to a set of questionnaires and display the respective impacts on the state of the microbiome. Since the overall prediction also depends on the individual's general characteristics (such as BMI, age, weight, ethnicity, etc.), recommendations for maintaining or improving the microbiome also depend on the individual's characteristics. In these embodiments, the system determines that one or more of these factors are harmful or beneficial to the microbiome of the individual for whom the recommendations are calculated. The systems and methods disclosed in these embodiments recommend either reducing or adding these modifiable factors, which can be adapted to the individual.

[0065] In various embodiments, the disclosed system provides a recommendation or advisory function, where the system suggests a combination of features or factors that result in an improved or optimal microbiome state. For example, if a user accesses the system after using antibiotics and indicates this in their responses to a set of questionnaires, the system of this disclosure may predict the microbiome state accordingly and provide the basis for that prediction with respect to individual factors, such as antibiotic use, that influence the final prediction. Thus, the system disclosed herein can function not only as a predictive system but also as a recommendation engine that provides personalized advice to help an individual achieve the goal of having a good microbiome state.

[0066] The term “feature” is used repeatedly herein. In some embodiments, the term “feature” as used herein refers to input parameters to a model. The term includes responses obtained from a set of questionnaires. As used herein, the term feature may include, for example, anthropometric measurements such as age, sex, height, and weight; lifestyle features such as exercise / physical activity, alcohol use, smoking status, anxiety levels, depressive states, and stress levels; travel; use of drug therapies such as antibiotics; and comprehensive information regarding disease conditions such as IBD and diabetes. These features are not necessarily mutually exclusive. For example, certain features such as age and exercise may be related, or the use of certain drug therapies may be related to several diseases, or certain diseases may coexist as comorbidities.

[0067] As is known in the art, anthropometric measurements are measurements of an object. In one embodiment, anthropometric measurements are selected from the group consisting of sex, age (years), weight (kilograms), height (meters), and body mass index (kg / m-2). Other anthropometric measurements will also be known to those skilled in the art.

[0068] The terms “ethnicity” or “race” can be used in the context of specific geographical groups to cluster different subgroups. For example, in the United States, categories include white, black or African American, American Indian or Alaskan Native, Asian, and Native Hawaiian or other Pacific Islander.

[0069] The term "lifestyle characteristics" refers to any lifestyle choice made by the subject, and includes all dietary intake data, activity scales, or data obtained from questionnaires regarding lifestyle, motivation, or preferences. In one embodiment, lifestyle characteristics relate to whether the subject is a drinker or a non-drinker. In another embodiment, lifestyle characteristics relate to whether the subject is a smoker or a non-smoker. In yet another embodiment, lifestyle characteristics relate to whether the subject exercises regularly.

[0070] The term "anxiety state" refers to feelings of unease, such as worry or fear, which may be mild or severe, affecting an individual. Anxiety is generally analyzed using effective questionnaires. For example, a general health questionnaire consists of 60 questions about generally moderate physical and anxiety symptoms. 30-item and 12-item questionnaires are also commonly used. The Patient Health Questionnaire (PHQ-9) and the Center for Epidemiological Research's Depression Scale (CES-D) are further examples of anxiety scales / scores. In another embodiment, anxiety scores may be self-assessed by the individual.

[0071] In a preferred embodiment, anxiety scores are measured using the DASS-21 methodology (Lovibond, SH and Lovibond, PF (1995). "Manual for the Depression Anxiety & Stress Scales." (2nd edition), Sydney: Psychology Foundation; Lovibond P: "Overview of the DASS and Its Uses." Obtained from http: / / www2.psy.unsw.edu.au / dass / over.htm). For MDD data, the final score was calculated by multiplying the score obtained with DASS-21 by 2 to obtain the recommended cutoff score for the conventional severity labels (normal, moderate, severe), and the DASS-21 questionnaire shown on the right (https: / / maic.qld.gov.au / wp-content / uploads / 2016 / 07 / DASS-21.pdf) was used.

[0072] The Depression Scale (DASS) is a set of three self-report scales designed to measure negative emotional states of depression, anxiety, and stress. Each of the three DASS scales contains 14 items, divided into subscales of 2 to 5 items each with similar content. The depression scale assesses dysphoria, hopelessness, devaluation of life, self-deprecation, decreased interest / involvement, anesthesia, and apathy. The anxiety scale assesses autonomic nervous system agitation, skeletal muscle effects, situational anxiety, and subjective experiences of anxious feelings. The stress scale is sensitive to levels of chronic, nonspecific agitation. The stress scale assesses difficulty relaxing, nervous agitation, and agitation / excitement, irritability / hyperresponsiveness, and restlessness. Participants are asked to use a 4-point severity / frequency scale to assess the extent to which they experienced each state in the past week. Scores for depression, anxiety, and stress are calculated by summing the scores of the relevant items. In addition to the basic 42-item questionnaire, a simplified version with 7 items per scale, the DASS-21, is available. Individuals who score highly on the DASS-21 anxiety scale are characterized by anxiety, panic attacks; tremors, dizziness (shaky); perceived dry mouth, shortness of breath, palpitations, sweating palms; and concern about potential performance and control impairments.

[0073] The term “depressed state” refers to a mood disorder that causes persistent feelings of sadness and loss of interest. Also known as major depressive disorder or clinical depression, it can affect how an individual feels, thinks, and behaves, and can cause a variety of emotional and physical problems. In one embodiment, the depression rating scale / score is completed by the individual. The Beck Depression Scale is a self-reported list of 21 questions covering symptoms such as irritability, fatigue, weight loss, decreased libido, and feelings of guilt, helplessness, or fear of punishment. In a preferred embodiment, the depression score is assessed using the DASS-21 method, as detailed in the previous paragraph. Characteristics of high scores on the DASS-21 depression scale include self-deprecation; dejection, sadness, pessimism; conviction that life has no meaning or value; pessimism about the future; inability to experience pleasure or satisfaction; inability to have interest or concern; and dslowness and decreased spontaneity.

[0074] Similarly, stress scores are also preferably measured using the DASS-21 method, as described above. Characteristics of individuals who score highly on the DASS-21 stress scale include excessive excitement, tension; inability to relax; irritability, agitation; irritability; easily startled; nervousness, anxiety, restlessness; and inability to tolerate interruption or delay.

[0075] In a particularly preferred embodiment, the user inputs their answers to questions about, for example, their health status, drug use, antibiotic use, location, age (years), BMI, recent travel, race, alcohol consumption (e.g., type, frequency, amount of alcohol), smoking status, weight (kg), height (cm), IBD, GI symptoms (e.g., abdominal distension, bowel movement quality), weight changes, depression, anxiety, stress levels, season, who they live with, and drinking water sources into the device. The device then processes this information according to the definitions above and provides predictions about the user's microbiome state in terms of whether it is "low" or "not low".

[0076] In a particularly preferred embodiment, the user inputs their responses to questions regarding, for example, age (years), location, health status, alcohol consumption (e.g., type of alcohol (red or white wine, unspecified); frequency of alcohol consumption, amount of alcohol), smoking status, use of drug therapy, use of antibiotics, weight (kg) on ​​the most recent trip, anxiety level, depressive state, stress level, varicella, vaccination status (e.g., influenza vaccination date, pneumococcal vaccination date), lactose intolerance, race, and frequency of makeup use into the device. The device processes this information and provides a prediction of the user's microbiome state in terms of "high" or "not high" according to the definitions above.

[0077] In a particularly preferred embodiment, the user inputs responses to questions on a questionnaire regarding, for example, health status, age (years), location, use of drug therapy, use of antibiotics, alcohol consumption (e.g., type of alcohol (red or white wine, unspecified); frequency of alcohol consumption, amount of alcohol), smoking status, vaccination status, race, BMI, recent travel, anxiety level, depressive state, stress level, varicella, IBD, GI symptoms (e.g., abdominal distension, quality of bowel movements), and weight (kg) into a computer-implemented device. The device then processes this information and provides a prediction of the user's microbiome state in terms of "low" or "high" according to the definitions above. Examples of preferred questionnaires are shown in Figures 12-1 to 12-3 and Figures 16-1 to 16-12.

[0078] The device may generally be a server on a network. However, any device may be used as long as it can process data such as biomarker data and / or anthropometric data and lifestyle data using a processor or central processing unit (CPU). The device may be, for example, a smartphone, tablet, or personal computer, and it outputs information indicating the state of the user's microbiome.

[0079] In further embodiments, the present invention provides a method for recommending lifestyle changes to a subject. Lifestyle changes to the subject may be any changes, such as changes in diet, more exercise / physical activity, a different work environment, and / or living environment. Changing the subject's lifestyle also includes, for example, indicating the subject's choice to change their lifestyle. The subject may prefer to exercise more or stop excessive alcohol consumption in each session. Therefore, when providing these recommendations to maintain or improve the state of the subject's gut microbiome, the subject's preferences or choices may be taken into consideration.

[0080] In one embodiment, the method further includes combining the levels of one or more biomarkers, such as those obtained by a subject in other health screenings or health examinations, with one or more anthropometric measurements and / or lifestyle characteristics of the subject, or other characteristics already used in a questionnaire. While each of such health biomarkers may have a predictive value in the method of the present invention, combining the values ​​of multiple biomarkers can improve the accuracy of the method and the recommendations. As just one example, anthropometric measurements are selected from a group of questionnaires consisting of sex, weight, height, age, and body mass index, and lifestyle characteristics are whether the subject is a drinker or non-drinker, or whether the subject is a smoker or non-smoker, which are then combined with other health biomarkers, such as cholesterol or blood pressure levels, to further enhance the performance of the predictive model.

[0081] The methods described herein may be implemented as computer programs that run on general-purpose hardware such as one or more computer processors. In some embodiments, the functions described herein may be implemented by devices such as smartphones, tablet devices, or personal computers. In one aspect, the present invention provides a computer program product that includes computer-implementable instructions for causing a programmable computer to predict the state of a microbiome based on the level of features obtained from the questionnaire or linked device described herein.

[0082] In another aspect, the present invention provides a computer program product that includes computer-implementable instructions for causing a device to predict the state of its microbiome, taking into account the levels of one or more biomarkers from a user. The computer program product may be input with anthropometric measurements and / or lifestyle characteristics obtained from the user. As described herein, anthropometric measurements include age, weight, height, sex, and body mass index, and lifestyle characteristics include smoking status, stress levels, anxiety levels, depressive states, and frequency of physical activity / exercise.

[0083] Referring here to Figure 1, a block diagram is shown illustrating an example of the electrical system of a host device 100 that can be used to implement at least a portion of the computerized prediction and recommendation system disclosed herein.

[0084] In one embodiment, the device 100 shown in Figure 1 corresponds to one or more servers and / or other computing devices that provide some or all of the following functions: (a) enabling remote users of the system to access the system of the disclosure; (b) providing one or more web pages that enable remote users to interface with the system of the disclosure; (c) storing and / or computing underlying data necessary to implement the system of the disclosure, such as predictive models and recommendation algorithms, features used by these models and algorithms, and feature values ​​used for determinations by these models and algorithms; (d) computing and displaying components; and / or (e) providing personalized recommendations and advice on various features that can be reduced or enhanced to help an individual reach an optimal microbiome state.

[0085] In the embodiment of the architecture shown in Figure 1, device 100 includes a main unit 104 which preferably includes one or more processors 106 electrically connected by an address / data bus 113 to one or more memory devices 108, other computer circuits 110, and / or one or more interface circuits 112. The one or more processors 106 can be any suitable processor, such as microprocessors of the INTEL PENTIUM® or INTEL CELERON® family. PENTIUM® and CELERON® are trademarks registered with Intel Corporation and refer to commercially available microprocessors. It should be understood that in other embodiments, other commercially available or specially designed microprocessors may be used as processor 106. In one embodiment, processor 106 is a system-on-a-chip ("SOC") specifically designed for use in the systems of this disclosure.

[0086] In one embodiment, device 100 further comprises memory 108, preferably including volatile memory and non-volatile memory. Memory 108 preferably stores one or more software programs that interact with the hardware of host device 100 and other devices in the system, as described later. In addition to or instead of this, programs stored in memory 108 can interact with one or more client devices, such as client device 102 (described in detail below), to provide those devices with access to media content stored in device 100. Programs stored in memory 108 can be executed by processor 106 in any preferred manner.

[0087] The interface circuit(s) 112 may be implemented using any preferred interface standard, such as an Ethernet® interface and / or a Universal Serial Bus (USB) interface. One or more input devices 114 may be connected to the interface circuit(s) 112 to input data and commands to the main unit(s) 104. For example, the input devices 114 may be a keyboard, mouse, touchscreen, trackpad, trackball, isopoint, and / or voice recognition system. In one embodiment, where device(s) 100 is designed to be operated or interacted with only by a remote device, device(s) 100 may not include the input devices 114. In other embodiments, the input devices 114 include one or more storage devices, such as one or more flash drives, hard disk drives, solid-state drives, cloud storage, or other storage devices or solutions, which provide data input to the host device(s) 100. In one embodiment, the system is configured to integrate with one or more input devices 114, which are personal mobile devices carried by the user. For example, a user wearing a pedometer or activity tracker can provide data from these devices to the system and thus input activity levels.

[0088] One or more storage devices 118 can also be connected to the main unit 104 via the interface circuit 112. For example, a hard drive, CD drive, DVD drive, flash drive, and / or other storage devices can be connected to the main unit 104. The storage device 118 can store any kind of data used by device 100, including data on preferred input features of a model and their decision ranges, data on possible answers for each of the questions in a set of questionnaires, data on the users of the system, data on previously generated microbiome assessment states, data on previously generated suggestions or recommendations, and individual user preferences for input features or sets of features, which may be willing to function for improvement or not, as shown in block 150, and for improving any other appropriate data required to implement the system of this disclosure.

[0089] In some embodiments, the recommendation system shown by block 150 may store different database models, which may include low-to-low models, high-to-high models, low-to-high models; a consensus model scoring module (e.g., for giving final predictions); and / or an optimization module (e.g., for providing the most reliable results); a recommendation module (e.g., for providing users with personalized advice on how to maintain or improve the state of their own microbiome); a constraint module (e.g., for incorporating user-side constraints); and a final recommendation module (e.g., for taking into account multiple model inputs and user constraints).

[0090] Alternatively, or in addition to the above, the storage device 118 may be implemented as cloud-based storage such that access to the storage 118 is via the Internet or other network connection circuits such as the Ethernet circuit 112.

[0091] One or more displays 120, and / or a printer, speaker, or other output device 119, may also be connected to the main unit 104 via an interface circuit 112. The displays 120 can be liquid crystal displays (LCDs), suitable projectors, or any other suitable type of display. The displays 120 generate visual representations of various data and functions of the host device 100 during its operation. For example, the displays 120 can be used to display information about the placement of an individual's microbiome state in a distribution observed for a reference population, deep analyses of an individual's responses to a set of questionnaires, and relevant recommendations for maintaining or improving the microbiome state. As mere examples, Figures 12-1 to 12-3 show information about an individual user in their responses to a set of questionnaires. Figure 13 shows the results regarding the prediction of the user's microbiome state. Figure 14 shows features that positively or negatively affect the user's microbiome state based on the user's responses. Figure 15 shows personalized recommendations and advice for maintaining or improving the state of the microbiome, as well as instructions on how to follow them.

[0092] In the illustrated embodiment, a user of the computerized recommendation system interacts with device 100 using a suitable client device, such as client device 102. In various embodiments, client device 102 is any device capable of accessing content provided or supplied by host device 100. For example, client device 102 may be any device capable of running a suitable web browser for accessing a web-based interface to host device 100. Alternatively or in addition to this, one or more applications or parts of applications that provide some of the functions described herein may run on client device 102, in which case client device 102 would only need to connect to host device 100 as described above in order to access data stored in host device 100.

[0093] In one embodiment, this connection of devices (i.e., device 100 and client device 102) is facilitated by a network connection via the Internet and / or other network illustrated in Figure 1 by the cloud 116. The network connection may be any suitable network connection, such as an Ethernet connection, a digital subscriber line (DSL), a Wi-Fi connection, a cellular data network connection, a telephone line-based connection, a connection over coaxial cable, or another suitable network connection.

[0094] In one embodiment, the host device 100 is a device that provides cloud-based services such as cloud-based authentication and access control, storage, streaming, and feedback. In this embodiment, the specific hardware details of the host device 100 are not important to the implementer of the System of the Disclosure, rather, in such embodiment, the implementer of the System of the Disclosure interacts with the host device 100 in a convenient manner using one or more application programmer interfaces (APIs), for example, by entering information about user input into a set of questionnaires to provide user preferences or constraints and other interactions, which are described in more detail below.

[0095] Access to device 100 and / or client device 102 can be managed by appropriate security software or security measures. Access for individual users can be defined by device 100 and may be restricted to certain data and / or actions, such as selecting various options for a question or viewing predicted values, depending on the individual's responses. Other users of either host device 100 or client device 102 may modify other data, such as weights, sensitivities, or feature range values, depending on the identity of those users. Therefore, users of the System may be required to register with device 100 before accessing the content provided by the System of this Disclosure.

[0096] In a preferred embodiment, each client device 102 has a structural or design configuration similar to that described above with respect to device 100. That is, in one embodiment, each client device 102 includes a display device, at least one input device, at least one memory device, at least one storage device, at least one processor, and at least one network interface device. It should be understood that by including such components common to well-known desktop, laptop, or mobile computer systems (including smartphones and tablet computers, etc.), the client device 102 facilitates interaction between multiple users of the corresponding system.

[0097] In various embodiments, device 100 and / or device 102, as illustrated in Figure 1, can actually be implemented as multiple different devices. For example, device 100 may actually be implemented as multiple server devices working together to implement the media content access system described herein. In various embodiments, one or more additional devices not shown in Figure 1 interact with device 100 to enable or facilitate access to the system disclosed herein. For example, in one embodiment, host device 100 communicates via network 116 with one or more public, private, or proprietary repositories, such as public, private, or proprietary repositories, for information such as microbiome-friendly advice or other advice, recommendations, nutritional information, nutrient content information, menu planners, recipe databases, information on healthy ranges, energy information, or environmental impact information.

[0098] In one embodiment, the system of the Disclosure does not include a client device 102. In this embodiment, the functions described herein are provided on a host device 100, and the user of the system interacts directly with the host device 100 using an input device 114, a display device 120, and an output device 119. In this embodiment, the host device 100 provides some or all of the functions described herein as user-facing functions.

[0099] In various embodiments, the systems disclosed herein are configured as a plurality of modules, each module performing a specific function or set of functions. The modules in these embodiments may be software modules executed by a general-purpose processor, software modules executed by a dedicated processor, firmware modules executed on appropriate dedicated hardware devices, or hardware modules (such as application-specific integrated circuits ("ASICs")) that fully perform the functions enumerated herein by circuitry. In embodiments where special hardware is used to perform some or all of the functions described herein, the systems disclosed herein may use one or more registers or other data input pins to control settings or to adjust the functions of such special hardware.

[0100] The user's goal of predicting the state of their microbiome can be examined over time to detect patterns of problems or improvements. The system can then be used to identify recommended shifts in habits, foods, supplements, menus, or recipes needed to move towards a better microbiome state. In some embodiments, the systems and methods disclosed herein can be used by dietitians, healthcare professionals, and individual users.

[0101] Figure 2 shows a microbiome recommendation system according to one embodiment of the present disclosure. System 200 comprises a user device 202 and a recommendation system 204. In another embodiment of the present disclosure, the recommendation system 204 may be an example of an embodiment of the recommendation system 150 in Figure 1. The user device 202 may be implemented as a computing device such as a computer, smartphone, tablet, smartwatch, or other wearable device on which the associated user can communicate with the recommendation system 204. The user device 202 may also be implemented as a voice assistant configured to receive voice requests from the user and process the requests locally on a computer device near the user or on a remote computing device (e.g., on a remote computing server).

[0102] The recommendation system 204 includes one or more of the following: a display 206, an attribute acquisition unit 208, an attribute comparison unit 210, an evidence-based evaluation and recommendation engine 212, an attribute analysis unit 214, an attribute storage unit 216, memory 218, and a CPU 220. It should be noted that in some embodiments, the display 206 may be additionally or alternatively located within the user device 202. In one example, the recommendation system 204 may be configured to acquire requests for multiple microbiome health recommendations 240. For example, a user may install an application on the user device 202 that prompts the user to register for the recommendation service. By registering for this service, the user device 202 may submit requests for microbiome health recommendations 240. In another example, a user may use the user device 202 to access a web portal using user-specific credentials. Through this web portal, the user may have the user device 202 request microbiome health recommendations from the recommendation system 204.

[0103] In another example, the recommendation system 204 may be configured to request and retrieve multiple user attributes 222. For example, the display 206 may be configured to present an attribute questionnaire 224 to the user. The attribute retrieval unit 208 may be configured to retrieve user attributes 222. In one example, the attribute retrieval unit 208 may retrieve multiple inputs 226 based on the attribute questionnaire 224 and determine multiple user attributes 222 based on those inputs. For example, the attribute retrieval unit 208 may receive input to the attribute questionnaire 224 indicating whether the user's habits are good for the microbiome and then suggest maintaining or improving user attributes 222 for the microbiome. In another example, the attribute retrieval unit 208 of a user device may retrieve user attributes 222 directly from the user device 102.

[0104] In another example, the attribute acquisition unit 208 may be configured to acquire results from a home testing kit, a standard health examination conducted by a healthcare professional, the results of this self-assessment tool used by the user, or the results of any external or third-party tests. The attribute acquisition unit 208 may be configured to determine user attributes 222 based on the results from any of these tests or tools. For example, the health of a user's microbiome may be determined by predicting the alpha diversity of microbial species before a microbiome health recommendation intervention. The same measurement may be predicted again during the period after the microbiome health intervention to determine whether there has been an improvement or maintenance of the user's microbiome health.

[0105] The recommendation system 204 may be further configured to compare multiple user attributes 222 with corresponding multiple evidence-based microbiome health criteria 228.

[0106] Furthermore, the attribute comparison unit 210 may be further configured to determine a microbiome reference set 232 based on the user's microbiome classification 230. For example, if the attribute comparison unit 210 determines, based on several user attributes 222, that the user falls into the obesity BMI classification 230, the attribute comparison unit 210 may select a microbiome reference set 232 that has been created and defined according to the specific needs of a healthy microbiome.

[0107] The comparison unit 210 may be further configured to select evidence-based microbiome criteria 128 from the determined microbiome criteria set 232 and compare the selected evidence-based microbiome criteria 228 with each of the corresponding user attributes 222. For example, once the microbiome criteria set 232 is determined, in response to that determination, the attribute comparison unit 210 may compare a user attribute 222 representing the user's antibiotic dosage with an evidence-based microbiome criterion 228 representing a baseline antibiotic dosage to determine whether the user's dosage is less than, at, or above the baseline antibiotic dosage from a microbiome perspective. While this example is based on a concrete numerical comparison, another example of a criterion comparison is qualitative and can vary from person to person. For example, a user attribute 222 might indicate that the exercise level the user is currently engaging in is lower than normal. An example of a criterion related to the user's exercise level might indicate that an average or high level of exercise is desirable, and therefore, a user attribute 222 indicating a lower exercise level would be judged to be below that criterion. Since different users are engaging in different levels of exercise, even under the same circumstances, such comparisons require a customized approach.

[0108] Furthermore, during the comparison in the example described above, the attribute comparison unit 210 may be configured to determine a user microbiome score 234 based on a comparison between the evidence-based microbiome criteria 228 and the user attribute 222. For example, if the user attribute 222 almost completely satisfies all or most of the corresponding evidence-based microbiome criteria 228, the attribute comparison unit 210 may determine a user microbiome score of 95 / 100. In another example, the score may be represented by letter grades, symbols, or other ranking systems, such as "low," "medium," or "high," so that the user can interpret how their current attribute ranks among the criteria. This user microbiome score 234 may be presented by the display 206.

[0109] The recommendation system 204 may be further configured to determine multiple microbiome support opportunities 238 based on a comparison of multiple user attributes 222 and corresponding multiple evidence-based microbiome criteria 228. For example, the attribute comparison unit 210 may determine a microbiome support opportunity 238 for all user attributes 222 that do not meet the corresponding evidence-based microbiome criteria. In this example, the corresponding evidence-based microbiome criterion 228 may require the user to take 2 ug / day of folate, but the user attribute may indicate that the user is taking only 1 ug / day of folate. Therefore, the attribute comparison unit 210 may determine an increase in folate intake as a microbiome support opportunity 238.

[0110] In another example, the attribute comparison unit 210 may be configured to identify a first set of user attributes 236, each consisting of several user attributes 222 that are below a corresponding evidence-based microbiome criterion 228, and a second set of user attributes 236, each consisting of several user attributes 222 that are equal to or greater than the corresponding evidence-based microbiome criterion 228. The first set of user attributes 236 is determined in the same way as in the example above, but the second set of user attributes 236 differs in that the user in question does not appear to be deficient, but there may be an opportunity to support microbiome health by recommending that the user maintain current practices, or an opportunity to further improve current practices. Thus, the recommendation system 204 may determine the opportunity to support microbiome health based on which attributes 222 fall into which set 236.

[0111] The recommendation system 204 may be further configured to identify multiple microbiome-friendly recommendations 240 based on multiple microbiome support opportunities 238. For example, the evidence-based diet and lifestyle recommendation engine 212 may be cloud-based. The recommendation engine 212 may include one or more of multiple databases 242, multiple dietary restriction filters 244, and optimization units 246. The recommendation engine 212 may identify multiple microbiome-friendly recommendations 240 based on multiple opportunities 238 according to one or more of the multiple databases 242, dietary restriction filters 244, and optimization units 246.

[0112] In another example, the recommendation system 204 may be configured to provide continuous recommendations based on previous user attributes. For example, the recommendation system 204 may include, in addition to the elements described above, an attribute storage unit 216 and an attribute analysis unit 214. The attribute storage unit 216 may be configured to add the acquired user attributes 222 to the attribute history database 248 as new entries based on when the user attributes 222 were acquired, in response to the attribute acquisition unit 108 acquiring multiple user attributes 222. For example, if the user attributes 222 are acquired by the attribute acquisition unit 208 on the first day, the attribute storage unit 216 adds the acquired user attributes 222 to the cumulative attribute history database 248 with the entry date, which in this example is the first day. Subsequently, if user attributes 222 are acquired by the attribute acquisition unit 208 on a second day, for example the following day, the attribute storage unit 216 further adds these new attributes to the attribute history database 248, noting that they were acquired on the second day, while also saving the attributes from the first day prior to that.

[0113] The attribute analysis unit 214 may be configured to analyze multiple user attributes 222 stored in the attribute history database 248, and analyzing the stored user attributes 222 may include conducting a long-term study 250. Continuing the above example, the attribute analysis unit 214 may conduct a long-term study of user attributes 222 from each of the sets of user attributes 222 found in the attribute history database 248, from day 1, from day 2, and all other sets of user attributes 222. The evidence-based diet and lifestyle recommendation engine 212 may be further configured to generate multiple microbiome-friendly recommendations 240 based on at least the stored user attributes 222 found in the attribute history database 248 and the analysis performed by the attribute analysis unit 214.

[0114] In one embodiment, the attribute analysis unit 214 is further configured to repeatedly analyze multiple user attributes 222 stored in the attribute history database 248 in response to the attribute storage unit 216 adding a new entry to the attribute history database 248, thereby effectively reanalyzing all the data in the attribute history database 248 immediately after a new user attribute 222 is acquired. Similarly, the evidence-based diet and lifestyle recommendation engine 212 may be further configured to repeatedly generate multiple microbiome health recommendations 240 in response to the attribute analysis unit 214 completing its analysis, thereby effectively generating new microbiome health recommendations 240 that take into account all past and present user attributes 222 whenever a new set of user attributes 222 is acquired.

[0115] In various embodiments, user-specific (or group-specific) inputs to the system of this disclosure are programmable and configurable and include, but are not limited to, gender, age, weight, height, physical activity level, obesity, etc.

[0116] In a further embodiment, an individual may provide its own weighted values, which are adjusted to its personal choices and health status. Using these individual-specific ranges and / or weighted values, the disclosed system can then calculate fully personalized advice for maintaining or improving the individual's microbiome state.

[0117] In one embodiment, the system of the Disclosure includes, or is connected to, a database containing food items, menus or recipes, and their respective nutrient content. In this embodiment, the system of the Disclosure includes a fuzzy search function that allows a user to input foods they have consumed (or plan to consume) and then search the database to find the item that most closely matches the user's input. In this embodiment, the system of the Disclosure uses stored nutritional information about the matching food items to determine whether they are microbiome-friendly, as described later.

[0118] In various embodiments, the system of the present disclosure further includes an interface (e.g., a graphical user interface) for displaying the amount of each nutrient available in each food item constituting a meal and the amount of energy that should be consumed. In some embodiments, this interface allows the user to modify the amount of different foods or energy that should be consumed. In other embodiments, the system is configured to determine the amount of food consumed or energy consumed using data not entered by the user, such as by scanning one or more barcodes, QR codes®, or RFID tags, an image recognition system, or by tracking items ordered from a menu or items purchased at a grocery store.

[0119] Various embodiments of the disclosed system display a dashboard or other appropriate user interface to the user, customized based on the user's needs. Conveniently, embodiments of the system disclosed herein provide a graphical user interface that, for the first time, allows the user to input data about their responses to a set of questionnaires and, based appropriately on predictions, see a display of a score that reflects the overall placement of the user's state within a commonly observed microbiome distribution.

[0120] All methods and procedures described herein can be implemented using one or more computer programs or components. These components may be provided as a series of computer instructions on any conventional computer-readable or machine-readable medium, including volatile and non-volatile memories such as RAM, ROM, flash memory, magnetic or optical disks, optical memory, or other storage media. The instructions may also be provided as software or firmware, and may be implemented in whole or in part in hardware components such as ASICs, FPGAs, DSPs, or any other similar devices. The instructions may be configured to be executed by one or more processors that perform or facilitate all or part of the performance of the disclosed methods and procedures when executing a series of computer instructions.

[0121] As described above, in some embodiments, the disclosed system relies on one or more modules (hardware, software, firmware, or a combination thereof) to perform the various functions described above.

[0122] The inventors have shown that it is possible to create a predictive tool that enables the prediction of the state of the gut microbiome, such as microbiome diversity, based on characteristics obtained from questionnaires such as health, lifestyle, dietary habits, or preferences. Furthermore, where known equivalents exist for a particular characteristic, such equivalents are incorporated as specifically referred to herein. Further advantages and aspects of the present invention are evident from the figures and non-limiting embodiments.

[0123] In some embodiments, the present invention is a method for determining the state of the intestinal microbiome, (i) A step to determine the state of the intestinal microbiome in the subject, (ii) A step of providing recommendations for improving or maintaining the state of the microbiome in the subject, This provides a method that includes [something].

[0124] In some embodiments, the present invention provides a determination of the state of the gut microbiome using a questionnaire for predicting the microbiome diversity of a target.

[0125] In one embodiment, the present invention additionally provides the determination of the state of the gut microbiome using a biological sample for quantifying the microbiome diversity of the subject.

[0126] In some embodiments, the method of the present invention is implemented using a computer.

[0127] In some embodiments, the method of the present invention evaluates characteristic parameters related to the state of the target gut microbiome.

[0128] In some embodiments, characteristic parameters related to the state of the gut microbiome are, (i) the geographical country of the subject residence, including its specific latitude and longitude; (ii) whether antibiotics have been used in the past year, and the use of antibiotics, including if antibiotics have been used in the past year; (iii) whether drug therapy has been used in the past 12 months, and the use of drug therapy, including if drug therapy has been used in the past 12 months; (iv) Anthropometric data including age, weight, height, body mass index, and sex; (v) Alcohol consumption, including the type of alcohol, the amount consumed, and the frequency of consumption; (vi) Smoking status and; (vii) Exercise and / or physical activity, including the location, frequency, and duration of indoor or outdoor exercise; (viii) ethnicity; (ix) seasons and; (x) Travel and including the place and duration of the trip; (xi) Sleep, including the duration and / or quality of sleep in hourly units; (xii) a state of stress, anxiety, and / or depression; (xiii) A medical history including blisters, irritable bowel disease, or diabetes; (xiv) A history of allergies, including seasonal allergies or food allergies and / or food intolerances; (xv) Vaccination status, including influenza vaccine or pneumococcal vaccine; (xvi) Use of dietary supplements, including vitamin or mineral supplements; (xvii) Sources of drinking water; (xviii) Personal hygiene, including tooth flossing, use of deodorants, and use of cosmetics; (xix) The type of food intake, the amount of food intake, and the frequency of food intake, including the intake of vegetables, fruits, fermented foods and / or whole grains. It is selected from the group consisting of the following.

[0129] In some embodiments, characteristic parameters related to the state of the gut microbiome are, (i) the geographical country of the subject residence, including its specific latitude and longitude; (ii) whether antibiotics have been used in the past year, and the use of antibiotics, including if antibiotics have been used in the past year; (iii) whether drug therapy has been used in the past 12 months, and the use of drug therapy, including if drug therapy has been used in the past 12 months; (iv) Anthropometric data including age, weight, height, body mass index, and sex; (v) Alcohol consumption, including the type of alcohol, the amount consumed, and the frequency of consumption; (vi) Smoking status and; (vii) Exercise and / or physical activity, including the location, frequency, and duration of indoor or outdoor exercise; (viii) ethnicity; (ix) seasons and; (x) Travel and including the place and duration of the trip; (xi) Sleep, including the duration and / or quality of sleep in hourly units; (xii) a state of stress, anxiety, and / or depression; (xiii) A medical history including blisters, irritable bowel disease, or diabetes; (xiv) A history of allergies, including seasonal allergies or food allergies and / or food intolerances; (xv) Vaccination status, including influenza vaccine or pneumococcal vaccine; (xvi) Use of dietary supplements, including vitamin or mineral supplements; (xvii) Sources of drinking water; (xviii) Personal hygiene, including tooth flossing, use of deodorants, and use of cosmetics; (xix) The type of food intake, the amount of food intake, and the frequency of food intake, including the intake of vegetables, fruits, fermented foods and / or whole grains. It is selected from the group consisting of the following.

[0130] In some embodiments of the present invention, this method (i) A step of determining at least one characteristic parameter related to the state of the gut microbiome in the subject; (ii) The step of comparing at least one characteristic parameter related to the state of the gut microbiome in the subject with a population database of subjects in the same geographical region; (iii) A step of determining whether the subject is low, medium, or high with respect to at least one characteristic parameter, They are involved in the process.

[0131] In some embodiments, subjects are notified of the state of their gut microbiome via a computer interface as shown in Figures 1 and 2.

[0132] In some embodiments, the systems and methods of the present invention contribute to maintaining and improving the state of the microbiome by providing recommendations for beneficial microbiome health, such as nutritional supplements, dietary recommendations, menu recommendations, and recipe recommendations, for improving or maintaining the α-diversity of microbial species in the gut.

[0133] In one embodiment of the present invention, by measuring the parameter diversity of microbial species in the gut, it is possible to determine whether the microbiome has improved or maintained its health from biological samples taken from a subject before and after recommending the dietary therapy of the present invention. Therefore, it is possible to determine over time whether an individual has made a positive improvement in their microbiome health after following the diet, menu, and recipe recommendations of the present invention.

[0134] In various embodiments, the systems disclosed herein provide recommendations for supplements, food items, menus, or recipes that demonstrate nutritional effects on the microbiome. In these embodiments, the system determines and stores one or more indicators of the individual's needs for the individual over a given period, such as a single meal, a full day, a week, or a month, for which recommendations are calculated.

[0135] In one embodiment of the present invention, the method and system of the present invention are (i) Whole grain foods, (ii) Beans and legumes, (iii) Fibers, (iv) Nuts and seeds, and (v) Includes recommendations for foods or nutrient groups selected from the group consisting of omega-3 fatty acids.

[0136] In one embodiment of the present invention, the method and system of the present invention are (i) Whole grain foods, (ii) Beans and legumes, (iii) Fibers, (iv) Nuts and seeds, and (v) Recommendations for foods or nutrient groups, including recommendations for meal plans or recipes that include omega-3 fatty acids.

[0137] Those skilled in the art will understand that all aspects of the invention disclosed herein can be freely combined without departing from the scope of the disclosed invention. Furthermore, aspects described for different embodiments of the invention may be combined. Although the invention has been described by examples, it should be understood that changes and modifications can be made without departing from the scope of the invention as defined in the claims.

[0138] As used herein, the words “comprises,” “comprising,” and similar words should not be interpreted as exclusive or exhaustive. In other words, they are intended to mean “includes, but not limited to.”

[0139] The above description illustrates aspects of the systems disclosed herein. As stated, the systems and methods of this disclosure can be used to predict the state of an individual's microbiome as defined by other microbiome indicators not mentioned herein, and to indicate the influence of other factors not mentioned herein, such as other endogenous, exogenous, or environmental factors, based on any suitable measurable characteristics. The systems and methods of this disclosure are not limited to determining only the microbiome state as defined herein, nor are they limited to describing the influence of factors on the microbiome enumerated herein. Furthermore, the functions of the systems described above are not limited to those shown herein. It should be understood that various changes and modifications to the embodiments described herein will be obvious to those skilled in the art. Such changes and modifications may be made without departing from the spirit and scope of this subject matter and without impairing the intended benefits. Accordingly, such changes and modifications are intended to be covered by the appended claims.

[0140] Next, various preferred features and embodiments of the present invention will be described by non-limiting examples. [Examples]

[0141] Example 1: Construction of a model for predicting the state of the microbiome The AGP data is a publicly available dataset containing microbiome data from 9,511 individuals, with relevant metadata features associated with common survey questions about personal characteristics, lifestyle, dietary habits, and medical status (approximately 200 questions in total). Further details of this study and a link to this dataset are available in the AGP publication (McDonald D et al., "mSystems.", 2018).

[0142] A predictive model was constructed to determine the state of the microbiome of individual subjects. In particular, the model predicted the alpha diversity of the microbiome using several feature parameters to determine whether the subject belonged to one of the categories defined above: "low" or "not low"; "high" or "not high"; or "low" or "high".

[0143] To build a classification model, the data was split into a training set ("training") and a test set ("holdout / test set"). For optimal model performance, the inventors used downsampling to balance any class imbalances that may occur based on the bin definition.

[0144] The training set was used by a machine learning algorithm to train the model. The training set involved finding the variables (i.e., features) and thresholds (or coefficients) to be used to classify groups. Learning from the data was done in a cross-validation manner, where some parts of the training data were used to train the model and other parts were used for internal testing (k-fold cross-validation, e.g., 3x), or this method was also repeated several times (repeated k-fold cross-validation, e.g., 10x, 10 iterations).

[0145] The holdout / test set was used solely to check the performance of the final trained model. Therefore, this holdout / test dataset was not used during the model training phase. The inventors evaluated multiple statistical models using readily available tools (R software, Python) and identified the best models for low vs. non-low, high vs. non-high, and low vs. high groups or categories.

[0146] Evaluating model performance was crucial at every stage of modeling. Once the model was trained, it was applied to holdout / test data that had not been used during the training phase. The model calculated the probability of belonging to each group (e.g., "low," "not low"). Final decisions were made based on these probabilities, and therefore the use of thresholds was necessary. These thresholds influenced the final classification of objects, regardless of whether they were correctly classified or not. Thus, errors were evaluated for different selections of thresholds. For each given threshold, a confusion matrix was calculated. This confusion matrix essentially counts the number of correctly classified and incorrectly classified objects. By using different thresholds, many confusion matrices were generated, which were then used to derive sensitivity and specificity at different thresholds. These two metrics (sensitivity and specificity) were generally presented in the form of receiver-operated curves (ROCs); these summarized the model performance across several thresholds.

[0147] For this model, receiver operation characteristic (ROC) curves were constructed. The inventors defined either a "low" group of subjects (and a "non-low" group) and predicted the probability that a subject would fall into this group; or the inventors defined a subject as falling into a "high" group (and a "non-high" group) and predicted the probability that a subject would fall into this group; or the inventors defined a subject as falling into a "low" group (and a "high" group) and predicted the probability that a subject would fall into this group.

[0148] The dataset used in the example prediction model comes from the American Gut Project (AGP) database (http: / / americangut.org).

[0149] Example 2: Low gut microbiome diversity model vs. high gut microbiome diversity model (I) For the low gut microbiota diversity model versus the gut microbiota diversity model, the run was performed with these parameters (80-20% training-holdout / test split, age _ years ≥ 20 and age _ years ≤ 70, BMI ≥ 18.5 and BMI ≤ 35, country == "USA"). These are the results for the holdout / test set: precision (0.687), 95% confidence interval (0.6375, 0.7335), equilibrium precision (0.6858), κ (0.37), low class sensitivity / prediction (0.6746), high class specificity / prediction (0.6971), positive predictor (0.6441), and negative predictor (0.7250). The top features used by this model to determine the low or high gut microbiota diversity status were ethnicity, antibiotic use, age, health status (no diabetes, IBD), and BMI.

[0150] Example 3: Low gut microbiome diversity model vs. high gut microbiome diversity model (II) This model was run with the following parameters: Country: USA, Age: 20–70 years, BMI: 18.5–30. The following results were obtained: For Random Forest (RF), precision (0.6928), 95% confidence interval (0.6132, 0.7648), equilibrium precision (0.6929), κ (0.3858), low-class sensitivity / prediction (0.7105), high-class specificity / prediction (0.6753), positive predictive value (0.6835), and negative predictive value (0.7027); for Support Vector Machine (SVM), precision (0.6667), 95% confidence interval (0.586, 0.7407), equilibrium precision (0.666 8) κ(0.3335), low-class sensitivity / prediction (0.6842), high-class specificity / prediction (0.6494), positive predictive value (0.6582), and negative predictive value (0.6757); for decision trees (DT), precision (0.6078), 95% confidence interval (0.5237, 0.6857), balanced precision (0.6082), κ(0.2162), low-class sensitivity / prediction (0.6579), high-class specificity / prediction (0.5584), positive predictive value (0.5952), and negative predictive value (0.6232). The top features of the algorithms (and diversity measures) were age, antibiotic use, alcohol consumption, travel, and exercise.

[0151] Example 4: Low gut microbiota diversity model vs. non-low gut microbiota diversity model (I) This model was run using the following parameters: training (80-20% holdout / test split, age_years=all ages, bmi=all BMIs, country=all countries), these are the results for the holdout / test set (n=1230): precision (0.679), 95% confidence interval (0.6511, 0.7041), balanced precision (0.62530), κ (0.1845), low-class sensitivity / prediction (0.54378), non-low-class specificity / prediction (0.70681), positive predictor (0.28434), and negative predictor (0.87853).

[0152] Example 5: High gut microbiota diversity model vs. non-high gut microbiota diversity model (I) This model was run using the following parameters: training (80-20% holdout / test split, age_years=all, bmi=all BMI, country=all countries). These are the results for the holdout / test set: precision (0.5984), 95% confidence interval (0.5709, 0.6255), balanced precision (0.6204), κ (0.1705), high-class sensitivity / prediction (0.6595), non-high-class specificity / prediction (0.5812), positive predictor (0.3072), and negative predictor (0.8584).

[0153] Example 6: Low gut microbiota diversity model vs. high gut microbiota diversity model (III) This model was run using the following parameters: training (80-20% holdout / trial split, age_years=all, bmi=all, country=all), these are the results for the holdout / trial set: precision (0.7117), 95% confidence interval (0.6696, 0.7512), balanced precision (0.7038), κ (0.4103), low-class sensitivity / prediction (0.6406), high-class specificity / prediction (0.7670), positive predictor (0.6814), and negative predictor (0.7329).

[0154] Example 7: Low gut microbiota diversity model vs. non-low gut microbiota diversity model (II) To define "low" - 1st / lower quartile for all three of these diversity measures, the following cutoffs were used in the American Gut Project (AGP) data: MEAN_OBSERVEDOTU≦88.30, MEAN_FAITHPD≦10.91, and MEAN_SHANNON≦4.38. To define "not low" for all three of these diversity measures, the cutoffs used in the AGP data were MEAN_OBSERVEDOTU>88.30, MEAN_FAITHPD>10.91, and MEAN_SHANNON>4.38.

[0155] This model was run with the following parameters: bin definition as the first / lowest quartile versus the rest defined for all three diversity measures, input AGP data with a survey minimum response rate of 0.65, input with no cutoff applied to any features using a random forest algorithm in 3x cross-validation training mode, a post-processing training size of 2370, and a holdout / test size of 1490. These are the results for the training set in cross-validation mode: sensitivity (0.65 ± 0.02), specificity (0.63 ± 0.01), precision (0.63 ± 0.01), AUC-ROC (0.7 ± 0.02). These are the results for the holdout / test set: sensitivity (0.67), specificity (0.62), precision (0.63), AUC-ROC (0.72). ROC curves and AUC values ​​are shown in Figures 3A-1 and 3A-2.

[0156] Example 8: Low gut microbiota diversity model vs. non-low gut microbiota diversity model (III) This model was run with the following parameters: bin definition as less than (mean - 1 * standard deviation) versus the rest defined for all three diversity measures, input AGP data with a survey minimum response rate of 0.85, input with no cutoff applied to any of the features using a random forest algorithm in a 3x, 3-repeat cross-validation training mode, post-processing training size of 1234, and holdout / test size of 1560. These are the results for the training set in cross-validation mode: sensitivity (0.62 ± 0.04), specificity (0.65 ± 0.02), precision (0.65 ± 0.02), AUC-ROC (0.7 ± 0.02). These are the results for the holdout / test set: sensitivity (0.62), specificity (0.71), precision (0.71), AUC-ROC (0.73).

[0157] The ROC curves and AUC values ​​are shown in Figures 3B-1 and 3B-2 for both (i) the training (cross-validation) and (ii) the holdout / test set. Figure 9 shows the improvement in the model's performance as features considered important to this model were added one by one through SHAP analysis (AUC values ​​of the ROC curves for the training (cross-validation) and holdout / test set).

[0158] Example 9: High gut microbiota diversity model vs. low gut microbiota diversity model (II) High-bin and non-high-bin are defined together by three diversity measures observed: OTU, Faith PD, and Shannon. To define "high" - 3rd / upper quartile for all three of these diversity measures, the following cutoffs were used on the American Gut Project (AGP) data: MEAN_OBSERVEDOTU>137.8, MEAN_FAITHPD>15.65, and MEAN_SHANNON>5.5. To define "non-high" for all three of these diversity measures, the cutoffs used on the AGP data were MEAN_OBSERVEDOTU≦137.8, MEAN_FAITHPD≦15.65, and MEAN_SHANNON≦5.5.

[0159] This model was run with the following parameters: bin definition as the third / upper quartile versus the rest defined on all three diversity measures, input AGP data with a survey minimum response rate of 0.65, input with no cutoff applied to any features, random forest algorithm in 3x cross-validation training mode, post-processing training size: 2564, holdout / test size: 1520. These are the results of training set in cross-validation mode: sensitivity (0.6±0.04), specificity (0.68±0.02), precision (0.67±0.01), AUC-ROC (0.71±0.01). These are the results of the holdout / test set: sensitivity (0.66), specificity (0.66), precision (0.66), AUC-ROC (0.73). ROC curves and AUC are shown in Figures 4A-1 and 4A-2.

[0160] Example 10: High gut microbiota diversity model vs. non-high gut microbiota diversity model (III) This model was run with the following parameters: bin definition as -(mean + 1*standard deviation) versus the remainder defined for all three diversity measures, input AGP data with a survey minimum response rate of 0.85, input with no cutoff applied to any of the features using a random forest algorithm in a 3x, 3-repeat cross-validation training mode, post-processed training size of 1408, and holdout / test size of 1585. These are the results for the training set in cross-validation mode: sensitivity (0.69±0.02), specificity (0.61±0.03), precision (0.68±0.01), AUC-ROC (0.71±0.02). These are the results for the holdout / test set: sensitivity (0.71), specificity (0.64), precision (0.70), AUC-ROC (0.72).

[0161] The ROC curves and AUC values ​​are shown in Figures 4B-1 and 4B-2 for both (i) the training (cross-validation) and (ii) the holdout / test set. Figure 10 shows the improvement in the model's performance as features considered important to this model were added one by one through SHAP analysis (AUC values ​​of the ROC curves for the training (cross-validation) and holdout / test set).

[0162] Example 11: Low gut microbiota diversity model vs. high gut microbiota diversity model (IV) Low and high bins are defined together using the three observed diversity measures: OTU, Faith PD, and Shannon. To define "low" (1st / lower quartile) for all three diversity measures, the following cutoffs were used on the American Gut Project (AGP) data: MEAN_OBSERVEDOTU≦88.30, MEAN_FAITHPD≦10.91, and MEAN_SHANNON≦4.38. To define "high" (3rd / upper quartile) for all three diversity measures, the cutoffs used on the AGP data were MEAN_OBSERVEDOTU>137.8, MEAN_FAITHPD>15.65, and MEAN_SHANNON>5.5.

[0163] This model was run with the following parameters: bin definition as the first / lowest quartile versus the third / upper quartile defined for all three diversity measures; input AGP data with a minimum survey response rate of 0.65; input with no cutoff applied to any features using a random forest algorithm in 3x cross-validation training mode; post-processing training size of 2370; and holdout / test size of 617. These are the results for the training set in cross-validation mode: sensitivity (0.73 ± 0.01), specificity (0.67 ± 0.03), precision (0.7 ± 0.02), AUC-ROC (0.78 ± 0.02). These are the results for the holdout / test set: sensitivity (0.71), specificity (0.71), precision (0.71), AUC-ROC (0.79). ROC curves and AUC values ​​are shown in Figures 5A-1 and 5A-2.

[0164] Example 12: Low gut microbiota diversity model vs. high gut microbiota diversity model (V) This model was run with the following parameters: bin definition as less than (mean - 1 * standard deviation) versus input AGP data with a minimum survey response rate of 0.85, defined as greater than (mean + 1 * standard deviation) for all three diversity measures; input with no cutoff applied to any features using a random forest algorithm in a 3x, 3-repeat cross-validation training mode; post-processing training size of 1232; and holdout / test size of 331. These are the results for the training set in cross-validation mode: sensitivity (0.73 ± 0.03), specificity (0.72 ± 0.03), precision (0.73 ± 0.02), AUC-ROC (0.81 ± 0.03). These are the results for the holdout / test set: sensitivity (0.72), specificity (0.74), precision (0.73), AUC-ROC (0.82).

[0165] The ROC curves and AUC values ​​are shown in Figures 5B-1 and 5B-2 for both (i) the training (cross-validation) and (ii) the holdout / test set. Figure 11 shows the improvement in the model's performance as features considered important to the model were added one by one through SHAP analysis (AUC values ​​of the ROC curves for the training (cross-validation) and holdout / test set).

[0166] Example 13: Key features and related recommendations of low-gut microbiome diversity models versus non-low-gut microbiome diversity models. For the model presented in Example 7, the top 30 features constituting the model are shown in Figures 6A-1 and 6A-2. Both (i) and (ii) were obtained by performing a SHAP algorithm analysis. (i) shows the average impact of each feature on the model output in order of importance from high to low. The primary / best feature was the top horizontal bar. The next best feature was the second horizontal bar, and so on. (ii) shows the impact of each instance / sample feature on the model output in more detail. The color gradient from gray to black shows the low to high values ​​of that feature. The vertical line of 0.00 defines the direction of the impact (left side is a negative impact on the model output, and right side is a positive impact on the model output). Here, the SHAP analysis output was for the reference class which was "low".

[0167] If a feature has a black value pointing to the right of the vertical line at 0.00, this indicates that a higher value of this feature contributes positively to the model output. Conversely, if a feature has a black value pointing to the left of the vertical line at 0.00, this indicates that a higher value of this feature contributes negatively to the model output. Similarly, if a feature has a gray value pointing to the right of the vertical line at 0.00, this indicates that a lower value of this feature contributes positively to the model output. Conversely, if a feature has a gray value pointing to the left of the vertical line at 0.00, this indicates that a lower value of this feature contributes negatively to the model output.

[0168] As can be seen from Figure 6A-1, for example, some of the key features that this model uses to predict low gut microbiota diversity versus non-low gut microbiota diversity were associated with antibiotic use, alcohol consumption, frequency of alcohol consumption, type of alcohol, country of residence such as the United States or the United Kingdom, frequency of vegetable consumption, and exercise frequency.

[0169] As can be seen from this, a history of antibiotic use negatively impacted the microbiome—this was one of the top features used by this model (Figure 6A-1); lower values ​​of antibiotic history were positively correlated with the impact on the model output for the "low" class when antibiotics were taken recently, meaning that recent antibiotic use pushed the microbiome state downwards. Similar results were observed for antibiotic use history in different low vs. non-low models shown in Example 8 (Figures 6B-1 and 6B-2).

[0170] In Figures 17-1 to 17-12, for each feature, the SHAP dependency plot shows a point on the x-axis with the feature value and the corresponding Shapley value on the y-axis for each data instance / sample. SHAP explained the prediction for each instance by calculating the contribution of each feature to the prediction. The explanation of the Shapley value was represented as a linear model, as an additive feature attribution method. Since the reference class here was "low," the positive coefficient of the SHAP value for the corresponding x-value of the feature indicates how much the model was influenced by this feature when predicting the "low" class.

[0171] Figure 17-1 shows a SHAP-dependent plot for antibiotic history. For all individuals with recent antibiotic use within the past year (all data points on the x-axis except for the value 500), the SHAP value was positive, indicating that this was associated with being in the "low" class of microbiome status. Similarly, only for individuals who have been using antibiotics for more than a year, the SHAP value was negative, indicating that this is associated with being in a "non-low" microbiome status, although not currently "low".

[0172] In summary, the SHAP analysis shown in Figures 6A-1, 6A-2, 6B-1, 6B-2, and 17-1 concluded that recent antibiotic use has negatively impacted the microbiome and pushed it into a "low" state. Therefore, the resulting advice or recommendation is that, if antibiotics must be taken based on a physician's prescription, the recommendation of this invention is to take other complementary solutions to improve the state of the microbiome both during and after antibiotic use.

[0173] Based on the same reasoning and explanation as above, a comprehensive examination of Figures 6A-1, 6A-2, 6B-1, 6B-2, and 17-3 suggests that red wine consumption was beneficial to the microbiome, as it was associated with a "non-low" microbiome state. Based on these results, the recommendation of this invention would be to drink red wine occasionally (1-2 times / week), regularly (3-5 times / week), or daily.

[0174] Based on the same reasoning and explanation as above, a comprehensive examination of Figures 6A-1, 6A-2, 6B-1, 6B-2, and 17-4 suggests that residents of the United States tend to have a "low" microbiome state, while residents of the United Kingdom tend to have a "non-low" microbiome state. It should be noted that this was a complex multivariate analysis attempting to explain all the data not only in relation to microbiome state but also in relation to other parts of the data. This demonstrated the importance of the geographical location of the population cohort. The recommendation of this invention would be to enhance microbiome state through supplemental dietary interventions, particularly for individuals located in geographically disadvantaged areas.

[0175] Based on the same reasoning and explanation as above, a comprehensive examination of Figures 6A-1, 6A-2, and 17-6 suggests that the frequency of vegetable consumption is related to the state of the microbiome. Specifically, increased vegetable consumption was associated with a "non-low" microbiome state. Therefore, the recommendation of this invention would be to consume at least 2-3 servings of vegetables per day, including potatoes (1 serving = vegetables / 1 / 2 cup potatoes; 1 cup raw leafy greens).

[0176] The final recommendations were the result of a complex multivariate analysis in which the characteristics were interrelated and had different ultimate impacts on the state of the individual's microbiome.

[0177] The system of the present invention, having a user-friendly digital interface, incorporates these recommendations to communicate directly with the user in order to improve the state of the user's own microbiome.

[0178] Example 14: Key features and related recommendations of high-gut microbiome diversity models versus non-high-gut microbiome diversity models. For the model presented in Example 9, the top 30 features constituting the model are shown in Figures 7A-1 and 7A-2. Both (i) and (ii) were obtained by performing SHAP algorithm analysis. For further technical explanations regarding the interpretation of different SHAP analysis plots, see Example 13. Briefly, (i) shows the average impact of each feature on the model output in order of importance from high to low. (ii) shows the impact of each instance / sample feature on the model output in more detail. Here, in Figures 7A-1 and 7A-2, the SHAP analysis output was for the reference class that was "high".

[0179] As can be seen in Figure 7A-1, some of the key features of this model, which predicts high-gluten-microbiota diversity versus low-gluten-microbiota diversity, were associated with age, country / location of residence, health status (no IBD, no diabetes), antibiotic use, exercise frequency, frequency of home-cooked meals, BMI, alcohol frequency, alcohol consumption, type of alcohol, vegetable frequency, fruit frequency, etc.

[0180] The country / location of residence, antibiotic use, frequency of alcohol, alcohol consumption, type of alcohol, and frequency of vegetable consumption were discussed in Example 13. Assumptions and recommendations regarding these characteristics for improving microbiome status are similar to those stated in Example 13. In short, the recommendations of the present invention would be to enhance microbiome status through supplemental dietary interventions, particularly for individuals located in geographically disadvantaged areas; to take other complementary solutions to enhance microbiome status both during and after antibiotic use; to drink red wine regularly (3-5 times / week) or daily; and to consume at least 2-3 servings of vegetables per day, including potatoes (1 serving = 1 / 2 cup vegetables / potato; 1 cup raw leafy greens).

[0181] Figures 7A-2 and 18-1 to 18-12 show details of how these characteristics influenced the microbiome to a "high" state. Since the reference class is set as "non-high" here, and therefore the association direction is "non-high," Figure 7B-2 shows details of how these characteristics influenced the microbiome to a "non-high" state. In short, not using antibiotics in the past year, regular or daily consumption of red wine, and daily consumption of vegetables were associated with a "high" microbiome state.

[0182] Next, we will discuss some of the key features unique to high-model versus non-high-model comparisons.

[0183] Being healthy was positively associated with a “high” microbiome state, as inferred from Figures 7A-2, 7B-2, and 18-2. Here, healthy was defined as an adult aged 20–69 years with a BMI of 18.5 kg / m2–30 kg / m2, no history of inflammatory bowel disease (IBD) or diabetes, and no antibiotic use in the previous year. BMI was calculated from height and weight. Information on BMI and antibiotic use was presented as independent features in the model as well. The recommendation was to consult a physician for appropriate medical advice according to the specific condition(s) if health status was affected by IBD, diabetes, advanced age, and an abnormal BMI range.

[0184] As can be interpreted from Figures 7A-2, 7B-2, 18-7, and 18-8, the frequency and location of exercise were associated with a "high" microbiome state. The recommendation was to exercise outdoors regularly (3-5 times / week) or daily.

[0185] Fruit consumption frequency was positively associated with a "high" microbiome state (Figures 7A-2 and 18-10). The recommendation was to consume at least 2-3 servings of fruit per day (1 serving = 1 / 2 cup of fruit; 1 medium-sized fruit; 4 ounces of 100% fruit juice).

[0186] Similarly, the frequency of home-cooked meals was positively associated with a "high" microbiome state, as inferred from Figures 7A-2 and 18-11. The recommendation was to cook and consume home-cooked meals daily (excluding instant foods such as boxed macaroni and cheese and ramen (ready-to-eat meals)).

[0187] The final recommendations were the result of a complex multivariate analysis in which the characteristics were interrelated and had different ultimate impacts on the state of the individual's microbiome.

[0188] The system of the present invention, having a user-friendly digital interface, incorporates these recommendations to communicate directly with the user in order to improve the state of their own microbiome.

[0189] Example 15: Key features and related recommendations of low-gut microbiome diversity models versus high-gut microbiome diversity models. For the model presented in Example 11, the top 30 features constituting the model are shown in Figures 8A-1 and 8A-2. Both (i) and (ii) were obtained by performing SHAP algorithm analysis. For further technical explanations regarding the interpretation of different SHAP analysis plots, please refer to Example 13. Briefly, (i) shows the average impact of each feature on the model output in order of importance from high to low. (ii) shows the impact of each instance / sample feature on the model output in more detail. Here, in Figures 8A-1 and 8A-2, the SHAP analysis output was for a reference class that was "low".

[0190] As can be seen in Figure 8A-1, for example, some of the key features that this model uses to predict low versus high gut microbiota diversity were associated with health status, age, antibiotic use, place of residence, alcohol consumption, frequency of alcohol consumption, type of alcohol, exercise, BMI, race, frequency of salty snack consumption, frequency of fruit consumption, frequency of home-cooked meals consumption, and frequency of vegetable consumption. The specific impact of these features on the model output was shown in Figure 8A-2 and Figures 19-1 to 19-15.

[0191] Examples 13 and 14 described above provide detailed interpretations of almost all of these features, including their contributions to the model and relevant recommendations and advice to follow regarding these features in order to maintain or improve the state of the microbiome.

[0192] In short, the recommendations of this invention would be to improve the state of the microbiome through supplemental dietary interventions, especially for individuals located in geographically disadvantaged areas; to take other complementary solutions to improve the state of the microbiome both during and after antibiotic use; to drink red wine regularly (3-5 times / week) or daily; to consume at least 2-3 servings of vegetables per day, including potatoes (1 serving = 1 / 2 cup of vegetables / potatoes; 1 cup of raw leafy greens); to consult a doctor for appropriate medical advice according to the specific condition(s) affected by IBD, diabetes, advanced age, and a non-normal BMI range; to exercise outdoors regularly (3-5 times / week) or daily; to consume at least 2-3 servings of fruit per day (1 serving = 1 / 2 cup of fruit; 1 medium-sized fruit; 4 ounces of 100% fruit juice); and to cook and consume home-cooked meals daily (excluding instant foods such as boxed macaroni and cheese and ramen).

[0193] The frequency of salty snack consumption is an additional feature found here that negatively impacts the microbiome (Figures 8A-2 and 19-9). The inventors recommended not consuming salty snacks (such as potato chips, nacho chips, corn chips, buttered popcorn, and french fries) or consuming them infrequently (less than once a week).

[0194] The final recommendations were the result of a complex multivariate analysis, where the characteristics were interrelated and the combination of factors had different ultimate impacts on the state of the individual's microbiome.

[0195] The system of the present invention, having a user-friendly digital interface, incorporates these recommendations to communicate directly with the user in order to improve the state of their own microbiome.

[0196] Example 16: Construction of a model for predicting the state of the microbiome in MDD data The MDD dataset consists of over 6000 metagenomic species profiles from fecal samples, paired with over 2000 relevant metadata features. The metadata used in this analysis included demographic, medical, physical activity, and dietary information. Shannon diversity (based on natural logarithm) and species richness were calculated.

[0197] The first step involved classifying α-diversity into low, non-low, medium, and high according to several of the definitions described above. Using both measures of α-diversity, metagenomic profiles were classified as having one of the diversity levels: low, non-low, medium, and high. These categories were then used in one of three different models to predict α-diversity as a categorical variable.

[0198] A list of 289 metadata features from MDD was evaluated as potential input features for the model. MDD was split into a discovery set and an evaluation set. For continuous variables, outliers were removed using a 95% (approximate Z = 1.96) confidence interval. For categorical variables, classes that accounted for less than 5% of the cohort were aggregated and re-labeled as "other", unless all classes would be aggregated if the feature was removed from the list of selectable features and other features used.

[0199] The MDD was divided into a discovery set (87.5% of the samples) and a hidden evaluation set (12.5% of the samples). The discovery set was used to train, optimize and select machine learning models. Next, the best-performing model on the discovery set was evaluated on the hidden evaluation set. Hold-out validation ensured that the dataset would not be overfitted by the model after many iterations and that all data would remain truly unseen. This set was stratified to have the same response distribution as the entire cohort after the aforementioned hygiene process. In an iterative approach, a large number of models were generated, with the discovery set randomly split into a training dataset and a test dataset in each iteration. However, the hidden evaluation set remained unchanged. This was to ensure that as the number of iterations approaches (theoretically) infinity, the possibility of the model generalizing randomly to data remains as close to 0 percent as possible.

[0200] Iteratively, 20 features were randomly selected from the set of 289 metadata features. MDD data was prepared for these 20 features, and the model discovery set was randomly split into a training set (75%) and a test set (25%). Samples were filtered based on the selected 20 features. Samples missing any metadata point from the selected feature pool were removed. Several models were trained on the 20 selected features (including neural network, random forest, gradient boosting, etc.) and optimized by cross-validation on the training set. The optimized models were evaluated on the test set. Identification of feature groups with similar predictive capabilities was performed using reinforcement learning. These steps were repeated for different sets of 20 features selected by reinforcement learning until metrics such as AUC converged to a steady-state value. Finally, evaluation of the best-performing model was performed against the hidden evaluation set.

[0201] Given that the dataset contains 289 features and 20 features were to be selected, approximately 3.46×10 30This means that there are unordered combinations of (i.e., nonirion) elements. To overcome this computational limitation, but still ensure that useful features are used, a reinforcement learning approach was employed. Accuracy recall AUC(AUC) PR The model was used with a reinforcement learning-based optimization method (multi-agent reinforcement learning) to identify feature sets with similar predictive capabilities. PR We maximized the reward function to generate clusters of models with similar feature inputs that performed well, and the reward function included the derivative of the mean importance of each feature, favoring robust models. The models were implemented in Haskell and primarily utilized the "reinforcement" framework. After integration, the well-performing models were manually checked for sensitive features that were not suitable for asking participants.

[0202] Models were trained and evaluated on classification tasks for low vs. non-low alpha diversity, low vs. high alpha diversity, and low vs. medium vs. high alpha diversity. The modeling method was repeated separately for each of these classification tasks. As described above, the discovery set (87.5% of the samples) was split into a training set (75%) and a test set (25%), which were used to train and evaluate the models. All models took 20 randomly selected features as input and predicted the answer as a categorical label (classified microbial diversity group). Model architectures included neural networks (NNs) implemented in PyTorch with Hyperopt and Optuna, distributed random forests (DRFs) optimized with H2O, and gradient-boosted machines (GBMs) optimized with H2O. All models were trained on the same split of data, trimmed 75%, and sanitized data (as the remainder depends on the number of samples removed due to sanitization). The remaining 25% was used for the test set.

[0203] Overall, gradient-boosted machines (GBMs) performed best across several feature sets, where neural networks often overfitted, while distributed random forests (DRFs) performed well. This was beneficial because GBMs provide better model explainability than neural network-based methods. Feature importance was extracted from GBM and DRF models. Locally interpretable model-agnostic explanations (LIMEs) were used to identify features important to the neural network models.

[0204] The best-performing models were evaluated on a hidden evaluation set. A total of 32 models were evaluated on this hidden data over the project span. Features were examined to ensure there was no overfitting. This also included investigating correlations between features and response variables, correlations between features and features, and correlations between response variables and response variables. This included the use of Pearson correlation, F-statistics, and chi-squared tests for different types of data.

[0205] Example 17: High-performance model vs. low-performance model (MDD data) For MDD data, "low" was defined as a sample with a Shannon index of 0.59–3.63 and richness of 14–150, based on the natural logarithm of the Shannon index. "High" was defined as a sample with a Shannon index of 4.05–4.88 and richness of 196–331, based on the natural logarithm of the Shannon index. AUC of 0.837 and AUC of 0.577. PR The best-performing model for the "high-low" task performed on (Figure 20). The mean error per class was 0.232, and the maximum precision was 0.827. Maximum precision is the maximum precision given specific true positive and false positive rate (TPR, FPR) thresholds. Here, FPR = 0.248, which maximized the F1 score (0.741). AUC and AUC PRThe values ​​were calculated using a hidden evaluation set. This model used 20 metadata features. The relative weights of each feature are shown in Figure 21. The features were a mixture of binary (antibiotic use), categorical (age), and continuous (BMI).

[0206] Example 18: Low-profile vs. Non-low-profile model (MDD data) For MDD data, "low" was defined as samples with a Shannon index of 0.59–3.50 and abundance of 14–138, based on the natural logarithm of the Shannon index. "Non-low" was defined as samples with a Shannon index of 2.09–4.88 and abundance of 57–331, based on the natural logarithm of the Shannon index. The best-performing model for the "low-non-low" task had an AUC-ROC of 0.902 (Figure 22). The precision was 0.883. The ratio of low to non-low samples was 1:3.59 out of a total of 422. Many of the 20 features (and their relative importance) overlapped with those determined to be optimal for the high-low model in Example 17, such as physical activity, height, weight, and alcohol consumption. Non-overlapping features in the low-non-low model included stress / anxiety, vegetable consumption, cat or dog ownership, overseas travel, smoking, and bloating. The relative weights of the 20 features are shown in Figure 23.

[0207] Example 19: Constructed ensemble model for predicting microbiome state in MDD data Individual models performed well on binary classification problems such as "high-low," but they performed poorly in solving imbalanced tasks like "low-medium-high." To overcome this, ensemble modeling (also known as layered modeling) was employed (Figure 24). Binary classification for "high vs. low" with a continuous output model using thresholds performed best. Other models investigated included multi-output models using consensus decisions from binary combinations of high, medium, and low, i.e., high / medium + high / low and medium / low models used together.

[0208] The binary model used a threshold for the decision that maximized F1 during training. The model's output was the probabilities of two classes, e.g., p0 and p1. The decision threshold was learned by the model, and for example, if p0 > 0.50, the result could be assigned to class 0 and class 1 when used with alternative probability p1. If both thresholds failed, the sample was said to be "low confidence" and passed to the continuous model (Figure 24).

[0209] The continuous model was trained with the same data as the "high-low" model, meaning all intermediate samples were discarded. Low samples were defined as when both alpha diversity measures fell below the third ternary, high as similarly the highest third ternary, and all intermediate samples were labeled intermediate. The features remained the same to ensure consistent sample size. Additionally, this kept the questionnaire small.

[0210] Example 20: Low-Medium-High Model (MDD Data) For our own MDD data, we defined "low" as samples with Shannon exponents ranging from 0.59 to 3.63 and richness from 14 to 150, based on the natural logarithm of the Shannon exponents. "Medium" was defined as samples with Shannon exponents ranging from 3.63 to 4.05 and richness from 150 to 196, based on the natural logarithm of the Shannon exponents. "High" was defined as samples with Shannon exponents ranging from 4.05 to 4.88 and richness from 196 to 331, based on the natural logarithm of the Shannon exponents. The ensemble modeling approach provides a model with an accuracy of 0.75 (Figure 25). The final previous hidden test set was imbalanced. Quantitative overlap between the two measures of diversity meant that most samples were moderate or intermediate (Figure 26). [Table 1]

Claims

1. A method for determining the state of the gut microbiome, which is performed by computer. A step of receiving characteristic parameters related to the state of the gut microbiome in a subject, wherein the characteristic parameters are (i) The geographical country of the subject residence, including its specific latitude and longitude, (ii) Whether antibiotics have been used in the past year, and the use of antibiotics, including if antibiotics have been used in the past year. (iii) Whether drug therapy has been used in the past 12 months, and the use of drug therapy, including if drug therapy has been used in the past 12 months, (iv) Body measurement data including age, weight, height, body mass index, and sex, (v) Alcohol consumption, including the type of alcohol, the amount consumed, and the frequency of consumption, (vi) Smoking status and, (vii) Exercise and / or physical activity, including the location, frequency, and duration of indoor or outdoor exercise, (viiii) national character, (ix) The trip, including the place and duration of the trip, (x) Sleep, including duration and / or quality of sleep in hourly units, (xi) stress, anxiety, and / or depressive states, (xi) Medical history including blisters, irritable bowel disease, diabetes, (xiiii) A history of allergies, including seasonal allergies or food allergies and / or food intolerances, Vaccination status, including influenza (xiv) vaccine or pneumococcal vaccine, (xv) Use of nutritional supplements, including vitamin or mineral supplements, (xvi) Sources of drinking water and (xvii) Personal hygiene, including tooth flossing, use of deodorants, and use of cosmetics, (xviiii) The type of food intake, the amount of food intake and the frequency of food intake, including the intake of vegetables, fruits, fermented foods and / or whole grains. A receiving process selected from the group consisting of, A step of determining the state of the intestinal microbiome in the subject in relation to the position of the subject within the distribution of the population, by using one or more predictive models based on the aforementioned characteristic parameters, A step of determining recommendations for improving or maintaining the state of the gut microbiome in the subject based on the characteristic parameters, A step of providing the aforementioned recommendation to the aforementioned target, Includes, The step of determining the state of the intestinal microbiome in the subject is: The process involves using the aforementioned prediction model to process the feature parameters and predict the alpha diversity score of the microbiome. The alpha diversity score of the aforementioned microbiome is compared with the distribution of a reference population, Based on the comparison results of the distribution, it is determined whether the state of the gut microbiome in the subject is low, medium, or high. including, method.

2. The method according to claim 1, wherein the characteristic parameters are obtained from the subject's responses to a questionnaire.

3. The method according to claim 1 or 2, wherein the determination of the state of the intestinal microbiome is further performed by a biological sample for quantifying the microbiome diversity of the subject.

4. The step of determining recommendations for improving or maintaining the state of the gut microbiome in the subject is: The characteristic parameters related to the state of the gut microbiome are compared with one or more corresponding benchmarks. Based on the comparison results with one or more benchmarks, the characteristic parameter is evaluated as low, medium, or high. The method according to any one of claims 1 to 3, including

5. The method according to claim 4, wherein the subject is notified of the state of its own intestinal microbiome via a computer interface.

6. A computer-implemented system for determining the state of the intestinal microbiome and providing recommendations for improving or maintaining the state of the microbiome in a subject, according to the method described in any one of claims 1 to 5.

7. The aforementioned recommendation is, (i) Whole grain foods, (ii) Beans and legumes, (iii) Fibers, (iv) Nuts and seeds, and The method according to any one of claims 1 to 5, comprising a recommendation for a group of foods or nutrients selected from the group consisting of (v) omega-3 fatty acids.

8. The above recommendations regarding foods or groups of nutrients, (i) Whole grain foods, (ii) Beans and legumes, (iii) Fibers, (iv) Nuts and seeds, and (v) Contains omega-3 fatty acids The method according to claim 7, including the recommendation of a meal plan or recipe.

9. The method according to any one of claims 1 to 5, 7, or 8, wherein the recommendations include recommendations for lifestyle habits relating to exercise, alcohol consumption, sleep, dental hygiene, and the use of nutritional supplements.

Citation Information

Patent Citations

  • Method for examining intestinal bacteria

    JP2019200687A

  • Method and system for microbiome-derived diagnostics and therapeutics

    US20170270272A1

  • Pet food recommend device and pet food recommend method, supplement recommend device and supplement recommend method, and bowel age calculation formula determination method and bowel age calculation method

    WO2020213667A1