A formula milk powder selection method based on neural network explainability to explore the relationship between bacterial interaction and metabolic products

By analyzing gut microbiota interactions and metabolite relationships through neural networks, personalized milk powder formulas are designed, solving the problem of lack of personalization in traditional milk powder formulas and achieving more effective improvement in health status.

CN119296657BActive Publication Date: 2025-12-16ICARBONX (ZHUHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411415846.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-11
Publication Date
2025-12-16
Estimated Expiration
2044-10-11

AI Technical Summary

Technical Problem

Traditional methods of selecting infant formula cannot effectively combine the interaction of individual gut microbiota and the relationship of metabolites, resulting in a lack of personalization and targeting in infant formula, making it difficult to meet the health needs of different groups.

Method used

Using a neural network-based approach, we established a specific population database, explored the relationships between gut microbiota interactions and metabolites, and combined this with environmental factors to design personalized milk powder formulas.

Benefits of technology

This enables targeted intervention based on individual gut health status, improves the personalized selection of milk powder formula, and enhances the effect on improving health status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119296657B_ABST
    Figure CN119296657B_ABST
Patent Text Reader

Abstract

The present application relates to a formula milk powder selection method based on neural network explainability to explore the relationship between bacterial interaction and metabolites, comprising: establishing a comprehensive database of specific population queues; building a global neural network; inputting key factors and lifestyle scales into the environmental factor analysis network in the secondary sub-network; extracting main factors in the multi-level interaction network according to network weights, and determining the strain and genus of several specific factor-related strains; extracting specific factors, metabolites and main strains and genera, inputting them into the secondary sub-network for classification, dividing the specific population into several categories, obtaining the intestinal health status of the specific population in each category, and designing and selecting specific milk powder formula according to the associated intestinal strains and genera. The present application uses AI neural network to explore the relationship between bacterial interaction and metabolites, combines with the living environmental factors, reveals the metabolite formation process and disease risk, and realizes the personalized selection of milk powder formula.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of neural networks, in particular to a formula milk powder selection method based on neural network explainability to explore the relationship between bacterial interaction and metabolic products. BACKGROUND

[0002] Human-related diseases are often caused by multiple factors, including environmental factors, living habits, daily diet, etc. Since human life activities are a metabolic process, food nutrition is obtained, absorbed and metabolized, and different nutrients have different effects. Therefore, milk powder is also a common choice. In addition to meeting the nutritional needs, it is also required to have a "diet therapy" effect to improve one's own health status based on one's own health awareness.

[0003] Traditional formula milk powder selection methods mainly include the following two methods:

[0004] 1. In vitro simulated digestion and absorption experiment: based on the research of related literature, different nutritional combinations are determined, a variety of protein combinations are designed in advance, and corresponding gradient solutions are configured. Add protease, lipase and other substances to simulate the gastric juice environment at 37℃, study the formula with the highest protein absorption efficiency, and further design in vitro cell absorption experiment to observe the protein absorption rate on the culture dish to obtain the best formula.

[0005] 2. Animal experiment: a variety of protein combinations are designed in advance, and young experimental animals are fed for a period of time, then the external indicators such as body length and weight of the experimental animals are compared, the body nutrient intake is measured, and further injection of foreign cells simulates invasion to measure the phagocytosis rate to reflect the immunity, thereby determining the best formula of milk powder. SUMMARY

[0006] The embodiment of the present application provides a formula milk powder selection method based on neural network explainability to explore the relationship between bacterial interaction and metabolic products, so as to realize personalized selection of milk powder formula.

[0007] According to the embodiment of the present application, a formula milk powder selection method based on neural network explainability to explore the relationship between bacterial interaction and metabolic products is provided, including the following steps:

[0008] Establish a comprehensive database of specific population cohorts;

[0009] Build a global neural network, input the multi-omics data of the specific population cohort in the comprehensive database for training, obtain the characteristic relationship between bacterial interaction and its metabolic products, and extract key factors therefrom;

[0010] The key factors and the life habit scale are input into the environment factor analysis network in the secondary sub-network, the interaction relationship between the environmental factors and the metabolites is analyzed, and the explanatory correlation relationship between the specific factors and the metabolites and the strain genus is obtained;

[0011] The main factors are extracted in the multi-level interaction network according to the network weight, and the strain genus associated with the specific factor is determined;

[0012] The specific factors, metabolites and main strain genus are extracted, input into the secondary sub-network for classification, the specific population is divided into several categories, the intestinal health status of the specific population in each category is obtained, and targeted milk powder formula design and selection are performed according to the associated intestinal strain genus.

[0013] Further, in the comprehensive database of the specific population queue, the user sample is collected and the basic information is collected for the population object meeting the condition combination, the corresponding strain, genus, gene and metabolism multi-omics data are obtained by performing macro-genome sequencing and mass spectrum analysis on the sample.

[0014] Further, the key data screened out by the correlation specific population queue is subjected to a Venn diagram intersection comparison, and the correlation screening result is confirmed to be effective.

[0015] Further, the global neural network is designed to train the omics data related to the specific population, including macro-genome data and metabolism data, and the corresponding interaction network probability relationship is obtained; wherein the global neural network includes but is not limited to LSTM, RNN, MLP, CNN, RF and XG commonly used algorithm components.

[0016] Further, the environment analysis network of the secondary sub-network is used to analyze the relationship between the environmental factors and the metabolites, so as to obtain the multi-level interaction network correlation relationship of the environmental factors, and the metabolites are further associated with the corresponding strain genus through the environmental factors.

[0017] Further, the classification network of the secondary sub-network automatically divides the main factors into several types according to the extracted main factors, including user information, metabolites and strain genus, analyzes the data characteristics of the specific strain genus in each type, and is used to guide the selection of nutritional substances in the milk powder formula, and specifically adjusts the intestinal health status, wherein the calculation method of classification includes but is not limited to manual grouping, automatic classification algorithm and automatic clustering algorithm.

[0018] Further, the comprehensive database is divided into four layers, the first layer is user information, living habits and environment, eating habits, the second layer is macrogenomic data, by collecting user fecal samples, by sequencing experiment to obtain intestinal related flora gene, flora abundance, genus abundance and other data reflecting the user intestinal health status, the third layer is metabolomics data, further by mass spectrometry analysis to obtain metabolite composition, gene expression amount metabolic data, and the fourth layer is multi-layer interaction relationship network data, by specific factors, the associated metabolites, strain and genus causal process are obtained from the interaction relationship network graph.

[0019] Further, the neural network framework is a logical container, which internally includes a global neural network and a secondary sub-network, receives queue topic data, macrogenomic data and metabolic data, wherein the queue topic data and the macrogenomic data are calculated with each other as input and output respectively, and the multi-level interaction relationship of the specific population is mined.

[0020] Further, the global neural network uses a nonlinear neural network algorithm to fit the interaction relationship between metabolites and intestinal flora, and bidirectionally trains the metabolites and intestinal flora, in which the metabolites are input to predict the intestinal flora in a forward direction, and the intestinal flora is input to predict the metabolites in a reverse direction, after multiple calculations, the interaction network relationship between the metabolites and the intestinal flora is obtained by extracting the weight.

[0021] Further, according to the network relationship weight, the action process, the action degree and the accurate explanation of the expression amount of any metabolite or environmental factor are obtained, the multi-level interaction network graph is drawn and the key path is identified, and the causal relationship explanation of the related factor cause is obtained.

[0022] The formula milk powder selection method based on neural network explainability to mine flora interaction and metabolite relationship in the embodiment of the application mines the flora interaction and metabolite relationship by AI neural network, combines the living environment factors, reveals the metabolite cause process and disease risk, and realizes personalized selection of milk powder formula. BRIEF DESCRIPTION OF DRAWINGS

[0023] The drawings described herein are used to provide further understanding of the application, and form a part of the application, the illustrative embodiments of the application and the description thereof are used to explain the application, and do not constitute improper limitation on the application. In the drawings:

[0024] Figure 1 The formula milk powder is customized based on the interaction relationship of flora and metabolism;

[0025] Figure 2 The basic information of 712 users is 5% screenshot;

[0026] Figure 3Figure 5% screenshot of metagenomic sequencing experimental data;

[0027] Figure 4 Figure 5% screenshot of metabolomic analysis experimental data;

[0028] Figure 5 Figure 4-level data structure diagram for specific cohort;

[0029] Figure 6 Figure computational scheme for milk powder formula selection based on non-linear multi-level interaction network relationship;

[0030] Figure 7 Figure comparison of metabolic features for different network pairs of global neural network;

[0031] Figure 8 Figure 10% screenshot of Bifidobacterium longum verification data;

[0032] Figure 9 Figure screening effect diagram based on Bifidobacterium longum verification model;

[0033] Figure 10 Figure 10% screenshot of Bifidobacterium longum verification data;

[0034] Figure 11 Figure screening effect diagram based on Bifidobacterium longum verification model;

[0035] Figure 12 Figure graph of mutual weight output relationship;

[0036] Figure 13 Figure single-factor multi-level interaction network graph;

[0037] Figure 14 Figure multi-factor multi-level interaction network graph;

[0038] Figure 15 Figure explanatory graph of living environment factors;

[0039] Figure 16 Figure explanatory graph of multi-level causal process of dietary factors;

[0040] Figure 17 Figure graph of overall interaction relationship network and main factor coverage;

[0041] Figure 18 Figure comparison graph of main factor interaction network graphs for different groups;

[0042] Figure 19 Figure comparison graph of theoretical population normal distribution;

[0043] Figure 20 Figure P-value check graph of theoretical population normal distribution;

[0044] Figure 21Design flowcharts for application implementation examples;

[0045] Figure 22 406 sample images were collected for the application examples;

[0046] Figure 23 A comparative analysis chart of feeding methods;

[0047] Figure 24 A comparative analysis chart of birth methods;

[0048] Figure 25 For interpretable multi-omics network interaction feature maps;

[0049] Figure 26 A comparison diagram of multi-omics network features for interpretability of feeding methods;

[0050] Figure 27 A process explanatory diagram of the multi-omics characteristic factors of breastfeeding;

[0051] Figure 28 Extract the main health factor map from the feature map based on different combinations of constraint domain condition values. Detailed Implementation

[0052] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0053] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0054] The technical problem solved by the present application is to reveal the metabolic process and disease risk of the target population by AI neural network mining of the relationship between the intestinal flora interaction and its metabolites, and combining the living environment factors, so as to realize the personalized selection of milk powder formula (such as Figure 1 ).

[0055] The method has open definition domain conditions, including age, gender, region, altitude, disease history, etc. Combined with the target population in market information, a number of definition domain conditions and their condition values are proposed, and multiple definition domain conditions are combined into a specificity. The target population that meets these definition domain conditions is described as a specific population. For example, the design theme is "chronic disease patients", and the definition domain conditions are three high, diabetes, BMI overweight, etc.

[0056] In the embodiment of the method, the definition domain conditions are "carrying 14 diseases", "0-30 years old, grouped into infants, children, teenagers, young people, middle-aged people", and the causes of metabolites are obtained by studying the living habits, work and rest, and dietary structure of the specific population. Further, the key strains and genera are found, and the specific population is classified, the intestinal health status of each classification is comprehensively evaluated, and a number of milk powder formulas are determined, and the key nutrients in the formula are selected, so as to realize the "diet therapy" intervention on human health. The process of the embodiment of the method includes the following steps:

[0057] A comprehensive database of specific population queues is established, and 712 user basic information meeting the specific conditions are collected, including name, gender, height, age, weight, work and rest, dietary habits, grouped according to age 1-6, 7-12, 13-18, 19-24, 25-30 (such as Figure 2 ), user fecal samples are collected, and multiple omics data including flora abundance, genus abundance, strain gene, and metabolite data are obtained through macro gene sequencing (such as Figure 3 ) and metabolic mass spectrometry experiment analysis (such as Figure 4 ).

[0058] The data model N-Tier is established, and in this embodiment, 4-Tier is set, wherein L1 is a queue topic data, and collected user basic information includes user basic information, life habit Q&A, dietary structure, etc. The related information of specific population is marked by using the authoritative literature PubMed in the biomedical and health field, such as smoking and drinking, irregular work and rest, dietary characteristics and the like in life habit, and the disease, risk and the like health status information associated with the disease are marked, L2 is colony abundance data of metagenomics collected by the queue, and the strain species and genus are marked by using the microorganism metabolite database MimeDB, so that the name of the strain and the genus is obtained, L3 is metabolite abundance data collected by the queue, and the identification InChIKey of the metabolite is obtained by marking by using the chemical database PubChem and the biological related chemical entity database ChEBI, and L4 is a multi-level interaction network relationship based on L1 / L2 and L2 / L3, and further classification is performed on the specific population selected by different characteristics, and the related function or application is designed according to the corresponding classification, such as the milk powder formula in the method embodiment (for example Figure 5 ).

[0059] Based on the above data model, the overall calculation scheme of the milk powder formula selection of the multi-level interaction network relationship based on the nonlinear model is designed (for example Figure 6 ).

[0060] Wherein S10 is a specific population queue after a limited domain condition is selected, the sample of the S10 queue is subjected to mass spectrum analysis to obtain L3 metabolite data S102, then PubChem / ChEBI is used in S103 to obtain the InChiKey of the S104 metabolite, the dimension of the metabolomics data itself is large, the commonly used machine learning algorithm random forest is used in S105 to perform feature mining on the data, then the linear Pearson algorithm is used to obtain the correlation between the metabolite components, the weight is extracted and the data is filtered, only the metabolites that can be annotated by HMDB are retained, and other metabolites are removed, and then the S301 in the global network is input.

[0061] Meanwhile, the sample of the S10 queue is subjected to macrogenomic sequencing to obtain L2 macrogenomic data S202, metaphla4 is used in S203 to obtain S204 abundance of the flora, including the strain, the genus and the expression amount, the machine learning algorithm random forest is used in S205 to perform feature mining on the data, then the linear Pearson algorithm is used to obtain the internal correlation between the strain and the genus, the weight is extracted and the data is filtered, the non-zero values less than 25% are removed, the genus with an average abundance less than 0.1% is removed, and the related main factors are retained, and then the S302 in the global network is input.

[0062] The neural network framework S30 is a logical container including the global neural network S300 and the secondary sub-network S600, which receives the L1 queue topic data, the L2 macro gene data, and the L3 metabolic data, wherein the L1 and the L2 respectively take each other as input and output to calculate the interaction relationship, and further mine the multi-level interaction relationship of the specific population.

[0063] The global neural network S300 uses a nonlinear neural network algorithm to fit the interaction relationship between the metabolites S301 and the intestinal flora S302. Unlike the traditional linear Pearson method, the nonlinear neural network algorithm can more accurately mine the correlation between each other. Here, the global neural network is bidirectionally trained on S301 and S302. The forward direction takes S301 as input to predict S302, and the reverse direction takes S302 as input to predict S301. After multiple calculations, the weights are extracted to obtain the interaction network relationship between S301 and S302.

[0064] Further description of the global neural network S300 design, S300 is a logical component container, integrating different algorithm components, which can load time series related algorithm components, or traditional feature representation algorithm components. The standard neural network types in it all belong to standard algorithm components, such as LSTM, RandomForest, XGBoost, CNN, MLP, etc., which are briefly described as LM, RF, XG, CNN, MLP.

[0065] For the representation ability of metabolic characteristics, the experimental data in the public literature is used to compare the representation ability of different algorithms for the interaction relationship of metabolites. Based on the indicators in the binary classification evaluation model, including F1 score (f1-score), accuracy (precision), and recall, the performance of several algorithm components is compared, including RF, XG, MLP-model1, MLP-model2, RF / XG combination, xSwiGLU / ReTa combination, etc. (such as Figure 7 ), and the model with the smallest residual, RF and XG, are selected as the algorithm components for metabolic representation in the embodiment of the method.

[0066] The algorithm combination selected by the global neural network S300 in this embodiment, wherein the timeline causal network S303 uses LM, and the traditional feature network S304 uses RF and XG. S301 and S302 are input into LM, RF, and XG, respectively, to obtain the LM, RF, and XG outputs, and the weight matrix of the hidden layer is extracted to obtain the interaction network relationship I lm , I rf , and I xg .

[0067] For the interaction network relationship I lmI rf I xg The weights of each factor are calculated and sorted according to S301 and S302 respectively. On the S301 side, associated bacterial strains and genera and their weight coefficients are extracted based on the interaction network relationships. For the LM subnet... Sort the LM outputs, where coef i To assign weights to metabolite i, the same applies to the outputs of the RF and XG subnets. The outputs of the RF or XG subnets are sorted, and the Top M major metabolites are extracted, thus establishing the interaction network relationship I. lm I rf I xg The three outputs R of the metabolites are obtained. lm R rf R xg Apply the same sorting operation to the S302 side to extract the Top N major bacterial strains and genera, and then analyze them in the interaction network I. lm I rf I xg The three outputs R′ of the strain and genus were obtained. lm 、R′ rf 、R′ xg .

[0068] The top N factors calculated by the global network S300 are summarized to obtain the comprehensive result S400. The calculation process of the comprehensive result is as follows: the intersection of LM, RF, and XG is taken, including the metabolite R. lm ∩R rf ∩R xg and strains and flora R′ lm ∩R′ rf ∩R′ xg Using the intersection results of metabolites R lm ∩R rf ∩R xg From Interaction Network Relationships I lm I rf I xg Obtain the set S′ of related bacterial strains, and use the intersection result R′ of the strains and genera. lm ∩R′ rf ∩R′ xg From Interaction Network Relationships I lm I rf I xg Obtain the set S of associated metabolites, thereby obtaining the screened metabolites (R). lm ∩R rf ∩R xg )∪S and bacterial strains (R′) lm ∩R ′ rf∩R ′ xg )∪S′, further extracting the correlation between the main factors in the metabolite and strain genus interaction network I xg And according to the Top N factors, the input of metabolites and the input of strain genus are obtained from S301 and S302 respectively to form the comprehensive output S400 of the global network.

[0069] In this embodiment, Bifidobacterium is taken as an example for illustration, S400 takes Top 100 strains, including strains and metabolites, and the strains are annotated using MiMeDB, and 489 items are obtained based on strain species name matching, and 826 items are obtained based on metabolite InChIKey matching, and further deduplication is performed on the data to obtain 385 annotated metabolites and 194 annotated strains. Since MiMeDB is a public authoritative dataset, the range of annotation is known knowledge, and the range of features learned by LM, RF and XG may exceed the annotation range of MiMeDB, which belongs to the current unknown knowledge. The intersection of the annotation range of MiMeDB and the learning results of the model is obtained, and it is confirmed that the learning results of the model are effective.

[0070] In this embodiment, s__Bifidobacterium_longum is taken as an example, MiMeDB is associated with 113 metabolites in sample data S301 (such as Figure 8 ), and the intersection analysis is performed on the Top 100 of the LM, RF and XG results of the global network output, wherein the intersection result of LM, RF and XG is 11 items, the intersection result of MiMeDB and RF is 26 items, the intersection result of MiMeDB and XG is 25 items, and the intersection result of MiMeDB and LM is 18 items (such as Figure 9 ), which proves that the model screening for Bifidobacterium longum is effective.

[0071] In this embodiment, s__Bifidobacterium_breve is taken as an example, MiMeDB is associated with 100 metabolites in sample data S301 (such as Figure 10 ), and the intersection analysis is performed on the Top 100 of the LM, RF and XG results of the global network output, wherein the intersection result of LM, RF and XG is 15 items, the intersection result of MiMeDB and RF is 24 items, the intersection result of MiMeDB and XG is 27 items, and the intersection result of MiMeDB and LM is 14 items (such as Figure 11 ), which proves that the model screening for Bifidobacterium breve is effective.

[0072] ​Then the comprehensive result S400 is input to the input end S602 of the two-layer sub-network. The two-layer sub-network S600 is also a logical component container, which can integrate different neural network algorithm components. The embodiment of the method inherits the environmental factor analysis network S603 and the classification sub-network S604.

[0073] In S501, the L1 queue topic data, including user basic information, life habit Q&A in the embodiment, is monitored and labeled using PubMeb in S502 to obtain the health status evaluation of the user, including disease-related unhealthy life habits, disease risks, such as smoking, drinking, etc. will cause changes in related intestinal flora, and may cause associated diseases, etc. Then the data is input to the two-level sub-network S601.

[0074] In the two-layer analysis network, the environmental factor analysis sub-network S603 uses S601 including life habits, food, environmental factors, etc. as input and uses S602 as output. Then S602 is used as input and S601 is used as output. After bidirectional training, the interaction network I of life habits and bacterial strains is obtained. env .

[0075] In the two-layer analysis network, the classification sub-network S604 combines the input of S601 and the input of S602. There are four grouping methods. The first method uses a neural network algorithm to realize sample grouping. The second method uses a clustering algorithm to form several natural categories according to the inherent characteristics of the data. Since the number of categories is large, adjacent categories are combined into several groups according to the Euclidean distance. The third method uses a classification algorithm to specify the number of categories. By continuously adjusting the position of each category center point in the multi-dimensional space of the data, the center point of each data point to the category it belongs to is realized. The fourth method specifies the grouping label when the sample is collected. The above four methods can be replaced according to the specific research target in the implementation path. In the embodiment, the fourth method is adopted, that is, the sample grouping specified by artificial is used. In the five age end grouping, the correlation between the strains and genera associated with 14 diseases, metabolites, and living environmental factors is comprehensively observed. Whether each group has obvious characteristics is analyzed to determine different milk powder formulas.

[0076] The weight relationships of environmental factors, metabolites, and strains and genera are summarized based on the above calculation results. The interaction network of metabolites and strains and genera uses I' xg In the embodiment, the XG network weight with the best metabolic characterization is used to obtain the mutual weight output S700 as I env ∪I′ xg (As Figure 12) In this embodiment, the relationship between strain genes and environmental factors is calculated to be 2073600 pairs, the relationship between genera and environmental factors is calculated to be 398161 pairs, the relationship between genera, metabolites, and environmental factors is calculated to be 36216324 pairs, and the relationship between strain genes, metabolites, and environmental factors is calculated to be 27133681 pairs.

[0077] Based on the pairs of different omics dimensions of S700, the correlation between each other is calculated, Top N' is set, and the weight W is extracted n , the multi-level interaction network relationship weight S800 is extracted, and the indicator is set In the exhaustive process, the of pairs is determined whether it is within the weight range of Top N, and the weight of pairs is extracted from S700 to obtain the multi-level interaction network relationship weight S800.

[0078] According to the network relationship weight S800, any metabolite or environmental factor can obtain the accurate explanation of the whole process, the degree of action, and the expression amount, and through drawing the multi-level interaction network diagram and identifying the key path, the causal relationship explanation S900 of the cause of the related factors can be obtained.

[0079] The explanation S900 can be represented by a multi-level interaction network diagram, and according to S800, different levels are distinguished by line color, and different weight sizes are distinguished by line thickness. Starting from a single environmental factor, the metabolites, strain genera can be associated to draw the action relationship network (such as Figure 13 ), and when multiple environmental factors interact, the metabolites, strain genera can be associated to draw a multi-factor interaction relationship network diagram (such as Figure 14 ). Taking a specific living environmental factor as an example, such as the effect of smoking on metabolism and intestinal flora, the relevant main metabolites and the correlation explanation of the intestinal strain genera state can be obtained from the interaction network relationship diagram (such as Figure 15 ), and further, the causal process explanation, such as meat in dietary habits, the relevant main metabolites and intestinal strain genera data are extracted according to S800 to draw a Sankey diagram, and the multi-level causal process explanation of the dietary factor is obtained (such as Figure 16 ).

[0080] In the embodiment of the present method, the top 30000 is taken, the overall interaction relationship network is obtained by drawing a multi-level interaction relationship network diagram, and then the top n (n = 1, 2, …, n) of the metabolites is taken according to the weight, and the same method is used for drawing, and the points and lines related to the associated life factors, metabolites and strains and flora are drawn using red, and whether the environmental factors and strains and genera associated with the circled metabolites are main factors is observed, so as to determine that the interaction network relationship range of the metabolites marked by n = 17 occupies the main part of the overall network (such as Figure 17 ).

[0081] Based on the above-mentioned 17 main metabolites, the multi-level main factors of the specific population cohort are further determined, and the corresponding multi-level interaction relationship network diagram is drawn for different age groups, including life and work, eating habits, gender, age, metabolites, strains and flora, and it is further found that the network diagrams of different age groups have obvious different characteristics (such as Figure 18 ). After it is determined that the selected specific population cohort related to 14 diseases has different characteristics of intestinal flora in different age groups, the health status of the intestinal flora related in different groups is evaluated, several kinds of milk powder formulations are designed, the nutrient molecules having a targeted intervention effect on intestinal strains and genera are selected, the health status of the intestinal flora is adjusted, and the improvement of the health status of the user is realized.

[0082] The P value check S1000 is performed for the explanatory health causes, the credibility of the explanation is confirmed, the index data extracted in S900 is shuffled using a random method, the theoretical population data conforming to the normal distribution of each factor index is generated, the mean square error, the average value and the CV value of each factor index are counted, whether the S900 sample data conforms to the theoretical distribution is compared, and whether the model generated screening result S900 is credible is judged according to the confidence interval ≤0.05. In the embodiment, 200 times of random and 10000 times of random are used, the x axis is the theoretical value, and the y is the Top 100 index sorting value. The probability graph of the S900 sample data and the theoretical population data is compared (such as Figure 19 ). It is reflected from the result that the S900 sample result screened out conforms to the normal distribution, and the middle section is relatively consistent. The standard deviation std, the mean value mean and the coefficient of variation CV can be obtained. Most of the S900 samples are in the confidence interval ≤0.05 (such as Figure 20 ). Therefore, it is confirmed that the nonlinear model of the present method conforms to the real world, and the above-mentioned calculation process is integrated to obtain the nonlinear model S1200.

[0083] In a specific application, the consumer S20 collects a fecal sample for mass spectrometry analysis S1301 and metagenomic sequencing S1302, and collects corresponding cohort theme data, such as the basic information and lifestyle questionnaire Q&A used in this embodiment, performs predictive analysis S1400 on the user data, calculates the multi-level interaction network characteristics of the user and the cohort similarity score S1500, and then selects a milk powder formula based on the user's intestinal flora characteristics S1600.

[0084] Based on the above method, a verification embodiment is designed, the theme cohort T1001 is set as "infant milk powder", the domain conditions T1002 include "feeding mode, birth mode", each domain condition is compared with the control group through the condition value, the characteristics such as gene, metabolism, and flora interaction are compared and analyzed, and several characteristic classifications are mined through the method, each classification represents a health condition, the health status in the condition is identified through health annotation, and the nutritional formula is determined to regulate the intestinal "flora" and promote the health of infants (such as Figure 21 ).

[0085] 406 infant fecal samples are collected, and the cohort T1003 registers information such as feeding mode, gender, birth mode, and birth day (such as Figure 22 ), wherein BF is breast feeding, FF is milk powder feeding, and the birth mode includes cesarean section and natural birth. The samples are subjected to metagenomic sequencing and mass spectrometry analysis, and T1004 uses the method to calculate the multi-omics multi-level interaction network characteristics.

[0086] The feeding mode includes breast feeding and milk powder feeding, the Top20 of the genus, species, and metabolites are extracted and analyzed, and through comparison, it can be found that, for example, the higher the Shannon abundance of the species, the greater the probability of milk powder feeding, and vice versa, the greater the probability of breast feeding, the higher the abundance of Clostridium B in the species gene, the greater the probability of milk powder feeding, and vice versa, the greater the probability of breast feeding, and by extracting metabolites based on metabolic pathways from the Top20 genes, it can also be found that there are obvious differences between the two groups (such as Figure 23 ).

[0087] The feeding methods include cesarean section and natural birth. The top 20 of the genus, species and metabolites are analyzed respectively. It is found through comparative analysis that, for example, the higher the Shannon abundance of the species Phocaeicola merdigallinarum, the greater the probability of natural birth, and vice versa, the greater the probability of cesarean section. Similarly, the higher the abundance of the species Bacteroides intestinigallinarum, the greater the probability of natural birth, and vice versa, the greater the probability of cesarean section. The higher the abundance of the species gene Ruminiclostridium, the greater the probability of natural birth, and vice versa, the greater the probability of cesarean section. By extracting metabolites based on metabolic pathways from the top 20 genes, it can also be found that there are obvious differences between the two groups (such as Figure 24 ).

[0088] In T1005, the method is further used to generate a multi-omics network interaction graph with explainability (such as Figure 25 ), and based on different defined domain condition values, a corresponding multi-omics network interaction graph can also be formed (such as Figure 26 ). By comparing breast-feeding and formula-feeding, the corresponding network interaction feature maps have obvious differences. The explainability of the specified defined domain condition value is obtained by using the explainable multi-omics network interaction graph, such as the species, species and metabolic processes affected by breast-feeding (such as Figure 27 ).

[0089] Based on the combination of the defined domain condition values, four network feature maps are formed, and the main health affecting factors are extracted and health labeled and evaluated (such as Figure 28 ). Here, if the number of combinations is too large, the number of classifications can be specified, and all samples are classified using a classification algorithm, and then the common features of each class of sample features are extracted, and then health labeling and evaluation are performed. After understanding the characteristics of each class, a number of standard model libraries T1006 are formed, such as A, B, C, etc., and based on health labeling and evaluation, the formula is formulated, and the infant intestinal flora is improved through nutrient absorption. Here, four defined domain condition value combinations are used.

[0090] Subsequently, any infant sample T1007 is sampled, and after macro-genome sequencing and mass spectrometry analysis using the method, the multi-omics multi-level interaction network features of the sample T1008 are calculated, and then based on the standard model library T1006, the distance between the networks T1009 is calculated, and the sample is matched to a specific standard template, such as A.

[0091] In T1010, according to the multi-omics multi-level interaction network characteristics of the sample, based on the health factor of the main factor and the matched limited domain condition value, an explanatory report is provided for the intestinal health of the infant, and milk powder formula A is selected according to the specific individual health state of the infant.

[0092] Compared with the prior art, the application has the advantages and positive effects:

[0093] 1. The innovation of the present scheme is to analyze the metabolite causes related to specific factors by using AI neural network, to sample, experiment and analyze the metabolites of the target population, to combine the external living environment factors such as dietary structure and living habits, to deeply mine the health status and possible risk factors of the target population, and to customize the milk powder formula in a targeted manner, which is closer to the real nutritional needs of individuals.

[0094] 2. The innovation of the present scheme is to deeply discover intestinal strains and genera by analyzing the metabolites of specific populations, and to determine the formula selection by comprehensively evaluating the health status, which is more targeted and personalized in improving health compared with the traditional positive selection method of milk powder formula for improving immunity or high absorption of specific nutrients based on the milk powder formula obtained by designing animal experiments.

[0095] The invention points of the present application are as follows:

[0096] 1. The innovation of the present scheme is the explainability of AI neural network, which collects metabolic samples for experiments for specific population cohorts, obtains multi-omics data such as corresponding flora, genes and metabolites, trains by using AI neural network, obtains specific factor interaction relationship network diagram, thereby knows the key processes related to specific factors, and further analyzes the interaction relationship between environmental factors and metabolites by using secondary sub-network, and finally obtains the explainability of metabolites.

[0097] 2. The innovation of the present scheme is a multi-level specific factor interaction network, which can clearly find the associated metabolites and intestinal strains and genera thereof by using the interaction network diagram for specific factors, extracts the main key factors for classification, comprehensively evaluates the intestinal health status and cause process of these categories, and customizes the milk powder formula in a targeted manner, and uses the nutritional substances in the milk powder formula to effectively intervene in the health of specific populations.

[0098] The above-mentioned embodiment numbers of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0099] In the above-mentioned embodiments of the present application, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0100] In several embodiments provided in the present application, it should be understood that the disclosed technology can be implemented by other ways. Among them, the system embodiments described above are only illustrative, for example, the division of units can be a logical function division, and actual implementation can have another division mode, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or modules shown or discussed can be indirect coupling or communication connection between the units or modules through some interfaces, and can be electrical or other forms.

[0101] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed to multiple units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0102] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0103] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The foregoing storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0104] The above is only the preferred embodiment of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.

Claims

1. A method for selecting formula milk powder based on the interpretability of neural networks to discover the relationships between microbial interactions and metabolites, characterized in that, Includes the following steps: Establish a comprehensive database for specific population queuing; A global neural network was constructed and trained using multi-omics data from a specific population cohort in a comprehensive database. The characteristic relationships between microbial interactions and their metabolites were obtained, and key factors were extracted from them. The key factors and lifestyle scales were input into the environmental factor analysis network in the second-level subnetwork to analyze the interaction between environmental factors and metabolites, and to obtain the explanatory association between specific factors and metabolites, and between bacterial strains and genera. Based on network weights, key factors are extracted from multi-level interaction networks, and strains and genera associated with several specific factors are identified. Specific factors, metabolites, and major bacterial strains are extracted and input into a secondary subnetwork for classification. The specific population is divided into several categories, and the intestinal health status of the specific population in each category is obtained. Based on the associated intestinal bacterial strains, targeted milk powder formula design and selection are carried out.

2. The formula milk powder selection method based on the interpretability of neural networks to discover the relationships between microbial interactions and metabolites according to claim 1, characterized in that, In the comprehensive database that establishes specific population groups, user samples are collected from populations that meet the criteria, and basic information is gathered. By performing metagenomic sequencing and mass spectrometry analysis on the samples, corresponding multi-omics data on bacterial species, genera, genes, and metabolism are obtained.

3. The formula milk powder selection method based on neural network interpretability for discovering microbial interactions and metabolite relationships, as described in claim 1, is characterized in that... Key data selected from the association-specific population cohort were compared using Venn diagram intersection to confirm the effectiveness of the association screening results.

4. The formula milk powder selection method based on neural network interpretability for discovering microbial interactions and metabolite relationships, as described in claim 1, is characterized in that... Global neural network design is used to train on omics data related to specific populations, including metagenomic data and metabolic data, to obtain their corresponding interaction network probability relationships; Global neural networks include, but are not limited to, commonly used algorithm components such as LSTM, RNN, MLP, CNN, RF, and XG.

5. The formula milk powder selection method based on neural network interpretability for discovering microbial interactions and metabolite relationships, as described in claim 1, is characterized in that... The environmental analysis network of the second-level sub-network is used to analyze the relationship between environmental factors and metabolites, thereby obtaining the multi-level interaction network relationship of environmental factors, and linking environmental factors to metabolites, and further linking them to the corresponding bacterial strains and genera.

6. The formula milk powder selection method based on the interpretability of neural networks to discover the relationships between microbial interactions and metabolites according to claim 1, characterized in that, The classification network of the second-level sub-network automatically classifies the main factors into several types based on the extracted main factors, including user information, metabolites, and bacterial strains and genera. It analyzes the data characteristics of specific bacterial strains and genera in each type to guide the selection of nutrients in milk powder formula and to regulate intestinal health in a targeted manner. The classification calculation methods include, but are not limited to, manual grouping, automatic classification algorithms, and automatic clustering algorithms.

7. The formula milk powder selection method based on neural network interpretability for discovering microbial interactions and metabolite relationships, as described in claim 1, is characterized in that... The comprehensive database is divided into four layers. The first layer contains user information, lifestyle habits and environment, and dietary habits. The second layer contains metagenomic data, which is obtained by collecting user fecal samples and sequencing experiments to obtain data on gut-related microbiome genes, microbiome abundance, and genus abundance, reflecting the user's gut health status. The third layer contains metabolomics data, which is further obtained through mass spectrometry analysis to obtain metabolic data on metabolite composition and gene expression levels. The fourth layer contains multilayer interaction network data, which uses specific factors to obtain the causal processes of related metabolites, bacterial strains, and genera from the interaction network diagram.

8. The formula milk powder selection method based on the interpretability of neural networks to discover the relationships between microbial interactions and metabolites according to claim 1, characterized in that, The neural network framework is a logical container that includes a global neural network and secondary sub-networks. It receives cohort topic data, metagenomic data, and metabolic data. The cohort topic data and metagenomic data use themselves as inputs and each other as outputs to calculate their interaction relationships, thereby mining multi-level interaction relationships specific to a particular population.

9. The formula milk powder selection method based on neural network interpretability for discovering microbial interactions and metabolite relationships, as described in claim 8, is characterized in that... The global neural network uses a nonlinear neural network algorithm to fit the interaction relationship between metabolites and gut microbiota. It is trained bidirectionally on metabolites and gut microbiota. The forward direction uses metabolites as input to predict gut microbiota, and the reverse direction uses gut microbiota as input to predict metabolites. After multiple calculations, the weights are extracted to obtain the interaction network relationship between metabolites and gut microbiota.

10. The formula milk powder selection method based on the interpretability of neural networks to discover the relationships between microbial interactions and metabolites according to claim 1, characterized in that, Based on network relationship weights, the entire process of action of any metabolite or environmental factor, as well as its degree of action and expression level, can be accurately explained. By drawing multi-level interaction network diagrams and identifying key paths, the causal relationships of related factors can be explained.

Citation Information

Patent Citations

  • Systems and methods for treating a dysbiosis using fecal-derived bacterial populations

    CA2995786A1

  • Metabolic flux prediction and health index evaluation system based on intestinal flora species spectrum

    CN117766144A