Natural food source efficacy formula screening method based on deep learning and NLP
By constructing a dynamic component and efficacy knowledge graph, combining graph neural network and Bayesian inference, the problem of difficult to capture the dynamic relationship between natural food ingredients and efficacy is solved, and the accuracy and personalized improvement of the group screening is achieved.
Patent Information
- Application Number
- CN202510031620.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
AI Technical Summary
The existing natural food source efficacy prescription screening methods are difficult to accurately reflect the dynamic changes and dose effects of natural food ingredients, and cannot accurately capture the dynamic relationship between food ingredients and efficacy over time and dose.
By constructing a method based on deep learning and NLP, a dynamic component knowledge graph and a efficacy knowledge graph of time dimension are introduced, and a graph neural network and Bayesian inferred formula are used, and a comprehensive score of the grouping group is calculated to obtain the best grouping plan.
The accurate analysis of the dynamic relationship between natural food ingredients and efficacy over time and dose is achieved, and the accuracy and personalization of the prescription screening is improved, which is in line with the timeliness requirements in actual use.
Smart Images

Figure CN119943283A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent screening of prescriptions, and in particular to a method for screening efficacy prescriptions derived from natural food sources based on deep learning and NLP. Background Art
[0002] The natural food-derived efficacy formula screening method based on deep learning and NLP aims to optimize the matching relationship between natural food ingredients and efficacy and improve the personalized accuracy of natural food formulas. By constructing a dynamic time-dimensional knowledge graph and combining it with a graph neural network to model the ingredient relationship, it controls the impact of the dosage effect on efficacy, and introduces Bayesian causal inference to model the causal relationship between ingredients and efficacy, thereby achieving accurate analysis of efficacy and personalized formula recommendations based on different time periods and different dosages.
[0003] Existing methods for screening efficacy formulas derived from natural foods are usually unable to accurately reflect the dynamic changes and dosage effects of natural food ingredients. In addition, since traditional knowledge graph models lack the time dimension combined with the analysis of efficacy dosage and the mutual influence between foods, it will lead to the problem of being unable to accurately capture the dynamic relationship between food ingredients and efficacy over time and dosage. Therefore, a method for screening efficacy formulas derived from natural foods based on deep learning and NLP is provided. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for screening efficacy formulas from natural food sources based on deep learning and NLP, so as to solve the problem raised in the above background technology that the traditional knowledge graph model lacks the time dimension combined with the analysis of efficacy dosage and the mutual influence between foods, which leads to the inability to accurately capture the dynamic relationship between food ingredients and efficacy that changes with time and dosage.
[0005] To achieve the above objectives, the present invention provides a method for screening effective formulas derived from natural foods based on deep learning and NLP, comprising the following steps:
[0006] S1. Collect ingredient and efficacy data of natural foods from ingredient databases, scientific research literature, and clinical data, and use NLP technology to clean and standardize the ingredient and efficacy data;
[0007] S2. Based on the composition data of natural foods, we introduce the time dimension to build a dynamic composition knowledge graph. We use graph neural network technology to convert the composition knowledge graph into a vector representation and build a natural food composition analysis model.
[0008] S3. Based on the efficacy data of natural foods, we introduce the time dimension to build an efficacy knowledge graph. Based on the ingredient data and efficacy data and introducing the dose effect, we use the Bayesian inference formula to calculate the probability of the causal influence of ingredient data on efficacy data.
[0009] S4. Based on the natural food ingredient analysis model, efficacy knowledge graph and causal influence probability formula, input the prescription constraint requirements to calculate the comprehensive prescription score and obtain the optimal prescription plan.
[0010] As a further improvement of this technical solution, in S1, the ingredient data and efficacy data of natural foods are collected from ingredient databases, scientific research literature, and clinical data, and the ingredient data and efficacy data are cleaned and standardized using NLP technology. The specific method steps are as follows:
[0011] S1.1. Collect the ingredient data of natural food C = {c1, c2…, c i ,…,c n} and efficacy data E={e1,e2…,e j ,…,e m};
[0012] Among them, c i is the i-th component, i=1,2,…,n; e j is the jth ingredient, j = 1, 2, ..., m; n is the total number of ingredient data; m is the total number of efficacy data;
[0013] S1.2. Use NLP technology to clean and standardize ingredient data and efficacy data.
[0014] As a further improvement of this technical solution, in S2, a dynamic ingredient knowledge graph is constructed based on the composition data of natural foods by introducing the time dimension. The ingredient knowledge graph is converted into a vector representation using graph neural network technology to construct a natural food composition analysis model. The specific method steps are as follows:
[0015] S2.1. Construct a dynamic ingredient knowledge graph based on natural food ingredient data and introduce the time dimension;
[0016] S2.2. Input the dynamic component knowledge graph into the graph neural network and calculate the vector representation of the dynamic component knowledge graph through graph convolution operation;
[0017] S2.3. Based on the vector representation of the dynamic ingredient knowledge graph, a natural food ingredient analysis model is constructed using a multi-layer fully connected neural network model.
[0018] As a further improvement of this technical solution, in S2.1, a dynamic ingredient knowledge graph is constructed based on the ingredient data of natural foods and the time dimension is introduced. The specific method steps are as follows:
[0019] S2.1.1. Set timestamp T = {t1, t2…, t i,…,t n} fused to component data C = {c1, c2…, c i ,…,c n}:
[0020] C time ={(c1,t1),(c2,t2),…,(c n ,t n )};
[0021] Among them, the i-th component c i Corresponding to the i-th timestamp t i ; is the component data with time dimension;
[0022] S2.1.2. Define an undirected graph:
[0023] G C =(V C ,E C )=({c1,c2,…,c n},{(c i ,c i ′),(c i ″,c i ″′),…});
[0024] Among them, V C ={c1,c2,…,c n} is a set of component nodes; E C ={(c i ,c i ′),(c i ″,c i ″′),…} is the set of relationship edges between components; (c i ,c i ′) represents the i-th component c i and the i′th component c i ′; (c i ″,c i ″′) represents the i″th component c i ″ and the i′′th component c i The relationship edge between ″′;
[0025] S2.1.3. Introduce the time dimension into the undirected graph and use time as the attribute of the relationship edge between components to obtain a dynamic component knowledge graph:
[0026] Dynamic component knowledge graph:
[0027]
[0028] in, is the dynamic component knowledge graph; t ii ′ is the i-th component c i and the i′th component c i The timestamp of the relationship edge between ′.
[0029] As a further improvement to this technical solution, in S2.2, the dynamic component knowledge graph is input into the graph neural network, and the vector representation of the dynamic component knowledge graph is calculated through graph convolution operation. The specific steps are as follows:
[0030] S2.2.1. Calculate the i-th component c through graph convolution operation i The feature representation of:
[0031]
[0032] in, is the i-th component c in the γth layer i The feature representation of ; σ is the ReLU nonlinear activation function; is the i-th component c i The set of neighbor nodes of is the i′th component c in the γ-1 layer i ′’s feature representation; d i is the i-th component c i degree; d i ′ is the i′th component c i ′ degree; W (γ) is the weight matrix of the γth layer; b (γ) is the bias term of the γth layer; αt ii ′ is the i-th component c i and the i′th component c i The weight coefficient of the timestamp of the relationship edge between ′;
[0033] S2.2.2, the i-th component c of the γ-th layer i Feature representation Get the i-th component c from the output of the graph neural network i The vector representation of :
[0034]
[0035] V=[v1,v2,…,v i ,…,v n ];
[0036] Among them, v i is the i-th component c i V is the vector representation set corresponding to the component data C.
[0037] As a further improvement of this technical solution, in S2.3, a natural food component analysis model is constructed using a multi-layer fully connected neural network model based on the vector representation of the dynamic component knowledge graph. The specific method steps are as follows:
[0038] S2.3.1. Construct a natural food composition analysis model based on a multi-layer fully connected neural network model:
[0039]
[0040] in, is the i-th component c i The vector representation v i The analysis results;
[0041] S2.3.2, according to the composition data C={c1,c2…,c i ,…,c n}Training a natural food composition analysis model.
[0042] As a further improvement of this technical solution, in S3, the time dimension is introduced based on the efficacy data of natural foods to construct an efficacy knowledge graph, and the dosage effect is introduced based on the ingredient data and efficacy data, and the Bayesian inference formula is used to calculate the probability of the causal influence of the ingredient data on the efficacy data. The specific method steps are as follows:
[0043] S3.1. Introducing the time dimension into the efficacy data of natural foods to build an efficacy knowledge graph:
[0044]
[0045] in, is the efficacy knowledge graph; t jj ′ is the jth efficacy e j and the j′th efficacy e j The timestamp of the relationship edge between ′;
[0046] S3.2. Based on the ingredient data and efficacy data and introducing the dose effect, use the Bayesian inference formula to calculate the probability of the causal effect of the ingredient data on the efficacy data.
[0047] As a further improvement of this technical solution, in S3.2, based on the ingredient data and efficacy data and introducing the dose effect, the Bayesian inference formula is used to calculate the probability of the causal effect of the ingredient data on the efficacy data. The specific method steps are as follows:
[0048] S3.2.1. The Bayesian inference formula based on ingredient data and efficacy data is as follows:
[0049]
[0050] Among them, P(e j ∣c i ) is a given i-th component c i When the jth efficacy e j The conditional probability of P(c i ∣e j ) is the given j-th efficacy e j When the i-th component c i The conditional probability of P(e j ) is the jth efficacy e j Prior probability; P(c i ) is the i-th component c i The prior probability of
[0051] S3.2.2. Introduce the dose effect and reconstruct the Bayesian dose inference formula:
[0052] Bayesian inference formula for dose:
[0053]
[0054] Among them, P(e j ∣c i ,d i ) is a given i-th component c i and dose d i When the jth efficacy e j The conditional probability of P(c i ,d i ∣e j ) is the given j-th efficacy e j When the i-th component c i and dose d i The conditional probability of P(c i ,d i ) is the i-th component c i and dose d i The joint probability of
[0055] Dose response function inference formula:
[0056] P(c i ,d i ∣e j )=P(c i ∣e j )·f(d i );
[0057] Among them, f(d i ) is the dose response function;
[0058] The dose effects are as follows:
[0059] The i-th component ci The dose is d i , and the dose d i For the jth efficacy e j The effect of the play has an impact, we can calculate the conditional probability P(c i ,d i ∣e j ) to consider the effect of dose on causality, and construct the dose-response function f(d i );
[0060] S3.2.3. Based on the Bayesian inference formula for dose and the dose-response function inference formula, the causal effect probability formula is obtained:
[0061] Causal influence probability formula:
[0062]
[0063] Among them, P(e j ∣c i ,d i ) is a given i-th component c i and dose d i When the jth efficacy e j The conditional probability of .
[0064] As a further improvement of this technical solution, in S4, based on the natural food component analysis model, efficacy knowledge graph and causal influence probability formula, the prescription constraint requirements are input to calculate the prescription comprehensive score and obtain the optimal prescription solution. The specific method is as follows:
[0065] S4.1. Input the formula constraints and calculate the comprehensive score for each formula using the natural food ingredient analysis model, efficacy knowledge graph, and causal influence probability formula;
[0066] S4.2. Find the highest comprehensive formula score. The formula corresponding to the highest comprehensive formula score is the best formula solution.
[0067] As a further improvement of this technical solution, in S4.1, the formula constraint conditions are input, and the comprehensive score of each formula is calculated using the natural food component analysis model, efficacy knowledge graph, and causal influence probability formula. The specific method steps are as follows:
[0068] S4.1.1. Set the formula constraints:
[0069] Efficacy constraints:
[0070] E target ={e j1 ,e j2 ,…,e jp};
[0071] Among them, e jp is the pth effect that is expected to be achieved; E tar get Efficacy constraint set;
[0072] Composition constraints:
[0073]
[0074] in, is the kth formula The number of components; ε min is the kth formula The minimum number of components; ε max is the kth formula The maximum number of ingredients;
[0075] Dose range setting:
[0076]
[0077] Among them, d i,min is the i-th component c i Minimum dose; d i,max is the i-th component c i The maximum dose limit;
[0078] S4.1.2. Calculate the comprehensive score for each formula using the natural food ingredient analysis model, efficacy knowledge graph, and causal influence probability formula:
[0079] From the efficacy knowledge graph In the experiment, extract the feature vector set U=[u1,u2…,u j ,…,u m ],u j is the jth efficacy e j The eigenvector of
[0080] Calculate the comprehensive score for each formula:
[0081]
[0082] Among them, S k is the kth formula The comprehensive score of the formula.
[0083] Compared with the prior art, the present invention has the following beneficial effects:
[0084] 1. In this natural food-derived efficacy formula screening method based on deep learning and NLP, a dynamic ingredient knowledge graph is constructed based on the time dimension to capture the time-varying relationship between food ingredients and efficacy. Combined with the dosage effect of food ingredients, it accurately analyzes the efficacy changes of different ingredient combinations under different time and dosage conditions, overcoming the limitations of traditional static knowledge graphs and making food formula screening more in line with the timeliness requirements in actual use.
[0085] 2. In this method for screening effective formulas derived from natural food sources based on deep learning and NLP, deep learning training is performed on the vector representation of the dynamic ingredient knowledge graph through graph neural networks to model the complex relationships between natural food ingredients. The functions of the ingredients are accurately analyzed through multi-layer fully connected neural networks to identify the synergistic and antagonistic effects between ingredients, further improving the accuracy and personalization of formula screening. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] Figure 1 The figure is a flow chart of the overall method of the present invention. DETAILED DESCRIPTION
[0087] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0088] Example:
[0089] See also Figure 1 As shown, this embodiment provides a method for screening natural food-derived efficacy formulas based on deep learning and NLP, comprising the following steps:
[0090] S1. Collect ingredient and efficacy data of natural foods from ingredient databases, scientific research literature, and clinical data, and use NLP technology to clean and standardize the ingredient and efficacy data;
[0091] In this embodiment S1, the composition data and efficacy data of natural foods are collected from the composition database, scientific research literature, and clinical data, and the composition data and efficacy data are cleaned and standardized using NLP technology. The specific method steps are as follows:
[0092] S1.1. Collect the ingredient data of natural food C = {c1, c2…, c i ,…,c n} and efficacy data E={e1,e2…,ej ,…,e m};
[0093] Among them, c i is the i-th component, i=1,2,…,n; e j is the jth ingredient, j = 1, 2, ..., m; n is the total number of ingredient data; m is the total number of efficacy data;
[0094] S1.2. Use NLP technology to clean and standardize ingredient data and efficacy data.
[0095] In this embodiment, S1.2, NLP technology is used to clean and standardize the ingredient data and efficacy data. The specific method is as follows:
[0096] S1.2.1. Denoising ingredient and efficacy data:
[0097] c i ′=RemoveNoise(c i )
[0098] e j ′=RemoveNoise(e j );
[0099] Among them, c i ′ is the denoising i-th component; e j ′ is the jth component of denoising;
[0100] S1.2.2. Perform word segmentation on the denoised component data and efficacy data:
[0101] Tokenized(c i ′)={to1,to2,…,to p};
[0102] Tokenized(e j ′)={to′1,to′2,…,to′ q};
[0103] Among them, to p is a component vocabulary unit; to′ q It is a functional vocabulary unit;
[0104] The result after word segmentation is:
[0105] POS(c i ′)={(to1,POS1),(to2,POS2),…,(to p ,POS p )};
[0106] POS(e j ′)={(to′1,POS1′),(to′2,POS2′),…,(to′ q ,POS′ q )};
[0107] Among them, POS p POS is the part-of-speech tag of the component vocabulary unit; POS′ q Part of speech tag for functional vocabulary unit; POS(c i ′) is the word segmentation result of the component data after denoising; POS(e j ') is the word segmentation result of efficacy data after denoising;
[0108] S1.2.3. Standardize the component data and efficacy data after denoising and word segmentation:
[0109]
[0110] Among them, c i ″ is the standard composition data; e j ″ is the standard efficacy data;
[0111] Then the standard ingredient data and standard efficacy data are unified as follows: ingredient data C = {c1, c2…, c i ,…,c n} and efficacy data E={e1,e2…,e j ,…,e m}.
[0112] S2. Based on the composition data of natural foods, we introduce the time dimension to build a dynamic composition knowledge graph. We use graph neural network technology to convert the composition knowledge graph into a vector representation and build a natural food composition analysis model.
[0113] In this embodiment S2, a dynamic ingredient knowledge graph is constructed based on the composition data of natural foods by introducing the time dimension. The ingredient knowledge graph is converted into a vector representation using graph neural network technology to construct a natural food composition analysis model. The specific steps are as follows:
[0114] S2.1. Construct a dynamic ingredient knowledge graph based on natural food ingredient data and introduce the time dimension;
[0115] S2.2. Input the dynamic component knowledge graph into the graph neural network and calculate the vector representation of the dynamic component knowledge graph through graph convolution operation;
[0116] S2.3. Based on the vector representation of the dynamic ingredient knowledge graph, a natural food ingredient analysis model is constructed using a multi-layer fully connected neural network model.
[0117] In this embodiment S2.1, a dynamic ingredient knowledge graph is constructed based on the ingredient data of natural foods and by introducing the time dimension. The specific steps are as follows:
[0118] S2.1.1. Set timestamp T = {t1, t2…, t i ,…,t n} fused to component data C = {c1, c2…, c i ,…,c n}:
[0119] C time ={(c1,t1),(c2,t2),…,(c n ,t n )};
[0120] Among them, the i-th component c i Corresponding to the i-th timestamp t i ; is the component data with time dimension;
[0121] S2.1.2. Define an undirected graph:
[0122] G C =(V C ,E C )=({c1,c2,…,c n},{(c i ,c i ′),(c i ″,c i ″′),…});
[0123] Among them, V C ={c1,c2,…,c n} is a set of component nodes; E C ={(c i ,c i ′),(c i ″,c i ″′),…} is the set of relationship edges between components; (c i ,c i ′) represents the i-th component c i and the i′th component c i ′; (c i ″,c i ″′) represents the i″th component c i ″ and the i′′th component c i The relationship edge between ″′;
[0124] S2.1.3. Introduce the time dimension into the undirected graph and use time as the attribute of the relationship edge between components to obtain a dynamic component knowledge graph:
[0125] Dynamic component knowledge graph:
[0126]
[0127] in, is the dynamic component knowledge graph; t ii ′ is the i-th component c i and the i′th component c i The timestamp of the relationship edge between ′.
[0128] In this embodiment S2.2, the dynamic component knowledge graph is input into the graph neural network, and the vector representation of the dynamic component knowledge graph is calculated through graph convolution operation. The specific method steps are as follows:
[0129] S2.2.1. Calculate the i-th component c through graph convolution operation i The feature representation of:
[0130]
[0131] in, is the i-th component c in the γth layer i The feature representation of ; σ is the ReLU nonlinear activation function; is the i-th component c i The set of neighbor nodes of is the i′th component c in the γ-1 layer i ′’s feature representation; d i is the i-th component c i degree; d i ′ is the i′th component c i ′ degree; W (γ) is the weight matrix of the γth layer; b (γ) is the bias term of the γth layer; αt ii ′ is the i-th component c i and the i′th component c i The weight coefficient of the timestamp of the relationship edge between ′;
[0132] S2.2.2, the i-th component c of the γ-th layer i Feature representation Get the i-th component c from the output of the graph neural network i The vector representation of :
[0133]
[0134] V=[v1,v2,…,v i ,…,v n ];
[0135] Among them, vi is the i-th component c i V is the vector representation set corresponding to the component data C.
[0136] In this embodiment S2.3, based on the vector representation of the dynamic ingredient knowledge graph, a multi-layer fully connected neural network model is used to construct a natural food ingredient analysis model. The specific steps are as follows:
[0137] S2.3.1. Construct a natural food composition analysis model based on a multi-layer fully connected neural network model:
[0138]
[0139] in, is the i-th component c i The vector representation v i The analysis results;
[0140] S2.3.2, according to the composition data C={c1,c2…,c i ,…,c n}Training a natural food composition analysis model.
[0141] In this embodiment S2.3.2, according to the composition data C = {c1, c2 ..., c i ,…,c n}Train the natural food composition analysis model as follows:
[0142] based on is the i-th component c i The vector representation v i The analysis results and the actual i-th component c i , construct the component loss function:
[0143]
[0144] Among them, L is the loss function;
[0145] The weight matrix W of the γth layer (γ) and the bias term b of the γth layer (γ) The set is represented as parameter set θ;
[0146] Use the gradient descent algorithm to minimize the component loss function L to optimize the parameter set θ:
[0147]
[0148] Among them, θ (γ′+1) is the parameter set for the γ′+1th iteration; θ (γ′) is the parameter set for the γ′th round of iteration; η is the learning rate; is the gradient of the loss function L with respect to the parameter set θ.
[0149] S3. Based on the efficacy data of natural foods, we introduce the time dimension to build an efficacy knowledge graph. Based on the ingredient data and efficacy data and introducing the dose effect, we use the Bayesian inference formula to calculate the probability of the causal influence of ingredient data on efficacy data.
[0150] In this embodiment S3, the time dimension is introduced into the efficacy data of natural foods to construct an efficacy knowledge graph. Based on the ingredient data and efficacy data and the dosage effect is introduced, the Bayesian inference formula is used to calculate the probability of the causal influence of the ingredient data on the efficacy data. The specific method steps are as follows:
[0151] S3.1. Introducing the time dimension into the efficacy data of natural foods to build an efficacy knowledge graph:
[0152]
[0153] in, is the efficacy knowledge graph; t jj ′ is the jth efficacy e j and the j′th efficacy e j The timestamp of the relationship edge between ′;
[0154] S3.2. Based on the ingredient data and efficacy data and introducing the dose effect, use the Bayesian inference formula to calculate the probability of the causal effect of the ingredient data on the efficacy data.
[0155] In this embodiment S3.2, based on the ingredient data and efficacy data and introducing the dose effect, the Bayesian inference formula is used to calculate the probability of the causal effect of the ingredient data on the efficacy data. The specific method steps are as follows:
[0156] S3.2.1. The Bayesian inference formula based on ingredient data and efficacy data is as follows:
[0157]
[0158] Among them, P(e j ∣c i ) is a given i-th component c i When the jth efficacy e j The conditional probability of P(c i ∣e j ) is the given j-th efficacy e j When the i-th component c i The conditional probability of P(e j ) is the jth efficacy e j Prior probability; P(c i ) is the i-th component c i The prior probability of
[0159] S3.2.2. Introduce the dose effect and reconstruct the Bayesian dose inference formula:
[0160] Bayesian inference formula for dose:
[0161]
[0162] Among them, P(e j ∣c i ,d i ) is a given i-th component c i and dose d i When the jth efficacy e j The conditional probability of P(c i ,d i ∣e j ) is the given j-th efficacy e j When the i-th component c i and dose d i The conditional probability of P(c i ,d i ) is the i-th component c i and dose d i The joint probability of
[0163] Dose response function inference formula:
[0164] P(c i ,d i ∣e j )=P(c i ∣e j )·f(d i );
[0165] Among them, f(d i ) is the dose response function;
[0166] The dose effects are as follows:
[0167] The i-th component c i The dose is d i , and the dose d i For the jth efficacy e j The effect of the play has an impact, we can calculate the conditional probability P(c i ,d i ∣e j ) to consider the effect of dose on causality, and construct the dose-response function f(d i );
[0168] In this embodiment, the dose response function f(d i) is a comprehensive function constructed by experts based on the ingredient data and efficacy data of natural foods.
[0169] S3.2.3. Based on the Bayesian inference formula for dose and the dose-response function inference formula, the causal effect probability formula is obtained:
[0170] Causal influence probability formula:
[0171]
[0172] Among them, P(e j ∣c i ,d i ) is a given i-th component c i and dose d i When the jth efficacy e j The conditional probability of .
[0173] S4. Based on the natural food ingredient analysis model, efficacy knowledge graph, and causal influence probability formula, input the formula constraint requirements to calculate the formula comprehensive score and obtain the optimal formula solution;
[0174] In this embodiment S4, based on the natural food component analysis model, efficacy knowledge graph and causal influence probability formula, the prescription constraint requirements are input to calculate the prescription comprehensive score and obtain the optimal prescription solution. The specific method is as follows:
[0175] S4.1. Input the formula constraints and calculate the comprehensive score for each formula using the natural food ingredient analysis model, efficacy knowledge graph, and causal influence probability formula;
[0176] S4.2. Find the highest comprehensive formula score. The formula corresponding to the highest comprehensive formula score is the best formula solution.
[0177] In this embodiment S4.1, the formula constraint conditions are input, and the comprehensive score of each formula is calculated using the natural food component analysis model, the efficacy knowledge map, and the causal influence probability formula. The specific method steps are as follows:
[0178] S4.1.1. Set the formula constraints:
[0179] Efficacy constraints:
[0180] E target ={e j1 ,e j2 ,…,e jp};
[0181] Among them, e jp is the pth effect that is expected to be achieved; E tar get Efficacy constraint set;
[0182] Composition constraints:
[0183]
[0184] in, is the kth formula The number of components; ε min is the kth formula The minimum number of components; ε max is the kth formula The maximum number of ingredients;
[0185] Dose range setting:
[0186]
[0187] Among them, d i,min is the i-th component c i Minimum dose; d i,max is the i-th component c i The maximum dose limit;
[0188] S4.1.2. Calculate the comprehensive score for each formula using the natural food ingredient analysis model, efficacy knowledge graph, and causal influence probability formula:
[0189] From the efficacy knowledge graph In the experiment, extract the feature vector set U=[u1,u2…,u j ,…,u m ],u j is the jth efficacy e j The eigenvector of
[0190] Calculate the comprehensive score for each formula:
[0191]
[0192] Among them, S k is the kth formula The comprehensive score of the formula.
[0193] The basic principles, main features, and advantages of the present invention are shown and described above. It should be understood by those skilled in the art that the present invention is not limited to the above-described embodiments. The above-described embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention, and such changes and modifications fall within the scope of the invention claimed.
Claims
1. A method for screening effective prescriptions from natural food sources based on deep learning and NLP, characterized in that: The following steps are involved: S1. Collect the ingredient data and efficacy data of natural foods from ingredient databases, scientific research literature, and clinical data, and use NLP technology to clean and standardize the ingredient data and efficacy data; S2. Based on the composition data of natural foods, the time dimension is introduced to build a dynamic composition knowledge graph, and the composition knowledge graph is converted into a vector representation using graph neural network technology to build a natural food composition analysis model; S3. Based on the efficacy data of natural foods, the time dimension is introduced to construct an efficacy knowledge graph. Based on the ingredient data and efficacy data, the dosage effect is introduced, and the Bayesian inference formula is used to calculate the probability of the causal influence of ingredient data on efficacy data. S4. Based on the natural food ingredient analysis model, efficacy knowledge graph and causal influence probability formula, input the formulation constraints to calculate the comprehensive formula score and obtain the best formulation plan.
2. The method for screening effective prescriptions from natural food sources based on deep learning and NLP according to claim 1, characterized in that: In S1, the ingredient data and efficacy data of natural foods are collected from ingredient databases, scientific research literature, and clinical data, and the ingredient data and efficacy data are cleaned and standardized using NLP technology. The specific method steps are as follows: S1.
1. Collect the ingredient data of natural foods C = {c1, c2…, c i ,…,c n } and efficacy data E={e1,e2…,e j ,…,e m }; Among them, c i is the i-th component, i=1,2,…,n; e j is the jth ingredient, j = 1, 2, ..., m; n is the total number of ingredient data; m is the total number of efficacy data; S1.
2. Use NLP technology to clean and standardize ingredient data and efficacy data.
3. The method for screening natural food-derived efficacy prescriptions based on deep learning and NLP according to claim 2, characterized in that: In S2, a dynamic ingredient knowledge graph is constructed based on the composition data of natural food by introducing the time dimension, and the ingredient knowledge graph is converted into a vector representation using graph neural network technology to construct a natural food composition analysis model. The specific method steps are as follows: S2.
1. Based on the ingredient data of natural foods and introducing the time dimension, a dynamic ingredient knowledge graph is constructed; S2.
2. Input the dynamic component knowledge graph into the graph neural network and calculate the vector representation of the dynamic component knowledge graph through graph convolution operation; S2.
3. Based on the vector representation of the dynamic ingredient knowledge graph, a natural food ingredient analysis model is constructed using a multi-layer fully connected neural network model.
4. The method for screening natural food-derived efficacy prescriptions based on deep learning and NLP according to claim 3, characterized in that: In S2.1, a dynamic ingredient knowledge graph is constructed based on the ingredient data of natural food and the time dimension is introduced. The specific method steps are as follows: S2.1.
1. Set the timestamp T = {t1, t2…, t i ,…,t n } is fused into component data C = {c1, c2…, c i ,…,c n }: C time ={(c1,t1),(c2,t2),…,(c n ,t n )}; Among them, the i-th component c i Corresponding to the i-th timestamp t i ; is the component data with time dimension; S2.1.2, define an undirected graph: G C =(V C ,E C )=({c1,c2,…,c n },{(c i ,c i ′),(c i ″,c i ″′),…}); Among them, V C ={c1,c2,…,c n } is a set of component nodes; E C ={(c i ,c i ′),(c i ″,c i ″′),…} is the set of relationship edges between components; (c i ,c i ′) represents the i-th component c i and the i′th component c i ′; (c i ″,c i ″′) represents the i″th component c i " and the i"th component c i The relationship edge between ″′; S2.1.
3. Introduce the time dimension in the undirected graph and use time as the attribute of the relationship edge between components to obtain a dynamic component knowledge graph: Dynamic component knowledge graph: in, is the dynamic component knowledge graph; t ii ′ is the i-th component c i and the i′th component c i The timestamp of the relationship edge between ′.
5. The method for screening effective prescriptions from natural food sources based on deep learning and NLP according to claim 4, characterized in that: In S2.2, the dynamic component knowledge graph is input into the graph neural network, and the vector representation of the dynamic component knowledge graph is calculated through the graph convolution operation. The specific method steps are as follows: S2.2.
1. Calculate the i-th component c through graph convolution operation i The characteristic representation of: in, is the i-th component c of the γ-th layer i The feature representation of ; σ is the ReLU nonlinear activation function; is the i-th component c i The set of neighbor nodes of is the i′th component c in the γ-1th layer i ′’s characteristic representation; d i is the i-th component c i degree; d i ′ is the i′th component c i ′ degree; W (γ) is the weight matrix of the γth layer; b (γ) is the bias term of the γth layer; αt ii ′ is the i-th component c i and the i′th component c i The weight coefficient of the timestamp of the relationship edge between ′; S2.2.
2. The i-th component c of the γ-th layer i The feature representation Output the i-th component c from the graph neural network i The vector representation of is: V=[v1,v2,…,v i ,…,v n ]; Among them, v i is the i-th component c i The vector representation of ; V is the vector representation set corresponding to the component data C.
6. The method for screening natural food-derived efficacy prescriptions based on deep learning and NLP according to claim 5, characterized in that: In S2.3, based on the vector representation of the dynamic ingredient knowledge graph, a multi-layer fully connected neural network model is used to construct a natural food ingredient analysis model. The specific method steps are as follows: S2.3.
1. Construct a natural food composition analysis model based on a multi-layer fully connected neural network model: in, is the i-th component c i The vector representation v i The analysis results; S2.3.2, according to the composition data C = {c1, c2…, c i ,…,c n }Training natural food composition analysis models.
7. The method for screening natural food-derived efficacy prescriptions based on deep learning and NLP according to claim 6, characterized in that: In S3, the time dimension is introduced based on the efficacy data of natural foods to construct an efficacy knowledge graph, and the dosage effect is introduced based on the ingredient data and efficacy data, and the causal influence probability of the ingredient data on the efficacy data is calculated using the Bayesian inference formula. The specific method steps are as follows: S3.
1. Introducing the time dimension based on the efficacy data of natural foods to build an efficacy knowledge graph: in, is the efficacy knowledge graph; t jj ′ is the jth efficacy e j and the j′th efficacy e j The timestamp of the relationship edge between ′; S3.
2. Based on the ingredient data and efficacy data and introducing the dose effect, the Bayesian inference formula is used to calculate the probability of the causal effect of the ingredient data on the efficacy data.
8. The method for screening natural food-derived efficacy prescriptions based on deep learning and NLP according to claim 7, characterized in that: In S3.2, based on the ingredient data and efficacy data and introducing the dose effect, the Bayesian inference formula is used to calculate the probability of the causal effect of the ingredient data on the efficacy data. The specific method steps are as follows: S3.2.
1. The Bayesian inference formula based on ingredient data and efficacy data is as follows: Among them, P(e j |c i ) is a given i-th component c i When the jth effect e j The conditional probability of i |e j ) is the given j-th efficacy e j When the i-th component c i The conditional probability of j ) is the jth efficacy e j The prior probability of i ) is the i-th component c i The prior probability of S3.2.
2. Introduce the dose effect and reconstruct the Bayesian dose inference formula: Dose Bayesian inference formula: Among them, P(e j |c i ,d i ) is a given i-th component c i and dose d i When the jth effect e j The conditional probability of i ,d i |e j ) is the given j-th efficacy e j When the i-th component c i and dose d i The conditional probability of i ,d i ) is the i-th component c i and dose d i The joint probability of Dose response function inference formula: P(c i ,d i |e j )=P(c i |e j )·f(d i ); Among them, f(d i ) is the dose response function; The dose effects are as follows: The i-th component c i The dose is d i , and the dose d i For the jth efficacy e j The influence of the play can be calculated by the conditional probability P(c i ,d i |e j ) to consider the effect of dose on causality, and construct the dose-response function f(d i ); S3.2.
3. Based on the dose Bayesian inference formula and the dose response function inference formula, the causal effect probability formula is obtained: Causal influence probability formula: Among them, P(e j |c i ,d i ) is a given i-th component c i and dose d i When the jth effect e j The conditional probability of .
9. The method for screening effective prescriptions from natural food sources based on deep learning and NLP according to claim 8, characterized in that: In S4, based on the natural food component analysis model, the efficacy knowledge graph and the causal influence probability formula, the prescription constraint requirements are input to calculate the prescription comprehensive score and obtain the best prescription scheme. The specific method is as follows: S4.
1. Input the formula constraints and use the natural food ingredient analysis model, efficacy knowledge graph and causal influence probability formula to calculate the comprehensive score of each formula; S4.
2. Find the highest comprehensive formula score. The formula corresponding to the highest comprehensive formula score is the best formula solution.
10. The method for screening effective prescriptions from natural food sources based on deep learning and NLP according to claim 9, characterized in that: In S4.1, the formula constraint conditions are input, and the comprehensive score of each formula is calculated using the natural food component analysis model, the efficacy knowledge graph, and the causal influence probability formula. The specific method steps are as follows: S4.1.
1. Set the formula constraints: Efficacy constraints: AND target ={and j1 ,And j2 ,…,And jp }; Among them, e jp is the pth effect that is expected to be achieved; E target Efficacy constraint set; Composition constraints: in, is the kth group The number of components; ε min is the kth group The minimum number of components; ε max is the kth group The maximum number of ingredients; Dose range setting: Among them, d i,min is the i-th component c i Minimum dose; d i,max is the i-th component c i The maximum dose limit; S4.1.
2. Calculate the comprehensive score of each formula using the natural food ingredient analysis model, efficacy knowledge graph and causal influence probability formula: From the efficacy knowledge graph In the above example, we extract the feature vector set U = [u1,u2…,u j ,…,u m ],u j is the jth efficacy e j The eigenvector of Calculate the overall score for each formula: Among them, S k is the kth group The comprehensive score of the formula.