A method and system for screening active ingredients and target points of a food as a medicine food by a double strategy
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JIANGNAN UNIV
- Filing Date
- 2025-10-13
- Publication Date
- 2026-07-24
AI Technical Summary
Existing molecular docking screening processes are inefficient, have biased or missed target predictions, and lack intelligent scoring mechanisms. The various stages of functional food and food-medicine homology research and development are fragmented and lack a collaborative mechanism, making it difficult to meet the needs of rapid, accurate, and efficient screening of active ingredients and targets in food-medicine homology foods.
A chemical composition matrix of food products with medicinal and edible properties was constructed by integrating multi-source databases. A rule engine-machine learning hybrid metabolite prediction model was used to identify metabolic sites. A low-energy three-dimensional ligand conformation library was generated by combining heterogeneous graph neural network scoring. Activity and toxicity were predicted by multi-task graph neural network. The NSGA-III multi-objective evolutionary algorithm was used to optimize the formulation, thus constructing a dual-strategy screening system for active ingredients and targets of food products with medicinal and edible properties.
It improved the biological relevance and accuracy of target prediction, significantly enhanced the model's generalization ability and prediction accuracy, shortened the development cycle, reduced R&D costs, and realized the modern development and systems pharmacology research of food resources that are both food and medicine.
Smart Images

Figure CN121641269B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a dual-strategy screening method and system for active ingredients and targets in food products that are both food and medicine, belonging to the fields of food science and computational chemistry. Background Technology
[0002] With the continuous development of computational biology, food science, pharmaceutical informatics, and artificial intelligence, the potential of natural active ingredients, especially plant-derived natural products, in drug development has received increasing attention. Currently, methods for screening active ingredient formulations and therapeutically relevant targets mainly rely on molecular docking technology. This technology simulates the interaction between ligand small molecules and target proteins to predict their binding affinity and infer their potential pharmacological effects. However, existing molecular docking methods for screening active ingredients and disease targets still have the following limitations: (1) The process is complex and cumbersome, and highly dependent on manual labor: The traditional molecular docking process for predicting active ingredients and targets involves multiple steps, is complex to operate, inefficient, and heavily dependent on manual processing. (2) Low efficiency of rapid screening of active ingredients: There are many types of ingredients in food that is both food and medicine. Traditional experimental screening of active ingredients is costly, time-consuming and lacks efficient prediction methods. (3) Lack of virtual metabolic prediction: Most screening methods ignore the metabolic transformation process of active ingredients in food and medicine homology in vivo, resulting in bias or omission in target prediction; (4) Target screening lacks intelligent scoring mechanism: Traditional molecular docking results mostly rely on binding energy values for sorting, lacking multi-dimensional scoring that combines biological background knowledge or artificial intelligence algorithms; (5) Lack of integrated system: The research and development of functional foods and food-medicine homology components presents a "fragmented" process, lacking an effective collaborative mechanism, and there is a lack of effective interconnection and collaboration between tools and databases in various links.
[0003] To overcome the above problems, there is an urgent need in this field to develop a standardized automated process system that integrates small molecule active metabolites, efficient molecular docking and intelligent scoring mechanisms, activity prediction, toxicity exclusion, and formulation design, so as to improve the accuracy, automation level and biological significance interpretation capabilities of virtual screening of plant active ingredients and targets and virtual screening of active ingredients. Summary of the Invention
[0004] [Technical Issues] The technical problem to be solved by this invention is that the existing molecular docking screening process has problems such as low screening efficiency, target prediction deviation or omission, lack of intelligent scoring mechanism, and lack of coordination in the various stages of functional food and food-medicine homology research and development, which makes it difficult to meet the needs of rapid, accurate and efficient screening of active ingredients and targets in food-medicine homology foods.
[0005] [Technical Solution] To address the above problems, this invention provides a dual-strategy screening method and system for active ingredients and targets in food products that are both food and medicine. In a first aspect, the present invention provides a dual-strategy screening method for active ingredients and targets in food products that are both food and medicine, the method comprising: Step 1: Integrate the BATMAN-TCM 2.0, CheEMBL-Food subset, TCMSP, HERB, TCM.ID, and HIT2.0 databases to construct a chemical composition matrix for food and medicine homology. The chemical composition matrix for food and medicine homology includes chemical composition structure information, molecular descriptors, and stereochemical information. The chemical composition structural information includes: molecular diagram, SMILES, InChlKey; the molecular descriptor includes: molecular weight, molecular formula, functional group, ring structure, polarity, hydrophilicity and hydrophobicity; the stereochemical information includes symmetry, isomers, and planar symmetry. Step 2: Implement the target screening strategy; The target screening strategy specifically involves: developing a rule engine-machine learning hybrid metabolite prediction model; identifying reaction sites and potential metabolic sites of molecules in the chemical composition matrix of food-medicine homology based on the chemical component structure information, molecular descriptors, and stereochemical information; calculating potential primary and secondary metabolites using a reaction rule system; integrating and deduplicating to construct a high-confidence virtual metabolite database; predicting virtual metabolites based on the compound target prediction database to obtain an active target dataset; integrating potential active targets related to physiological states based on a bioinformatics database; constructing a target set by taking the intersection of the active target dataset with the target set; downloading its crystal structure to generate target proteins; performing high-throughput molecular docking between the virtual metabolite database and target proteins; integrating the physical prior knowledge between target proteins and virtual metabolites using a heterogeneous graph neural network; characterizing their interaction in isovariant geometric space; and using a training scoring model to score and rank the high-throughput molecular docking results, screening the optimal ligand-receptor complex configuration, and outputting the best target related to physiological states, i.e., the target protein. The reaction rule system includes reaction rules, reaction mechanisms, a reaction database, and a reaction prediction algorithm; the reaction rules include phase I metabolic reactions and phase II metabolic reactions; the reaction mechanisms include enzyme catalysis mechanisms; the reaction database includes MetaCyc and KEGG; and the reaction prediction algorithm is an expert system model. Step 3: Implement the functional component screening strategy; The functional component screening strategy is as follows: Based on the chemical component structure information in the chemical component matrix of food homology, a graph neural network (GNN) is used to extract chemical priors, including functional group categories, number of rotatable bonds, aromaticity characteristics, electron cloud density distribution, and stereochemical constraint information; based on the chemical priors and the target binding pocket structure in the optimal ligand-receptor complex configuration, a conditional diffusion generation model is used to generate a low-energy three-dimensional ligand conformation library; a high-throughput parallel docking engine is used to dock the low-energy three-dimensional ligand conformation library to the active pocket of the target protein, and a multi-task graph neural network (Food-GNN) is used to predict the activity and toxicological indicators of the compounds, outputting candidate molecules; Step 4: Based on candidate molecules, prohibited substances are automatically excluded through the compliance engine Rule-Food; the NSGA-III multi-objective evolutionary algorithm is used to optimize the formulation of candidate molecules based on multiple dimensions of activity, safety, cost, and sensory factors, generating Pareto's cutting-edge functional food formulation combinations.
[0006] Optionally, the method of integrating multi-source databases in step 1 includes: merging repetitive compounds from multiple source databases using InChIKey as the unique identifier; standardizing and normalizing SMILES using RDKit and OpenBabel; completing missing values using structure prediction and physicochemical parameter calculation; performing molecular descriptor calculation and matrix processing; and generating 3D structural coordinates from SMILES and SDF using RDKitETKDG and OpenBabel to form an integrated database of two-dimensional and three-dimensional structures.
[0007] Optionally, the rule engine-machine learning hybrid metabolite prediction model in the target screening strategy is based on a multi-model joint prediction strategy, combining the PharmMapper structure reverse docking method, the DeepDTA deep learning method, and the SEA similarity analysis method to improve the accuracy and reliability of target prediction. The rule engine-machine learning hybrid metabolite prediction model is trained and optimized based on metabolic reaction data from the MetaCyc, BRENDA, KEGG, HumanCyc, ECNumbers, SMARTS structure matching rules, ChEMBL, PubChem, and Cytochrome P450 databases. It is used to simulate the reaction process catalyzed by metabolic enzymes in vivo, and to predict the potential metabolites of small molecule drugs or natural products in vivo. The implementation method is as follows:
[0008]
[0009]
[0010]
[0011] In the formula, This represents the input compound; This indicates the probability that compound M is metabolized by the cytochrome P450 enzyme CPY isoform; This represents the probability score output by the trained random forest classifier when it accepts a molecule as input. This represents the probability that a specific atom, bond, and position is a metabolic target. A neural network model representing the prediction of metabolic sites; Characterized by the local chemical environment of atoms and bonds; This represents the molecular feature vector generated by the GNN; Indicates the input molecule; Represents a graph neural network; Indicates the Transformer decoder; This indicates the output metabolic product molecules.
[0012] Optionally, the target screening strategy includes the following compound target prediction databases: SEA, PharmMapper, SwissTargetPrediction, and DeepDTA; the bioinformatics databases include: Digenet, DrugBank, GeneCards, and TTD databases; and the high-throughput molecular docking is performed in parallel on a Linux system using AutodockVina and Autodock GPU.
[0013] Optionally, in the target screening strategy, the heterogeneous graph neural network is modeled based on graph structure data of different conformations of target proteins and virtual metabolites, and is trained on a training set that has undergone data augmentation and redundancy removal. The heterogeneous graph neural network integrates the physical prior knowledge between proteins and ligands to screen the optimal ligand-receptor complex conformation and output the best target related to physiological state in the following way:
[0014]
[0015]
[0016] in, Represents the raw attention score. Representing the geometric characteristics of the edges, This represents converting geometric features into scalar weight factors, and ⊙ represents the element-wise multiplication of two vectors and the element-wise multiplication of two matrices. Characterizing the structural features of the edges, This represents mapping discrete data into a continuous high-dimensional vector representation. This indicates that the structure is embedded in the projection as an offset value. This represents element-wise addition and biased addition. This represents the final attention score; The feature vector of the ligand; Indicates the first ligand The atom in the first The node feature vectors generated by the layer; This represents summing the eigenvectors of all nodes of the ligand; This indicates that the node features are weighted; Indicates the affinity score; This represents a multilayer perceptron.
[0017] Optionally, in the functional component screening strategy, the high-throughput parallel docking engine utilizes the parallel computing of Autodock vina and Autodock GPU to calculate the binding affinity between the ligand and the target protein; the target binding pocket structure is the active site on the target protein that binds to the ligand.
[0018] Optionally, the conditional diffusion generation model in the functional component screening strategy is the PocketDiffusion-MFH model; The PocketDiffusion-MFH model generates a three-dimensional ligand conformation library of active compounds in food-derived medicinal materials. Using two-dimensional molecular diagrams of the compounds as input, and based on structural priors and the target binding pocket structure in the optimal complex configuration, it generates low-energy three-dimensional ligand conformations through multiple rounds of progressive denoising. The implementation method is as follows:
[0019] in, Indicates the first The intermediate state of a step; Indicates the first The state of the step; Indicates the current time step; Indicates conditional information; This represents the distribution of the inverse process after model parameterization; Indicates model parameters; The model predicts the first The mean of the step Gaussian distribution; The model predicts the first The covariance matrix of the step Gaussian distribution; This indicates a normal distribution.
[0020] Optionally, in the functional component screening strategy, the multi-task graph neural network Food-GNN constructs a molecular graph, treating each compound's atoms as nodes and chemical bonds as edges, and iteratively propagates feature information between nodes in the network to obtain the embedded representation of atoms. The embedded representation of atoms is calculated by the following formula:
[0021] in, Represents atoms In iteration The amount of embedding; Represents atoms In iteration The embedding vector; Represents atoms The set of adjacent atoms; Representing adjacent atoms In iteration The embedding vector; Represents atoms and The characteristic vector of the chemical bond between them; Represents a message function; Indicates the update function; The resulting atomic embedding representation is input into the multi-task classification head and regression head to predict the activity and toxicity indicators of the compound, respectively. The prediction formulas are as follows:
[0022] in, This represents the input feature vector; This represents the weight matrix of the shared hidden layer; (•) represents the activation function; Indicates the first A weight matrix specific to each detection task; Indicates the first Bias terms for each detection task; (•) represents the Sigmoid function; Indicates the first The predicted probability of a detection; A cross-modal fusion approach is employed to integrate multi-source information, including chemical structure features (CS), morphological features (MO), and gene expression features (GE), ultimately leading to a predicted probability. The expression is as follows:
[0023] in, This indicates that the first [item] is predicted using only chemical structure. The probability of a detection; This represents the probability predicted using only morphological features; This represents the probability predicted using only gene expression. This represents the final predicted probability after fusion.
[0024] Optionally, step 4 further includes using the Prolog language to implement the compliance engine Rule-Food to determine the legality of candidate molecules. The compliance engine Rule-Food includes ChinaFoodDB, the CIRS regulatory database, national standard GN, FDA, and EFSA. The prohibited substances are substances banned by regulatory authorities for safety, toxicity, or other reasons. The objective function of the NSGA-III multi-objective evolutionary algorithm for formula optimization of candidate molecules is:
[0025]
[0026]
[0027]
[0028] in, This indicates the total number of candidate food and medicine homologous ingredients in the formula, i.e., the number of candidate molecules; Indicates the first The amount or proportion of each food ingredient; No. The functional activity of each component; Indicates the first The risk of toxicity or adverse reactions of each component; Indicates the first The unit price of each component; Indicates the first The usage amount of each ingredient; Indicates the first The contribution of each component to the overall sensory score; This represents the weighting factor for each component in the three corresponding functions.
[0029] Secondly, this invention provides a dual-strategy screening system for active ingredients and targets in food products that are both food and medicine, the system comprising: The data construction module is used to integrate multi-source databases and construct a chemical composition matrix; The target screening module is used to identify metabolic sites based on chemical composition matrix, construct target set and output candidate ligand-receptor combinations through high-throughput molecular docking and heterogeneous graph neural network scoring, and output the best target related to physiological state. The functional component screening module is used to extract structural priors based on chemical component information in the chemical component matrix, generate a low-energy three-dimensional ligand conformation library based on structural priors and target binding pocket structures, and predict the activity and toxicity of compounds through a multi-task graph neural network Food-GNN, and output candidate molecules. The decision optimization module integrates the compliance engine Rule-Food with the NSGA-III multi-objective evolutionary algorithm to exclude prohibited substances based on candidate molecules and optimize the generation of the optimal formulation combination.
[0030] [Beneficial Effects] (1) Step 2 of the present invention systematically considers the metabolic transformation process of food homology in vivo, simulates the multi-stage metabolites formed by food homology compounds in vivo, and constructs a virtual metabolic prediction model, thereby improving the biological relevance and accuracy of target prediction; In step 2, a heterogeneous graph neural network is introduced to realize high-dimensional interactive scoring, which can comprehensively evaluate the real binding ability between ligand and target, which is more comprehensive and intelligent than traditional docking scoring, and reduces misjudgment.
[0031] (2) Step 3 of the present invention constructs a multi-task graph neural network Food-GNN, which is based on molecular structure graph representation learning, comprehensively integrates the structural information of food components, and realizes high-precision prediction of their bioactivity, functional performance, production cost, toxicity risk and other multi-dimensional characteristics, significantly improving the generalization ability and prediction accuracy of the model.
[0032] (3) Step 4 of the present invention adopts the NSGA-III multi-objective optimization algorithm, which can simultaneously consider the synergistic effect of each component in the formula on activity enhancement and toxicity reduction, obtain the optimal balance between multiple objectives, improve the functionality, safety and market value of food products, while reducing the number of experiments, shortening the product development cycle and reducing R&D costs.
[0033] In summary, this invention integrates several key technologies, including a data matrix of chemical components from food and medicinal plants, a virtual metabolism prediction model, a low-energy three-dimensional ligand conformation library generated based on a conditional diffusion model, high-throughput molecular docking, heterogeneous graph neural network scoring, and a multi-task graph neural network (Food-GNN) for activity and toxicity prediction. This constructs a dual-strategy automated, high-throughput screening process that significantly reduces manual operation and greatly improves the accuracy and efficiency of drug candidate target screening. This invention is particularly suitable for functional mining of active ingredients from food and medicinal plants, effectively promoting the modern development and systems pharmacology research of food and medicinal resources, and providing new tools for the mechanism research and precise development of food and medicinal components. Attached Figure Description
[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0035] Figure 1 The flowchart illustrates the dual-strategy system for screening active ingredients and targets in food products that are both food and medicine provided by this invention.
[0036] Figure 2 The diagram illustrates the specific implementation steps of the target screening strategy provided by this invention.
[0037] Figure 3 This is a diagram of the virtual metabolism prediction model architecture provided by the present invention.
[0038] Figure 4 This is a schematic diagram of the process of the heterogeneous graph neural network scoring module provided by the present invention.
[0039] Figure 5 A comparison chart showing the number of compounds predicted by different metabolic prediction models provided in this invention.
[0040] Figure 6 The optimal active metabolic small molecules and their protein targets are shown in Example 2 of this invention. Detailed Implementation
[0041] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0042] Example 1: This embodiment uses ten food-medicine homologous ingredients—bitter melon, mulberry leaf, astragalus, lotus leaf, hemp seed, ganoderma lucidum, sea buckthorn, perilla, licorice, and sophora japonica—as examples to systematically demonstrate the application of the dual-strategy screening method described in this invention in the development of functional foods, specifically targeting the physiological state of type 2 diabetes. See [link / reference]. Figure 1 The specific implementation steps are as follows: Step 1: Integrate multi-source databases to construct a chemical composition matrix of food and medicine homology. This study focused on ten food-grade medicinal ingredients: bitter melon, mulberry leaf, astragalus, lotus leaf, hemp seed, ganoderma lucidum, sea buckthorn, perilla, licorice, and sophora japonica. It integrated multiple databases including BATMAN-TCM 2.0, ChemBL-Food subset, TCMSP, HERB, TCM.ID, and HIT2.0, and used Python scripts to automatically retrieve and download their active ingredient information. A total of 8,456 natural small molecules were obtained, including their chemical composition structure information, molecular descriptors, and stereochemical information. The chemical composition structure information included molecular diagrams, SMILES, and InChIKey; the molecular descriptors included molecular weight, molecular formula, functional groups, ring structure, polarity, hydrophilicity, and hydrophobicity; and the stereochemical information included symmetry, isomers, and planar symmetry. The InChIKey was used as the unique identifier for deduplication. SMILES was standardized and normalized using RDKit and Open Babel. Missing values were filled using structural prediction and physicochemical parameter calculation. SMILES and SDF were then used to generate 3D structural coordinates using RDKit ETKDG and Open Babel, ultimately forming a chemical composition matrix that integrates two-dimensional and three-dimensional structures.
[0043] Step 2: Implement the target screening strategy, see Figure 2 : (1) A rule engine-machine learning hybrid metabolite prediction model trained on databases such as MetaCyc, BRENDA, KEGG, and Cytochrome P450 is used to identify the reaction sites and potential metabolic sites of molecules in the chemical composition matrix constructed in step 1. The potential metabolites are calculated and output by combining the reaction rule system covering phase I and phase II metabolic reactions. After integration and deduplication, a high-confidence virtual metabolite database is constructed. (2) Based on compound target prediction databases such as SEA, PharmMapper, SwissTargetPrediction, and DeepDTA, the above virtual metabolites were predicted to obtain an active target dataset. (3) Based on bioinformatics databases such as DisGeNET, DrugBank, GeneCards, and TTD, integrate potential active targets related to type 2 diabetes (including AMPKα2, PPARγ, α-glucosidase, GLUT4, etc., totaling 1,247), take the intersection with the active target dataset obtained in step (2), construct a target set, and download its crystal structure to generate target proteins; (4) High-throughput molecular docking was performed between the virtual metabolite database and the target protein. This involved parallel computation using AutodockVina and Autodock GPU on a Linux system. A heterogeneous graph neural network was used to integrate prior physical knowledge between the target protein and the virtual metabolite, characterizing their interaction in an isovariant geometric space. A training scoring model was used to score and rank the docking results, selecting the optimal ligand-receptor complex configuration. Cluster analysis of the top 1% of compounds by docking energy revealed that saponins (such as momordicin from bitter melon) exhibited the best binding affinity to AMPKα2, with an average docking score <-10.5 kcal / mol.
[0044] See Figure 3 The rule-based engine-machine learning hybrid metabolite prediction model is based on a multi-model joint prediction strategy, combining the PharmMapper structure reverse docking method, DeepDTA deep learning method, and SEA similarity analysis method to improve the accuracy and reliability of target prediction. This model is trained and optimized using metabolic reaction data from the MetaCyc, BRENDA, KEGG, HumanCyc, ECNumbers, SMARTS structure matching rules, ChEMBL, PubChem, and Cytochrome P450 databases. It simulates in vivo reaction processes catalyzed by metabolic enzymes to predict potential metabolites of small molecule drugs or natural products in vivo. The implementation method is as follows:
[0045]
[0046]
[0047]
[0048] In the formula, This represents the input compound; This indicates the probability that compound M is metabolized by the cytochrome P450 enzyme CPY isoform; This represents the probability score output by the trained random forest classifier when it accepts a molecule as input. This represents the probability that a specific atom, bond, and position is a metabolic target. A neural network model representing the prediction of metabolic sites; Characterized by the local chemical environment of atoms and bonds; This represents the molecular feature vector generated by the GNN; Indicates the input molecule; Represents a graph neural network; Indicates the Transformer decoder; This indicates the output metabolic product molecules.
[0049] See Figure 4 The heterogeneous graph neural network (HBR) models target proteins and virtual metabolites based on graph structure data of different conformations. It is trained on a training set that has undergone data augmentation and redundancy removal. The HBR integrates prior physical knowledge between proteins and ligands to screen for the optimal ligand-receptor complex conformation and outputs the best target related to physiological state in the following way:
[0050]
[0051]
[0052] in, Represents the raw attention score. Representing the geometric characteristics of the edges, This represents converting geometric features into scalar weight factors, and ⊙ represents the element-wise multiplication of two vectors and the element-wise multiplication of two matrices. Characterizing the structural features of the edges, This represents mapping discrete data into a continuous high-dimensional vector representation. This indicates that the structure is embedded in the projection as an offset value. This represents element-wise addition and biased addition. This represents the final attention score; The feature vector representing the ligand; Indicates the first ligand The atom in the first The node feature vectors generated by the layer; This represents summing the eigenvectors of all nodes of the ligand; This indicates that the node features are weighted; Indicates the affinity score; This represents a multilayer perceptron.
[0053] Step 3: Implement the functional component screening strategy (1) Based on the chemical composition matrix in step 1, the graph neural network (GNN) is used to extract the structural priors of the chemical composition, including functional group categories, number of rotatable bonds, aromaticity characteristics, electron cloud density distribution and stereochemical constraint information. (2) Based on the above chemical priors and the binding pocket structure of AMPKα2 and other targets in the optimal ligand-receptor complex configuration obtained in step 2, a low-energy three-dimensional ligand conformation library is generated using the PocketDiffusion-MFH conditional diffusion generation model. (3) The generated ligand conformation library is docked to the active pocket of the target protein through a high-throughput parallel docking engine, and the activity (pIC50) and toxicological indicators (including 7 ADMET parameters such as protein binding rate PPB, hepatotoxicity DILI, and CYP450 inhibitory potential) of the compound are predicted by the multi-task graph neural network Food-GNN, and candidate molecules are output.
[0054] The high-throughput parallel docking engine utilizes the parallel computing of Autodock vina and Autodock GPU to calculate the binding affinity between ligands and target proteins; the target binding pocket structure is the active site on the target protein that binds to the ligand.
[0055] Furthermore, the PocketDiffusion-MFH conditional diffusion generation model generates a three-dimensional ligand conformation library of active compounds in food-derived medicinal materials. Using two-dimensional molecular diagrams of the compounds as input, and based on structural priors and target-binding pocket structures in the optimal complex configuration, low-energy three-dimensional ligand conformations are generated through multiple rounds of progressive denoising. The implementation method is as follows:
[0056] in, Indicates the first The intermediate state of a step; Indicates the first The state of the step; Indicates the current time step; Indicates conditional information; This represents the distribution of the inverse process after model parameterization; Indicates model parameters; The model predicts the first The mean of the step Gaussian distribution; The model predicts the first The covariance matrix of the step Gaussian distribution; Represents a normal distribution; Furthermore, the multi-task graph neural network Food-GNN constructs a molecular graph, treating each compound's atoms as nodes and chemical bonds as edges. Iteratively propagating feature information between nodes within the network yields the embedded representation of atoms, which is calculated using the following formula:
[0057] in, Represents atoms In iteration The amount of embedding; Represents atoms In iteration The embedding vector; Represents atoms The set of adjacent atoms; Representing adjacent atoms In iteration The embedding vector; Represents atoms and The characteristic vector of the chemical bond between them; Represents a message function; Indicates the update function; The resulting atomic embedding representation is input into the multi-task classification head and regression head to predict the activity and toxicity indicators of the compound, respectively. The prediction formulas are as follows:
[0058] in, This represents the input feature vector; This represents the weight matrix of the shared hidden layer; (•) represents the activation function; Indicates the first A weight matrix specific to each detection task; Indicates the first Bias terms for each detection task; (•) represents the Sigmoid function; Indicates the first The predicted probability of a detection; A cross-modal fusion approach is employed to integrate multi-source information, including chemical structure features (CS), morphological features (MO), and gene expression features (GE), ultimately leading to a predicted probability. The expression is as follows:
[0059] in, This indicates that the first [item] is predicted using only chemical structure. The probability of a detection; This represents the probability predicted using only morphological features; This represents the probability predicted using only gene expression. This represents the final predicted probability after fusion.
[0060] To evaluate the ability of the multi-task graph neural network Food-GNN in screening highly active samples, an enrichment factor (EF) is introduced for assessment:
[0061] in, This indicates the number of truly active compounds in the top 1% of compounds predicted by the model. This represents the total number of compounds predicted to be in the top 1%; This represents the total number of actual active compounds in the entire compound library; Total number of compounds in the library; EF represents the enrichment factor;
[0062] The multi-task loss function of the Food-GNN multi-task graph neural network consists of classification loss and regression loss:
[0063]
[0064]
[0065] in Indicates sample In the mission The real labels on it; Indicates sample In the mission The original output of the model; This represents the probability of being predicted as positive. Weights of positive samples (toxic molecules); The number of regression tasks is represented; loss represents the total loss function. Indicates the number of molecules; Indicates the number of classification tasks; Indicates the index of the current molecule; An index representing the classification task; Indicates the index of the regression task.
[0066] As predicted, the predicted results for bitter melon saponins are: pIC50 (AMPKα2) = 7.3, PPB = 22%, CYP450 inhibition risk: negative, DILI risk: low.
[0067] The Rule-Food compliance engine, which includes ChinaFoodDB, CIRS regulatory database, national standard GN, FDA and EFSA, is implemented using Prolog language. Based on the candidate molecules output in step 3, such as bitter melon saponins, the Rule-Food engine automatically excludes prohibited substances that are banned by regulatory authorities for safety, toxicity or other reasons. Step 4: Implement the Rule-Food compliance engine in Prolog, which includes ChinaFoodDB, CIRS regulatory database, national standard GN, FDA, and EFSA. Based on the candidate molecules output in Step 3, such as momordicin, Rule-Food automatically excludes prohibited substances banned by regulatory authorities for safety, toxicity, or other reasons. It has been verified that momordicin is a permitted ingredient in the "List of Food and Drug Homologous Ingredients," with a predicted daily acceptable intake (ADI) >500 mg / day, and no genotoxicity or sensitizing risk. The NSGA-III multi-objective evolutionary algorithm is used to optimize the formulation of candidate molecules with the objectives of "maximizing activity, minimizing toxicity, minimizing cost, and optimizing sensory perception." The objective function is as follows:
[0068]
[0069]
[0070]
[0071] in, This indicates the total number of candidate food and medicine homologous ingredients in the formula, i.e., the number of candidate molecules; Indicates the first The amount or proportion of each food ingredient; No. The functional activity of each component; Indicates the first The risk of toxicity or adverse reactions of each component; Indicates the first The unit price of each component; Indicates the first The usage amount of each ingredient; Indicates the first The contribution of each component to the overall sensory score; This represents the weighting factor for each component in the three corresponding functions.
[0072] After optimization, the optimal formulation combination obtained at the Pareto frontier is: bitter melon saponin freeze-dried powder: 15%; mulberry leaf polysaccharide extract: 35%; astragalus liposome preparation: 50%. This formulation simultaneously meets multiple objectives such as functional activity, safety, regulatory compliance, and economy, and has the potential to be developed into functional foods such as oral solid beverages and capsules.
[0073] Example 2: This embodiment uses lotus leaf as a typical plant with both medicinal and edible uses, and demonstrates in detail the screening of the optimal ligand-receptor complex configuration using steps 1 and 2 of the method of this invention. The specific implementation steps are as follows: Step 1: Integrate multi-source databases to construct a chemical composition matrix for foods that are both food and medicine; Using lotus leaves as an example of a food and medicinal ingredient, we integrated multiple databases including BATMAN-TCM 2.0, CheEMBL-Food subset, TCMSP, HERB, TCM.ID, and HIT2.0, and used Python automated scripts to retrieve and download their active ingredients. A total of 166 compounds were obtained, including typical components such as quercetin, kaempferol, and isoquercetin, some of which are shown in Table 1.
[0074] Table 1. Information on ten typical active ingredients in lotus leaves
[0075] Using InChIKey as the unique identifier for deduplication, SMILES were normalized using RDKit and Open Babel to supplement missing molecular descriptors and stereochemical parameters. The ETKDG algorithm was then used to generate a three-dimensional structure, ultimately constructing a chemical composition matrix containing chemical composition structural information, molecular descriptors, and stereochemical information. The chemical composition structural information includes the molecular diagram, SMILES, and InChIKey; the molecular descriptors include molecular weight, molecular formula, functional groups, ring structure, polarity, hydrophilicity, and hydrophobicity; and the stereochemical information includes symmetry, isomers, and planar symmetry.
[0076] Step 2: Implement target screening strategy (1) A rule-based engine-machine learning hybrid metabolic prediction model trained on databases such as MetaCyc, BRENDA, KEGG, and Cytochrome P450 was used. The smiles and three-dimensional structures of lotus leaf components were input, and metabolic sites (SOMs) were identified. The model was then combined with a reaction rule system covering phase I and phase II metabolic reactions to calculate and output potential primary and secondary metabolic transformation products. Four metabolic prediction tools—Biotransformer 3.0, MetaPredictor, MetaTrans, and Sygma—were used to generate initial metabolites of 929, 522, 335, and 951 respectively. (See...) Figure 5 The outputs of each model were integrated, and the SMILES outputs of each model were converted into InChIKey for deduplication using RDKit. 91 high-confidence metabolites that were predicted by at least three models were retained to generate a virtual metabolite database.
[0077] (2) The SMILES of 91 high-confidence metabolites were input into target prediction tools such as SEA, PharmMapper, SwissTargetPrediction, and DeepDTA to obtain 827 potential active target genes. At the same time, 1492 NAFLD-related targets were obtained from disease databases such as DisGeNET, DrugBank, GeneCards, and TTD. The intersection was taken to obtain 200 core targets. The crystal structures of the 200 core targets were downloaded from the PDB database. The three-dimensional structures of targets without crystal structures were predicted by AlphaFold3. Then, preprocessing such as dehydration, hydrogenation, and charge distribution was performed to generate target protein structure files that can be used for molecular docking.
[0078] (3) AutoDock Vina and AutoDock GPU were used to perform high-throughput docking of metabolite-target protein pairs to generate multi-conformation complexes. A heterogeneous graph neural network was used to intelligently score the docking results, integrate geometric features and structural priors, calculate binding affinity (prob), and screen out the optimal ligand-receptor complex conformation, as shown in Table 2.
[0079] Table 2. Top 10 binding scores of virtual metabolites of lotus leaf chemical components to NAFLD-related targets.
[0080] Among them, the binding conformation score of the metabolite of myricetin, the main component of lotus leaf, to protein 3WI2 was the highest (EquiScore = 0.99337304). The visualization results are shown below. Figure 6 .
[0081] The above-mentioned virtual metabolite prediction, target intersection screening, and heterogeneous graph neural network intelligent scoring, combined with high-throughput molecular docking and multi-dimensional interaction analysis, can obtain the optimal ligand-receptor complex with stable binding conformation and high affinity, as well as its corresponding target, providing a reliable target basis for subsequent functional component screening and formulation optimization.
[0082] This embodiment uses the Autodock GPU method for parallel computing, running hundreds of docking tasks simultaneously at a speed of 3 seconds per task. This is nearly 8 times faster than the traditional molecular docking Autodock Vina method (24.4 seconds per task), significantly reducing computation time. This advantage is particularly evident for screening large-scale compound libraries.
[0083] Example 3: This embodiment provides a dual-strategy screening system for active ingredients and targets in food products that are both food and medicine, including: The data construction module is used to integrate multi-source databases and construct a chemical composition matrix; The target screening module is used to identify metabolic sites based on chemical composition matrix, construct target set and output candidate ligand-receptor combinations through high-throughput molecular docking and heterogeneous graph neural network scoring, and output the best target related to physiological state. The functional component screening module is used to extract structural priors based on chemical component information in the chemical component matrix, generate a low-energy three-dimensional ligand conformation library based on structural priors and target binding pocket structures, and predict the activity and toxicity of compounds through a multi-task graph neural network Food-GNN, and output candidate molecules. The decision optimization module integrates the compliance engine Rule-Food with the NSGA-III multi-objective evolutionary algorithm to exclude prohibited substances based on candidate molecules and optimize the generation of the optimal formulation combination.
[0084] This invention constructs an integrated automated screening process, significantly reducing manual operation and greatly improving the screening efficiency from plant extracts to drug candidate targets, solving the problems of complex and cumbersome traditional processes and high dependence on manual labor. It systematically considers the metabolic transformation process of food-medicine homology foods in vivo, simulating the multi-stage metabolites formed by compounds catalyzed by metabolic enzymes such as CYP450, thus overcoming the deficiency of existing screening methods that ignore metabolic transformation links and improving the realism and biological significance of target prediction. A multi-task graph neural network, Food-GNN, based on molecular structure graph representation learning, is constructed to comprehensively integrate the structural information of food-derived components, enabling high-throughput and accurate prediction of multi-dimensional characteristics such as bioactivity, functional performance, and toxicity risk, significantly improving the model's generalization ability and prediction accuracy. The NSGA-III multi-objective evolutionary algorithm is adopted, considering multiple objectives such as enhanced activity, reduced toxicity, and cost control, to obtain the optimal balance between functional activity, safety, and economy in formulation design, reducing the number of experiments, shortening the development cycle, and lowering R&D costs. This invention encompasses two strategies: "target screening" and "component screening," enabling end-to-end virtual screening, activity verification, and functional food formulation design of food-medicine homologous active ingredients.
[0085] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A dual-strategy screening method for active ingredients and targets in food products that are both food and medicine, characterized in that, The method includes: Step 1: Integrate the BATMAN-TCM 2.0, CheEMBL-Food subset, TCMSP, HERB, TCM.ID, and HIT2.0 databases to construct a chemical composition matrix for food and medicine homology. The chemical composition matrix for food and medicine homology includes chemical composition structure information, molecular descriptors, and stereochemical information. The chemical composition structural information includes: molecular diagram, SMILES, InChlKey; the molecular descriptor includes: molecular weight, molecular formula, functional group, ring structure, polarity, hydrophilicity and hydrophobicity; the stereochemical information includes symmetry, isomers, and planar symmetry. Step 2: Implement the target screening strategy; The target screening strategy specifically involves: developing a rule engine-machine learning hybrid metabolite prediction model; identifying reaction sites and potential metabolic sites of molecules in the chemical composition matrix of food-medicine homology based on the chemical component structure information, molecular descriptors, and stereochemical information; calculating potential primary and secondary metabolites using a reaction rule system; integrating and deduplicating to construct a high-confidence virtual metabolite database; predicting virtual metabolites based on the compound target prediction database to obtain an active target dataset; integrating potential active targets related to physiological states based on a bioinformatics database; constructing a target set by taking the intersection of the active target dataset with the target set; downloading its crystal structure to generate target proteins; performing high-throughput molecular docking between the virtual metabolite database and target proteins; integrating the physical prior knowledge between target proteins and virtual metabolites using a heterogeneous graph neural network; characterizing their interaction in isovariant geometric space; and using a training scoring model to score and rank the high-throughput molecular docking results, screening the optimal ligand-receptor complex configuration, and outputting the best target related to physiological states, i.e., the target protein. The reaction rule system includes reaction rules, reaction mechanisms, a reaction database, and a reaction prediction algorithm; the reaction rules include phase I metabolic reactions and phase II metabolic reactions; the reaction mechanisms include enzyme catalysis mechanisms; the reaction database includes MetaCyc and KEGG; and the reaction prediction algorithm is an expert system model. Step 3: Implement the functional component screening strategy; The functional component screening strategy is as follows: Based on the chemical component structure information in the chemical component matrix of food homology, a graph neural network (GNN) is used to extract chemical priors, including functional group categories, number of rotatable bonds, aromaticity characteristics, electron cloud density distribution, and stereochemical constraint information; based on the chemical priors and the target binding pocket structure in the optimal ligand-receptor complex configuration, a conditional diffusion generation model is used to generate a low-energy three-dimensional ligand conformation library; a high-throughput parallel docking engine is used to dock the low-energy three-dimensional ligand conformation library to the active pocket of the target protein, and a multi-task graph neural network (Food-GNN) is used to predict the activity and toxicological indicators of the compounds, outputting candidate molecules; Step 4: Based on candidate molecules, prohibited substances are automatically excluded through the compliance engine Rule-Food; the NSGA-III multi-objective evolutionary algorithm is used to optimize the formulation of candidate molecules based on multiple dimensions of activity, safety, cost, and sensory factors, generating Pareto's cutting-edge functional food formulation combinations.
2. The dual-strategy screening method for active ingredients and targets in food products that are both food and medicine according to claim 1, characterized in that, The method of integrating multi-source databases in step 1 includes: merging duplicate compounds from multiple source databases using InChIKey as the unique identifier; standardizing and normalizing SMILES using RDKit and OpenBabel; completing missing values using structure prediction and physicochemical parameter calculation; performing molecular descriptor calculation and matrix processing; and generating 3D structural coordinates from SMILES and SDF using RDKitETKDG and OpenBabel to form an integrated database of two-dimensional and three-dimensional structures.
3. The dual-strategy screening method for active ingredients and targets in food and medicine homology according to claim 1, characterized in that, The target screening strategy employs a rule-engine-machine learning hybrid metabolite prediction model based on a multi-model joint prediction strategy. This model combines the PharmMapper structure reverse docking method, DeepDTA deep learning method, and SEA similarity analysis method to improve the accuracy and reliability of target prediction. The rule-engine-machine learning hybrid metabolite prediction model is trained and optimized using metabolic reaction data from the MetaCyc, BRENDA, KEGG, HumanCyc, ECNumbers, SMARTS structure matching rules, ChEMBL, PubChem, and Cytochrome P450 databases. This model simulates the reaction process catalyzed by metabolic enzymes in vivo, enabling the prediction of potential metabolites of small molecule drugs or natural products in vivo. The implementation method is as follows: In the formula, This represents the input compound; This indicates the probability that compound M is metabolized by the cytochrome P450 enzyme CPY isoform; This represents the probability score output by the trained random forest classifier when it accepts a molecule as input. This represents the probability that a specific atom, bond, and position is a metabolic target. A neural network model representing the prediction of metabolic sites; Characterized by the local chemical environment of atoms and bonds; This represents the molecular feature vector generated by the GNN; Indicates the input molecule; Represents a graph neural network; Indicates the Transformer decoder; This indicates the output metabolic product molecules.
4. The dual-strategy screening method for active ingredients and targets in food products that are both food and medicine according to claim 1, characterized in that, The target screening strategy includes the following compound target prediction databases: SEA, PharmMapper, SwissTargetPrediction, and DeepDTA; the bioinformatics databases include: Digenet, DrugBank, GeneCards, and TTD databases; and the high-throughput molecular docking is performed in parallel on a Linux system using AutodockVina and Autodock GPU.
5. The dual-strategy screening method for active ingredients and targets in food and medicine homology according to claim 1, characterized in that, In the target screening strategy, the heterogeneous graph neural network is modeled based on graph structure data of different conformations of target proteins and virtual metabolites, and the heterogeneous graph neural network is trained based on a training set that has undergone data augmentation and redundancy removal. The following method utilizes heterogeneous graph neural networks to integrate prior physical knowledge between proteins and ligands, screens for the optimal ligand-receptor complex conformation, and outputs the best target associated with physiological state: in, Represents the raw attention score. Representing the geometric characteristics of the edges, This represents converting geometric features into scalar weight factors, and ⊙ represents the element-wise multiplication of two vectors and the element-wise multiplication of two matrices. Characterizes the structural features of the edges. This represents mapping discrete data into a continuous high-dimensional vector representation. This indicates that the structure is embedded in the projection as an offset value. This represents element-wise addition and bias addition. This represents the final attention score; The feature vector of the ligand; Indicates the first ligand The atom in the first The node feature vectors generated by the layer; This represents summing the eigenvectors of all nodes of the ligand; This indicates that the node features are weighted; Indicates the affinity score; This represents a multilayer perceptron.
6. The dual-strategy screening method for active ingredients and targets in food and medicine homology according to claim 1, characterized in that, The high-throughput parallel docking engine in the functional component screening strategy utilizes the parallel computing of Autodock vina and Autodock GPU to calculate the binding affinity between ligands and target proteins. The target binding pocket structure is the active site on the target protein that binds to the ligand.
7. The dual-strategy screening method for active ingredients and targets in food products that are both food and medicine according to claim 1, characterized in that, The conditional diffusion generation model in the functional component screening strategy is the PocketDiffusion-MFH model; The PocketDiffusion-MFH model generates a three-dimensional ligand conformation library of active compounds in food-derived medicinal materials. Using two-dimensional molecular diagrams of the compounds as input, and based on structural priors and the target binding pocket structure in the optimal complex configuration, it generates low-energy three-dimensional ligand conformations through multiple rounds of progressive denoising. The implementation method is as follows: in, Indicates the first The intermediate state of a step; Indicates the first The state of the step; Indicates the current time step; Indicates conditional information; This represents the distribution of the inverse process after model parameterization; Indicates model parameters; The model predicts the first The mean of the step Gaussian distribution; The model predicts the first The covariance matrix of the step Gaussian distribution; This represents a normal distribution.
8. The dual-strategy screening method for active ingredients and targets in food products that are both food and medicine according to claim 1, characterized in that, The functional component screening strategy utilizes a multi-task graph neural network, Food-GNN, which constructs a molecular graph, treating each compound's atoms as nodes and chemical bonds as edges. Iteratively propagating feature information between nodes within the network yields the embedded representation of atoms. This embedded representation is calculated using the following formula: in, Represents atoms In iteration The amount of embedding; Represents atoms In iteration The embedding vector; Represents atoms The set of adjacent atoms; Representing adjacent atoms In iteration The embedding vector; Represents atoms and The characteristic vector of the chemical bond between them; Represents a message function; Indicates the update function; The resulting atomic embedding representation is input into the multi-task classification head and regression head to predict the activity and toxicity indicators of the compound, respectively. The prediction formulas are as follows: in, This represents the input feature vector; This represents the weight matrix of the shared hidden layer; (•) represents the activation function; Indicates the first A weight matrix specific to each detection task; Indicates the first Bias terms for each detection task; (•) represents the Sigmoid function; Indicates the first The predicted probability of each detection; A cross-modal fusion approach is employed to integrate multi-source information, including chemical structure features (CS), morphological features (MO), and gene expression features (GE), ultimately leading to a predicted probability. The expression is as follows: in, This indicates that the prediction of the first [item] is based solely on chemical structure. The probability of a detection; This represents the probability predicted using only morphological features; This represents the probability predicted using only gene expression. This represents the final predicted probability after fusion.
9. The dual-strategy screening method for active ingredients and targets in food and medicine homology according to claim 1, characterized in that, Step 4 further includes using the Prolog language to implement the compliance engine Rule-Food to determine the legality of candidate molecules. The compliance engine Rule-Food includes ChinaFoodDB, the CIRS regulatory database, the national standard GN, and the FDA and EFSA. The prohibited substances are those banned by regulatory authorities for safety, toxicity, or other reasons. The objective function of the NSGA-III multi-objective evolutionary algorithm for formula optimization of candidate molecules is: in, This indicates the total number of candidate food and medicine homologous ingredients in the formula, i.e., the number of candidate molecules; Indicates the first The amount or proportion of each food ingredient; No. The functional activity of each component; Indicates the first The risk of toxicity or adverse reactions of each component; Indicates the first The unit price of each component; Indicates the first The usage amount of each ingredient; Indicates the first The contribution of each component to the overall sensory score; This represents the weighting factor for each component in the three corresponding functions.
10. A dual-strategy screening system for active ingredients and targets in food products that are both food and medicine, characterized in that, The system is used to perform the dual-strategy screening method for active ingredients and targets in food homology as described in any one of claims 1 to 9, the system comprising: The data construction module is used to integrate multi-source databases and construct a chemical composition matrix; The target screening module is used to identify metabolic sites based on chemical composition matrix, construct target set and output candidate ligand-receptor combinations through high-throughput molecular docking and heterogeneous graph neural network scoring, and output the best target related to physiological state. The functional component screening module is used to extract structural priors based on chemical component information in the chemical component matrix, generate a low-energy three-dimensional ligand conformation library based on structural priors and target binding pocket structures, and predict the activity and toxicity of compounds through a multi-task graph neural network Food-GNN, and output candidate molecules. The decision optimization module integrates the compliance engine Rule-Food with the NSGA-III multi-objective evolutionary algorithm to exclude prohibited substances based on candidate molecules and optimize the generation of the optimal formulation combination.
Citation Information
Patent Citations
CN120319354A
CN120600158A