Natural product anti-aging component screening method and system based on multi-modal deep learning

By combining multimodal deep learning methods with Chinese and Western medicine data, a dual-channel transfer learning model was constructed, which solved the problems of high cost and difficult prediction of traditional anti-aging drug screening, and achieved efficient and accurate screening of natural product anti-aging ingredients and key information output.

CN120636607APending Publication Date: 2025-09-12XIAN INT UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510786159.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Traditional anti-aging drug screening has high costs and low throughput, data on Chinese and Western medicines are fragmented, and the anti-aging mechanisms of Chinese herbal ingredients are unclear, resulting in difficult predictions and insufficient model generalization capabilities.

Method used

A multimodal deep learning method was used to combine traditional Chinese medicine and Western medicine data to construct a dual-channel transfer learning model. The anti-aging activity of natural products was predicted through standardization processing, multimodal feature extraction and model training.

Benefits of technology

It improves data utilization, reduces the demand for traditional Chinese medicine data, and enhances the model's prediction accuracy and generalization ability. It can successfully predict new anti-aging traditional Chinese medicine ingredients and output key atomic fragments and target association rules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636607A_ABST
    Figure CN120636607A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a natural product anti-aging component screening method and system based on multi-modal deep learning, and relates to the technical field of artificial intelligence assisted drug discovery. The method comprises the following steps: collecting multi-source data, and carrying out standardization processing on the multi-source data; performing multi-modal feature processing on the multi-source data subjected to the standardization processing; constructing a dual-channel transfer learning model according to the multi-source data after the multi-modal feature processing, and training and finely adjusting the dual-channel transfer learning model; and predicting the anti-aging activity of the natural product to be tested according to the fine-tuned dual-channel transfer learning model. According to the natural product anti-aging component screening method based on multi-modal deep learning, the data utilization rate can be increased, and important atom fragments and key targets can be output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of artificial intelligence-assisted drug discovery, and in particular to a method and system for screening anti-aging ingredients from natural products based on multimodal deep learning. Background Art

[0002] Traditional anti-aging drug screening relies on in vitro experiments, which are costly and low-throughput. Furthermore, existing computational methods utilize only single molecular fingerprints or target data. Furthermore, existing techniques fragment data between Chinese and Western medicines, failing to leverage the transfer learning value of known anti-aging knowledge from Western medicines. Finally, the unclear anti-aging mechanisms of traditional Chinese medicine ingredients make prediction difficult, while models lack generalization capabilities under small sample sizes and lack multidimensional modeling of molecular target pathway associations. Therefore, a method for screening natural product anti-aging ingredients is needed. Summary of the Invention

[0003] The multimodal deep learning-based natural product anti-aging ingredient screening method provided in the present disclosure can improve data utilization and output important atomic fragments and key targets.

[0004] According to a first aspect of an embodiment of the present disclosure, a method for screening anti-aging ingredients from natural products based on multimodal deep learning is provided, the method comprising:

[0005] Collecting multi-source data and performing standardization processing on the multi-source data;

[0006] performing multimodal feature processing on the standardized multi-source data;

[0007] Constructing a dual-channel transfer learning model based on the multi-source data after multimodal feature processing, and training and fine-tuning the dual-channel transfer learning model;

[0008] The anti-aging activity of the natural products to be tested was predicted based on the fine-tuned dual-channel transfer learning model.

[0009] In one embodiment, collecting multi-source data and performing standardization processing on the multi-source data includes:

[0010] Acquire a traditional Chinese medicine data source and a western medicine data source, and preprocess the traditional Chinese medicine data source and the western medicine data source;

[0011] Screening positive samples of traditional Chinese medicine and positive samples of Western medicine, and constructing a data set based on the positive samples of traditional Chinese medicine and the positive samples of Western medicine;

[0012] The data set is divided into a training set and a test set; wherein the training set is used for model training, and the test set is used for model verification.

[0013] In one embodiment, performing multimodal feature processing on the standardized multi-source data includes:

[0014] generating a molecular 3D conformation and extracting molecular geometric features based on the molecular 3D conformation;

[0015] Constructing target network features and determining target feature vectors based on the target network features;

[0016] The molecular geometric features and the target feature vector are concatenated and input into a fully connected layer, and a fused multimodal feature vector is outputted after dimensionality reduction processing.

[0017] In one embodiment, constructing a dual-channel transfer learning model based on the multi-source data after multimodal feature processing, and training and fine-tuning the dual-channel transfer learning model includes:

[0018] Pre-train channels based on the multimodal feature matrix of Western medicine;

[0019] Fine-tune the channel according to the multimodal feature matrix of traditional Chinese medicine;

[0020] Use Western medicine data to pre-train the model to optimize the anti-aging activity classification task, and stop training when the validation set AUC reaches a preset value;

[0021] Freeze the first two layers of the pre-trained model and fine-tune the last layer using a learning rate decay strategy.

[0022] In one embodiment, predicting the anti-aging activity of the natural product to be tested based on the fine-tuned dual-channel transfer learning model includes:

[0023] The SMILES structure of the natural product to be tested is input into the fine-tuned dual-channel transfer learning model to generate a 3D conformation and extract molecular geometric features;

[0024] Query the target information of the natural product to be tested and calculate the target characteristics;

[0025] The molecular geometric features and the target features are spliced ​​and input into a fully connected layer for calculation, and a fused multimodal feature vector is output;

[0026] Inputting the multimodal feature vector into the fine-tuned dual-channel transfer learning model to output the anti-aging activity probability;

[0027] Output the attention weights of key atomic fragments and visualize the association rules between the key atomic fragments and key targets.

[0028] In one embodiment, obtaining a traditional Chinese medicine data source and a western medicine data source, and preprocessing the traditional Chinese medicine data source and the western medicine data source includes:

[0029] The molecular structures, target information, and anti-aging activity data of 33,643 traditional Chinese medicine ingredients were extracted from the Herb database as the traditional Chinese medicine data source;

[0030] Extract 17,458 Western medicines from the DrugBank database, and obtain the corresponding molecular structures, targets, and efficacy annotations as Western medicine data sources; wherein the anti-aging drugs include positive control drugs;

[0031] Remove entries with incomplete molecular structures and / or unlabeled targets from the traditional Chinese medicine data source and the Western medicine data source;

[0032] Performing unified target naming on entries in the traditional Chinese medicine data source and the Western medicine data source;

[0033] The screening of positive samples of traditional Chinese medicine and positive samples of western medicine includes:

[0034] Drugs that target core aging pathways, drugs marked as "senolytic" or "antiaging" in DrugBank, or drugs with clear records in the DrugAge or Senolytic Compounds databases were selected as positive Western medicine samples;

[0035] Drugs that share at least one target with the Western medicine positive sample, or drugs with ingredients that are clearly reported in literature to have anti-aging activity are selected as traditional Chinese medicine positive samples.

[0036] In one embodiment, generating a molecular 3D conformation and extracting molecular geometric features based on the molecular 3D conformation includes:

[0037] Use OpenBabel to perform conformational search on each molecule to optimize it to the lowest energy state and output a 3D molecular structure file containing atomic coordinates;

[0038] Inputting the 3D molecular structure file into a pre-trained three-dimensional graph attention network to perform atomic-level feature encoding and interatomic spatial distance weight calculation; wherein the atomic-level feature encoding includes encoding of atom type, hybridization state, and formal charge;

[0039] Output the geometric characteristics of each molecule.

[0040] In one embodiment, constructing target network features and determining target feature vectors based on the target network features includes:

[0041] Construct a PPI network of targets based on the STRING database;

[0042] Calculate the node centrality score using NetworkX; the node centrality score is used to measure the hubness of the target in the aging pathway and characterize the efficiency of target information transmission;

[0043] Integrate the co-expression intensity data of target genes and aging markers in the GTEx database;

[0044] The target pathway enrichment p-value was obtained through the Reactome API;

[0045] The centrality scores, co-expression intensity data and pathway p-values ​​were concatenated into a target feature vector.

[0046] In one embodiment, the pre-training channel based on the multimodal feature matrix of Western medicine includes:

[0047] Inputting the multimodal feature matrix of western medicine into the network structure is a 3-layer fully connected neural network, the loss function is a cross entropy loss function, the training parameters are a learning rate of 0.001, a batch size of 128, and the network is trained for 50 rounds to obtain the pre-training channel;

[0048] The fine-tuning of the channel according to the multimodal feature matrix of traditional Chinese medicine includes:

[0049] Parameter migration, to load the weights of the pre-trained model as initialization parameters;

[0050] Stratified sampling to ensure that each target family is evenly distributed in the training set;

[0051] Comparative loss function to constrain the similarity of molecular features within the same target family.

[0052] According to a second aspect of the embodiments of the present disclosure, a natural product anti-aging ingredient screening system based on multimodal deep learning is provided, the system comprising: an acquisition module, a processing module, a construction module and a prediction module; wherein,

[0053] The acquisition module collects multi-source data and performs standardization processing on the multi-source data;

[0054] The processing module performs multimodal feature processing on the multi-source data after the standardization process;

[0055] The construction module constructs a dual-channel transfer learning model based on the multi-source data after multimodal feature processing, and trains and fine-tunes the dual-channel transfer learning model;

[0056] The prediction module predicts the anti-aging activity of the natural product to be tested based on the fine-tuned dual-channel transfer learning model.

[0057] The present disclosure provides a method for screening anti-aging ingredients from natural products based on multimodal deep learning, and a system and method for screening anti-aging ingredients from natural products based on multimodal deep learning provided in an embodiment of the present disclosure. On an independent test set, the AUC can reach 0.92, and the F1 can reach 0.85 (in the traditional method, the AUC is 0.78, and the F1 is 0.71. This traditional method is a combination of molecular fingerprints and random forests), and can successfully predict three new anti-aging Chinese medicine ingredients (such as vitamin K2). At the same time, the system and method for screening anti-aging ingredients from natural products based on multimodal deep learning provided in an embodiment of the present disclosure can improve data utilization. Through transfer learning, it reduces the demand for Chinese medicine data by 60%, and can output association rules between important atomic fragments and key targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 A flowchart of the method for screening anti-aging ingredients from natural products based on multimodal deep learning provided for the implementation of the present disclosure.

[0059] Figure 2 A flowchart of the method for screening anti-aging ingredients from natural products based on multimodal deep learning provided for the implementation of the present disclosure.

[0060] Figure 3 A flowchart of the method for screening anti-aging ingredients from natural products based on multimodal deep learning provided for the implementation of the present disclosure.

[0061] Figure 4 A flowchart of the method for screening anti-aging ingredients from natural products based on multimodal deep learning provided for the implementation of the present disclosure.

[0062] Figure 5 A flowchart of the method for screening anti-aging ingredients from natural products based on multimodal deep learning provided for the implementation of the present disclosure.

[0063] Figure 6 Architectural diagram of the multimodal deep learning-based natural product anti-aging ingredient screening system provided for the implementation of the present disclosure. DETAILED DESCRIPTION

[0064] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of systems consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0065] Figure 1 This is a flow chart of a method for screening anti-aging ingredients from natural products based on multimodal deep learning provided in the embodiments of the present disclosure. Figure 1As shown, the method includes:

[0066] Step 101: Collect multi-source data and perform standardization on the multi-source data;

[0067] In one embodiment, Figure 2 As shown, the collecting of multi-source data and standardizing the multi-source data include:

[0068] Step 1011: Acquire a traditional Chinese medicine data source and a western medicine data source, and pre-process the traditional Chinese medicine data source and the western medicine data source;

[0069] In one embodiment, obtaining a traditional Chinese medicine data source and a western medicine data source, and preprocessing the traditional Chinese medicine data source and the western medicine data source includes:

[0070] The molecular structures (SMILES format), target information (UniProt ID) and anti-aging activity data of 33,643 traditional Chinese medicine ingredients were extracted from the Herb database as the traditional Chinese medicine data source;

[0071] 17,458 Western medicines were extracted from the DrugBank database, and the corresponding molecular structures, targets, and efficacy annotations were obtained as Western medicine data sources;

[0072] Remove entries with incomplete molecular structures and / or unlabeled targets from the traditional Chinese medicine data source and the Western medicine data source;

[0073] Through UniProt ID standardization, unified target naming is performed on the entries in the traditional Chinese medicine data source and the Western medicine data source;

[0074] Step 1012: screening positive samples of traditional Chinese medicine and positive samples of western medicine, and constructing a data set based on the positive samples of traditional Chinese medicine and the positive samples of western medicine;

[0075] In one embodiment, screening positive samples of traditional Chinese medicine and positive samples of western medicine includes:

[0076] Drugs that target core aging pathways, drugs marked as "senolytic" or "antiaging" in DrugBank, or drugs with clear records in the DrugAge or Senolytic Compounds databases were selected as positive Western medicine samples;

[0077] Specifically, in this step, the Western medicine positive sample screening rules must meet at least one of the following:

[0078] Target Match: Directly matches UniProt IDs targeting core aging pathways (mTOR / SIRT1 / AMPK).

[0079] Drug efficacy labeling: drugs labeled as "senolytic" or "antiaging" in DrugBank.

[0080] Database validation: drugs that have been clearly recorded in the DrugAge or Senolytic Compounds database.

[0081] Positive sample screening rules for traditional Chinese medicine:

[0082] Drugs that share at least one target with the Western medicine positive sample, or drugs with ingredients that are clearly reported in literature to have anti-aging activity are selected as traditional Chinese medicine positive samples.

[0083] Specifically, drugs that share at least one target with the positive Western medicine samples, such as UniProt ID matching, are selected, and drugs with ingredients that are clearly reported in the literature to have anti-aging activity, such as astragaloside IV, are selected.

[0084] It should be noted that, in this embodiment, molecules that do not meet the above conditions need to be randomly selected as negative samples to ensure class balance.

[0085] Step 1013 divides the data set into a training set and a test set; wherein the training set is used for model training, and the test set is used for model verification.

[0086] In this step, 80% of the data were selected for model training, including 26,914 kinds of traditional Chinese medicine and 14,038 kinds of Western medicine; 20% of the data were selected as an independent test set for final verification, including 6,729 kinds of traditional Chinese medicine and 3,420 kinds of Western medicine.

[0087] Step 102: performing multimodal feature processing on the standardized multi-source data;

[0088] In one embodiment, Figure 3 As shown, the multimodal feature processing of the standardized multi-source data includes:

[0089] Step 1021: Generate a 3D conformation of the molecule, and extract molecular geometric features based on the 3D conformation of the molecule;

[0090] In one embodiment, generating a molecular 3D conformation and extracting molecular geometric features based on the molecular 3D conformation includes:

[0091] The one-hot encoding of the atomic type and hybridization state is concatenated for each atom, and the formal charge is normalized to form the atomic feature vector.

[0092] The exponential weights are calculated based on the inter-atomic distances to generate the attention matrix.

[0093] The atomic feature matrix and attention matrix are input to a three-layer graph attention network (GAT). Each layer aggregates neighbor features by calculating the attention coefficient and is activated by ReLU. After the multi-head outputs are spliced, the third layer is compressed into a 256-dimensional molecular geometric feature vector through global average pooling.

[0094] The specific implementation process of this step is:

[0095] Input: SMILES string or molecular 3D conformation (SDF format), generated and optimized by OpenBabel (MMFF94 force field).

[0096] 1. Atomic feature encoding:

[0097] (1) Atom type (One-Hot encoding):

[0098] Let the set of atomic types be , a total of K types.

[0099] The type of atom i , whose encoding vector for:

[0100] ;

[0101] in, is an indicator function, which is 1 if the condition is met and 0 otherwise.

[0102] (2) Hybrid state (One-Hot encoding):

[0103] Assume that the hybrid state set is , the hybridization state of atom i , whose encoding vector for:

[0104] .

[0105] (3) Formal charge normalization:

[0106] Formal charge on atom i , normalized to [-1, 1]:

[0107] ;

[0108] in and are the minimum and maximum values ​​of formal charge in the data set, respectively.

[0109] (4) Atomic feature splicing:

[0110] The final eigenvector of atom i is:

[0111] .

[0112] 2. Spatial attention weight calculation:

[0113] For atoms i and j, calculate the distance weight:

[0114] ;

[0115] Constructing the attention matrix , where N is the number of atoms.

[0116] 3. Graph Attention Network (GAT) forward propagation:

[0117] Input: atomic feature matrix and attention matrix.

[0118] Network structure: 3 GAT layers, where the hidden layer dimension is 128 and the number of multi-head attention heads is 4.

[0119] Output: Molecular geometry feature vector, dimension 256.

[0120] (1) Single-head attention mechanism:

[0121] For the Layer GAT, input feature matrix , output features :

[0122] a. Calculate the attention coefficient:

[0123] ;

[0124] in, is the weight matrix, is the attention parameter vector.

[0125] b. Feature aggregation:

[0126] ;

[0127] in, is the neighbor set of atom i, is an activation function, such as ReLU.

[0128] (2) Multi-head attention mechanism:

[0129] Assume the number of heads is M (in this example, M=4), and the output dimension of each head is , and finally spliced ​​together:

[0130] ;

[0131] The total dimensions are: .

[0132] (3) Multi-layer GAT structure:

[0133] Input layer: atomic features ;

[0134] Layer 1: 4-head GAT, output ;

[0135] Layer 2: 4-head GAT, output ;

[0136] Layer 3: Global average pooling, output molecular geometric feature vector ,

[0137] .

[0138] Step 1022: construct target network features, and determine target feature vectors based on the target network features;

[0139] In one embodiment, constructing target network features and determining target feature vectors based on the target network features includes:

[0140] The PPI network of the target was constructed based on the STRING database. In this embodiment, the PPI network of the target was constructed based on the STRING database, covering interactions with a confidence level > 0.7.

[0141] Calculate the node centrality score using NetworkX; the node centrality score is used to measure the hubness of the target in the aging pathway and characterize the efficiency of target information transmission;

[0142] Integrate the co-expression intensity data of target genes and aging markers in the GTEx database; in this step, integrate the co-expression data of target genes and aging markers (such as CDKN2A) in the GTEx database (Pearson correlation coefficient>0.5).

[0143] Obtain target pathway enrichment p-values ​​through the Reactome API. In this step, obtain target pathway enrichment p-values ​​(e.g., mTOR signaling pathway, p < 0.05) through the Reactome API.

[0144] The centrality scores, co-expression intensity data and pathway p-values ​​were concatenated into a target feature vector for feature fusion.

[0145] The specific implementation process of this step is:

[0146] 1. Calculate node centrality score

[0147] Perform minimum and maximum normalization on the Betweenness Centrality (BC) of each target:

[0148] ;

[0149] and are the minimum and maximum centrality values ​​of all targets in the PPI network, respectively.

[0150] 2 Reactome pathway enrichment p-value transformation:

[0151] Take a negative logarithmic transformation of pathway enrichment p-values: ;

[0152] like , it is set to the preset maximum value.

[0153] 3 GTEx co-expression intensity retains the original value:

[0154] Directly use the Pearson correlation coefficient As co-expression intensity: ;

[0155] 4. Feature concatenation and dimension mapping:

[0156] Concatenate the three processed features into the original feature vector:

[0157] .

[0158] The 3D original features are mapped to a 128-dimensional target feature vector through a fully connected layer:

[0159] ;

[0160] in, is the weight matrix, is the bias vector; is the activation function.

[0161] Step 1023: concatenate the molecular geometric features and the target feature vector and input them into a fully connected layer, and output a fused multimodal feature vector after dimensionality reduction processing.

[0162] In this step, the molecular geometric features (256 dimensions) and target network features (128 dimensions) are used as input and reduced to a unified representation (dimension: 192) through the fully connected layer (FC layer). ;in, , is a learnable parameter. Finally, the fused multimodal feature vector (dimension 192) is output for model training.

[0163] Step 103: constructing a dual-channel transfer learning model based on the multi-source data after multimodal feature processing, and training and fine-tuning the dual-channel transfer learning model;

[0164] In one embodiment, Figure 4 As shown, the step of constructing a dual-channel transfer learning model based on the multi-source data after multimodal feature processing, and training and fine-tuning the dual-channel transfer learning model includes:

[0165] Step 1031: pre-train the channel according to the multimodal feature matrix of western medicine;

[0166] In one embodiment, the pre-training channel based on the multimodal feature matrix of Western medicine includes:

[0167] The multimodal feature matrix of Western medicine is input into the network structure of a 3-layer fully connected neural network, the loss function is the cross entropy loss function, the training parameters are a learning rate of 0.001, a batch size of 128, and the network is trained for 50 rounds to obtain the pre-training channel.

[0168] Step 1032: fine-tune the channel according to the multimodal feature matrix of traditional Chinese medicine;

[0169] The fine-tuning of the channel according to the multimodal feature matrix of traditional Chinese medicine includes:

[0170] Parameter migration, to load the weights of the pre-trained model as initialization parameters;

[0171] Stratified sampling to ensure that each target family is evenly distributed in the training set;

[0172] Comparative loss function to constrain the similarity of molecular features within the same target family.

[0173] In this step, the molecular feature similarity of the same target family is constrained by the following formula.

[0174] ;

[0175] in, is the Euclidean distance, is the target similarity, .

[0176] Step 1033: Use Western medicine data to pre-train the model to optimize the anti-aging activity classification task, and stop training after the validation set AUC reaches a preset value;

[0177] In this step, only Western medicine data was used to train the model, optimize the anti-aging activity classification task, and stop training after the validation set AUC reached 0.89.

[0178] Step 1034: Freeze the first two layers of the pre-trained model and fine-tune the last layer using a learning rate decay strategy.

[0179] In this step, the first two layers of the pre-trained model are frozen, and only the last layer is fine-tuned with a learning rate decay strategy (initial lr = 0.0001, decaying by 50% every 10 rounds).

[0180] The specific implementation process of this step is:

[0181] In the pre-training stage (Western medicine data):

[0182] Network structure: 3-layer fully connected network (input 192 → 128 → 1).

[0183] Loss function: Cross entropy loss: .

[0184] Optimization parameters: Adam (lr=0.001, batch_size=128), early stopping strategy (validation set AUC ≥ 0.89).

[0185] In the fine-tuning stage (TCM data):

[0186] Parameter migration: load pre-trained weights and freeze the first two layers.

[0187] Contrastive loss function: .

[0188] Stratified sampling: Group by target family to ensure even coverage of each target type.

[0189] Learning rate decay: initial lr=0.0001, decaying by 50% every 10 rounds

[0190] Step 104: predict the anti-aging activity of the natural product to be tested based on the fine-tuned dual-channel transfer learning model.

[0191] In one embodiment, Figure 5 As shown, the anti-aging activity of the natural product to be tested is predicted based on the fine-tuned dual-channel transfer learning model, including:

[0192] Step 1041: Input the SMILES structure of the natural product to be tested into the fine-tuned dual-channel transfer learning model to generate a 3D conformation and extract molecular geometric features;

[0193] The specific implementation process of this step is as follows:

[0194] 1. Molecule input verification:

[0195] SMILES validity check: Use the Chem.MolFromSmiles() function of RDKit to verify the validity of SMILES. If invalid, an error code (such as Error 101) is returned.

[0196] from rdkit import Chem

[0197] mol = Chem.MolFromSmiles(smiles)

[0198] Supports SMILES string or SDF file format input.

[0199] Target completion:

[0200] If the target was not specified, the Herb database API was called to retrieve related targets based on molecular name or structural similarity (similarity threshold > 0.7).

[0201] 2. Molecular geometric feature extraction (calling the model in step 102):

[0202] 3D conformation generation:

[0203] Execute the following command using OpenBabel:

[0204] obabel -ismi input.smi -O output.sdf --gen3D --minimize --ff MMFF94;

[0205] gen3D: Generate initial 3D conformation.

[0206] minimize: Use the MMFF94 force field for energy minimization.

[0207] Output file: Generates an SDF file containing atomic coordinates, recording the spatial position (XYZ coordinates) of each atom.

[0208] Feature calculation: Input the pre-trained 3DGAT network and output 256-dimensional features.

[0209] Step 1042: query the target information of the natural product to be tested and calculate the target characteristics;

[0210] In this step, the model in step 102 is called to calculate the 128-dimensional target features.

[0211] Step 1043: Input the multimodal feature vector into the fine-tuned dual-channel transfer learning model to output the anti-aging activity probability;

[0212] The specific implementation process of this step is as follows:

[0213] 1. Calculation of the probability of anti-aging activity:

[0214] Input: Load the fine-tuned two-channel transfer learning model (PyTorch format .pt file) and fused feature vector (192 dimensions).

[0215] Probability prediction: The anti-aging activity probability is output through the last layer of Sigmoid function:

[0216] ;

[0217] in, is the Sigmoid function.

[0218] Activity determination: If , labeled as “anti-aging active ingredient”.

[0219] Step 1044: output the attention weight of the key atomic fragment, and visualize the association rules between the key atomic fragment and the key target.

[0220] The specific implementation process of this step is as follows:

[0221] 1. Identification of key atomic fragments:

[0222] Atom weight sorting:

[0223] For each atom i, calculate the global attention weight mean: extract the atomic level attention weight matrix from the 3DGAT layer .

[0224] Atom weight calculation:

[0225] ;

[0226] Wherein, N is the number of atoms.

[0227] Sort by weight from high to low and filter the top 10% atoms.

[0228] Substructure clustering: Use the GetSubstructMatches method of RDKit to cluster adjacent high-weight atoms into chemical substructures (such as benzene rings and glycosidic bonds).

[0229] Sample code:

[0230] from rdkit import Chem

[0231] mol = Chem.MolFromSmiles(smiles)

[0232] core = Chem.MolFromSmarts('[#6]1:[#6]:[#6]:[#6]:[#6]:[#6]:1') # Benzene ring pattern

[0233] matches = mol.GetSubstructMatches(core).

[0234] 2. Results visualization:

[0235] Molecule highlighting: Use RDKit's Draw.MolToImage() function to mark high-weight atoms in red.

[0236] Target pathway association diagram: Use Cytoscape to plot the location of targets in the PPI network and highlight core targets (such as SIRT1).

[0237] Final output:

[0238] {

[0239] "status": "success",

[0240] "probability": 0.93,

[0241] "prediction": "Active (Anti-aging)",

[0242] "key_fragments": [

[0243] {

[0244] "smarts": "[O][C](=[O])C1=CC=CC=C1",

[0245] "atom_indices": [5,6,7,8],

[0246] "average_weight": 0.91,

[0247] "structure_image": "base64_encoded_image"

[0248] }

[0249] ],

[0250] "critical_targets": [

[0251] {

[0252] "uniprot_id": "Q96EB6",

[0253] "name": "SIRT1",

[0254] "pathway": "NAD-dependent protein deacetylation",

[0255] "betweenness_centrality": 0.85

[0256] } ]

[0258] }.

[0259] The present disclosure provides a method for screening anti-aging ingredients from natural products based on multimodal deep learning, and a system and method for screening anti-aging ingredients from natural products based on multimodal deep learning provided in an embodiment of the present disclosure. On an independent test set, the AUC can reach 0.92, and the F1 can reach 0.85 (in the traditional method, the AUC is 0.78, and the F1 is 0.71. This traditional method is a combination of molecular fingerprints and random forests), and can successfully predict three new anti-aging Chinese medicine ingredients (such as vitamin K2). At the same time, the system and method for screening anti-aging ingredients from natural products based on multimodal deep learning provided in an embodiment of the present disclosure can improve data utilization. Through transfer learning, it reduces the demand for Chinese medicine data by 60%, and can output association rules between important atomic fragments and key targets.

[0260] It should be noted that the AUC in the embodiments of the present disclosure stands for Area Under the ROC Curve (or AUC for short), and its English name is Area Under the ROC Curve. The ROC (Receiver Operating Characteristic) curve is a graphical tool for evaluating the performance of a binary classification model.

[0261] The AUC value represents the area enclosed by the ROC curve and the coordinate axis, and its value range is between 0.5 (random guessing) and 1 (perfect classification). The larger the value, the better the model performance.

[0262] The full name of F1 in the embodiments of the present disclosure is F1 score (or F1 value) in Chinese, and F1 score in English means the harmonic mean of precision and recall, which is used to comprehensively evaluate the performance of the classification model.

[0263] formula:

[0264] ;

[0265] Among them, the closer the F1 value is to 1, the better the model balances precision and recall.

[0266] Figure 6 This is an architecture diagram of a natural product anti-aging ingredient screening system based on multimodal deep learning provided by the present disclosure. Figure 6 As shown, the system includes: an acquisition module 601, a processing module 602, a construction module 603 and a prediction module 604; wherein the acquisition module 601 is used to acquire multi-source data and standardize the multi-source data; the processing module 602 is used to perform multimodal feature processing on the multi-source data after standardization; the construction module 603 is used to construct a dual-channel transfer learning model based on the multi-source data after multimodal feature processing, and train and fine-tune the dual-channel transfer learning model; the prediction module 601 is used to predict the anti-aging activity of the natural product to be tested based on the fine-tuned dual-channel transfer learning model.

[0267] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be pre-stored in a computer-readable storage medium. When executed, the program performs the steps in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0268] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the disclosure herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0269] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A method for screening anti-aging ingredients from natural products based on multimodal deep learning, characterized in that: The method comprises: Collecting multi-source data and performing standardization processing on the multi-source data; performing multimodal feature processing on the standardized multi-source data; Constructing a dual-channel transfer learning model based on the multi-source data after multimodal feature processing, and training and fine-tuning the dual-channel transfer learning model; The anti-aging activity of the natural products to be tested was predicted based on the fine-tuned dual-channel transfer learning model.

2. The method according to claim 1, characterized in that The collecting of multi-source data and standardizing the multi-source data includes: Acquire a traditional Chinese medicine data source and a western medicine data source, and preprocess the traditional Chinese medicine data source and the western medicine data source; Screening positive samples of traditional Chinese medicine and positive samples of Western medicine, and constructing a data set based on the positive samples of traditional Chinese medicine and the positive samples of Western medicine; The data set is divided into a training set and a test set; wherein the training set is used for model training, and the test set is used for model verification.

3. The method according to claim 1, characterized in that The performing multimodal feature processing on the standardized multi-source data includes: generating a molecular 3D conformation and extracting molecular geometric features based on the molecular 3D conformation; Constructing target network features and determining target feature vectors based on the target network features; The molecular geometric features and the target feature vector are spliced ​​and input into the fully connected layer, and the fused multimodal feature vector is output after dimensionality reduction processing.

4. The method according to claim 1, wherein The step of constructing a dual-channel transfer learning model based on the multi-source data after multimodal feature processing, and training and fine-tuning the dual-channel transfer learning model includes: Pre-train channels based on the multimodal feature matrix of Western medicine; Fine-tune the channel according to the multimodal feature matrix of traditional Chinese medicine; Use Western medicine data to pre-train the model to optimize the anti-aging activity classification task, and stop training when the validation set AUC reaches a preset value; Freeze the first two layers of the pre-trained model and fine-tune the last layer using a learning rate decay strategy.

5. The method according to claim 1, characterized in that The method of predicting the anti-aging activity of the natural product to be tested based on the fine-tuned dual-channel transfer learning model includes: The SMILES structure of the natural product to be tested is input into the fine-tuned dual-channel transfer learning model to generate a 3D conformation and extract molecular geometric features; Query the target information of the natural product to be tested and calculate the target characteristics; The molecular geometric features and the target features are spliced ​​and input into a fully connected layer for calculation, and a fused multimodal feature vector is output; Inputting the multimodal feature vector into the fine-tuned dual-channel transfer learning model to output the anti-aging activity probability; Output the attention weights of key atomic fragments and visualize the association rules between the key atomic fragments and key targets.

6. The method according to claim 2, characterized in that The obtaining of a traditional Chinese medicine data source and a western medicine data source, and preprocessing the traditional Chinese medicine data source and the western medicine data source include: The molecular structures, target information, and anti-aging activity data of 33,643 traditional Chinese medicine ingredients were extracted from the Herb database as the traditional Chinese medicine data source; 17,458 Western medicines were extracted from the DrugBank database, and the corresponding molecular structures, targets, and efficacy annotations were obtained as Western medicine data sources; Remove entries with incomplete molecular structures and / or unlabeled targets from the traditional Chinese medicine data source and the Western medicine data source; Performing unified target naming on entries in the traditional Chinese medicine data source and the Western medicine data source; The screening of positive samples of traditional Chinese medicine and positive samples of western medicine includes: Select drugs that target core aging pathways, drugs marked as "senolytic" or "antiaging" in DrugBank, or drugs with clear records in the DrugAge or Senolytic Compounds database as positive Western medicine samples; Drugs that share at least one target with the Western medicine positive sample, or drugs with ingredients that are clearly reported in literature to have anti-aging activity are selected as traditional Chinese medicine positive samples.

7. The method according to claim 3, characterized in that The generating of the molecular 3D conformation and extracting the molecular geometric features according to the molecular 3D conformation comprises: Use OpenBabel to perform conformational search on each molecule to optimize it to the lowest energy state and output a 3D molecular structure file containing atomic coordinates; Inputting the 3D molecular structure file into a pre-trained 3D graph attention network to perform atomic-level feature encoding and interatomic spatial distance weight calculation; wherein the atomic-level feature encoding includes encoding of atom type, hybridization state, and formal charge; Output the geometric characteristics of each molecule.

8. The method according to claim 3, characterized in that The step of constructing target network features and determining target feature vectors based on the target network features includes: Construct a PPI network of targets based on the STRING database; Calculate the node centrality score using NetworkX; the node centrality score is used to measure the hubness of the target in the aging pathway and characterize the efficiency of target information transmission; Integrate the co-expression intensity data of target genes and aging markers in the GTEx database; The target pathway enrichment p-value was obtained through the Reactome API; The centrality scores, co-expression intensity data and pathway p-values ​​were concatenated into a target feature vector.

9. The method according to claim 4, characterized in that The pre-training channel according to the multimodal feature matrix of Western medicine includes: Inputting the multimodal feature matrix of western medicine into the network structure is a 3-layer fully connected neural network, the loss function is a cross entropy loss function, the training parameters are a learning rate of 0.001, a batch size of 128, and the network is trained for 50 rounds to obtain the pre-training channel; The fine-tuning of the channel according to the multimodal feature matrix of traditional Chinese medicine includes: Parameter migration, to load the weights of the pre-trained model as initialization parameters; Stratified sampling to ensure that each target family is evenly distributed in the training set; Comparative loss function to constrain the similarity of molecular features within the same target family.

10. A natural product anti-aging ingredient screening system based on multimodal deep learning, characterized by: The system includes: an acquisition module, a processing module, a construction module and a prediction module; wherein, The acquisition module collects multi-source data and performs standardization processing on the multi-source data; The processing module performs multimodal feature processing on the multi-source data after the standardization process; The construction module constructs a dual-channel transfer learning model based on the multi-source data after multimodal feature processing, and trains and fine-tunes the dual-channel transfer learning model; The prediction module predicts the anti-aging activity of the natural product to be tested based on the fine-tuned dual-channel transfer learning model.