Rapid screening method for active ingredients of apocynum venetum based on deep learning
By constructing a deep learning-based graph neural network model and combining graph structure modeling and data augmentation strategies, the problem of rapid, low-cost, and automated screening of active ingredients in Apocynum venetum was solved, achieving efficient new drug discovery.
Patent Information
- Application Number
- CN202510975089.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-11-07
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies are insufficient for rapidly, cost-effectively, and automatically screening the active ingredients of Chinese medicinal plants such as Apocynum venetum. This results in traditional methods being cumbersome, time-consuming, and requiring large amounts of reagents, and makes it difficult to achieve large-scale and rapid screening of components in complex plant extracts.
A deep learning-based graph neural network model was constructed, combining graph structure modeling and data augmentation mechanisms. Molecular structure data was converted into feature vectors, and a deep learning model was built to predict active ingredients. Highly active candidate ingredients were verified through in vitro experiments.
This enables rapid, low-cost, and automated screening of active ingredients from Apocynum venetum, improving the efficiency of new drug discovery, shortening the screening cycle, and enhancing the robustness and generalization ability of the model.
Smart Images

Figure CN120913692A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of drug screening methods, in particular to a rapid screening method for active ingredients of Apocynum venetum based on deep learning. BACKGROUND
[0002] Apocynum venetum is a traditional Chinese medicinal material, widely distributed in northwest China and Inner Mongolia, etc., with heat-clearing and tranquilizing, antihypertensive and antioxidant activities. Modern pharmacological studies have shown that Apocynum venetum is rich in flavonoids, phenylpropanoids, organic acids and various trace elements, etc. active ingredients, with good free radical scavenging ability and neuroprotective effect, especially in antioxidant, hypolipidemic, anti-inflammatory and other aspects with important development potential.
[0003] Currently, the research on active ingredients of Apocynum venetum mainly relies on traditional means such as laboratory extraction, separation and purification and in vitro activity screening. This method is tedious, long cycle and large reagent consumption, and it is difficult to realize large-scale and rapid screening of components in complex plant extracts, which significantly restricts the research on active ingredients and mechanism of natural medicines.
[0004] With the development of artificial intelligence, especially deep learning technology, chemical informatics and machine learning methods have been widely used in small molecule drug screening, activity prediction and efficacy evaluation. In recent years, graph neural network (GNN) has gradually become one of the mainstream models for predicting the properties of drug molecules, because it can effectively model the complex relationship between nodes and edges in the structure graph of molecules. However, most of the existing methods are applied to synthetic drug molecule libraries, and the adaptability to the structural heterogeneity and data sparsity of natural plant complex components is insufficient, and there is a lack of systematic modeling and specificity enhancement strategies for active ingredients of traditional Chinese medicine.
[0005] Therefore, it is urgent to construct an efficient screening method for natural products of traditional Chinese medicine, which can combine graph structure modeling, deep learning prediction and data enhancement mechanism, especially for rapid, low-cost and automated prediction and optimization of active ingredients in medicinal plants such as Apocynum venetum, in order to improve the efficiency of new drug discovery and support the research on the material basis of drug efficacy. SUMMARY
[0006] The present application aims to provide a rapid screening method for active ingredients of Apocynum venetum based on deep learning, to solve the problems in the background art.
[0007] To achieve the above-mentioned purpose, the present application provides the following technical scheme:
[0008] A rapid screening method for active ingredients of Apocynum venetum based on deep learning, comprising the following steps:
[0009] (1) Data collection and preprocessing: Obtain the chemical composition set of Apocynum venetum C={m1,m2,…,m N Molecular structure data and bioactivity data of Y = {y} m |m∈C}, the molecular structure is converted into a feature vector F using cheminformatics tools. m , where m∈C,y m This represents the true bioactivity value of molecule m.
[0010] (2) Deep learning model construction: Constructing a deep learning model f(F) m ;θ), the input is the molecular feature vector F m The output is the predicted activity value. Where θ represents the model parameters;
[0011] (3) Model training and optimization: Training dataset based on publicly available chemical databases Train the model and adjust the parameter θ by optimizing the loss function;
[0012] (4) Screening of active ingredients: The feature vector F of Apocynum venetum chemical components is used to screen the active ingredients. m Input model to predict activity value Screening a set of highly active candidate ingredients
[0013] (5) Experimental verification: In vitro experiments were conducted on molecules in the candidate component set S to verify their bioactivity y. m .
[0014] As a preferred technical solution of the present invention, the molecular structure data is represented by SMILES encoding or molecular fingerprint, preferably Morgan fingerprint, defined as:
[0015] F m ={f i |F i ∈{0,1},i=1,2,…,K}
[0016] Among them, F m Let f be the eigenvector of the molecule m∈C. i Let K represent the i-th fingerprint feature, and K be the fingerprint length.
[0017] As a preferred technical solution of the present invention, the deep learning model adopts a convolutional neural network, and its feature extraction process is as follows:
[0018] h l =σ(W l ·h l-1 +b l )
[0019] Among them, hl is the output of the first hidden layer, h0=F m , m is the eigenvector of the molecule m∈C, W l is the weight matrix of the first layer, b l is the bias vector, and σ is the ReLU activation function, and the output h L is used to generate the predicted activity value
[0020] As a preferred technical solution of the present application, the deep learning model can be replaced by a graph neural network, and the node update formula is:
[0021]
[0022] wherein, is the eigenvector of node i in the kth layer, and the initial node eigenvector is derived from F m , is the neighbor node set of node i, AGGREGATE is the mean aggregation function, W k and b k are the weight matrix and bias, respectively, σ is the ReLU function, and the final node eigenvector is generated by aggregation
[0023] As a preferred technical solution of the present application, the model training adopts a mean square error loss function, which is defined as:
[0024]
[0025] wherein, L(θ) is the loss function, M is the number of samples in the training set D, y m is the true activity value of the molecule m, is the model predicted activity value, and the optimization objective is to minimize L(θ), and the parameters θ are updated by the Adam optimizer.
[0026] As a preferred technical solution of the present application, the training process introduces L2 regularization, and the optimized loss function is:
[0027]
[0028] wherein, L reg (θ) is the regularization loss, L(θ) is the loss function, λ is the regularization coefficient (preferably 0.01), W l is the weight matrix of the first layer, which prevents the model from overfitting to the training set D.
[0029] As a preferred technical solution of the present application, the active ingredient screening step screens high-activity ingredients by threshold τ:
[0030]
[0031] wherein S is a candidate component set, is the predicted activity value in claim 1, C is a chemical component set of Apocynum venetum, and τ is set according to the type of biological activity (such as IC50<50μM), and the screening result is directly used for in vitro verification.
[0032] As a preferred technical solution of the present application, the in vitro experimental verification includes one or more of the following tests: antioxidant activity test, based on DPPH free radical scavenging rate, the calculation formula is:
[0033]
[0034] wherein A control , A sample are the absorbance of the control group and the sample, respectively; or antihypertensive activity test, based on ACE inhibition rate; or anti-inflammatory activity test, based on NO production inhibition rate, to verify the biological activity y m of the candidate component m∈S.
[0035] As a preferred technical solution of the present application, the method includes a data enhancement step to generate a virtual molecule structure to expand the training set:
[0036] D aug =D∪{(F m′ ,y m′ )∣m′=T(m),m∈C}
[0037] wherein D aug is the enhanced training set, D is the training data set in the public chemical database, T is a chemical transformation function, F m′ , y m′ are the feature vector and the predicted activity value of the virtual molecule m', respectively.
[0038] Compared with the prior art, the present application has the following beneficial effects:
[0039] The present application uses a graph neural network model to deeply model the structural features of various chemical components in Apocynum venetum, which can fully extract high-order information contained in atoms and their adjacent relationships in molecules, improve the expression ability of complex natural product structures, and has stronger structure representation ability and prediction accuracy compared with traditional fingerprint features or one-dimensional molecular description methods.
[0040] The present application introduces a supervised learning mechanism, trains the prediction model combined with existing biological activity data, can quickly evaluate the activity strength of unknown components without relying on a large number of in vitro experiments, greatly shortens the activity screening period, reduces the consumption of experimental resources, and improves the efficiency of natural product research.
[0041] The application further introduces a molecular data enhancement strategy based on structural disturbance, effectively alleviates the problems of insufficient traditional Chinese medicine ingredient samples and sparse labels by constructing virtual molecules and activity pairs to enhance the training sample size, and improves the model robustness and generalization ability, and is suitable for rapid high-throughput activity screening of diverse compounds in natural plants such as apocynum venetum. BRIEF DESCRIPTION OF DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the description of the embodiments of the present application or the prior art will be briefly introduced and explained below. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creating any creative labor.
[0043] Figure 1 is a flowchart of a rapid screening method for active ingredients of apocynum venetum based on deep learning of the present application. DETAILED DESCRIPTION
[0044] In order to further illustrate the rapid screening method for active ingredients of apocynum venetum based on deep learning of the present application, the specific implementation method will be described in detail below. These embodiments are only used to illustrate the technical solutions of the present application, and do not limit the protection scope. Equivalent replacements or improvements made by those skilled in the art without departing from the technical solutions of the present application all belong to the protection scope of the present application.
[0045] The present application will be further described in detail below in combination with the drawings and embodiments, but the protection scope of the present application is not limited to the following contents.
[0046] Embodiment: The rapid screening method for active ingredients of apocynum venetum based on deep learning provided by the present application is to accurately predict the activity value of the chemical components of apocynum venetum by constructing an efficient graph neural network model, and then screen potential high-activity ingredients to improve the efficiency of new drug development. The present embodiment details the complete technical process from data construction, model training to activity screening, as follows:
[0047] First, the molecular set C = {m1, m2,..., m n} of the chemical components of apocynum venetum is structure-encoded, the standardized molecular structure representation method (such as SMILES) is used, and the graph learning tool is used to convert it into a node feature vector and an adjacency matrix to construct a feature data set Y = {y m | m e C}, wherein y m is the biological activity experimental value of the compound m, indicating the inhibition or activation effect on the target protein target point. Each compound structure is converted into a graph representation G m = (C m , Am ), where represents the node feature matrix, A m ∈{0, 1} n×n is the adjacency matrix. The graph convolution operation extracts the structural semantic features by the following formula:
[0048] h1=σ(W1·h0+b1);
[0049] where h0=F m is the input molecular structure feature vector; W1is the weight matrix, b1is the bias term, and the activation function uses ReLU.
[0050] Further, a graph attention mechanism is introduced to model the different contributions of each atomic neighbor node to the center node, and the update method is:
[0051]
[0052] where, is the last layer feature of the adjacent node, is the neighbor set of node v, and AGGREGATE represents the neighbor feature aggregation operation.
[0053] In the above network structure training process, the model loss function is defined as the mean square error loss:
[0054]
[0055] where, is the model predicted activity value, and M is the number of samples in the training set. The Adam optimizer is used to iteratively update the model parameters.
[0056] To improve the generalization ability of the model, an L2 regularization term is introduced into the loss function, and the total loss function is as follows:
[0057]
[0058] where λ is the regularization coefficient, L is the number of network layers, and W l is the weight parameter of the lth layer.
[0059] The inhibition rate is used as an indicator for activity evaluation criteria, and the in vitro experimental results are verified, and the inhibition rate calculation formula is as follows:
[0060]
[0061] where, control , A sample represent the absorbance values of the control group and the sample group, respectively.
[0062] In addition, in order to expand the training data and improve the robustness of the model, a data enhancement strategy is introduced. Through a chemical transformation function T(·), a virtual structure m' is generated from the original molecular structure m e C, forming an enhanced data set:
[0063] D aug =D∪{(F m ,y m )∣m′=T(m),m∈C};
[0064] Where F m , y m are the feature representation and predicted activity value of the virtual molecule m' respectively.
[0065] Finally, Morgan fingerprint is used to encode the features of all compounds, and the fingerprint length K is set. The feature representation of each molecule is:
[0066] F m ={f i |f i ∈{0,1},i=1,2,...,K};
[0067] After the model training is completed, the trained deep learning model is applied to the activity prediction of new compound structures. For unknown molecules m * Input its graph structure, output its predicted activity value If it meets the activity threshold judgment standard (such as IC50<50μm), it is determined as a potential active ingredient, and enters the subsequent verification stage.
[0068] Through the implementation of the present application, the screening efficiency and accuracy of active compounds in apocynum venetum components can be effectively improved, the experimental cost is reduced, and the new drug development process is accelerated.
[0069] In order to verify the effectiveness of the method of the present application, the chemical components extracted from the leaves of apocynum venetum are taken as the research object, and a neural network model is constructed to predict the antioxidant activity, and the target is to quickly screen out candidate components with high DPPH free radical scavenging ability.
[0070] 1, original data construction: collect the SMILES structural formula of 30 main chemical components in apocynum venetum leaves, denoted as set C
[0071] C={m1,m2,...,m 30};
[0072] Use RDKit tool to convert SMILES representation to molecular graph structure G m =(C m ,A m ), where is the node (atom) feature matrix, and Am ∈{0,1} n×n is the adjacency matrix (bond connection). At the same time, the corresponding antioxidant activity experimental value y (quantified as the reciprocal of IC50 as the activity value) is extracted from the literature and experimental database to construct the training data set m
[0073] 2. Feature extraction and model training: Use Morgan fingerprint (radius = 2, length K = 2048 bits) to extract the structure fingerprint vector of each molecule:
[0074] F m ={f i ∈{0,1},i=1,...,2048};
[0075] Input the features into the graph neural network, and the first layer is a graph convolution layer:
[0076] h1=σ(W1·h0+b1),h0=F m ;
[0077] The second layer adopts the graph attention network mechanism:
[0078]
[0079] 3. Loss function and optimization strategy: The model loss function adopts mean square error loss + L2 regularization:
[0080]
[0081] Where, the number of training set samples M = 30, and the Adam optimizer is used for parameter iterative update.
[0082] 4. Data augmentation processing: Use chemical transformation rules T(·) (such as hydroxymethylation, halogen substitution, etc.) to generate virtual structures for the original structure, and construct an enhanced training set:
[0083] D aug =D∪{(F m ,y m )∣m′=T(m),m∈C};
[0084] The final training set is expanded to 90 samples.
[0085] 5. Activity prediction and screening verification: After the model training is completed, the newly discovered 5 compounds m * ∈C new in the apocynum venetum extract are predicted, and the activity value is output. If the normalized activity is higher than the threshold value, the candidate high-activity component is determined. Compounds M15, M22 and M29 are selected for DPPH free radical scavenging experiment in vitro, and the inhibition rate is calculated as follows:
[0086]
[0087] Wherein:
[0088] The experimental result of M22 is A control = 0.72, The result is highly consistent with the model prediction value which verifies the reliability and practicability of the model.
[0089] The contents not described in detail in the specification belong to the prior art known to those skilled in the art, although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments or make equivalent replacement for part of the technical features, any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A deep learning-based rapid screening method for active ingredients of Apocynum venetum, characterized in that, Comprising the following steps: (1) Data collection and pretreatment: Obtain the molecular structure data and biological activity data Y = {y m ∣m∈C} of the Apocynum venetum chemical component set C = {m1,m2,…,m m n}, and convert the molecular structure into a feature vector F m by using a chemical informatics tool, where m∈C,y m is the true biological activity value of the molecule m; (2) Deep learning model construction: construct a deep learning model f(F m ; θ) with the input of molecular feature vector F m and the output of predicted activity value where θ is the model parameter; (3) Model training and optimization: based on the training dataset in the public chemical database Model training is performed, and the parameters θ are adjusted by optimizing the loss function; (4) Active ingredient screening: the characteristic vector F of the chemical components of Apocynum venetum L. is inputted into the model to predict the active value m Inputting the model, predicting the active value Screening the high active candidate component set (5) Experimental verification: in vitro experiments are performed on the molecules in the candidate component set S to verify their biological activity y m .
2. The method of claim 1, wherein, The molecular structure data is defined by SMILES encoding or molecular fingerprint representation, preferably Morgan fingerprint, as follows: F m = {f i | f i ∈ {0,1}, i = 1,2,…,K} where F m is the feature vector of the molecule m e C, f i represents the i-th bit of the fingerprint feature, K is the length of the fingerprint, preferably 2048, and the feature vector F m is the input of the deep learning model of claim 1.
3. The method according to claim 1, wherein, The deep learning model adopts a convolutional neural network, The feature extraction process is as follows: h l = σ(W l · h l-1 + b l ) where h l is the output of the l-th hidden layer, h0= F m , F m is the feature vector of molecule m e C, W l is the l-th layer weight matrix, b l is the bias vector, and s is the ReLU activation function, with output h L is used to generate the predicted activity value 4. The method according to claim 1, wherein the method is characterized by, The deep learning model can be replaced by a graph neural network, and the node update formula is as follows: where, is the feature vector of node i in the k-th layer, the initial node features are from F m derived, is the set of neighbor nodes of node i, AGGREGATE is the mean aggregation function, W k , b k are the weight matrix and bias respectively, σ is the ReLU function, and the final node features are summarized to generate 5. The method according to claim 1, wherein the method is characterized by, The model training adopts a mean square error loss function, defined as follows: where L(θ) is the loss function, M is the number of samples in the training set D, y m is the true activity value of molecule m, is the model predicted activity value, and the optimization objective is to minimize L(θ) by updating the parameters θ with an Adam optimizer.
6. The method according to claim 5, wherein the method is characterized by, The training process introduces L2 regularization, and the optimized loss function is as follows: wherein L reg (θ) is a regularization loss, L(θ) is a loss function, λ is a regularization coefficient, W l is the weight matrix of the l-th layer, which prevents the model from overfitting the training set D.
7. The method according to claim 1, wherein the method is characterized by, The active ingredient screening step screens high-activity ingredients by a threshold τ: Wherein, S is the candidate component set, The activity value predicted by the model in claim 1, C is the chemical component set of Apocynum venetum L, τ is set according to the type of biological activity (such as IC50<50 μM), and the screening result is directly used for in vitro verification.
8. The method according to claim 1, wherein the method is characterized by, The in vitro experimental verification includes one or more of the following tests: antioxidant activity test, based on DPPH free radical scavenging rate, and the calculation formula is as follows: where A control , A sample are the absorbances of the control and sample, respectively; or the antihypertensive activity test based on ACE inhibition rate; or the anti-inflammatory activity test based on NO production inhibition rate, verifying the biological activity y m of the candidate component m ∈ S.
9. The method according to claim 1, wherein the method is characterized by, The method comprises a data enhancement step to generate virtual molecular structures to expand the training set: D aug = D U {(F m′ , y m′ ) | m' = T(m), m e C} where D aug is the augmented training set, D is the training dataset in the public chemical database, T is the chemical transformation function, F m′ , y m′ are the feature vector and the predicted activity value of the virtual molecule m ′ , respectively.