A pathway-embedded traditional Chinese medicine efficacy substance synergistic effect prediction method and system
By constructing a pathway-embedded predictive model for the synergistic effects of active substances in traditional Chinese medicine (TCM), and utilizing graph convolutional neural networks and deep Q-network algorithms, the problem of predicting the synergistic effects of active substances in TCM was solved, achieving efficient and accurate drug screening and prediction, and improving the accuracy and efficiency of screening and prediction of active substances in TCM.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies are insufficient to efficiently and accurately predict the synergistic effects of active substances in traditional Chinese medicine (TCM). They lack systematic data analysis and mining capabilities, fail to fully utilize modern information technology, and the complexity of TCM components leads to time-consuming and costly experimental methods, making it difficult to fully reveal the interactions between components and their comprehensive impact on organisms.
A drug efficacy prediction model is constructed using pathway embedding technology, graph convolutional neural networks, and deep Q-network algorithms. By preprocessing drug molecular structure data, constructing the drug-target correlation relationship, fusing the adjacency matrix of target pathways, and training with deep learning, a drug synergistic effect prediction model is generated.
It has enabled accurate, efficient and comprehensive prediction of the synergistic effects of active substances in traditional Chinese medicine, improved the accuracy of molecular structure characterization and information retention, enhanced the ability to predict target interactions, and realized the automation and intelligence of drug screening and prediction.
Smart Images

Figure CN120148904B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of predicting the synergistic effects of active substances in traditional Chinese medicine, specifically to a method and system for predicting the synergistic effects of active substances in traditional Chinese medicine with pathway embedding. Background Technology
[0002] In the field of traditional Chinese medicine (TCM) research and development, the compatibility of TCM herbs is the most unique aspect of the TCM theoretical system, representing its strength and characteristic. How to combine and match medicinal substances to achieve efficacy equal to or better than the original formula, and to produce synergistic effects from the original medicinal substances, is crucial. Traditional research on the combination of medicinal substances in TCM mainly relies on laboratory bioactivity testing. While these methods have a certain experimental basis, the process of combining various substances is time-consuming, costly, and difficult to cover the chemical spectrum of TCM medicinal substances.
[0003] Existing methods for studying the synergistic effects of active substances in traditional Chinese medicine (TCM) face several major challenges: First, the components of TCM are complex, and single experimental methods are insufficient to fully reveal the interactions between components and their comprehensive effects on organisms; second, traditional methods typically lack systematic data analysis and mining capabilities, making it difficult to extract useful information from massive experimental data, thus affecting the accuracy and efficiency of efficacy prediction; third, the interactions between active substances in TCM and biological targets are complex, and there is a lack of effective computational models to support the prediction of these complex interactions; finally, existing technologies often fail to fully utilize modern information technologies, such as artificial intelligence and machine learning, to optimize the screening and prediction process of active substances.
[0004] Therefore, in response to the above problems and technical challenges, there is an urgent need to provide a new method and system for predicting the synergistic effects of active substances in traditional Chinese medicine, so as to achieve accurate and efficient screening of active substances and prediction of their effects. Summary of the Invention
[0005] To address the problems existing in the prior art, this invention provides a method and system for predicting the synergistic effects of pharmacodynamic substances in traditional Chinese medicine through pathway embedding. This invention integrates pathway embedding technology, graph convolutional neural networks, deep Q-network algorithms, and deep learning training technologies to construct a more accurate and efficient efficacy prediction model, achieving accurate, efficient, and comprehensive prediction of the synergistic effects of pharmacodynamic substances in traditional Chinese medicine.
[0006] To achieve the above technical objectives, the present invention adopts the following technical solution:
[0007] This invention provides a method for predicting the synergistic effect of pathway-embedded pharmacodynamic substances in traditional Chinese medicine, the method comprising the following steps:
[0008] Step S1: Preprocess the drug molecule structure data to generate a representation vector of the drug molecule;
[0009] Step S2: Based on the target database information and the representation vector of drug molecules, construct the association relationship between the drug and each target;
[0010] Step S3: Use the collected data on the relationship between targets and biological pathways to generate an adjacency matrix of multiple target pathways. Then, fuse the adjacency matrix with the obtained association between the drug and each target to obtain the embedding feature of each target.
[0011] Step S4: Use the obtained target embedding features for deep learning training to generate a drug synergistic effect prediction model;
[0012] Step S5: Use the drug synergy prediction model to directly present the synergy prediction results of the drug combination.
[0013] Preferably, step S1 specifically includes the following steps:
[0014] Step S11: Convert the drug molecule structure data represented by SMILES into a graph topology to form a molecular graph;
[0015] Step S12: Using the node degree heuristic, prioritize the nodes in the molecular graph using the deep Q-network algorithm;
[0016] Step S13: Based on the node priority ranking results and the node features of the molecular graph, selectively update the state of the nodes using a graph convolutional neural network;
[0017] Step S14: Extract features from nodes based on graph convolutional neural networks to generate representation vectors of drug molecules.
[0018] Preferably, step S2 is as follows: based on the information in the target database, using the feature vector of the drug, the feature association and regularization constraint of the target features are calculated through the self-attention mechanism to generate the interaction probability between the drug and the specific biological target, so as to construct the association relationship between the drug and each target.
[0019] Preferably, step S3 specifically includes the following steps:
[0020] Step S31: Collect the relationship data of target points in different biological pathways, regard multiple target points as graph nodes, construct the association relationship between target points under each pathway between graph nodes, and map the association relationship data into an adjacency matrix of multiple target pathways;
[0021] Step S32: The adjacency matrix of multiple target pathways and the obtained drug-target association relationship are fused into a multi-level graph embedding representation. In this process, a graph convolutional neural network is used to embed pathway information to obtain target embedding features.
[0022] Preferably, step S4 specifically includes the following steps:
[0023] Step S41: Fuse the target embedding features of a pair of molecular structures and map them to the synergistic effect score space of the molecular pair, and calculate the synergistic effect score.
[0024] Step S42: Based on the obtained synergy score, train the model using a deep learning training algorithm to generate a drug synergy prediction model.
[0025] The present invention also provides a pathway-embedded synergistic effect prediction system for traditional Chinese medicine active substances, which includes a molecular structure characterization module, a drug-target interaction module, a target-pathway information embedding module, an end-to-end training module, and a visualization module.
[0026] The molecular structure characterization module is used for preprocessing drug molecular structure data;
[0027] The drug-target interaction module generates representation vectors of drug molecules and is used to construct the association between drugs and various targets.
[0028] The target-pathway information embedding module is used to generate multiple target-pathway adjacency matrices to describe the relationships between targets.
[0029] An end-to-end training module is used to generate a drug synergy prediction model to predict the synergistic effects of drug combinations.
[0030] The visualization module presents the predicted results of the synergistic effects of drug combinations.
[0031] Preferably, the molecular structure characterization module includes a data input unit and a node priority sorting unit;
[0032] The data input unit converts the input drug molecule SMILES data into a graph topology structure, forming a molecular graph;
[0033] The node priority sorting unit uses the node degree heuristic index to sort the nodes in the molecular graph by priority.
[0034] Preferably, the drug-target interaction module includes a graph convolution calculation unit, a vector generation unit, and an interaction prediction unit;
[0035] The graph convolution computation unit updates the state of each node using a selective graph convolution algorithm;
[0036] Vector generation unit, used to generate the final representation vector of drug molecules, providing molecular structure characterization;
[0037] The interaction prediction unit predicts the probability of interaction between the drug and the target by combining information from the target database.
[0038] Preferably, the target-pathway information embedding module includes a pathway adjacency matrix generation unit and a multi-graph fusion convolution unit;
[0039] The pathway adjacency matrix generation unit maps the relationships between biological targets to an adjacency matrix of target pathways, describing the correlation of targets in different pathways.
[0040] The multi-graph fusion convolutional unit uses multiple target points as graph nodes, constructs a graph structure using the adjacency matrix of multiple target point pathways, and embeds pathway information using multi-graph convolution operations to obtain target point embedding features, thereby achieving multi-level information fusion of the association between target points.
[0041] Preferably, the end-to-end training module includes a target-disease association mapping unit and a deep learning training unit;
[0042] The target-disease association mapping unit fuses the target embedding features of a pair of molecular structures and maps them to the synergistic effect score space of the molecular pair.
[0043] The deep learning training unit trains the model using deep learning training algorithms based on the data output by the target-disease association mapping unit, generating the final predictive model of the synergistic effect of traditional Chinese medicine active substances.
[0044] Compared with the prior art, the beneficial effects of the present invention are:
[0045] 1) To achieve efficient input and preprocessing of traditional Chinese medicine molecular data, this invention converts SMILES data into a graph topology structure, prioritizes the nodes in the molecule, and uses a graph convolutional neural network algorithm to extract features from the nodes, generating representation vectors of drug molecules. This effectively captures long-range dependencies and provides efficient and rich molecular structure representation. Compared with previous methods, it can significantly improve the accuracy of molecular structure representation and information retention.
[0046] 2) Enhanced target interaction prediction capability: Based on the generated drug molecule representation vector and combined with information from the target database, this invention accurately predicts the interaction probability between drugs and specific biological targets, providing high-quality data support for subsequent pathway analysis, thereby enhancing the effectiveness of predicting the synergistic effect of pharmacodynamic substances.
[0047] 3) Achieving multi-pathway information embedding: This invention constructs an adjacency matrix of multiple target pathways to effectively integrate the information of targets and biological pathways, establishes the association between targets and pathways, and uses a graph convolutional neural network to embed pathway information to obtain target embedding features. By embedding multi-pathway information, a comprehensive description of the association and effects between targets can be achieved. This can effectively integrate pathway-related multi-target information, improve the screening and prediction accuracy of pharmacodynamic substances, and thus improve the depth and accuracy of synergistic effect prediction.
[0048] 4) Intelligent and automated data analysis: This invention integrates graph convolutional neural networks and deep Q-network algorithms to achieve full-process automation from data input and structural representation to synergistic effect prediction, which greatly improves the efficiency of drug screening and prediction, reduces the need for human intervention, and realizes the automation and intelligence of screening of active substances in traditional Chinese medicine.
[0049] 5) The system is refined and intelligent. By combining a variety of advanced information technologies, such as graph convolutional neural networks, multi-graph fusion convolutions, and multi-disease deep learning, this invention provides a refined and intelligent method and system for predicting drug efficacy. It provides strong technical support for screening the synergistic effects of active substances in traditional Chinese medicine and has significant application value and promotion potential. Attached Figure Description
[0050] Figure 1 This is a flowchart of a method for predicting the synergistic effect of pathway-embedded pharmacodynamic substances in traditional Chinese medicine according to an embodiment of the present invention;
[0051] Figure 2 This is a schematic diagram of a pathway-embedded synergistic effect prediction system module for traditional Chinese medicine active substances, according to an embodiment of the present invention. Detailed Implementation
[0052] The following will refer to the appendices in the embodiments of the present invention. Figure 1 and 2 The technical solutions in the embodiments of the present invention are clearly and completely described herein. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0053] In the description of this invention, it should be understood that the terms "coaxial," "bottom," "one end," "top," "middle," "other end," "upper," "side," "top," "inner," "front," "center," and "both ends," etc., indicate the orientation or positional relationship based on the appendix. Figure 1 and 2 The orientations or positional relationships shown are for the purpose of facilitating and simplifying the description of the present invention, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as limiting the present invention.
[0054] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "setting," "connection," "fixing," "screw connection," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal connection of two components or the interaction between two components. Unless otherwise explicitly limited, those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0055] Example 1
[0056] Combination Figure 1 As shown in the figure, this embodiment provides a method for predicting the synergistic effect of pathway-embedded pharmacodynamic substances in traditional Chinese medicine, including the following steps:
[0057] Step S1: Preprocess the drug molecule structure data to generate a representation vector of the drug molecule;
[0058] Step S2: Based on the target database information and the representation vector of drug molecules, construct the association relationship between the drug and each target;
[0059] Step S3: Use the collected data on the relationship between targets and biological pathways to generate an adjacency matrix of multiple target pathways. Then, fuse the adjacency matrix with the obtained association between the drug and each target to obtain the embedding feature of each target.
[0060] Step S4: Use the obtained target embedding features for deep learning training to generate a drug synergistic effect prediction model;
[0061] Step S5: Use the drug synergy prediction model to directly present the synergy prediction results of the drug combination.
[0062] This invention enhances the predictive ability of target interactions by accurately constructing the correlation between drugs and targets, providing data support for subsequent pathway analysis, thereby enhancing the effectiveness of predicting the synergistic effects of pharmacodynamic substances. Furthermore, by constructing an adjacency matrix of multiple target pathways, this invention establishes the correlation between targets and pathways, achieving multi-pathway information embedding and providing a comprehensive description of the correlations and effects between targets. It effectively integrates pathway-related multi-target information, improving the screening and prediction accuracy of pharmacodynamic substances, thereby enhancing the depth and accuracy of synergistic effect prediction.
[0063] Example 2
[0064] Combination Figure 1 As shown in the figure, this embodiment provides a method for predicting the synergistic effect of pathway-embedded pharmacodynamic substances in traditional Chinese medicine, including the following steps:
[0065] Step S1: Preprocess the drug molecule structure data to generate a representation vector of the drug molecule;
[0066] Step S2: Based on the target database information and the representation vector of drug molecules, construct the association relationship between the drug and each target;
[0067] Step S3: Use the collected data on the relationship between targets and biological pathways to generate an adjacency matrix of multiple target pathways. Then, fuse the adjacency matrix with the obtained association between the drug and each target to obtain the embedding feature of each target.
[0068] Step S4: Use the obtained target embedding features for deep learning training to generate a drug synergistic effect prediction model;
[0069] Step S5: Use the drug synergy prediction model to directly present the synergy prediction results of the drug combination.
[0070] In this embodiment, the specific steps of step S1 are as follows:
[0071] Step S11: Convert the drug molecule structure data represented by SMILES into a graph topology to form a molecular graph. In specific implementation, input drug molecule structure data in SMILES format. A chemical library (such as RDKit) can be used to parse the SMILES string, extract each atom and the bonds between them. Atoms are abstracted as graph nodes, and the bonds between atoms are abstracted as graph edges, thereby converting the SMILES data into a graph topology to obtain the molecular graph structure G(X,A), where X represents the node feature matrix and A represents the adjacency matrix of the molecular graph.
[0072] Step S12: Using a node degree heuristic, a deep Q-network algorithm is employed to prioritize nodes in the molecular graph to determine their importance. The node degree heuristic, based on the number of edges connecting a node to other nodes, is an empirically designed metric for quickly evaluating node importance and other characteristics. Node priority is calculated using node degree, and nodes are ranked according to their priority, placing important nodes at the end of the graph sequence to ensure they receive more contextual information in subsequent feature extraction. This priority-based ranking ensures that more important nodes receive priority attention in later processing steps.
[0073] Step S13: Based on the node priority ranking and node features of the molecular graph, the state of nodes is selectively updated using a Graph Convolutional Neural Network (GCN). This process provides basic features for subsequent predictions. In the specific implementation, a GCN with selection gating is constructed, that is, a gating mechanism is added to the GCN to select which node features are important and need to be retained, and which node features can be ignored and allowed to enter the hidden state. Important, high-priority nodes are updated, while low-priority nodes are restricted from updating. This allows the model to process graph structure data more flexibly, improving the model's performance and efficiency. The formula for determining whether each node should be updated can be expressed as:
[0074]
[0075] in This is the new feature vector of node i after being updated by the Graph Convolutional Neural Network (GCN). This vector represents the updated state of node i. P(d(v) i The layer ) represents the degree feature extraction layer for node i. By inputting the degree d(vi) of node i into a fully connected layer, which consists of a set of weights and biases, it learns how to map the node degree to a new feature space. The degree feature extraction layer is a linear transformation, implemented by a linear layer and an activation function (ReLU). δ represents the hyperparameter. Let A represent the l-th layer features of node i, and let A represent the adjacency matrix. I represents the identity matrix. Let represent the normalized form of the adjacency matrix, and D represent the degree matrix. l Let represent the parameters of the l-th layer, and σ represent the activation function.
[0076] In this way, information from high-priority nodes is selectively propagated during GCN hierarchical updates, thereby increasing the influence of important nodes and avoiding the propagation of irrelevant or redundant information. Information propagation from low-priority nodes is suppressed, achieving a sparsity effect. Then, a gating mechanism selectively controls which node information enters the hidden state. By selectively propagating node information in a Graph Convolutional Network (GCN), important long-term dependencies are effectively preserved and propagated, compressing and transmitting long-distance dependencies.
[0077] Step S14: Extract features from nodes using a graph convolutional neural network to generate representation vectors for drug molecules. Specifically, node features X are further extracted using 1D convolution and activation functions of the graph convolutional neural network. v And generate the final representation vector X of the drug molecule. mThese representation vectors not only contain contextual information of key nodes, but also retain the long-range dependency characteristics filtered out in graph convolution, providing efficient and rich molecular structure characterization for subsequent drug-target interaction modules.
[0078] This invention enables efficient input and preprocessing of traditional Chinese medicine molecular data. It converts SMILES data into a graph topology, prioritizes nodes in the molecule, and uses a graph convolutional neural network algorithm to extract features from the nodes, generating representation vectors for drug molecules. This effectively captures long-range dependencies and provides efficient and rich molecular structure representation. Compared with previous methods, it can significantly improve the accuracy of molecular structure representation and information retention.
[0079] In this embodiment, step S2 specifically involves: based on information from the target database, using the drug's feature vector, calculating feature associations and regularizing target features through a self-attention mechanism to generate interaction probabilities between the drug and specific biological targets, thereby constructing associations between the drug and each target and providing foundational data for subsequent pathway embedding analysis. In a specific implementation, the feature associations of molecular representation vectors can be calculated using a self-attention mechanism, and the output of the self-attention mechanism can be normalized to provide an additional association adjacency matrix A for subsequent analysis. coor Then, consistency regularization is used to constrain the target features output by the self-attention mechanism during training. Similarity regularization is applied to self-supervised labels (such as prediction results) generated with the same target features, causing the model to automatically cluster the same target features in the feature space. Different target features are kept discrete through differential processing of the self-supervised task. In anti-pneumonia applications, data on the targets of the novel coronavirus (such as TNF-α, MAPK1, etc.) are used to predict the possible antiviral effects of traditional Chinese medicine molecules. This invention enhances the target interaction prediction capability, accurately predicts the interaction probability between drugs and specific biological targets, provides high-quality data support for subsequent pathway analysis, and thus enhances the effectiveness of predicting the synergistic effect of pharmacodynamic substances.
[0080] In this embodiment, step S3 specifically includes the following steps:
[0081] Step S31: Treat multiple targets as graph nodes, construct the association relationships between targets under each pathway between graph nodes, and map the association data into an adjacency matrix of multiple target pathways to represent the association of targets in different biological pathways. In specific implementation, relationship data of targets in different biological pathways are collected from multiple biological databases (such as KEGG, Reactome, etc.). Each pathway describes the biological interactions between different targets, such as metabolic pathways, signal transduction pathways, etc.
[0082] For each pathway, an adjacency matrix is generated based on the interaction relationships between target points. The adjacency matrix has a dimension of N×N, where N is the number of target points involved in the system. The elements in the matrix represent the relationship between target point pairs; for example, 1 indicates a direct interaction, and 0 indicates no direct interaction. The adjacency matrix is then standardized to ensure that each matrix has a consistent scale for subsequent multi-graph convolution processing.
[0083] Step S32: The adjacency matrices of multiple target pathways and the obtained drug-target association relationships are fused into a multi-level graph embedding representation. During this process, a graph convolutional neural network (GCN) is used to embed pathway information, obtaining target embedding features. These embedding features contain complex association information of targets in multiple pathways. The GCN utilizes these adjacency matrices to perform graph convolution, generating embedding features for each target, achieving multi-level information fusion of associations between targets. For specific applications in anti-pneumonia, this module can combine disease-related biological network data, including host-pathogen interaction pathways and pneumonia-related biological pathways.
[0084] The specific implementation steps of step S32 include:
[0085] Step S321: The feature information of the targets contained in the drug-target association data output from step S2 is used as input. Each target initially has a low-dimensional feature vector. This vector contains basic information about the target, such as molecular structure and biological function. Generally, the generation process of target features usually involves extracting basic information about the target (such as amino acid sequence, function, structure, etc.) and encoding and representing it using deep learning methods (graph convolutional neural networks).
[0086] Step S322: Utilize graph convolution and adjacency matrix A B (Adjacency matrix of the i-th path) and adjacency matrix A coor This allows information from adjacent nodes to be aggregated. A multi-head mechanism is employed, allowing each convolutional layer to independently compute features from different pathways, and then these features are concatenated or averaged to obtain a richer representation.
[0087] Step S323: After multiple convolutions, fuse the outputs of each convolutional layer. Weighted averaging or attention mechanisms can be used to aggregate the results of multiple convolutions into a single comprehensive feature. Attention mechanisms can assign different weights to each pathway based on its relevance to the current task, thereby enhancing the information of important pathways. After the above steps, embed the output target points into feature X. bThis target embedding feature representation contains multi-level association information of the target in different pathways. This representation not only reflects the basic characteristics of the target but also includes its complex network relationships in multiple biological pathways. The representation of the target embedding feature for pneumonia will reflect its role in disease transmission-related pathways.
[0088] This invention embeds multi-pathway information, improving the accuracy of drug substance screening and prediction by establishing correlations between targets and pathways. It achieves a comprehensive description of the relationships and effects between targets, effectively integrating pathway-related multi-target information, thereby enhancing the depth and accuracy of synergistic effect prediction.
[0089] In this embodiment, the specific steps of step S4 are as follows:
[0090] Step S41: Fuse the target embedding features of a pair of molecular structures and map them to the synergy score space of the molecular pair, then calculate the synergy score. In specific implementation, the target embedding feature X output from step S32 is received. b A multilayer perceptron (MLP) is used to map the target embedding features of a pair of disease-associated molecular structures to a synergistic score space, and the synergistic score is calculated using the MLP. The calculation method can be expressed as follows:
[0091]
[0092] Where Linear represents a linear layer, and cat represents a feature cascade. and This is represented as a target embedding feature of a pair of molecular structures.
[0093] In the specific implementation, the paired molecular graph structure of the paired molecular structure data is obtained through the previous step S1. A two-stream parallel network structure is constructed to learn the target embedding features of each molecular structure data. Finally, the paired target embedding features obtained from the paired molecular structure data are fused and mapped to the synergistic effect score space of the molecular pair, and the synergistic effect score is output.
[0094] Step S42: Based on the obtained synergistic effect score, a model is trained using a deep learning training algorithm to generate a drug synergistic effect prediction model. This synergistic effect score is compared with the target label, and the model parameters are optimized using gradient descent to achieve end-to-end training. After training, a synergistic effect prediction model of traditional Chinese medicine active substances is obtained. During training, antiviral activity data of single drugs related to pneumonia and synergistic data of drug combinations are introduced to improve the antiviral prediction ability for the novel coronavirus. Finally, the model, trained on pneumonia data, can be used to predict potential synergistic anti-pneumonia treatment combinations.
[0095] In this embodiment, in step S5, when using the drug synergy prediction model for prediction, the prediction results are displayed in the form of graphics or virtual reality through data visualization and virtual reality technology, directly presenting the synergy prediction results of the drug combination, which is convenient for users to understand and further analyze.
[0096] Example 3
[0097] This embodiment, based on Embodiment 2, specifically illustrates the steps involved in prioritizing nodes in the molecular graph using a Deep Q-Network algorithm in step S12. The Deep Q-Network (DQN) is used to identify the importance of nodes in the molecular graph and prioritize them.
[0098] In its implementation, DQN learns the optimal strategy for propagating the influence of nodes across multiple neighborhoods, thereby effectively identifying nodes in the molecular structure that significantly impact the overall properties. In a molecular graph, each node (e.g., an atom) is connected to other nodes (e.g., neighboring atoms), and different nodes have different effects on the overall molecular properties (e.g., activity, stability, polarity). DQN learns a strategy model that allows the model to select and evaluate the importance of nodes layer by layer in the neighborhood, thus ranking key nodes. The specific learning and training process of DQN is as follows:
[0099] Step S121: In the reinforcement learning framework, define the state: state s v This represents the current node characteristics and graph structure information of the molecular graph. State s v Features X containing the target node v and its neighborhood features N v ,Right now:
[0100] s v ={X v N v}
[0101] Step S122: Define actions: Each action 'a' represents whether to move the current node v up one position in the importance ranking. The action space can be defined as either selecting node v or skipping node v.
[0102] Step S123: Define the reward: The reward r is designed based on the impact of node selection on the overall molecular structure and properties. For example, if a node contributes significantly to a certain target property of the molecule, a high reward is given.
[0103] Step S124: The goal of DQN is to approximate the Q-value function Q(s,a; θ) using a neural network, where θ are the parameters of the Q-network. The update process is as follows:
[0104] (1) For each node v, define state sv ={X v N v} and select an action a using the current strategy.
[0105] (2) Perform action a and observe the reward r and the next state s′.
[0106] (3) Store the experience (s,a,r,s′) into the experience replay buffer.
[0107] (4) Randomly sample a batch of experiences (s,a,r,s′) from the buffer and update the formula using the following Q value:
[0108]
[0109] Where θ - This represents the parameters of the target network, and the parameters of the Q network are periodically copied to the target network.
[0110] (5) Minimize the loss function to update the Q network parameters θ:
[0111] L(θ)=E((yQ(s,a;θ)) 2 )
[0112] Step S125: After training, calculate the final Q-value for each node, and sort the nodes from high to low based on the Q-values to obtain node v. i Importance ranking P(v) i ).
[0113] Step S126: A filtering mechanism controls which nodes' contextual information is retained and passed on. By selectively updating hidden states, the model can better filter out irrelevant node information when modeling long-range dependencies in a graph. For the current node in a node sequence, the selection mechanism allows it to receive only the influence of previously important nodes, thus preserving key graph structure information.
[0114] Example 4
[0115] Based on the pathway-embedded synergistic effect prediction method for pharmacodynamic substances in traditional Chinese medicine provided in Examples 1-3 above, this embodiment further provides a pathway-embedded synergistic effect prediction system for pharmacodynamic substances in traditional Chinese medicine, combined with... Figure 2 As shown, it includes a molecular structure characterization module, a drug-target interaction module, a target-pathway information embedding module, an end-to-end training module 4, and a visualization module.
[0116] The molecular structure characterization module is used for preprocessing drug molecular structure data;
[0117] The drug-target interaction module generates representation vectors of drug molecules and is used to construct the association between drugs and various targets.
[0118] The target-pathway information embedding module is used to generate multiple target-pathway adjacency matrices to describe the relationships between targets.
[0119] An end-to-end training module is used to generate a drug synergy prediction model to predict the synergistic effects of drug combinations.
[0120] The visualization module provides an intuitive presentation of the synergistic effect prediction results of drug combinations.
[0121] Example 5
[0122] Combination Figure 2 As shown, this embodiment provides a pathway-embedded synergistic effect prediction system for traditional Chinese medicine active substances, including a molecular structure characterization module, a drug-target interaction module, a target-pathway information embedding module, an end-to-end training module, and a visualization module.
[0123] The molecular structure characterization module is used for preprocessing drug molecular structure data;
[0124] The drug-target interaction module generates representation vectors of drug molecules and is used to construct the association between drugs and various targets.
[0125] The target-pathway information embedding module is used to generate multiple target-pathway adjacency matrices to describe the relationships between targets.
[0126] An end-to-end training module is used to generate a drug synergy prediction model to predict the synergistic effects of drug combinations.
[0127] The visualization module provides an intuitive presentation of the synergistic effect prediction results of drug combinations.
[0128] In this embodiment, the molecular structure characterization module includes a data input unit and a node priority sorting unit;
[0129] The data input unit is responsible for receiving molecular structure data of drug molecules and performing preliminary data preprocessing. It converts the input drug molecule data (in SMILES format) into a graph topology, forming a molecular graph.
[0130] The node priority ranking unit uses a node degree heuristic to rank the nodes (i.e., atoms) in the molecular graph to determine their importance. The unit places key nodes at the end of the graph sequence to ensure they receive more contextual information in subsequent feature extraction.
[0131] In this embodiment, the drug-target interaction module includes a graph convolution calculation unit, a vector generation unit, and an interaction prediction unit.
[0132] The graph convolution computation unit updates the state of each node through a selective graph convolution algorithm, providing basic features for subsequent predictions. Based on the node priority ranking results, it selectively controls which node information will enter the hidden state, compressing and transmitting long-distance dependencies.
[0133] The vector generation unit extracts features from nodes to generate the final representation vector of drug molecules, providing efficient and rich molecular structure characterization for subsequent predictions, serving as the basis for subsequent predictions.
[0134] The interaction prediction unit predicts the interaction probability between drugs and targets by combining information from the target database. It uses a self-attention mechanism to calculate the relationship between corresponding targets in molecular structural features, and incorporates consistency regularization to constrain the molecular structural feature representation, providing fundamental support for the pathway information embedding module.
[0135] In this embodiment, the target-path information embedding module includes a path adjacency matrix generation unit and a multi-graph fusion convolution unit.
[0136] The pathway adjacency matrix generation unit maps the relationships between biological targets into an adjacency matrix of target pathways based on the roles of biological targets in different biological pathways. The adjacency matrix is used to describe the correlation of targets in different pathways.
[0137] The multi-graph fusion convolutional unit uses target features output from the drug-target interaction module as graph nodes. It constructs a graph structure using multiple pathway adjacency matrices and embeds pathway information using multi-graph convolution operations to obtain target embedding features, achieving multi-level information fusion between target associations. For specific applications in anti-pneumonia, this module can combine disease-related biological network data, including host-pathogen interaction pathways and pneumonia-related biological pathways.
[0138] In this embodiment, the end-to-end training module includes a target-disease association mapping unit and a deep learning training unit;
[0139] The target-disease association mapping unit fuses the target embedding features of a pair of molecular structures and maps them to the synergistic effect score space of the molecular pair. This unit predicts the synergistic effect of drug combinations on specific diseases through feature fusion technology. In specific implementation, the pairwise molecular graph structure of the paired molecular structure data is obtained by using the molecular structure characterization module. A two-stream parallel network structure is constructed to learn the target embedding features of each molecular structure data. The two-stream parallel network consists of two parameter-sharing drug-target interaction modules and a target-pathway information embedding module. Finally, the paired target embedding features obtained from the paired molecular structure data are fed into the target-disease association mapping unit, which outputs the synergistic effect score.
[0140] The deep learning training unit trains the model using deep learning algorithms based on the data output by the target-disease association mapping unit, generating the final predictive model for the synergistic effects of traditional Chinese medicine active substances. During training, antiviral activity data of single drugs and synergistic data of drug combinations related to pneumonia are incorporated to improve the predictive ability against the novel coronavirus.
[0141] In this embodiment, the visualization module can intuitively present the prediction results of the drug-active substance synergistic effect prediction system through data visualization and virtual reality technology.
[0142] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.
Claims
1. A method for predicting the synergistic effect of pharmacodynamic substances in traditional Chinese medicine through pathway embedding, characterized in that, Includes the following steps: Step S1: Preprocess the drug molecule structure data to generate a representation vector of the drug molecule; Step S2: Based on the target database information and the representation vector of drug molecules, construct the association relationship between the drug and each target; Step S3: Generate an adjacency matrix of multiple target pathways using the collected data on the relationship between targets and biological pathways. Then, fuse the adjacency matrix with the obtained association between the drug and each target to obtain the embedding feature of each target. Step S4: Use the obtained target embedding features for deep learning training to generate a drug synergistic effect prediction model; Step S5: Using a drug synergy prediction model, directly present the synergy prediction results of the drug combination; Step S1 specifically includes the following steps: Step S11: Convert the drug molecule structure data represented by SMILES into a graph topology to form a molecular graph; Step S12: Using the node degree heuristic, prioritize the nodes in the molecular graph using the deep Q-network algorithm; Step S13: Based on the node priority ranking results and the node features of the molecular graph, selectively update the state of the nodes using a graph convolutional neural network; Step S14: Extract features from nodes based on graph convolutional neural networks to generate representation vectors of drug molecules; Step S3 specifically includes the following steps: Step S31: Collect the relationship data of target points in different biological pathways, regard multiple target points as graph nodes, construct the association relationship between target points under each pathway between graph nodes, and map the association relationship data into an adjacency matrix of multiple target pathways; Step S32: The adjacency matrix of multiple target pathways and the obtained drug-target association relationship are fused into a multi-level graph embedding representation. In this process, a graph convolutional neural network is used to embed pathway information to obtain target embedding features.
2. The method for predicting the synergistic effect of pathway-embedded pharmacodynamic substances in traditional Chinese medicine according to claim 1, characterized in that, Step S2 is as follows: Based on the information in the target database, the feature vector of the drug is used to calculate the feature association and regularization constraint of the target features through the self-attention mechanism, and the interaction probability between the drug and the specific biological target is generated to construct the association relationship between the drug and each target.
3. The method for predicting the synergistic effect of pathway-embedded pharmacodynamic substances in traditional Chinese medicine according to claim 2, characterized in that, Step S4 specifically includes the following steps: Step S41: Fuse the target embedding features of a pair of molecular structures and map them to the synergistic effect score space of the molecular pair, and calculate the synergistic effect score. Step S42: Based on the obtained synergy score, train the model using a deep learning training algorithm to generate a drug synergy prediction model.
4. A pathway-embedded synergistic effect prediction system for traditional Chinese medicine active substances, comprising using the pathway-embedded synergistic effect prediction method for traditional Chinese medicine active substances according to any one of claims 1-3, characterized in that, It includes a molecular structure characterization module, a drug-target interaction module, a target-pathway information embedding module, an end-to-end training module, and a visualization module; The molecular structure characterization module is used for preprocessing drug molecular structure data; The drug-target interaction module generates representation vectors of drug molecules and is used to construct the association between drugs and various targets. The target-pathway information embedding module is used to generate multiple target-pathway adjacency matrices to describe the relationships between targets. An end-to-end training module is used to generate a drug synergy prediction model to predict the synergistic effects of drug combinations. The visualization module presents the predicted results of the synergistic effects of drug combinations.
5. The pathway-embedded synergistic effect prediction system for traditional Chinese medicine active substances according to claim 4, characterized in that, The molecular structure characterization module includes a data input unit and a node priority sorting unit; The data input unit converts the input drug molecule SMILES data into a graph topology structure, forming a molecular graph; The node priority sorting unit uses the node degree heuristic index to sort the nodes in the molecular graph by priority.
6. The pathway-embedded synergistic effect prediction system for traditional Chinese medicine active substances according to claim 5, characterized in that, The drug-target interaction module includes a graph convolution calculation unit, a vector generation unit, and an interaction prediction unit; The graph convolution computation unit updates the state of each node using a selective graph convolution algorithm; Vector generation unit, used to generate the final representation vector of drug molecules, providing molecular structure characterization; The interaction prediction unit predicts the probability of interaction between the drug and the target by combining information from the target database.
7. The pathway-embedded synergistic effect prediction system for traditional Chinese medicine active substances according to claim 6, characterized in that, The target-path information embedding module includes a path adjacency matrix generation unit and a multi-graph fusion convolution unit; The pathway adjacency matrix generation unit maps the relationships between biological targets to an adjacency matrix of target pathways, describing the correlation of targets in different pathways. The multi-graph fusion convolutional unit uses multiple target points as graph nodes, constructs a graph structure using the adjacency matrix of multiple target point pathways, and embeds pathway information using multi-graph convolution operations to obtain target point embedding features, thereby achieving multi-level information fusion of the association between target points.
8. The pathway-embedded synergistic effect prediction system for traditional Chinese medicine active substances according to claim 7, characterized in that, The end-to-end training module includes a target-disease association mapping unit and a deep learning training unit; The target-disease association mapping unit fuses the target embedding features of a pair of molecular structures and maps them to the synergistic effect score space of the molecular pair. The deep learning training unit trains the model using deep learning training algorithms based on the data output by the target-disease association mapping unit, generating the final predictive model of the synergistic effect of traditional Chinese medicine active substances.
Citation Information
Patent Citations
Drug combination network based drug combined action predicting method
CN103065066A
Drug target binding affinity prediction method based on graph neural network
CN119132386A