Path-embedded traditional Chinese medicine efficacy substance synergistic effect prediction method and system
Through path embedding technology and deep learning methods, an efficient prediction model for synergistic effects of Chinese medicine drug-efficacy substances was constructed, solving the complexity and inefficiency of the research on synergistic effects of traditional Chinese medicine drug-efficacy substances in the existing technology, and achieving accurate and efficient prediction of drug-efficacy.
Patent Information
- Application Number
- CN202510210767.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-25
AI Technical Summary
The prior art faces challenges such as complex components, lack of systematic data analysis, complex biological target interactions and failure to utilize modern information technology in the study of synergistic effects of Chinese medicine drugs, resulting in inaccuracy and inefficiency of drug efficacy prediction.
The path embedding method is adopted, combined with graph convolutional neural network, deep Q network algorithm and deep learning training, an accurate and efficient drug efficacy prediction model is built to realize efficient input and preprocessing of drug molecular structure data, enhance target interaction prediction capabilities, carry out multi-path information embedding, and realize full process automation.
It significantly improves the screening and prediction accuracy of Chinese medicine-efficacy substances, realizes accurate, efficient and comprehensive prediction of the synergistic effects of Chinese medicine-efficacy substances, and improves the depth and accuracy of the prediction of synergistic effects of drug combinations.
Smart Images

Figure CN120148904A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of predicting the synergistic effect of traditional Chinese medicine active ingredients, and particularly to a method and system for predicting the synergistic effect of traditional Chinese medicine active ingredients with pathway embedding. Background Art
[0002] In the field of traditional Chinese medicine research and development, the compatibility of traditional Chinese medicine is the most unique link in the theoretical system of traditional Chinese medicine, and it is the advantage and characteristic of traditional Chinese medicine. How to combine the active ingredients of traditional Chinese medicine and make their efficacy equal to or better than that of the original formula and the original active ingredients to produce a synergistic effect is crucial. The traditional research on the combination of traditional Chinese medicine active ingredients mainly relies on laboratory bioactivity tests. Although these methods have a certain experimental basis, the combination traversal is time-consuming and costly, and it is difficult to cover the chemical space of traditional Chinese medicine active ingredients.
[0003] The existing methods for studying the synergistic effect of traditional Chinese medicine active ingredients face several major challenges: First, the components of traditional Chinese medicine are complex, and a single experimental method is difficult to comprehensively reveal the interactions between components and their comprehensive effects on organisms; Second, traditional methods usually lack systematic data analysis and mining capabilities, and it is difficult to extract useful information from a large amount of experimental data, which affects the accuracy and efficiency of efficacy prediction; Third, the interaction between the active ingredients of traditional Chinese medicine and biological targets is complex, and there is a lack of effective computational models to support the prediction of the complex interactions between active ingredients and targets; Finally, the existing technologies usually fail to make full use of modern information technologies, such as artificial intelligence and machine learning, to optimize the screening and prediction processes of active ingredients.
[0004] Therefore, in view of the above problems and technical challenges, there is an urgent need to provide a new method and system for predicting the synergistic effect of traditional Chinese medicine active ingredients to achieve accurate and efficient screening and effect prediction of active ingredients. Summary of the Invention
[0005] In view of the problems existing in the prior art, the present invention provides a method and system for predicting the synergistic effect of traditional Chinese medicine active ingredients with pathway embedding. The present invention integrates technologies such as pathway embedding, graph convolutional neural network, deep Q-network algorithm, and deep learning training to construct a more accurate and efficient efficacy prediction model, realizing accurate, efficient, and comprehensive prediction of the synergistic effect of traditional Chinese medicine active ingredients.
[0006] To achieve the above technical objectives, the present invention adopts the following technical solutions:
[0007] The present invention provides a method for predicting the synergistic effect of traditional Chinese medicine active ingredients with pathway embedding, the method comprising the following steps:
[0008] Step S1: Preprocess the drug molecular structure data to generate a representation vector of the drug molecule;
[0009] Step S2: Based on the information in the target database and the representation vectors of drug molecules, construct the association relationships between drugs and each target;
[0010] Step S3: Use the collected relationship data between targets and biological pathways to generate the adjacency matrices of multiple target pathways, and fuse the adjacency matrices with the obtained association relationships between drugs and each target to obtain the embedding features of each target;
[0011] Step S4: Use the obtained embedding features of targets for deep learning training to generate a drug synergy prediction model;
[0012] Step S5: Use the drug synergy prediction model to directly present the synergy prediction results of drug combinations.
[0013] Preferably, step S1 specifically includes the following steps:
[0014] Step S11: Convert the drug molecular structure data represented by SMILES into a graph topological structure to form a molecular graph;
[0015] Step S12: Adopt the node degree heuristic index to prioritize the nodes in the molecular graph through the deep Q-network algorithm;
[0016] Step S13: Based on the node prioritization results and the node features of the molecular graph, selectively update the states of the nodes through the graph convolutional neural network;
[0017] Step S14: Extract features from the nodes based on the graph convolutional neural network to generate the representation vectors of drug molecules.
[0018] Preferably, the specific process of step S2 is as follows: Based on the information in the target database, use the feature vectors of drugs, calculate the feature associations and regularize the target features through the self-attention mechanism, generate the interaction probabilities between drugs and specific biological targets, so as to construct the association relationships between drugs and each target.
[0019] Preferably, step S3 specifically includes the following steps:
[0020] Step S31: Collect the relationship data of targets in different biological pathways, regard multiple targets as graph nodes, construct the association relationships between targets under each pathway between the graph nodes, and map the association relationship data into the adjacency matrices of multiple target pathways;
[0021] Step S32: Fuse the adjacency matrices of multiple target pathways and the obtained association relationships between drugs and each target into a multi-level graph embedding representation, and use the graph convolutional neural network for pathway information embedding in this process to obtain the embedding features of targets.
[0022] Preferably, step S4 specifically includes the following steps:
[0023] Step S41: Fuse the target embedding features of a pair of molecular structures, map them to the synergistic effect score space of the molecular pair, and calculate the synergistic effect score;
[0024] Step S42: According to the obtained synergistic effect score, perform model training through a deep learning training algorithm to generate a drug synergistic effect prediction model.
[0025] The present invention also provides a traditional Chinese medicine pharmacodynamic substance synergistic effect prediction system based on pathway embedding, which includes a molecular structure characterization module, a drug-target interaction module, a target-pathway information embedding module, an end-to-end training module, and a visualization module;
[0026] The molecular structure characterization module is used to preprocess drug molecular structure data;
[0027] The drug-target interaction module generates a representation vector of the drug molecule and is used to construct the association relationship between the drug and each target;
[0028] The target-pathway information embedding module is used to generate multiple target pathway adjacency matrices to describe the association relationship between targets;
[0029] The end-to-end training module is used to generate a drug synergistic effect prediction model to predict the synergistic effect of drug combinations;
[0030] The visualization module presents the prediction results of the synergistic effect of drug combinations.
[0031] Preferably, the molecular structure characterization module includes a data input unit and a node priority sorting unit;
[0032] The data input unit converts the input drug molecular SMILES data into a graph topological structure to form a molecular graph;
[0033] The node priority sorting unit uses a node degree heuristic index to sort the nodes in the molecular graph by priority.
[0034] Preferably, the drug-target interaction module includes a graph convolution calculation unit, a vector generation unit, and an interaction prediction unit;
[0035] The graph convolution calculation unit updates the state of each node through a selective graph convolution algorithm;
[0036] The vector generation unit is used to generate the final representation vector of the drug molecule to provide molecular structure characterization;
[0037] The interaction prediction unit predicts the interaction probability between the drug and the target by combining the information in the target database.
[0038] Preferably, the target-pathway information embedding module includes a pathway adjacency matrix generation unit and a multi-graph fusion convolution unit;
[0039] The pathway adjacency matrix generation unit maps the relationships of biological targets into an adjacency matrix of target pathways, describing the relevance of targets in different pathways;
[0040] The multi-graph fusion convolution unit takes multiple targets as graph nodes, constructs a graph structure using multiple target pathway adjacency matrices, embeds pathway information through multi-graph convolution operations, obtains target embedding features, and realizes multi-level information fusion of the associations between targets.
[0041] Preferably, the end-to-end training module includes a target-disease association mapping unit and a deep learning training unit;
[0042] The target-disease association mapping unit fuses the target embedding features of a pair of molecular structures and maps them to the synergistic effect score space of the molecular pair;
[0043] The deep learning training unit performs model training through a deep learning training algorithm based on the data output by the target-disease association mapping unit to generate a final prediction model for the synergistic effect of traditional Chinese medicine active ingredients.
[0044] Compared with the prior art, the beneficial effects produced by the present invention are as follows:
[0045] 1) It realizes the efficient input and preprocessing of traditional Chinese medicine molecular data. The present invention converts SMILES data into a graph topology structure, prioritizes the nodes in the molecule, and uses the graph convolutional neural network algorithm to extract features of the nodes, generating a representation vector of the drug molecule, effectively capturing long-range dependencies, providing an efficient and rich molecular structure representation, and significantly improving the accuracy and information retention effect of the molecular structure representation compared with previous methods.
[0046] 2) It enhances the ability to predict target interactions. Based on the generated representation vector of the drug molecule and combined with the information in the target database, the present invention accurately predicts the interaction probability between the drug and specific biological targets, providing high-quality data support for subsequent pathway analysis, thereby enhancing the effectiveness of the prediction of the synergistic effect of active ingredients.
[0047] 3) It realizes the embedding of multi-pathway information. The present invention effectively integrates the information of targets and biological pathways by constructing adjacency matrices of multiple target pathways, establishes the association relationship between targets and pathways, and uses the graph convolutional neural network for the embedding of pathway information to obtain target embedding features, performing multi-pathway information embedding, realizing a comprehensive description of the associations and interactions between targets, being able to effectively integrate multi-target information related to pathways, improving the screening and prediction accuracy of active ingredients, and thus enhancing the depth and accuracy of the prediction of the synergistic effect.
[0048] 4) Intelligent and automated data analysis. The present invention integrates a graph convolutional neural network and a deep Q-network algorithm to achieve full-process automation from data input, structural representation to synergistic effect prediction, greatly improving the efficiency of drug screening and prediction, reducing the need for human intervention, and realizing the automation and intelligence of screening for effective components of traditional Chinese medicine.
[0049] 5) System refinement and intelligence. By combining a variety of advanced information technologies, such as graph convolutional neural networks, multi-graph fusion convolution, and deep learning for multiple diseases, the present invention provides a refined and intelligent method and system for predicting drug efficacy, providing strong technical support for screening the synergistic effects of effective components of traditional Chinese medicine, and having significant application value and promotion potential. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is a flowchart of a method for predicting the synergistic effect of effective components of traditional Chinese medicine with pathway embedding according to an embodiment of the present invention;
[0051] Figure 2 It is a schematic diagram of system modules for predicting the synergistic effect of effective components of traditional Chinese medicine with pathway embedding according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention. Figure 1 and 2 shown in the accompanying drawings, which are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.
[0053] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "coaxial", "bottom", "one end", "top", "middle", "the other end", "upper", "one side", "top", "inner", "front", "center", "both ends", etc. is based on the orientation or positional relationship shown in the accompanying drawings Figure 1 and 2 shown in the accompanying drawings, which are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.
[0054] In the present invention, unless otherwise clearly defined and limited, terms such as "installation", "setting", "connection", "fixation", "swivel connection", etc. shall be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral one; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components or the interaction relationship between two components. Unless otherwise clearly defined, for those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0055] Example 1
[0056] Combined with Figure 1 As shown, this embodiment provides a method for predicting the synergistic effect of traditional Chinese medicine efficacy substances with pathway embedding, including the following steps:
[0057] Step S1: Preprocess the drug molecular structure data to generate a representation vector of the drug molecule;
[0058] Step S2: Based on the target database information and the representation vector of the drug molecule, construct the association relationship between the drug and each target;
[0059] Step S3: Use the collected relationship data between the target and the biological pathway to generate the adjacency matrix of multiple target pathways, and fuse the adjacency matrix with the obtained association relationship between the drug and each target to obtain the embedding feature of each target;
[0060] Step S4: Use the obtained embedding feature of the target for deep learning training to generate a drug synergistic effect prediction model;
[0061] Step S5: Use the drug synergistic effect prediction model to directly present the prediction result of the synergistic effect of the drug combination.
[0062] By accurately constructing the association relationship between the drug and the target, the present invention enhances the ability to predict target interactions, provides data support for subsequent pathway analysis, and thus enhances the effectiveness of predicting the synergistic effect of efficacy substances; by constructing the adjacency matrix of multiple target pathways and establishing the association between the target and the pathway, the present invention realizes the embedding of multi-pathway information, comprehensively describes the association and interaction between targets, can effectively integrate multi-target information related to the pathway, improves the screening and prediction accuracy of efficacy substances, and thus enhances the depth and accuracy of predicting the synergistic effect.
[0063] Example 2
[0064] Combined with Figure 1 As shown, this embodiment provides a method for predicting the synergistic effect of traditional Chinese medicine efficacy substances with pathway embedding, including the following steps:
[0065] Step S1: Preprocess the drug molecular structure data to generate a representation vector of the drug molecule;
[0066] Step S2: Based on the target database information and the representation vector of the drug molecule, construct the association relationship between the drug and each target;
[0067] Step S3: Use the collected relationship data between the target and the biological pathway to generate the adjacency matrix of multiple target pathways, and fuse the adjacency matrix with the obtained association relationship between the drug and each target to obtain the embedding feature of each target;
[0068] Step S4: Use the obtained embedding feature of the target for deep learning training to generate a drug synergy prediction model;
[0069] Step S5: Use the drug synergy prediction model to directly present the synergy prediction result of the drug combination.
[0070] In this embodiment, the specific procedure of step S1 is as follows:
[0071] Step S11: Convert the drug molecular structure data represented by SMILES into a graph topology structure to form a molecular graph. In a specific implementation, input the drug molecular structure data in the SMILES format. The chemical library (such as RDKit) can be used to parse the SMILES string, extract each atom and the bonds between them. The atoms are abstracted as graph nodes, and the bonds between atoms are abstracted as the edges of the graph, so as to convert the SMILES data into a graph topology structure to obtain the molecular graph structure G(X,A), where X represents the node feature matrix and A represents the adjacency matrix of the molecular graph.
[0072] Step S12: Adopt the node degree heuristic index to prioritize the nodes in the molecular graph through the deep Q-network algorithm to determine the importance of the graph nodes. The node degree heuristic index is based on the number of edges connecting a node to other nodes and is designed based on experience. It is a measurement method used to quickly evaluate the characteristics such as the importance degree of nodes. Calculate the node priority through the node degree, and sort the nodes according to their priorities, arranging the important nodes at the end of the graph sequence to ensure that they obtain more context information in subsequent feature extraction. This sorting by priority can ensure that the nodes with higher importance are given priority attention in subsequent processing steps.
[0073] Step S13: Based on the priority sorting result of nodes and the node features of the molecular graph, selectively update the states of nodes through a Graph Convolutional Network (GCN). This process can provide basic features for subsequent predictions. In a specific implementation, a GCN combined with a selection gate is constructed, that is, a gating mechanism is added to the GCN. Through gating, it is selected which node features are important and need to be retained, and which node features can be ignored and allowed to enter the hidden state. Important nodes with high priority are updated, while nodes with low priority are updated in a restricted manner. This allows the model to more flexibly process graph-structured data and improve the performance and efficiency of the model. The formula for determining whether each node is updated can be expressed as:
[0074]
[0075] where is the new feature vector of node i after being updated by the Graph Convolutional Network (GCN). This vector is the state of node i after update. P(d(v i )) represents the degree feature extraction layer of node i. By inputting the degree d(vi) of node i into a fully connected layer, which consists of a set of weights and biases, a way to map the node degree to a new feature space can be learned. The degree feature extraction layer is a linear transformation, implemented by including a linear layer plus an activation function (ReLU). δ represents a hyperparameter, represents the l-th layer feature of node i, A represents the adjacency matrix, I represents the identity matrix. represents the normalized form of the adjacency matrix, D represents the degree matrix. w l represents the l-th layer parameter, and σ represents the activation function.
[0076] In this way, for nodes with higher priority, their information will be selectively propagated in the GCN hierarchical update, thereby enhancing the influence of important nodes and avoiding the propagation of irrelevant or redundant information. The information propagation of low-priority nodes will be suppressed, thus achieving a sparsification effect. Then, through the gating mechanism, it can selectively control which node information will enter the hidden state. By selectively propagating node information in the Graph Convolutional Network (GCN), it is ensured that important long-term dependencies can be effectively retained and propagated, realizing the compression and transmission of long-range dependencies.
[0077] Step S14: Based on the Graph Convolutional Network, perform feature extraction on nodes to generate a representation vector of the drug molecule. Specifically, further extract the node feature X v through the 1D convolution and activation function of the Graph Convolutional Network, and generate the final representation vector X mThese representation vectors not only contain the context information of key nodes but also retain the long-range dependence characteristics filtered out in graph convolution, providing an efficient and rich molecular structure representation for the subsequent drug-target interaction module.
[0078] The present invention realizes the efficient input and preprocessing of traditional Chinese medicine molecular data, converts SMILES data into a graph topology structure, prioritizes the nodes in the molecule, and uses a graph convolutional neural network algorithm to extract features of the nodes to generate a representation vector of the drug molecule, effectively capturing long-range dependence relationships, providing an efficient and rich molecular structure representation, and significantly improving the accuracy and information retention effect of the molecular structure representation compared with previous methods.
[0079] In this embodiment, the specific process of step S2 is as follows: Based on the information in the target database, using the feature vector of the drug, calculate the feature correlation and regularize the target features through the self-attention mechanism to generate the interaction probability between the drug and a specific biological target, so as to construct the association relationship between the drug and each target, providing basic data for subsequent pathway embedding analysis. In specific implementation, the feature correlation of the molecular representation vector can be calculated by using the self-attention mechanism, and the result output by the self-attention mechanism can be normalized to provide an additional association adjacency matrix A for subsequent use. coor Then, through consistency regularization, the target features output by the self-attention mechanism are constrained during the training process, and similarity regularization is imposed on the self-supervised labels (such as prediction results) generated by the same target features, so that the model automatically clusters the same target features in the feature space. Different target features remain discrete through the differential processing of the self-supervised task. In the anti-pneumonia application, the data of the targets of the new coronavirus (such as TNF-α, MAPK1, etc.) will be used to predict the possible antiviral effects of traditional Chinese medicine molecules. The present invention enhances the target interaction prediction ability, accurately predicts the interaction probability between the drug and a specific biological target, provides high-quality data support for subsequent pathway analysis, and thus enhances the effectiveness of predicting the synergistic effect of active pharmaceutical ingredients.
[0080] In this embodiment, the specific implementation steps of step S3 include:
[0081] Step S31: Regard multiple targets as graph nodes, construct the association relationship between the targets under each pathway between the graph nodes, and map the association relationship data into the adjacency matrix of multiple target pathways to represent the relevance of the targets in different biological pathways. In specific implementation, the relationship data of the targets in different biological pathways are collected from multiple biological databases (such as KEGG, Reactome, etc.). Each pathway describes the biological interactions between different targets, such as metabolic pathways, signal transduction pathways, etc.
[0082] For each pathway, an adjacency matrix is generated based on the interaction relationships between the targets. The dimension of the adjacency matrix is N×N, where N is the number of targets involved in the system. The elements in the matrix represent the relationships between target pairs. For example, 1 indicates the existence of a direct interaction, and 0 indicates no direct interaction. Then, the adjacency matrix is normalized to ensure that each matrix has a consistent scale for subsequent multi-graph convolution processing.
[0083] Step S32: Integrate the adjacency matrices of multiple target pathways and the obtained drug-target association relationships into a multi-level graph embedding representation. In this process, a graph convolutional neural network is used to embed the pathway information to obtain target embedding features, which contain the complex association information of the targets in multiple pathways. The graph convolutional neural network (GCN) uses these adjacency matrices for graph convolution to generate the embedding features of each target, realizing the multi-level information fusion of the associations between targets. For specific application to anti-pneumonia, this module can combine disease-related biological network data, including host-pathogen interaction pathways and pneumonia-related biological pathways.
[0084] The specific implementation steps of step S32 include:
[0085] Step S321: Use the feature information of the targets contained in the drug-target association relationship data output in step S2 as the input of the target features. Each target has a low-dimensional feature vector representation initially. This vector contains the basic information of the target, such as molecular structure and biological function, etc. Generally, the generation process of target features is usually by extracting the basic information of the targets (such as amino acid sequence, function, structure, etc.) and encoding and representing them through deep learning methods (graph convolutional neural network).
[0086] Step S322: Use graph convolution and the adjacency matrix A B (the adjacency matrix of the i-th pathway) and the adjacency matrix A coor to aggregate the information of adjacent nodes. The multi-head mechanism is adopted to allow each layer of convolution to independently calculate the features of different pathways, and then connect or average these features to obtain a richer representation.
[0087] Step S323: After multi-layer convolution, fuse the outputs of each layer of convolution. Weighted average or attention mechanism can be used to aggregate the multi-graph convolution results into a comprehensive feature. The attention mechanism can assign different weights to each pathway according to its relevance in the current task to enhance the information of important pathways. After the above steps, the target embedding feature X b。The representation of the target embedding features contains multi-level association information of the target in different pathways. This representation of the target embedding features can not only reflect the basic characteristics of the target, but also contain its complex network relationships in multiple biological pathways. The representation of the target embedding features for the pneumonia target will reflect its role in the pathways related to disease transmission.
[0088] The present invention performs multi-pathway information embedding. By establishing the association between the target and the pathway, the screening and prediction accuracy of the pharmaceutically active substances are improved. It can comprehensively describe the association and interaction between targets, effectively integrate multi-target information related to the pathway, and thus improve the depth and accuracy of the synergistic effect prediction.
[0089] In this embodiment, the specific step process of step S4 is as follows:
[0090] Step S41: Fuse the target embedding features of a pair of molecular structures and map them to the synergistic effect score space of the molecular pair to calculate the synergistic effect score. In a specific implementation, receive the target embedding feature X output from step S32 b , and use a multi-layer perceptron to map the target embedding features of a pair of molecular structures associated with the disease to the synergistic effect score space, and use the multi-layer perceptron to calculate the synergistic effect score. Its calculation method can be expressed as:
[0091]
[0092] Among them, Linear represents the linear layer, cat represents feature concatenation, and represent the target embedding features of a pair of molecular structures.
[0093] In a specific implementation, through the previous step S1, the paired molecular graph structures of the paired molecular structure data are obtained, a two-stream parallel network structure is constructed to learn the target embedding features of each molecular structure data respectively, and finally the paired target embedding features obtained from the paired molecular structure data are fused and mapped to the synergistic effect score space of the molecular pair to output the synergistic effect score.
[0094] Step S42: According to the obtained synergistic effect score, perform model training through a deep learning training algorithm to generate a drug synergistic effect prediction model. Compare the synergistic effect score with the target label, and optimize the model parameters by the gradient descent method to achieve end-to-end training. After training, a traditional Chinese medicine pharmaceutically active substance synergistic effect prediction model is obtained. During the training process, the antiviral activity data of a single drug related to anti-pneumonia and the drug combination synergistic data are introduced to improve the antiviral prediction ability for the novel coronavirus. The final model is trained with pneumonia data and can be used to predict potential anti-pneumonia synergistic treatment combinations.
[0095] In this embodiment, in step S5, when using the drug synergy effect prediction model for prediction, through data visualization and virtual reality technologies, the prediction results are presented in the form of graphs or virtual reality, directly presenting the drug combination's synergy effect prediction results, which is convenient for users to understand and further analyze.
[0096] Embodiment 3
[0097] This embodiment is based on Embodiment 2, and specifically illustrates the specific step process of prioritizing the nodes in the molecular graph through the deep Q-network algorithm in step S12. The importance of the nodes in the molecular graph is identified through the Deep Q-Network (DQN) and prioritized.
[0098] In specific implementation, DQN is used to learn the optimal strategy for nodes to spread influence in multiple layers of neighborhoods, thereby effectively identifying the nodes in the molecular structure that have important impacts on the overall properties. In the molecular graph, each node (such as an atom) is connected to other nodes (such as adjacent atoms), and different nodes have different impacts on the overall properties of the molecule (such as activity, stability, polarity, etc.). Through DQN, a policy model can be learned, enabling the model to select and evaluate the importance of nodes layer by layer in multiple layers of neighborhoods to achieve the ranking of key nodes. The specific learning and training process through DQN is as follows:
[0099] Step S121: Under the reinforcement learning framework, define the state: state s v represents the current node features and graph structure information of the molecular graph. State s v contains the features X v of the target node and its neighborhood features N v , that is:
[0100] s v = {X v , N v}
[0101] Step S122: Define the action: Each action a represents choosing whether to promote the current node v by one position in the importance ranking. The action space can be defined as choosing node v or skipping node v.
[0102] Step S123: Define the reward: The design of the reward r is based on the impact of node selection on the overall structure and properties of the molecule. For example, if a node makes a greater contribution to a certain target property of the molecule, a high reward is given.
[0103] Step S124: The goal of DQN is to approximate the Q-value function Q(s, a; θ) through a neural network, where θ is the parameter of the Q-network. The update process is as follows:
[0104] (1) For each node v, define the state sv = {X v , N v}, and select an action a using the current policy.
[0105] (2) Execute the action a and observe the reward r and the next state s'.
[0106] (3) Store the experience (s, a, r, s') in the experience replay buffer.
[0107] (4) Randomly sample a batch of experiences (s, a, r, s') from the buffer and use the following Q-value update formula:
[0108]
[0109] where θ - represents the parameters of the target network, and the parameters of the Q-network are copied to the target network periodically.
[0110] (5) Minimize the loss function to update the Q-network parameters θ:
[0111] L(θ) = E((y - Q(s, a; θ)) 2 )
[0112] Step S125: After the training is completed, calculate the final Q-value for each node, sort the nodes based on the Q-value from high to low, and obtain the importance ranking P(v i ) of the node v i .
[0113] Step S126: Control which nodes' context information is retained and transmitted through a screening mechanism. By selectively updating the hidden state, the model can better filter out irrelevant node information when modeling long-range dependencies in the graph. For the current node in the node sequence, the selection mechanism allows it to only receive the influence of previous important nodes, thus retaining the key graph structure information.
[0114] Example 4
[0115] Based on the method for predicting the synergistic effect of traditional Chinese medicine efficacy substances with pathway embedding provided in the above Examples 1 to 3, on this basis, this example provides a system for predicting the synergistic effect of traditional Chinese medicine efficacy substances with pathway embedding. As shown in combination with Figure 2 , it includes a molecular structure characterization module, a drug-target interaction module, a target-pathway information embedding module, an end-to-end training module 4, and a visualization module.
[0116] The molecular structure characterization module is used to preprocess the drug molecular structure data.
[0117] Drug - target interaction module, which generates a representation vector of the drug molecule and is used to construct the association relationship between the drug and each target;
[0118] Target - pathway information embedding module, which is used to generate multiple target - pathway adjacency matrices to describe the association relationship between targets;
[0119] End - to - end training module, which is used to generate a drug synergistic effect prediction model to predict the synergistic effect of drug combinations;
[0120] Visualization module, which intuitively presents the prediction results of the synergistic effect of drug combinations.
[0121] Example 5
[0122] Combined Figure 2 As shown, this example provides a traditional Chinese medicine effective substance synergistic effect prediction system with pathway embedding, including a molecular structure characterization module, a drug - target interaction module, a target - pathway information embedding module, an end - to - end training module, and a visualization module.
[0123] Molecular structure characterization module, which is used to pre - process the drug molecular structure data;
[0124] Drug - target interaction module, which generates a representation vector of the drug molecule and is used to construct the association relationship between the drug and each target;
[0125] Target - pathway information embedding module, which is used to generate multiple target - pathway adjacency matrices to describe the association relationship between targets;
[0126] End - to - end training module, which is used to generate a drug synergistic effect prediction model to predict the synergistic effect of drug combinations;
[0127] Visualization module, which intuitively presents the prediction results of the synergistic effect of drug combinations.
[0128] In this example, the molecular structure characterization module includes a data input unit and a node priority sorting unit;
[0129] Data input unit, which is responsible for receiving the molecular structure data of the drug molecule and completing the preliminary pre - processing of the data. It converts the input drug molecule in the SMILES data format into a graph topological structure to form a molecular graph.
[0130] Node priority sorting unit, which uses the node degree heuristic index to sort the nodes (i.e., atoms) in the molecular graph to determine the importance of the nodes. The node priority sorting unit arranges the key nodes at the end of the graph sequence to ensure that they obtain more context information in subsequent feature extraction.
[0131] In this embodiment, the drug-target interaction module includes a graph convolution calculation unit, a vector generation unit, and an interaction prediction unit.
[0132] The graph convolution calculation unit updates the state of each node through a selective graph convolution algorithm, providing basic features for subsequent predictions; according to the node priority sorting result, it selectively controls which node information will enter the hidden state, compresses and transmits long-range dependencies.
[0133] The vector generation unit extracts features from the nodes to generate the final representation vector of the drug molecule, providing an efficient and rich molecular structure characterization for subsequent predictions and serving as the basic data for subsequent predictions.
[0134] The interaction prediction unit predicts the interaction probability between the drug and the target by combining the information in the target database. By combining the information in the target database, it uses the self-attention mechanism to calculate the relationship of the corresponding target in the molecular structure features, and combines the consistency regularization to constrain the molecular structure feature representation, providing basic support for the pathway information embedding module.
[0135] In this embodiment, the target-pathway information embedding module includes a pathway adjacency matrix generation unit and a multi-graph fusion convolution unit.
[0136] The pathway adjacency matrix generation unit maps the relationship of biological targets into the adjacency matrix of the target pathway according to the role relationship of biological targets in different biological pathways. The adjacency matrix is used to describe the relevance of targets in different pathways.
[0137] The multi-graph fusion convolution unit uses the target features output by the drug-target interaction module as graph nodes, constructs a graph structure using multiple pathway adjacency matrices, and embeds the pathway information through multi-graph convolution operations to obtain target embedding features, realizing multi-level information fusion of the associations between targets. For specific application to anti-pneumonia, this module can combine disease-related biological network data, including host-pathogen interaction pathways and pneumonia-related biological pathways.
[0138] In this embodiment, the end-to-end training module includes a target-disease association mapping unit and a deep learning training unit;
[0139] The target-disease association mapping unit fuses the target embedding features of a pair of molecular structures and maps them to the synergistic effect score space of the molecular pair. This unit predicts the synergistic effect of drug combinations on specific diseases through feature fusion technology. In a specific implementation, the paired molecular graph structure of paired molecular structure data is obtained by using the molecular structure characterization module, and a two-stream parallel network structure is constructed to learn the target embedding features of each molecular structure data respectively. The two-stream parallel network consists of two drug-target interaction modules with shared parameters and a target-pathway information embedding module. Finally, the paired target embedding features obtained from the paired molecular structure data are sent into the target-disease association mapping unit to output the synergistic effect score.
[0140] The deep learning training unit performs model training through a deep learning training algorithm according to the data output by the target-disease association mapping unit, and generates a final prediction model for the synergistic effect of traditional Chinese medicine active ingredients. During the training process, single-drug antiviral activity data related to anti-pneumonia and drug combination synergistic data are introduced to improve the antiviral prediction ability against the novel coronavirus.
[0141] In this embodiment, the visualization module can intuitively present the prediction results of the active ingredient synergistic effect prediction system through data visualization and virtual reality technology.
[0142] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A pathway-embedded method for predicting the synergistic effect of Chinese medicine active substances, characterized in that: The following steps are involved: Step S1: preprocessing the drug molecule structure data to generate a representation vector of the drug molecule; Step S2: Based on the target database information and the representation vector of the drug molecule, the association relationship between the drug and each target is constructed; Step S3: Generate an adjacency matrix of multiple target pathways using the collected target-biological pathway relationship data, fuse the adjacency matrix with the obtained drug-target association relationship, and obtain the embedding feature of each target; Step S4: using the obtained target embedding features for deep learning training to generate a drug synergistic effect prediction model; Step S5: Using the drug synergistic effect prediction model, directly present the synergistic effect prediction results of the drug combination.
2. The method for predicting the synergistic effect of Chinese medicine active substances by pathway embedding according to claim 1, characterized in that: Step S1 specifically includes the following steps: Step S11: converting the drug molecular structure data represented by SMILES into a graph topology structure to form a molecular graph; Step S12: using the node degree heuristic indicator to prioritize the nodes in the molecular graph through the deep Q network algorithm; Step S13: Based on the priority ranking results of the nodes and the node features of the molecular graph, the states of the nodes are selectively updated through the graph convolutional neural network; Step S14: Extract features of nodes based on the graph convolutional neural network to generate representation vectors of drug molecules.
3. The method for predicting the synergistic effect of Chinese medicine active substances by pathway embedding according to claim 2, characterized in that: The specific process of step S2 is: based on the information of the target database, using the feature vector of the drug, calculating the feature association and regularization constraint target features through the self-attention mechanism, generating the interaction probability between the drug and the specific biological target, so as to construct the association relationship between the drug and each target.
4. The method for predicting the synergistic effect of Chinese medicine active substances by pathway embedding according to claim 3, characterized in that: Step S3 specifically includes the following steps: Step S31: collecting the relationship data of the targets in different biological pathways, treating the multiple targets as graph nodes, constructing the association relationship between the targets under each pathway between the graph nodes, and mapping the association relationship data into an adjacency matrix of the multiple target pathways; Step S32: The adjacency matrices of multiple target pathways and the obtained associations between drugs and targets are fused into a multi-level graph embedding representation. In this process, a graph convolutional neural network is used to embed pathway information to obtain target embedding features.
5. The method for predicting the synergistic effect of Chinese medicine active substances by pathway embedding according to claim 4, characterized in that: Step S4 specifically includes the following steps: Step S41: fusing the target embedding features of a pair of molecular structures and mapping them to the synergy score space of the molecular pair to calculate the synergy score; Step S42: Based on the obtained synergy score, model training is performed using a deep learning training algorithm to generate a drug synergy effect prediction model.
6. A pathway-embedded Chinese medicine pharmacological substance synergistic effect prediction system, using a pathway-embedded Chinese medicine pharmacological substance synergistic effect prediction method according to any one of claims 1 to 5 for prediction, characterized in that: It includes molecular structure characterization module, drug-target interaction module, target-pathway information embedding module, end-to-end training module, and visualization module; Molecular structure characterization module, used to pre-process drug molecular structure data; The drug-target interaction module generates a representation vector of the drug molecule and is used to construct the association between the drug and each target; The target-pathway information embedding module is used to generate multiple target-pathway adjacency matrices to describe the association relationship between targets; An end-to-end training module for generating a drug synergy prediction model to predict the synergistic effect of drug combinations; Visualization module presents the prediction results of synergistic effects of drug combinations.
7. The pathway-embedded Chinese medicine active substance synergistic effect prediction system according to claim 6, characterized in that: The molecular structure characterization module includes a data input unit and a node priority sorting unit; The data input unit converts the input drug molecule SMILES data into a graph topology to form a molecular graph; The node prioritization unit uses the node degree heuristic metric to prioritize the nodes in the molecular graph.
8. The pathway-embedded Chinese medicine active substance synergistic effect prediction system according to claim 7, characterized in that: The drug-target interaction module includes a graph convolution calculation unit, a vector generation unit and an interaction prediction unit; Graph convolution computation unit, which updates the state of each node through a selective graph convolution algorithm; A vector generation unit, used to generate a final representation vector of the drug molecule and provide a molecular structure representation; The interaction prediction unit predicts the interaction probability between drugs and targets by combining the information in the target database.
9. The pathway-embedded Chinese medicine active substance synergistic effect prediction system according to claim 8, characterized in that: The target-pathway information embedding module includes a pathway adjacency matrix generation unit and a multi-graph fusion convolution unit; The pathway adjacency matrix generation unit maps the relationship between biological targets into the adjacency matrix of target pathways, describing the association between targets in different pathways; The multi-graph fusion convolution unit takes multiple targets as graph nodes, uses multiple target pathway adjacency matrices to build a graph structure, and uses multi-graph convolution operations to embed pathway information to obtain target embedding features, thereby realizing multi-level information fusion of associations between targets.
10. The pathway-embedded Chinese medicine active substance synergistic effect prediction system according to claim 9, characterized in that: The end-to-end training module includes a target-disease association mapping unit and a deep learning training unit; The target-disease association mapping unit fuses the target embedding features of a pair of molecular structures and maps them to the synergy score space of the molecular pair; The deep learning training unit performs model training through a deep learning training algorithm based on the data output by the target-disease association mapping unit to generate the final prediction model for the synergistic effect of Chinese medicine active substances.
Citation Information
Patent Citations
Drug combination network based drug combined action predicting method
CN103065066A
Drug-drug interaction prediction method based on biological network global structure
CN115458044A
Drug target positioning evaluation system
CN115641909A
Drug-target interaction prediction method fusing multi-dimensional features
CN116206775A
Prediction model construction method, prediction method and device for drug combination synergistic effect
CN118412146A