Method and system for predicting potential drug interaction based on biological network

By adopting biological network methods in drug interaction prediction, combining natural language processing, knowledge graph embedding and graph neural networks, dynamically fusing the characteristics of drug structure, biological network and functional information, the problems of insufficient information integration and generalization capabilities in the existing methods are solved, and more efficient drug interaction prediction is achieved.

CN119943207AInactive Publication Date: 2025-05-06YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510430403.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing drug interaction prediction methods have problems such as insufficient information integration, difficult static fusion strategies to adapt to complex scenarios, and insufficient generalization ability of models to new drugs and sparse interaction relationships.

Method used

A biological network-based method is adopted to obtain drug interaction data sets and drug feature data, and extract drug structural features using natural language processing algorithms. The knowledge graph embedding method is used to learn drug biological networks. The graph neural network aggregates drug functional information, and dynamic feature fusion is carried out through cross matrix operations or self-attention mechanisms to generate fusion feature vectors for training prediction models.

Benefits of technology

It improves the accuracy and generalization ability of drug interaction prediction, can more effectively integrate multimodal information, adapt to complex scenarios, handle new drugs and sparse interaction relationships, reduces information redundancy and loss, and improves prediction accuracy and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943207A_ABST
    Figure CN119943207A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a system for predicting potential drug interaction based on a biological network. The method comprises the following steps: acquiring a drug interaction data set and drug characteristic data; the medicine characteristic data comprises medicine structure information, a medicine biological network and medicine function information; performing feature extraction on the drug structure information based on a natural language processing algorithm to generate drug structure characterization; learning the drug biological network data based on a knowledge graph embedding method, and generating drug network characterization containing heterogeneous nodes and edge relationships; feature aggregation is carried out on the multi-scale drug function information based on a graph neural network, and drug function characterization is generated; performing dynamic fusion on the drug structure characterization, the drug network characterization and the drug function characterization through cross matrix operation or a self-attention mechanism to generate a fusion feature vector; and training a prediction model based on the fused feature vector, and outputting a potential drug interaction prediction result. According to the scheme, multi-source data are integrated in an optimal mode, and the model generalization ability is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of bioinformatics, and in particular relates to a method and system for predicting potential drug interactions based on biological networks. Background Art

[0002] Prediction of drug interactions is a key link in clinical medicine and new drug development. Traditional methods mainly rely on clinical trials and in vitro experiments, but the high cost and long cycle severely limit their application efficiency. With the advancement of bioinformatics and computing technology, predictive models based on biological networks have gradually become a research focus. Early methods focused on homogeneous network modeling, such as predicting drug interactions through non-negative matrix factorization (such as the DDINMF model) or combining drug properties (such as the BioChemDDI model). This type of method regards network nodes as homogeneous entities and only uses topological structure information for inference, but ignores the heterogeneity of entities such as drugs, targets, and pathways in biological networks, making it difficult for the model to capture complex biological associations. More importantly, homogeneous network models cannot handle new drugs that do not appear in the training data, which seriously restricts their generalization ability.

[0003] In recent years, heterogeneous network-based methods (such as KGNN and KG2ECapsule) have significantly improved the semantic expression ability of features by introducing knowledge graph technology, integrating node types and relationship types into modeling. However, existing heterogeneous network methods still have shortcomings: the integration of multimodal data (such as drug molecular structure, functional properties, and biological networks) relies on simple operations (such as splicing or weighted averaging), and fails to fully explore the complementarity of different information sources; static fusion strategies are difficult to adapt to complex scenarios, resulting in information redundancy or loss, affecting prediction accuracy; the model's generalization ability for new drugs or sparse interactions is insufficient, limiting its practical application value. For example, although the KGNN model can utilize the heterogeneity of the knowledge graph, its feature fusion module lacks a dynamic weight allocation mechanism, making it difficult to balance the contribution of structural information and functional properties. Summary of the invention

[0004] The purpose of the present invention is to propose a drug interaction prediction method that effectively combines multimodal information and takes into account accuracy and generalization ability in view of the problems existing in the prior art. First, the limitations of the isomorphic network method are overcome, and the heterogeneous network information is fully utilized to integrate the node type and relationship type characteristics into the interaction network to improve the expression ability of the model. Secondly, biomedical features such as molecular structure, drug function and biological interaction network are fully utilized, and structural information and biochemical property information are taken into account to improve the generalization ability of the model. Thirdly, an adaptive information fusion mechanism is introduced to dynamically weight different modal information, integrate multi-source data in an optimal way, and improve the accuracy and stability of the prediction. Finally, predictions are made and the prediction results are applied in clinical medicine and new drug development to avoid further harm to patients caused by potential drug interactions and reduce the cost investment in drug development.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions: A method for predicting potential drug interactions based on biological networks, the method comprising: Obtain drug interaction datasets and drug characteristic data; The drug characteristic data includes drug structure information, drug biological network and drug function information; Extract drug structure information based on natural language processing algorithms to generate drug structure representations; Based on the knowledge graph embedding method, drug biological network data is learned to generate drug network representations containing heterogeneous nodes and edge relationships; Based on graph neural network, feature aggregation of multi-scale drug function information is performed to generate drug function representation; Dynamically fuse the drug structure representation, drug network representation, and drug function representation through cross matrix operations or self-attention mechanisms to generate a fused feature vector; The prediction model is trained based on the fused feature vector and the prediction results of potential drug interactions are output.

[0006] The cross matrix constructs feature interaction patterns through cross multiplication operations, uses convolution operations to extract local associations, and fuses multi-source information layer by layer. The self-attention mechanism dynamically assigns feature weights through learnable query, key, and value parameters to achieve context-aware modal complementary optimization. Experiments have shown that the prediction performance of the cross matrix and self-attention fusion methods are significantly better than the direct splicing method.

[0007] In the above-mentioned method for predicting potential drug interactions based on biological networks, extracting features from drug structure information based on natural language processing algorithms to generate drug structure representations specifically includes: The drug structure information includes SMILES structure data; the natural language processing algorithm is a CBOW model; The CBOW model embeds the characters in the SMILES structure data, calculates the probability of the central word through the context window, optimizes the CBOW model parameters by maximizing the probability of the actual central word appearing, and takes the mean of all character embedding vectors as the drug structure representation.

[0008] In the above-mentioned method for predicting potential drug interactions based on biological networks, the drug-biological network is knowledge graph network data; the knowledge graph embedding method is the CompIEx model; The CompIEx model models the triple relationship of the knowledge graph network through complex tensor decomposition and uses the loss function Optimize the embedding vectors of heterogeneous nodes and edge relationships, where Represents the model parameters, i.e., the embedding vector , w ri is the relation embedding representation of the i-th relation, e hi The i-th head node embedding representation, e ti The i-th tail node embedding represents, the above — represents e ti The conjugate vector of a vector, represents the predicted label, is the scoring function, ( h i, t i ,r i ) is in triple form, h i ,t i Represent entities h i and entities t i , r i Representing Entities h i and entities t i The type of biological relationship between represents the coefficient, and N represents the number of nodes. Entities are all individuals that appear in the biological network, and nodes are the graph structure representations of all these individuals.

[0009] In the above-mentioned method of predicting potential drug interactions based on biological networks, the graph neural network performs feature aggregation on multi-scale drug function information. For different functional relationship types, the features of heterogeneous neighbor nodes are mapped through learnable weights. After normalization and activation function processing, the node's own features are spliced ​​to generate drug function representation.

[0010] In the above-mentioned method for predicting potential drug interactions based on biological networks, the drug function representation is generated by feature aggregation of multi-scale drug function information based on graph neural networks, specifically including: Different relationship matrices are established for functional relationships from different perspectives, where each element is defined as , 1 and 0 respectively indicate the existence and non-existence of the relationship. Represents an entity member in an entity type. express v The heterogeneous neighbors of express u and v The type of functional relationship between them; For a specific function, first randomly initialize a d dimensional learnable matrix, representing the representation of all entities under this function f(u) ; Then, set different learnable weights for different features W r , representing all entities under the corresponding function f(u) Map to the corresponding functional space; Normalize the mapped representations and use learnable weights W c The standardized representations in each functional space are mapped to the common representation space, and then fused by neighborhood aggregation in the common representation space. v Neighbor entity information under different functions is obtained ; The fusion information of the entity Its characterization f(v) The final fused representation is concatenated and used to reconstruct the adjacency matrix of all functional relationships to optimize the graph neural network.

[0011] In the above-mentioned method for predicting potential drug interactions based on biological networks, the cross matrix operation method includes generating a matrix by cross multiplication of any two features, cross multiplying the matrix with the third feature again after dimensionality reduction by convolution, and finally generating a fused feature vector.

[0012] In the above-mentioned method for predicting potential drug interactions based on biological networks, the self-attention mechanism method includes using three features to calculate their corresponding importance scores, and then fusing them based on the weighted average of the scores: in (Q w ,K w ,V w )=x(Wq ,W k ,W v ) , W q , W k and W v Represent three learnable parameters respectively x represents the original feature vector.

[0013] In the above-mentioned method for predicting potential drug interactions based on biological networks, the prediction model is at least one selected from deep neural networks, random forests or support vector machines.

[0014] In the above-mentioned method for predicting potential drug interactions based on biological networks, the natural language processing algorithm is trained using all drugs downloaded from DrugBank to enable it to have drug structure characterization capabilities; The knowledge graph embedding method learns the structural relationship between heterogeneous entities based on the drug biological network knowledge graph containing several entities involving several entity types and several relationships involving several edge relationship types, so as to enable it to have the ability to generate drug network representation containing heterogeneous nodes and edge relationships; Graph neural networks take a variety of functional relationships as learning objects, and organically combine different functional information through domain aggregation to have the ability to generate drug functional representations.

[0015] In the above-mentioned method for predicting potential drug interactions based on biological networks, three characteristics of drug network representation, drug function representation, and drug structure representation and drug interaction data are used as input to train the cross-matrix operation / self-attention mechanism and the prediction model, so that the prediction model has the ability to output the prediction results of potential drug interactions.

[0016] A system for predicting potential drug interactions based on biological networks, comprising a drug structure feature extraction module, a drug biological network feature extraction module, a drug function feature extraction module, a feature fusion module and a prediction module; A drug structure feature extraction module is used to extract drug structure features based on drug structure information; A drug biological network feature extraction module is used to extract drug biological network features based on drug biological network data; A drug function feature extraction module is used to perform feature aggregation on multi-scale drug function information to extract drug function features; The feature fusion module is used to dynamically fuse drug structure features, drug biological network features, and drug functional features to obtain fusion features and input them into the prediction module; A prediction module, used to output relevant drug interaction prediction results based on fusion features; The drug structure feature extraction module, drug biological network feature extraction module, and drug function feature extraction module are obtained by training using their respective data; The feature fusion module and the prediction module are trained based on the outputs of the drug structure feature extraction module, the drug biological network feature extraction module, the drug function feature extraction module and the drug interaction dataset.

[0017] When making predictions, the drug structure feature extraction module, the drug biological network feature extraction module, and the drug function feature extraction module respectively output corresponding features based on the relevant data of the drug pair to be predicted. The feature fusion module outputs fused features based on the above three features, and the prediction module then makes a prediction of the interaction between the drug pair to be predicted based on the fused features.

[0018] The advantages of the present invention are: 1) This solution integrates three types of heterogeneous data: drug molecular structure, biological network topology, and functional attributes. Through multiple feature extraction modules, multimodal representation learning can improve the integrity of drug representation and make the prediction effect more accurate. 2) The generalization ability of the model is enhanced through multiple features. By jointly modeling biological networks, drug functions and drug structures, the model can handle drugs with missing information; 3) This solution uses a natural language processing algorithm to perform context-aware encoding on SMILES strings, converts chemical symbols into low-dimensional dense vectors, and generates drug structure representations that break through the limitations of traditional molecular descriptors; uses knowledge graph embedding methods to capture the antisymmetric relationships between heterogeneous entities such as drugs, targets, and pathways to generate network representations rich in semantic information; dynamically aggregates drug function information through multi-scale graph neural networks, and normalizes weight distribution to fuse heterogeneous neighbor features of different functional relationships. Through the above three methods, a multimodal feature extraction framework covering the chemical, topological, and biological properties of drug effects is realized, solving the problems of single information and insufficient generalization capabilities of traditional solutions; 4) By constructing a new model architecture based on graph neural networks, it can effectively learn node representations in heterogeneous graphs and integrate multi-scale information, solving the problems existing in homogeneous graph learning models, and has good scalability; 5) This scheme optimizes the information fusion strategy, enhances the complementarity of different modal information, solves the problem of information loss caused by simple splicing or weighted summation, and enables the model to be integrated more efficiently. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 The system block diagram of the system for predicting potential drug interactions based on biological networks of the present invention; Figure 2 A flow chart of the method for predicting potential drug interactions based on biological networks of the present invention; Figure 3 A flowchart of a drug structure information extraction module in the method for predicting potential drug interactions based on biological networks of the present invention; Figure 4 A flowchart of a drug function information extraction module in the method for predicting potential drug interactions based on biological networks of the present invention; Figure 5 This is a flow chart of an information fusion module based on a self-attention mechanism in the method for predicting potential drug interactions based on biological networks of the present invention; Figure 6 The ROC curve and PR curve of the self-attention mechanism feature fusion module based on five-fold cross validation in the embodiment of the present invention; Figure 7 The ROC curve and PR curve of the cross matrix feature fusion module based on five-fold cross validation in the embodiment of the present invention; Figure 8 The ROC curve and PR curve of the traditional direct splicing feature fusion module based on five-fold cross validation are used in the embodiment of the present invention; Fig. 9 Ablation experiments on three feature extraction modules are carried out for the present invention; Fig.10 ROC curve and PR curve based on five-fold cross validation for using the DecisionTreeClassifier model on the same data; Fig.11 ROC curve and PR curve based on five-fold cross validation for using GaussianNB model on the same data; Fig.12 The ROC curve and PR curve based on five-fold cross validation using the LogisticRegression model on the same data; Fig.13 The figure shows the comparison of the results of the embodiment of the present invention and the three models of DecisionTreeClassifier, GaussianNB and LogisticRegression under all evaluation indicators when trained and predicted on the same data; Fig.14 This is a case study of dexamethasone, a drug related to a viral disease, according to an embodiment of the present invention.

[0020] Figure numerals: drug structure feature extraction module 1; drug biological network feature extraction module 2; drug function feature extraction module 3; feature fusion module 4; prediction module 5. DETAILED DESCRIPTION

[0021] This protocol provides a method for predicting potential drug interactions based on biological networks, such as Figure 1 and Figure 2 As shown, it mainly includes three feature extraction modules: drug structure feature extraction module 1, drug biological network feature extraction module 2, drug function feature extraction module 3, and a feature fusion module 4. Through the feature extraction and integration of drug biological network, drug function information and drug structure information, a prediction module 5 is used to realize the identification of potential drug interactions. The specific process is as follows: S1. Collect and download drug interaction datasets for training prediction and data for extracting drug features, including drug-related drug biological networks, drug functional information, and drug structure information.

[0022] In this example, the drug interaction dataset is obtained from the public database TWOSIDES, including 548 drugs and 48,584 drug interaction data. In the feature file, the drug biological network data is obtained from the DRKG public knowledge graph dataset; drug function information including chemical structure, target, transporter, enzyme, biological pathway, indication, side effect and non-side effect is collected from DrugBank, SIDER, PubChem, KEGG and OFFSIDES databases; drug self-structure information is SMILES data collected from DrugBank.

[0023] S2. Figure 3 As shown in the figure, the SMILES data of all collected drugs are regarded as biological short sentences, and the natural language processing algorithm CBOW is used to represent the characteristics of drug SMILES characters. All symbol words in the biological short sentences are encoded, and each word is iteratively regarded as the central word, and the probability of its own occurrence is calculated using its context to optimize the symbol encoding. Specifically, each character is first encoded as x i , and then through the learnable weights W Construct the embedding vector of each character, and then calculate the average value of the embedding vector of the context character with a pre-set window size to obtain the hidden layer vector, that is, Finally, the probability of all words being the central word is calculated through the learnable output layer, and the model is optimized to complete the training by maximizing the probability of the actual central word appearing. The loss function is Among them, w o Indicates the central word; w I1 ,w I2 , …,w I Represents the context word, v’wj represents the embedding vector of the context word, v’ wo Represents the embedding vector of the central word.

[0024] In this process, in order to effectively learn the representation of drug SMILES, all drugs downloaded from DrugBank were used for training, and then applied to the drugs designed in the dataset of this project. CBOW was used to convert each drug SMILES into a matrix of stacked character embedding vectors, and the mean of all character embedding vectors was used as the representation of the drug's own structural information.

[0025] S3. The drug-bionetwork is mainly knowledge graph network data, and learning is performed based on the collected knowledge graph network data to construct features that take into account heterogeneous nodes and edge relationships. The knowledge graph used in this embodiment contains 97,238 entities, involving 13 entity types, and 5,874,261 relationships, involving 28 edge relationship types, and does not contain drug interaction relationships to avoid label leakage. The structural relationship between heterogeneous entities is captured through the knowledge graph embedding model. The graph covers all the drugs used in the project to construct a representation of the drug-bionetwork information. This process uses the knowledge graph embedding method to effectively learn rich multi-entity relationship type information to enhance drug representation.

[0026] Furthermore, considering that in large-scale medical knowledge graphs, drug-related action relationships are not symmetric, this solution uses the ComplEx method to capture antisymmetric relationships. Specifically, given a biomedical knowledge graph ,in E represents a set of biological entities, R Represents the type of biological relationship between them, expressed as a triple ( h i ,t i, r i ). ComplEx performs low-rank decomposition of the three-dimensional tensor of the heterogeneous graph based on complex numbers to obtain the vector representation of each entity and relationship. Each slice of the tensor represents the graph adjacency matrix under different relationship types. h i With entity t i The probability of interaction between ,in represents the sigmoid function, Represents the model parameters, i.e., the embedding vector , w ri is the relation embedding representation of the i-th relation, e hi The i-th head node embedding representation, eti The i-th tail node embedding represents, the above — represents e ti The conjugate vector of a vector, represents the predicted label, is its scoring function. Finally, through the loss function Optimize the drug embedding vector containing heterogeneous node information and edge relationship information.

[0027] As mentioned above, entities are all individuals appearing in the biological network, and nodes are the graph structure representation of all these individuals.

[0028] S4. Learning based on the collected multi-scale drug function information. Different types of functions represent the capabilities and roles of drugs under specific types. This solution designs a learning model based on the graph neural network algorithm that can be applied to various scales at the same time, and calculates a unified drug function representation based on considerations at each scale.

[0029] Specifically, Figure 4 As shown in the figure, different relationship matrices are established for functional relationships under different perspectives, where each element is defined as , 1 and 0 respectively indicate the existence and non-existence of the relationship. Represents an entity member in an entity type. express v The heterogeneous neighbors of express u and v The type of functional relationship between them.

[0030] For a specific function, first randomly initialize a d dimensional learnable matrix, representing the representation of all entities under this function Then, different learnable weights are set for different functions W r , mapping all entity representations under the corresponding function to the corresponding function space. After further standardization, the learnable weights are used W c Map them into a common representation space, where they are fused by neighborhood aggregation v Neighbor entity information under different functions is obtained , and its specific implementation process is defined as: (1) in Representation Node v In the relationship type r The set of adjacent nodes under represents the ReLU activation function, The representative normalization term is defined as (2) Finally, the fusion information of the entity is combined with the representation of the entity The final fusion representation is concatenated and used to reconstruct the adjacency matrix of all functional relationships to optimize the model. The fusion process is formulated as: (3) Among them, W' represents the weight of constructing the final node embedding vector, Concat represents directly connecting the two vectors in the brackets; b' represents the bias of constructing the final node embedding vector The loss function is defined as: (4) in is the projection matrix associated with the edge type.

[0031] This solution uses multiple functional relationships as learning objects and cleverly combines different functional information organically instead of simply adding and averaging. In addition, the model designed in this solution is scalable and can effectively enrich drug features by adding more functional information in specific problems. In addition, the feature representation of drug-related entities also incorporates multi-scale relationship information, which is calculated by the model as an accessory and can be used as a pre-training feature and applied to other computational experiments.

[0032] S5. For the three extracted features, a fusion strategy is designed to fuse all constructed feature information and finally input it into the prediction module 5 for prediction. This embodiment constructs two information integration strategies: 1) Based on the cross matrix, the two features are constructed into a matrix through the cross multiplication operation. The operation process is as follows: (5) C i ‘ Represents the cross matrix obtained after fusing two features; Then use the convolution operation on this matrix to reduce its dimension to a fused vector c i , the fused vector is further fused with the third feature by performing the above operation, and finally used for model prediction.

[0033] 2) Based on the self-attention mechanism, such as Figure 5 As shown in the figure, the attention module uses the three features to calculate their corresponding importance scores, and then performs weighted averaging based on the scores for fusion. The calculation process is as follows: (6) in ,W q , W k and W v Represent three learnable parameters respectively x represents the original feature vector.

[0034] When put into use, use any of the above strategies to achieve feature fusion.

[0035] All constructed feature information is fused through the above fusion strategy and finally input into the prediction module 5 for prediction. The prediction module 5 is used to implement the prediction model. Figure 2 As shown, the prediction model can be selected from support vector machine SVM, deep neural network, random forest, rotation forest, etc.

[0036] In order to verify the effect of this solution, this example evaluates the prediction performance of the model based on five-fold cross validation, and conducts a case study on dexamethasone, a drug related to the treatment of a viral disease: The dataset was divided into five parts, each of which was used as a test set in turn, and the remaining four parts were used as training sets. The results of five training and testing were used as the final evaluation results of the model. The training negative samples were negatively sampled to ensure the balance of positive and negative samples. For the case study, we used the entire dataset to train the model, and then the test set was the data in the original dataset that did not have annotated dexamethasone-related drug pairs, ensuring that the training data would not be used for testing. Finally, the results with high model prediction scores and samples that were not annotated in the dataset were searched in the literature to determine whether there was indeed a potential interaction.

[0037] Figure 6 The feature fusion module 4 uses the self-attention mechanism. Based on the ROC curve and PR curve of five-fold cross-validation, we can see from the figure that mean-AUC (average area under the curve) and mean-AP (average precision) are as high as 0.9663 and 0.9623, respectively, indicating that the prediction effect is very good.

[0038] Figure 7 The feature fusion module 4 uses a cross matrix, based on the ROC curve and PR curve of five-fold cross validation. From the figure, we can see that the mean-AUC (average area under the curve) and mean-AP (average precision) are as high as 0.9673 and 0.9648 respectively, and the prediction effect is better than that of the self-attention mechanism feature fusion module.

[0039] Figure 8 The embodiment of the present invention uses the traditional direct splicing feature fusion module 4, and the ROC curve and PR curve based on the five-fold cross validation are not as good as Figure 6 and Figure 7The two methods show that the fusion method of this scheme is more effective than the traditional direct splicing. However, it can be seen from the figure that the prediction effect of this method is also good, which shows that the feature extraction module of this scheme is effective.

[0040] Fig. 9 This is the result of the ablation experiment on three feature extraction modules implemented in the present invention. It can be seen that the best performance can be achieved only when all modules are used, indicating that each module is important, but the gap is not particularly obvious, indicating that our method can effectively utilize the complementarity between different information, making the generalization ability very strong. In the figure, auc_roc represents the area under the curve, auc_pr represents the area under the precision-recall curve; nobehavior represents that this solution removes the drug biological network module, nosimi represents that this solution removes the drug function information module, and nostr represents that this solution removes the drug structure information module.

[0041] Fig.10 In order to use the DecisionTreeClassifier model under the same data, the ROC curve and PR curve based on the five-fold cross validation are compared with this solution. It can be seen that this solution has a great performance improvement over the DecisionTreeClassifier model method, which reflects the advantages of this solution method.

[0042] Fig.11 The ROC curve and PR curve based on five-fold cross validation are shown in the figure using the GaussianNB model under the same data. Compared with the method of this solution, it can be seen that this solution also has a significant performance improvement over the GaussianNB model, which reflects the advantages of this solution.

[0043] Fig.12 In order to use the LogisticRegression model under the same data, the ROC curve and PR curve based on the five-fold cross validation are shown. Compared with this solution, it can be seen that this solution also has a significant performance improvement over the LogisticRegression model, which reflects the advantages of this solution.

[0044] Fig.13 The results of the embodiment of the present invention and the three models of DecisionTreeClassifier, GaussianNB and LogisticRegression under the same data training and prediction under all evaluation indicators are compared. It can be seen that the model provided by this solution obtains the highest score under all indicators, which illustrates the advantages of the method of this solution.

[0045] Fig.14This is a case study of dexamethasone, a drug related to a viral disease, in an embodiment of the present invention. The figure shows the names of the top 20 potential interacting drugs with the highest prediction scores, and checks the literature to see whether the predicted drugs do have potential interactions with dexamethasone. It can be seen from the figure that the model of this scheme successfully predicted 20 potential interacting drugs, 18 of which were confirmed by the literature, which basically proves the success of this scheme in the case study of dexamethasone, a drug for a viral disease.

[0046] The specific embodiments described herein are merely examples of the spirit of the present invention. Those skilled in the art may make various modifications or additions to the specific embodiments described or replace them in similar ways, but they will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

Claims

1. A method for predicting potential drug interactions based on biological networks, characterized in that: The method includes, Obtain drug interaction datasets and drug characteristic data; The drug characteristic data includes drug structure information, drug biological network and drug function information; Extract drug structure information based on natural language processing algorithms to generate drug structure representations; Based on the knowledge graph embedding method, drug biological network data is learned to generate drug network representations containing heterogeneous nodes and edge relationships; Based on graph neural network, feature aggregation of multi-scale drug function information is performed to generate drug function representation; Dynamically fuse the drug structure representation, drug network representation, and drug function representation through cross matrix operations or self-attention mechanisms to generate a fused feature vector; The prediction model is trained based on the fused feature vector and the prediction results of potential drug interactions are output.

2. The method for predicting potential drug interactions based on biological networks according to claim 1, characterized in that: The feature extraction of drug structure information based on natural language processing algorithm to generate drug structure representation specifically includes: The drug structure information includes SMILES structure data; the natural language processing algorithm is a CBOW model; The CBOW model embeds the characters in the SMILES structure data, calculates the probability of the central word through the context window, optimizes the CBOW model parameters by maximizing the probability of the actual central word appearing, and takes the mean of all character embedding vectors as the drug structure representation.

3. The method for predicting potential drug interactions based on biological networks according to claim 1, characterized in that: The drug biological network is a knowledge graph network data; the knowledge graph embedding method is a CompIEx model; The CompIEx model models the triple relationship of the knowledge graph network through complex tensor decomposition and uses the loss function Optimize the embedding vectors of heterogeneous nodes and edge relationships, where Represents the model parameters, i.e., the embedding vector , w ri is the relation embedding representation of the i-th relation, e hi The i-th head node embedding representation, e ti The i-th tail node embedding represents, the above — represents e ti The conjugate vector of a vector, represents the predicted label, is the scoring function, ( h i ,t i ,r i ) is in triple form, h i ,t i Represent entities h i and entities t i , r i Representing Entities h i and entities t i The type of biological relationship between represents the coefficient, and N represents the number of nodes.

4. The method for predicting potential drug interactions based on biological networks according to claim 1, characterized in that: In the feature aggregation of multi-scale drug function information by graph neural network, the features of heterogeneous neighbor nodes are mapped through learnable weights for different functional relationship types. After standardization and activation function processing, the node's own features are spliced ​​to generate drug function representation.

5. The method for predicting potential drug interactions based on biological networks according to claim 4, characterized in that: The feature aggregation of multi-scale drug function information based on graph neural network generates drug function representation, which specifically includes: Different relationship matrices are established for functional relationships from different perspectives, where each element is defined as , 1 and 0 respectively indicate the existence and non-existence of the relationship. Represents an entity member in an entity type. express v The heterogeneous neighbors of express u and v The type of functional relationship between them; For a specific function, first randomly initialize a d dimensional learnable matrix, representing the representation of all entities under this function f(u) ; Then, set different learnable weights for different features W r , representing all entities under the corresponding function Map to the corresponding functional space; Normalize the mapped representations and use learnable weights W c The standardized representations in each functional space are mapped to the common representation space, and then fused by neighborhood aggregation in the common representation space. v Neighbor entity information under different functions is obtained ; The fusion information of the entity Its characterization f(v) The final fused representation is concatenated and used to reconstruct the adjacency matrix of all functional relationships to optimize the graph neural network.

6. The method for predicting potential drug interactions based on biological networks according to claim 1, characterized in that: The cross matrix operation method includes generating a matrix by cross-multiplying any two features, cross-multiplying it with the third feature again after dimensionality reduction by convolution, and finally generating a fused feature vector.

7. The method for predicting potential drug interactions based on biological networks according to claim 1, characterized in that: The self-attention mechanism method includes using three features to calculate their corresponding importance scores, and then performing weighted averaging based on the scores for fusion: (6) in (Q w ,K w ,V w )=x(W q ,W k ,W v ) , W q , W k and W v Represent three learnable parameters respectively x represents the original feature vector.

8. The method for predicting potential drug interactions based on biological networks according to claim 1, characterized in that: The prediction model is at least one selected from a deep neural network, a random forest or a support vector machine; The natural language processing algorithm is trained using all drugs downloaded from DrugBank to enable it to have drug structure characterization capabilities; The knowledge graph embedding method learns the structural relationship between heterogeneous entities based on the drug biological network knowledge graph containing several entities involving several entity types and several relationships involving several edge relationship types, so as to enable it to have the ability to generate drug network representation containing heterogeneous nodes and edge relationships; Graph neural networks take a variety of functional relationships as learning objects, and organically combine different functional information through domain aggregation to have the ability to generate drug functional representations.

9. The method for predicting potential drug interactions based on biological networks according to claim 8, characterized in that: Taking drug network representation, drug function representation, drug structure representation and drug interaction data as input, the cross-matrix operation / self-attention mechanism and prediction model are trained to enable the prediction model to have the ability to output potential drug interaction prediction results.

10. A system for predicting potential drug interactions based on biological networks, characterized in that: It includes a drug structure feature extraction module (1), a drug biological network feature extraction module (2), a drug function feature extraction module (3), a feature fusion module (4) and a prediction module (5); A drug structure feature extraction module (1), used to extract drug structure features based on drug structure information; A drug biological network feature extraction module (2), used for extracting drug biological network features based on drug biological network data; A drug function feature extraction module (3), used for performing feature aggregation on multi-scale drug function information to extract drug function features; A feature fusion module (4) is used to dynamically fuse drug structure features, drug biological network features, and drug functional features to obtain fused features, and input the fused features into a prediction module (5); A prediction module (5), used for outputting relevant drug interaction prediction results based on the fusion features; The drug structure feature extraction module (1), the drug biological network feature extraction module (2), and the drug function feature extraction module (3) are obtained by training using their respective data; The feature fusion module (4) and the prediction module (5) are trained based on the outputs of the drug structure feature extraction module (1), the drug biological network feature extraction module (2), the drug function feature extraction module (3) and the drug interaction dataset.

Citation Information

Cited By

  • Drug synergistic effect prediction method, system and equipment based on heterograph tensor decomposition and medium

    CN120913697A

  • Multi-scale feature fusion drug interaction prediction method based on biological knowledge graph

    CN121789787A

  • A multi-scale feature fusion drug interaction prediction method based on a biological knowledge graph

    CN121789787B