A multi-scale feature fusion drug interaction prediction method based on a biological knowledge graph

CN121789787BActive Publication Date: 2026-08-11BEIJING INST OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0002]在现代临床医疗中,药物-药物相互作用(DDI)可能导致药效降低或引发严重不良反应

Benefits of technology

[0010]根据本发明提供的基于生物知识图谱的多尺度特征融合药物相互作用预测方法,获取待预测药物的分子结构数据、生化代谢通路数据及副作用数据,其中,待预测药物包括第一待预测药物与第二待预测药物两个被研究对象。基于分子结构数据,确定第一待预测药物的第一节点特征和第二待预测药物的第二节点特征;其中,第一节点特征和第二节点特征均用于表征待预测药物的分子结构特征;基于生化代谢通路数据、第一节点特征和第二节点特征,确定第一待预测药物与第二待预测药物的联合生物学特征;该联合生物学特征有效捕捉了两种药物在体内代谢过程中的协同或拮抗关联,例如二者在共同代谢通路中的竞争关系、代谢产物的相互影响等。同时,基于副作用数据,确定第一待预测药物的第一副作用特征和第二待预测药物的第二副作用特征。基于第一节点特征、第二节点特征、联合生物学特征、第一副作用特征和第二副作用特征,构建融合特征;通过融合特征,确定药物与药物相互作用预测结果,如此,本发明能够提高药物与药物相互作用预测的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure SMS_1
    Figure SMS_1
Patent Text Reader

Abstract

This invention relates to the field of medical information processing technology, and more particularly to a multi-scale feature fusion method for predicting drug interactions based on biological knowledge graphs. The method includes: acquiring molecular structure data, biochemical metabolic pathway data, and side effect data of the drug to be predicted; determining a first node feature of a first drug to be predicted and a second node feature of a second drug to be predicted based on the molecular structure data; determining the combined biological characteristics of the first and second drugs to be predicted based on the biochemical metabolic pathway data, the first node feature, and the second node feature; determining a first side effect feature and a second side effect feature based on the side effect data; determining a fusion feature based on the first node feature, the second node feature, the combined biological feature, the first side effect feature, and the second side effect feature; and determining the drug-drug interaction prediction result based on the fusion feature. Thus, this invention can improve the accuracy of drug-drug interaction prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical information processing technology, and in particular to a method for predicting drug interactions based on multi-scale feature fusion of biological knowledge graphs. Background Technology

[0002] In modern clinical medicine, drug-drug interactions (DDIs) can lead to reduced drug efficacy or serious adverse reactions.

[0003] Existing deep learning-based drug interaction (DDI) prediction methods suffer from the following technical drawbacks: 1. Lagging feature fusion (delayed fusion problem): Existing methods typically extract features from two drugs independently, only splicing them at the fully connected layer. This "late fusion" ignores the "atomic-level induced fit" nature of biochemical reactions, causing the model to lose cooperative signals between molecules and atoms (e.g., failing to capture how the oxidizing group of one drug induces the reducing group of another). 2. Graph noise and sparsity (static graph bottleneck): Existing knowledge graph-based methods typically perform reasoning on a fixed full-graph structure. However, biomedical KGs contain a large number of generalized noisy connections (such as high-frequency nodes like ATP and water molecules leading to over-smoothing), while new drugs often lack connections (cold start problem). Static graphs cannot dynamically remove noise or complete potential implicit associations for specific tasks, resulting in low accuracy in drug-drug interaction prediction.

[0004] Based on this, the present invention proposes a multi-scale feature fusion drug interaction prediction method based on biological knowledge graph to solve the above-mentioned technical problems. Summary of the Invention

[0005] This invention describes a multi-scale feature fusion method for predicting drug interactions based on biological knowledge graphs, which can improve the accuracy of drug-drug interaction prediction.

[0006] According to a first aspect, the present invention provides a multi-scale feature fusion method for predicting drug interactions based on biological knowledge graphs, comprising: Acquire molecular structure data, biochemical metabolic pathway data, and side effect data of the drug to be predicted; wherein, the drug to be predicted includes a first drug to be predicted and a second drug to be predicted; Based on the molecular structure data, the first node features of the first drug to be predicted and the second node features of the second drug to be predicted are determined. Based on the biochemical metabolic pathway data, the first node features, and the second node features, the combined biological characteristics of the first and second drugs to be predicted are determined. Based on the side effect data, a first side effect characteristic of the first drug to be predicted and a second side effect characteristic of the second drug to be predicted are determined. Based on the first node feature, the second node feature, the combined biological feature, the first side effect feature, and the second side effect feature, the fusion feature is determined; Based on the fusion features, the predicted drug-drug interactions are determined.

[0007] According to a second aspect, the present invention provides a multi-scale feature fusion drug interaction prediction device based on biological knowledge graphs, comprising: The acquisition unit is configured to acquire molecular structure data, biochemical metabolic pathway data, and side effect data of the drug to be predicted; wherein the drug to be predicted includes a first drug to be predicted and a second drug to be predicted. The first data processing unit is configured to determine the first node features of the first drug to be predicted and the second node features of the second drug to be predicted based on the molecular structure data. The second data processing unit is configured to determine the combined biological characteristics of the first drug to be predicted and the second drug to be predicted based on the biochemical metabolic pathway data, the first node characteristics, and the second node characteristics. The third data processing unit is configured to determine the first side effect characteristics of the first drug to be predicted and the second side effect characteristics of the second drug to be predicted based on the side effect data. The fourth data processing unit is configured to determine the fusion feature based on the first node feature, the second node feature, the combined biological feature, the first side effect feature, and the second side effect feature; The fifth data processing unit is configured to determine the drug-drug interaction prediction results based on the fusion features.

[0008] Thirdly, embodiments of this specification also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the method described in any embodiment of this specification.

[0009] Fourthly, embodiments of this specification also provide a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the methods described in any embodiment of this specification.

[0010] According to the multi-scale feature fusion drug interaction prediction method based on biological knowledge graph provided by the present invention, molecular structure data, biochemical metabolic pathway data, and side effect data of the drug to be predicted are obtained. The drug to be predicted includes two research objects: a first drug to be predicted and a second drug to be predicted. Based on the molecular structure data, a first node feature of the first drug to be predicted and a second node feature of the second drug to be predicted are determined; both the first and second node features are used to characterize the molecular structure features of the drug to be predicted. Based on the biochemical metabolic pathway data, the first node feature, and the second node feature, a joint biological feature of the first and second drugs to be predicted is determined. This joint biological feature effectively captures the synergistic or antagonistic relationship between the two drugs in the in vivo metabolic process, such as the competitive relationship between them in a common metabolic pathway, and the mutual influence of metabolites. Simultaneously, based on the side effect data, a first side effect feature of the first drug to be predicted and a second side effect feature of the second drug to be predicted are determined. Based on the first node feature, the second node feature, the joint biological feature, the first side effect feature, and the second side effect feature, a fusion feature is constructed. Through the fusion feature, the drug-drug interaction prediction result is determined. Thus, the present invention can improve the accuracy of drug-drug interaction prediction. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A flowchart illustrating a drug interaction prediction method according to one embodiment is shown; Figure 2 A schematic block diagram of a drug interaction prediction device according to one embodiment is shown. Detailed Implementation

[0013] The solution provided by the present invention will now be described with reference to the accompanying drawings.

[0014] Figure 1 A flowchart illustrating a drug interaction prediction method according to one embodiment is shown. It will be understood that this method can be executed by any device, apparatus, platform, or cluster of devices with computing and processing capabilities. Figure 1 As shown, the method includes: Step 100: Obtain molecular structure data, biochemical metabolic pathway data, and side effect data of the drug to be predicted; wherein, the drug to be predicted includes a first drug to be predicted and a second drug to be predicted; Step 102: Based on molecular structure data, determine the first node features of the first drug to be predicted and the second node features of the second drug to be predicted; Step 104: Based on biochemical metabolic pathway data, first node characteristics, and second node characteristics, determine the combined biological characteristics of the first and second predictable drugs. Step 106: Based on the side effect data, determine the first side effect characteristics of the first drug to be predicted and the second side effect characteristics of the second drug to be predicted; Step 108: Determine the fusion features based on the first node features, the second node features, the combined biological features, the first side effect features, and the second side effect features; Step 110: Based on the fusion features, determine the predicted results of drug-drug interactions.

[0015] In this embodiment, molecular structure data, biochemical metabolic pathway data, and side effect data of the drugs to be predicted are acquired. The drugs to be predicted include two research subjects: a first drug to be predicted and a second drug to be predicted. Based on the molecular structure data, first node characteristics of the first drug to be predicted and second node characteristics of the second drug to be predicted are determined. Both the first and second node characteristics are used to characterize the molecular structure of the drugs to be predicted. Based on the biochemical metabolic pathway data, the first node characteristics, and the second node characteristics, the combined biological characteristics of the first and second drugs to be predicted are determined. These combined biological characteristics effectively capture the synergistic or antagonistic relationships between the two drugs during their metabolism in vivo, such as their competitive relationship in a common metabolic pathway and the mutual influence of their metabolites. Simultaneously, based on the side effect data, first side effect characteristics of the first drug to be predicted and second side effect characteristics of the second drug to be predicted are determined.

[0016] Based on the first node feature, the second node feature, the combined biological feature, the first side effect feature, and the second side effect feature, a fusion feature is constructed; through the fusion feature, the prediction result of drug-drug interaction is determined. Thus, the present invention can improve the accuracy of drug-drug interaction prediction.

[0017] In one embodiment of the present invention, determining a first node feature of a first drug to be predicted and a second node feature of a second drug to be predicted based on molecular structure data includes: Based on molecular structure data, determine the molecular structure diagram; Based on the molecular structure diagram, the first node features of the first drug to be predicted and the second node features of the second drug to be predicted are determined. The molecular structure diagram includes atomic type, atomic degree, formal charge, hybridization mode, aromaticity and chirality labels.

[0018] In this embodiment, a molecular structure map is determined based on molecular structure data; based on the molecular structure map, the first node features of the first drug to be predicted and the second node features of the second drug to be predicted are determined; the drug SMILES sequence is converted into a molecular map using the RDKit tool, which includes atomic type (one-hot encoding, such as C, N, O, etc.), atomic degree, formal charge, hybridization (such as SP2, SP3), aromaticity (Bool value), and chirality.

[0019] In one embodiment of the present invention, fusion features are determined based on first node features, second node features, combined biological features, first side effect features, and second side effect features, including: Based on the features of the first node and the features of the second node, determine the interaction information; The first node features and the second node features are updated based on the interaction information to obtain the updated first node features and the updated second node features. The updated first node features and the updated second node features are read out and aggregated to obtain the microstructure characterization of the first drug to be predicted. Based on microstructural characterization, combined biological characteristics, first side effect characteristics, and second side effect characteristics, fusion characteristics were determined.

[0020] In this embodiment, cross-drug molecule interaction analysis is performed on the first and second node features. An attention mechanism is used to capture potential correlation signals between the two types of features, generating interaction information that characterizes the potential for microscopic interactions between drug molecules. Next, this interaction information is used to adaptively update the original first and second node features, strengthening the feature dimensions related to drug interactions and suppressing invalid noise information, resulting in updated first and second node features. Subsequently, a graph readout aggregation strategy is used to integrate the updated two types of node features, providing a first microstructure characterization of the drug to be predicted that comprehensively reflects the microscopic properties of drug molecules. Finally, this microstructure characterization is multimodally fused with joint biological features, first side effect features, and second side effect features. Through feature alignment and adaptive weight allocation, a fused feature is obtained.

[0021] In one embodiment of the present invention, based on biochemical metabolic pathway data, first node features, and second node features, the combined biological characteristics of a first predictable drug and a second predictable drug are determined, including: Construct a biochemical metabolic pathway knowledge graph based on biochemical metabolic pathway data; The biochemical metabolic pathway knowledge graph was structurally pruned and completed to obtain an optimized biochemical metabolic pathway knowledge graph. In the optimized biochemical metabolic pathway knowledge graph, information about multiple adjacent nodes related to the first and second drugs to be predicted is extracted to obtain a joint knowledge subgraph. Based on joint knowledge subgraphs and microstructure characterization, joint biological characteristics are determined; The node information includes targets, enzymes, pathways, diseases, and connection edges.

[0022] In this embodiment, firstly, using biochemical metabolic pathway data as the data source, multiple types of biomedical information, such as drug target associations, enzyme catalytic reactions, and pathway regulatory relationships, are integrated to construct a biochemical metabolic pathway knowledge graph. This graph is used to characterize the biological association network in the in vivo metabolism of drugs. Secondly, to address the redundant noise connections (such as invalid associations formed by universal metabolic molecules) and sparsity issues in the original knowledge graph, nodes and edges that are meaningless for predicting drug interactions are removed, and potential associations between drugs, pathways, and targets are supplemented, resulting in an optimized biochemical metabolic pathway knowledge graph. Subsequently, information on multiple adjacent nodes related to the first and second drugs to be predicted is extracted from the optimized biochemical metabolic pathway knowledge graph to obtain a joint knowledge subgraph. Based on the joint knowledge subgraph and microstructural characterization, joint biological characteristics are determined; the node information includes targets, enzymes, pathways, diseases, and connecting edges.

[0023] In one embodiment of the present invention, the combined biological characteristics are determined by the following formula:

[0024] In the formula, Let be the retention probability of the edges in the joint knowledge subgraph. Let be the node feature vector of one endpoint connected to the edge to be evaluated in the joint knowledge subgraph. Let be the node feature vector of the other endpoint connected to the edge to be evaluated in the joint knowledge subgraph. The embedding type vector corresponding to the edge. For microstructure characterization, For cosine similarity, Let be the feature vector of any non-adjacent node pair in the joint knowledge subgraph. Let be the feature vector of another arbitrary non-adjacent node pair in the joint knowledge subgraph. The importance score of the node. As the first learnable parameter, As the second learnable parameter, For the final embedding of each node in the reconstructed joint knowledge subgraph, For the aforementioned combined biological characteristics, It is the set of all valid nodes in the joint knowledge subgraph centered on the drug pair to be predicted, after structural pruning and completion optimization.

[0025] In this embodiment, an MLP scorer is used to calculate the retention probability of edges in the joint knowledge subgraph. If the retention probability of an edge is less than a preset retention probability, the edge is cut off, dynamically filtering out general nodes (such as generalized ATP connections) that are irrelevant to the current DDI task, preventing oversmoothing and improving accuracy. If the cosine similarity is greater than a preset similarity, a virtual edge of type Latent_Relation is created. This allows new or rare drugs without known connections to use pathways of similar entities for inference, solving the cold start problem. Through node importance scores, the model can automatically determine which nodes (such as specific target proteins) are most critical to the current DDI prediction and assign them high weights; and which nodes (such as general background) are unimportant and assign them low weights.

[0026] In one embodiment of the present invention, the interaction information is determined by the following formula:

[0027] In the formula, The first atom of the drug to be predicted Atoms of the second drug to be predicted The potential tendency to react Features of the first node Features of the second node The weight matrix is ​​a learnable bilinear weight matrix. For the first drug to be predicted, atoms The maximum response sensitivity affected by the overall influence of the second drug to be predicted The total number of atoms, For interactive information, For the similarity of internal structure, The dimension of the feature vector. For adjustable weights, The first drug to be predicted contains atoms Maximum response sensitivity induced by the second predictable drug.

[0028] In this embodiment, It is a learnable bilinear weight matrix that does not just calculate similarity, but learns atomic matching patterns across molecules (e.g., matching of positively charged atoms with negatively charged atoms). Atoms representing drug A The maximum response sensitivity affected by drug B as a whole. If A large value indicates that the atom is a potential active site; standard attention only calculates... (Internal structural similarity). This invention incorporates... If atoms and The bias values ​​of all drugs A and B are very high (i.e., all are susceptible to the influence of drug B), and the attentional weights between them are forcibly amplified. This makes the representation of drug A context-aware. The weights are adjustable. Effect: This causes the information aggregation process within the first predictor drug to focus on the active regions that "respond to the second predictor drug".

[0029] In one embodiment of the present invention, the fusion feature is constructed using the following formula:

[0030] In the formula, This is the original fusion vector; This is a vector concatenation operation; Characterization of microstructure; To combine biological characteristics, These are the first and second side effect characteristics. For the gated weight vector, For the weights of the gated network For the bias of the gated network. It is the Sigmoid activation function. As a feature of fusion, This indicates an element-wise multiplication operation.

[0031] In this embodiment, the weights of each mode are dynamically determined by a gating network, and the model can automatically determine whether the current DDI is mainly caused by chemical structure (high microstructure characterization weight) or by metabolic pathway (high microstructure characterization weight).

[0032] The foregoing has described specific embodiments of the invention. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0033] According to another embodiment, the present invention provides a multi-scale feature fusion drug interaction prediction device based on biological knowledge graph. Figure 2A schematic block diagram of a drug interaction prediction device according to one embodiment is shown. It will be understood that this device can be implemented by any apparatus, device, platform, or cluster of devices with computing and processing capabilities. Figure 2 As shown, the device includes: an acquisition unit 200, a first data processing unit 202, a second data processing unit 204, a third data processing unit 206, a fourth data processing unit 208, and a fifth data processing unit 210. The main functions of each component are as follows: The acquisition unit 200 is configured to acquire molecular structure data, biochemical metabolic pathway data, and side effect data of the drug to be predicted; wherein the drug to be predicted includes a first drug to be predicted and a second drug to be predicted. The first data processing unit 202 is configured to determine the first node features of the first drug to be predicted and the second node features of the second drug to be predicted based on the molecular structure data. The second data processing unit 204 is configured to determine the combined biological characteristics of the first drug to be predicted and the second drug to be predicted based on the biochemical metabolic pathway data, the first node characteristics and the second node characteristics. The third data processing unit 206 is configured to determine, based on the side effect data, a first side effect characteristic of the first drug to be predicted and a second side effect characteristic of the second drug to be predicted. The fourth data processing unit 208 is configured to determine fusion features based on the first node features, the second node features, the combined biological features, the first side effect features, and the second side effect features; The fifth data processing unit 210 is configured to determine the drug-drug interaction prediction results based on the fusion features.

[0034] In one embodiment of the present invention, the first data processing unit 202 is configured to perform the following operations: Based on the molecular structure data, a molecular diagram of the molecular structure is determined; Based on the molecular structure diagram, the first node features of the first drug to be predicted and the second node features of the second drug to be predicted are determined. The molecular structure diagram includes atomic type, atomic degree, formal charge, hybridization mode, aromaticity and chirality labels.

[0035] In one embodiment of the present invention, the fourth data processing unit 208 is configured to perform the following operations: Based on the features of the first node and the features of the second node, the interaction information is determined; The first node features and the second node features are updated based on the interaction information to obtain the updated first node features and the updated second node features. The updated first node features and the updated second node features are readout and aggregated to obtain the microstructure characterization of the first drug to be predicted. Based on the microstructure characterization, the combined biological characteristics, the first side effect characteristics, and the second side effect characteristics, the fusion characteristics are determined.

[0036] In one embodiment of the present invention, the second data processing unit 204 is configured to perform the following operations: Based on the biochemical metabolic pathway data, a biochemical metabolic pathway knowledge graph is constructed. The biochemical metabolic pathway knowledge graph was structurally pruned and completed to obtain an optimized biochemical metabolic pathway knowledge graph. Information on multiple adjacent nodes related to the first and second drugs to be predicted is extracted from the optimized biochemical metabolic pathway knowledge graph to obtain a joint knowledge subgraph. Based on the joint knowledge subgraph and the microstructure representation, joint biological characteristics are determined; The node information includes targets, enzymes, pathways, diseases, and connection edges.

[0037] In one embodiment of the present invention, the combined biological characteristics are determined by the following formula:

[0038] In the formula, Let be the retention probability of the edges in the joint knowledge subgraph. Let be the node feature vector of one endpoint connected to the edge to be evaluated in the joint knowledge subgraph. Let be the node feature vector of the other endpoint connected to the edge to be evaluated in the joint knowledge subgraph. The embedding type vector corresponding to the edge. For microstructure characterization, For cosine similarity, Let be the feature vector of any non-adjacent node pair in the joint knowledge subgraph. Let be the feature vector of another arbitrary non-adjacent node pair in the joint knowledge subgraph. The importance score of the node. As the first learnable parameter, As the second learnable parameter, For the final embedding of each node in the reconstructed joint knowledge subgraph, For the aforementioned combined biological characteristics, It is the set of all valid nodes in the joint knowledge subgraph centered on the drug pair to be predicted, after structural pruning and completion optimization.

[0039] In one embodiment of the present invention, the interaction information is determined by the following formula:

[0040] In the formula, The atoms of the first drug to be predicted With the atoms of the second drug to be predicted The potential tendency to react This is the feature of the first node. This is the feature of the second node. The weight matrix is ​​a learnable bilinear weight matrix. For the atoms in the first drug to be predicted The maximum response sensitivity affected by the overall effect of the second drug to be predicted. The total number of atoms, The interactive information, For the similarity of internal structure, The dimension of the feature vector. For adjustable weights, The first drug to be predicted contains atoms Maximum response sensitivity induced by the second predictable drug.

[0041] In one embodiment of the present invention, the fusion feature is constructed using the following formula:

[0042] In the formula, This is the original fusion vector; This is a vector concatenation operation; Characterization of the aforementioned microstructure; For the aforementioned combined biological characteristics, The first side effect feature and the second side effect feature are, For the gated weight vector, For the weights of the gated network For the bias of the gated network. It is the Sigmoid activation function. For the fusion feature, This indicates an element-wise multiplication operation.

[0043] According to another embodiment, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed in a computer, causes the computer to perform a combination Figure 1The method described.

[0044] According to another embodiment, an electronic device is also provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, it implements a combination... Figure 1 The method described.

[0045] The various embodiments in this invention are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0046] Those skilled in the art will recognize that, in one or more of the examples above, the functions described in this invention can be implemented using hardware, software, firmware, or any combination thereof. When implemented in software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium.

[0047] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of the present invention should be included within the scope of protection of the present invention.

Claims

1. A multi-scale feature fusion method for predicting drug interactions based on biological knowledge graphs, characterized in that, include: Acquire molecular structure data, biochemical metabolic pathway data, and side effect data of the drug to be predicted; wherein, the drug to be predicted includes a first drug to be predicted and a second drug to be predicted; Based on the molecular structure data, the first node features of the first drug to be predicted and the second node features of the second drug to be predicted are determined. Based on the biochemical metabolic pathway data, the first node features, and the second node features, the combined biological characteristics of the first and second drugs to be predicted are determined. Based on the side effect data, a first side effect characteristic of the first drug to be predicted and a second side effect characteristic of the second drug to be predicted are determined. Based on the first node feature, the second node feature, the combined biological feature, the first side effect feature, and the second side effect feature, the fusion feature is determined; Based on the fusion features, the drug-drug interaction prediction results are determined; The step of determining the first node features of the first drug to be predicted and the second node features of the second drug to be predicted based on the molecular structure data includes: Based on the molecular structure data, a molecular diagram of the molecular structure is determined; Based on the molecular structure diagram, the first node features of the first drug to be predicted and the second node features of the second drug to be predicted are determined. The molecular structure diagram includes atomic type, atomic degree, formal charge, hybridization mode, aromaticity and chirality labels; The determination of fusion features based on the first node features, the second node features, the combined biological features, the first side effect features, and the second side effect features includes: Based on the features of the first node and the features of the second node, the interaction information is determined; The first node features and the second node features are updated based on the interaction information to obtain the updated first node features and the updated second node features. The updated first node features and the updated second node features are readout and aggregated to obtain the microstructure characterization of the first drug to be predicted. Based on the microstructure characterization, the combined biological characteristics, the first side effect characteristics, and the second side effect characteristics, the fusion characteristics are determined; The determination of the combined biological characteristics of the first and second predictable drugs based on the biochemical metabolic pathway data, the first node characteristics, and the second node characteristics includes: Based on the biochemical metabolic pathway data, a biochemical metabolic pathway knowledge graph is constructed. The biochemical metabolic pathway knowledge graph was structurally pruned and completed to obtain an optimized biochemical metabolic pathway knowledge graph. Information on multiple adjacent nodes related to the first and second drugs to be predicted is extracted from the optimized biochemical metabolic pathway knowledge graph to obtain a joint knowledge subgraph. Based on the joint knowledge subgraph and the microstructure representation, joint biological characteristics are determined; The node information includes targets, enzymes, pathways, diseases, and connection edges; The combined biological characteristics are determined by the following formula: In the formula, Let be the retention probability of the edges in the joint knowledge subgraph. Let be the node feature vector of one endpoint connected to the edge to be evaluated in the joint knowledge subgraph. Let be the node feature vector of the other endpoint connected to the edge to be evaluated in the joint knowledge subgraph. Let c be the embedding type vector corresponding to the edge, and c be the microstructure representation. For cosine similarity, Let be the feature vector of any non-adjacent node pair in the joint knowledge subgraph. Let be the feature vector of another arbitrary non-adjacent node pair in the joint knowledge subgraph. The importance score of the node. As the first learnable parameter, As the second learnable parameter, For the final embedding of each node in the reconstructed joint knowledge subgraph, For the aforementioned combined biological characteristics, It is the set of all valid nodes in the joint knowledge subgraph centered on the drug pair to be predicted, after structural pruning and completion optimization; The interactive information is determined by the following formula: In the formula, The potential tendency for atom i of the first drug to be predicted to react with atom j of the second drug to be predicted. This is the feature of the first node. This is the feature of the second node. The weight matrix is ​​a learnable bilinear weight matrix. N represents the maximum reaction sensitivity of atom i in the first drug to be predicted to the overall influence of the second drug to be predicted. B The total number of atoms, The interactive information, For internal structural similarity, d is the dimension of the feature vector. For adjustable weights, The maximum responsiveness of atom k in the first drug to be predicted to the second drug to be predicted is induced by the second drug to be predicted. The fusion feature is constructed using the following formula: In the formula, The first line represents the original fused vector; the second line represents the vector concatenation operation. Characterization of the aforementioned microstructure; For the aforementioned combined biological characteristics, The first side effect feature and the second side effect feature are, For the gated weight vector, For the weights of the gated network For the bias of the gated network. It is the Sigmoid activation function. For the fusion feature, This indicates an element-wise multiplication operation.

2. A multi-scale feature fusion drug interaction prediction device based on biological knowledge graphs, characterized in that, For performing the method as described in claim 1, comprising: The acquisition unit is configured to acquire molecular structure data, biochemical metabolic pathway data, and side effect data of the drug to be predicted; wherein the drug to be predicted includes a first drug to be predicted and a second drug to be predicted. The first data processing unit is configured to determine the first node features of the first drug to be predicted and the second node features of the second drug to be predicted based on the molecular structure data. The second data processing unit is configured to determine the combined biological characteristics of the first drug to be predicted and the second drug to be predicted based on the biochemical metabolic pathway data, the first node characteristics, and the second node characteristics. The third data processing unit is configured to determine the first side effect characteristics of the first drug to be predicted and the second side effect characteristics of the second drug to be predicted based on the side effect data. The fourth data processing unit is configured to determine the fusion feature based on the first node feature, the second node feature, the combined biological feature, the first side effect feature, and the second side effect feature; The fifth data processing unit is configured to determine the drug-drug interaction prediction results based on the fusion features.

3. An electronic device, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method as described in claim 1.

4. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed in a computer, causes the computer to perform the method of claim 1.

Citation Information

Patent Citations

  • Method and system for predicting potential drug interaction based on biological network

    CN119943207A