Drug interaction prediction method based on drug multi-dimensional feature fusion

Through multi-dimensional feature fusion and multi-head cross-attention mechanism, the problems of incomplete features and insufficient fusion in drug interaction prediction are solved, and efficient and accurate drug interaction prediction is achieved.

CN120767010APending Publication Date: 2025-10-10HARBIN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510909411.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing drug interaction prediction methods fail to comprehensively consider the multidimensional characteristics of drugs, resulting in insufficient prediction performance and insufficient fusion of multidimensional features.

Method used

By obtaining multi-source feature information of drugs, the Jaccard similarity algorithm is used to generate one-dimensional similarity features, which are then combined with the DrugBank database to extract two-dimensional chemical space structure features. Edge coding, centrality coding, and spatial coding are then used to extract chemical space structure features. After dimensionality reduction using a lightweight self-attention mechanism, the one-dimensional similarity features and the two-dimensional chemical space structure features are fused through a multi-head cross-attention mechanism.

Benefits of technology

It significantly improves the accuracy and generalization ability of drug interaction prediction, reduces computational complexity, and improves the robustness and scalability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120767010A_ABST
    Figure CN120767010A_ABST
Patent Text Reader

Abstract

The invention relates to the field of drug interaction prediction, in particular to a drug interaction prediction method based on drug multi-dimensional feature fusion. The method comprises the following steps: acquiring drug multi-source feature information, chemical structure chart information and drug interaction information; then constructing a one-dimensional drug similarity feature and a two-dimensional chemical space structure feature based on drug multi-source feature information and chemical structure chart information so as to obtain a multi-dimensional feature of the drug; secondly, fusing multidimensional features of drugs based on multiple attention mechanisms; and finally predicting drug interaction by using MLP. The method comprises the following steps: calculating one-dimensional similarity features of medicines by adopting Jaccard similarity, splicing a plurality of matrixes to obtain comprehensive one-dimensional features, acquiring two-dimensional structural information of the medicines by utilizing RDkit, generating features through edges, centrality and space coding, fusing the features after lightweight self-attention dimensionality reduction, and finally combining the multi-dimensional features and a medicine interaction matrix to obtain a medicine interaction model. The drug interaction is predicted through the multi-layer perceptron model, and the drug interaction prediction performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of drug interaction prediction, and in particular to a drug interaction prediction method based on the fusion of multidimensional drug features. Background Art

[0002] Treating a disease with multiple drugs together is a common treatment method, but taking two or more drugs at the same time or over a continuous period of time may cause drug interactions, leading to adverse drug reactions, which in turn endanger health and may even cause death. Early drug interactions were mainly verified through experiments. This experimental method is not only very time-consuming but also expensive. With the advancement of computer technology, more and more machine learning-based computational methods are used to predict potential drug interactions. At present, drug interaction prediction methods based on machine learning still have some limitations. For example: some drug interaction prediction methods only use the one-dimensional characteristics or two-dimensional chemical space characteristics of drugs to construct prediction models, without comprehensively considering the multidimensional characteristics of drugs; although other methods use the multidimensional characteristics of drugs to construct drug interaction prediction models, there is a problem of insufficient multidimensional feature fusion in the use of multidimensional features, which in turn affects the performance of drug interaction prediction. Therefore, it is urgent to design a drug interaction prediction method based on the fusion of multidimensional drug features to solve the above problems. Summary of the Invention

[0003] The purpose of the present invention is to provide a drug interaction prediction method based on the fusion of multidimensional drug features, so as to solve the problem that the existing prediction methods in the above-mentioned background technology use the one-dimensional features or two-dimensional chemical space features of the drugs to construct a prediction model, do not comprehensively consider the multidimensional features of the drugs, and have the problem of insufficient multidimensional feature fusion in the use of multidimensional features, which in turn affects the performance of drug interaction prediction.

[0004] To achieve the above objectives, the present invention provides the following technical solution: a drug interaction prediction method based on multi-dimensional feature fusion of drugs, comprising the following steps:

[0005] Step S1: Acquire multi-source characteristic information of a drug, wherein the multi-source characteristic information includes drug property information, drug side effect information, and drug chemical structure diagram information;

[0006] Step S2: Calculate the drug feature similarity matrix based on the Jaccard similarity algorithm to generate a one-dimensional similarity feature of the drug;

[0007] Step S3: Obtain drug chemical structure information from the DrugBank database and extract the two-dimensional chemical space structural characteristics of the drug;

[0008] Step S4: extracting the chemical space structural features of the drug based on edge coding; extracting the chemical space structural features of the drug based on centrality coding; extracting the chemical space structural features of the drug based on space coding;

[0009] Step S5: Reduce the dimension of chemical space structural features based on lightweight self-attention mechanism;

[0010] Step S6: The one-dimensional similarity features and the two-dimensional chemical space structure features are fused through a multi-head cross-attention mechanism to generate multi-dimensional fusion features and predict drug interactions.

[0011] Preferably, in step S1, the drug attribute information includes: drug chemical structure characteristics (548×881), drug target characteristics (548×780), drug transporter characteristics (548×78), drug enzyme characteristics (548×129), drug indication characteristics (548×4897), and drug pathway characteristics (548×253); drug side effect information: SIDER side effect characteristics (548×4897) and OFFSIDES side effect characteristics (548×9496).

[0012] Preferably, in step S2, eight drug feature similarity matrices are generated by Jaccard similarity calculation, including: drug chemical structure similarity matrix, drug target similarity matrix, drug transporter similarity matrix, drug enzyme similarity matrix, drug indication similarity matrix, drug pathway similarity matrix, SIDER side effect similarity matrix and OFFSIDES side effect similarity matrix, wherein the Jaccard calculation formula is:

[0013]

[0014] Among them, P and Q are the feature sets of two drugs in the same drug feature matrix, respectively. The 8 drug feature similarity matrices are converted into column feature vectors, and the 8 column feature vectors are spliced ​​column by column using the feature splicing method to obtain the 300304×8-dimensional drug one-dimensional similarity feature D S .

[0015] Preferably, in step S3, the chemical spatial structure information of the drug is obtained by: downloading 548 SDF files from the DrugBank database, using the Chem.MolFromMolFile() function of the RDKit library to read the SDF file, and generating the Mol object of the RDKit; traversing each Mol object, and obtaining the atom and bond information through the mol.GetAtoms() and mol.GetBonds() functions.

[0016] Preferably, in step S4, edge coding is implemented in the following manner: an atomic feature vector is composed of 9 elements, including the number of atoms, chirality information, degree, formal charge, number of connected hydrogen atoms, number of free radical electrons, hybridization state, whether it participates in aromatic bonds, and whether it is located in a ring structure; an edge feature vector represents the bond type, stereochemical bond, and conjugation of adjacent atomic pairs, and an edge code is composed of the feature vector x of atom i. i and edge (i,j) constitute;

[0017] Formula (2) represents the characteristic vector x of atom i in the drug i :

[0018] x i =[α1,α2,α3,α4,α5,α6,α7,α8,α9] (2)

[0019] Among them: eigenvector x i It is composed of 9 elements α1α2α3,α4,α5,α6,α7,α8,α9, which respectively represent the number of the atom in the drug, chirality information (unspecified type, r-type or s-type), degree, formal charge, number of connected hydrogen atoms, number of free radical electrons, hybridization state, whether it participates in aromatic bonds and whether it is located in a ring structure;

[0020] Formula (3) represents the edge e of the bonded pair of adjacent atoms i and j. (i,j) :

[0021] e (i,j) =[β1,β2,β3] (3)

[0022] Among them: β1 represents the type of bond, β2 represents the stereochemistry of the bond, and β3 represents whether the bond is conjugated.

[0023] Preferably, in step S4, centrality coding is implemented in the following manner: based on the atomic feature vector x in formula (2) i , combined with the centrality characteristics of drug atom i, generate the centrality code of drug atom i and obtain the centrality code:

[0024]

[0025] Where: W X x i represents the atomic characteristics of drug atom i; O i represents the out-degree centrality vector of drug atom i, I i represents the in-degree centrality vector of drug atom i, W out O i +W in I i is the central feature of drug atom i, W X 、Wout and W in It is a learnable shared weight matrix that shares the weights of the embedding layer and attention block between drugs through the weight-sharing Siamese architecture to learn the characteristics of atoms during the parameter update process.

[0026] Preferably, in step S4, based on the edge feature vector e in formula (3) (i,j) To obtain the spatial encoding, first find the shortest path P from node i to node j, and then use the shortest path distance to represent the relationship between the positions of the two nodes in the spatial structure. If the node pair i, j is connected (can be non-adjacent), the shortest path distance is used as the edge e (i,j) The spatial position s (i,j) , for a node pair (i, j), if the node pair i, j is not connected, its edge e (i,j) The spatial position s (i,j) Set to -1, and after obtaining the representation of the edge and spatial structure, the spatial encoding of the drug is expressed by formula (5)

[0027]

[0028] in: It is the side of the drug (i,j) The embedding is done by taking the edge feature P in the shortest path of the atom pair (i, j) in the drug l and the shared weight matrix W edge The dot product is averaged to obtain, l is the lth edge between the atom pair (i, j) in the drug, and the total number of edges is k. (i,j) It is edge e (i,j) The spatial position of W edge and W spat are two learnable shared weight moments.

[0029] Preferably, the lightweight self-attention mechanism in step S5 is a stacked encoder based on sparse self-attention, including three probabilistic sparse self-attention blocks, wherein the first two self-attention blocks are followed by a convolution module and a maximum pooling module for distillation, and the sparse self-attention formula is:

[0030]

[0031] in: is a sparse matrix of the same size as the query, containing only the top-u queries, K is the key vector, V is the numeric vector, is the dimension scaling factor of the key vector;

[0032] After the features obtained in step S4 are input into the encoder, the probabilistic sparse self-attention calculation formula is:

[0033]

[0034] wherein: W K and W V are learnable shared weight matrices;

[0035] The first two attention blocks are distilled to preferentially extract important mappings, generating a more focused feature in the next level, and the dimension reduction feature calculation formula is:

[0036] X j+1 = MaxPool(ELU(Conv1d([X j ] op )))(8)

[0037] wherein: [X j ] op represents the ProbSparse self-attention mechanism input to the jth attention block, Conv1d() represents a one-dimensional convolution layer, ELU() is an activation function, and MaxPool() is a maximum pooling layer.

[0038] The embedding vector of the drug is processed by the encoder to generate a corresponding low-dimensional feature vector, and the low-dimensional feature vectors of the drugs are spliced to obtain the chemical space structure feature D C ;

[0039] Preferably, in the step S6, a multi-head cross-attention mechanism is used to realize the fusion of the drug similarity feature D S and the chemical space structure feature D C , the chemical space structure feature D C is projected onto two independent matrices V C and K C to obtain a set of values and keys, and the similarity feature D S is projected into another single matrix Q S ;

[0040] V C = D C W V , K C = D C W K , Q S = D S W Q (9)

[0041] wherein: W V , W K , W Q are weight matrices;

[0042] Preferably, the correlation matrix and the vector V CMultiply to obtain vector Z, and use cross-feature correlation to complement chemical structure features;

[0043]

[0044] in: Used to represent D C With D S The correlation between S is the query vector, K C is the key vector, V C is a numeric vector, is the dimension scaling factor of the key vector, which is transformed into Z by reprojecting the vector Z into the original space through nonlinear transformation C , and perform residual connection through formula (11), where W O is the output weight matrix before the feedforward neural network layer (FFN).

[0045] D′=α·Z C W O +β·D C (11)

[0046] Finally, a feedforward neural network (FFN) is applied to further refine the global information, thereby improving the robustness and accuracy of the model, and outputting the fusion feature D″ through formula (12);

[0047] D″=γ·D′+δ·FFN(D′) (12)

[0048] A learnable coefficient is applied to each residual connection branch in Equations (11) and (12) to adaptively learn data from different branches to achieve performance improvement, where α, β, γ, δ are learnable parameters that are initialized to 1 during training;

[0049] The fusion features and drug interaction matrix are input into a multi-layer perceptron with 3 layers to predict drug interactions, and the drug interaction prediction result Y is calculated by formula (13): pred ;

[0050] Y pred =MLP(D″,θ MLP ) (13)

[0051] Where: θ MLP Refers to the training parameters of MLP.

[0052] Compared with the prior art, the present invention has the following beneficial effects:

[0053] 1. This prediction method systematically integrates the multi-dimensional features of drugs (including drug properties, side effect information, and chemical space structure features) and uses a multi-head cross-attention mechanism to achieve deep feature interaction, effectively solving the problem of insufficient prediction performance caused by one-sided feature representation or insufficient fusion of traditional methods. Specifically, by extracting attribute information such as the chemical structure, target, enzyme, indication, and other attribute information of drugs from multiple source databases such as DrugBank, PubChem, and KEGG, as well as side effect data such as SIDER and OFFSIDES, and chemical structure information of drugs, a multi-dimensional feature covering the molecular properties, pharmacological properties, and two-dimensional chemical structure of drugs is constructed; at the same time, the three complementary methods of edge coding, centrality coding, and spatial coding are combined to extract the two-dimensional chemical space structure features of drugs and construct a high-dimensional chemical space representation. Further, the multi-head cross-attention mechanism is used to achieve the fusion of one-dimensional similarity features and two-dimensional chemical space features, ultimately significantly improving the accuracy and generalization ability of drug interaction prediction.

[0054] 2. This prediction method introduces a lightweight self-attention mechanism, which effectively reduces computational complexity while ensuring model performance, solving the problems of dimensionality explosion and low computational efficiency in traditional multi-dimensional feature fusion methods. Specifically, the probabilistic sparse self-attention block is combined with the convolution-pooling distillation module to efficiently reduce the dimensionality of chemical space structural features and reduce redundant calculations. At the same time, the multi-head cross-attention mechanism optimizes the feature interaction process through dynamic weight allocation to avoid information loss. This design not only reduces the model's dependence on computing resources, but also enhances its ability to process large-scale drug data, enabling the model to maintain high prediction accuracy while having good scalability, providing an efficient and reliable solution for drug combination screening and safety assessment in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is a schematic diagram of the overall structural scheme of the present invention;

[0056] Figure 2 Schematic diagram of the lightweight self-attention mechanism of the present invention;

[0057] Figure 3 Schematic diagram of the cross-attention mechanism of the present invention. DETAILED DESCRIPTION

[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0059] See also Figure 1-3 , an embodiment provided by the present invention:

[0060] A drug interaction prediction method based on multi-dimensional feature fusion of drugs includes the following steps:

[0061] Step S1: Acquire multi-source feature information of the drug, which includes drug property information, drug side effect information, and drug chemical structure diagram information;

[0062] Step S2: Calculate the drug feature similarity matrix based on the Jaccard similarity algorithm to generate a one-dimensional similarity feature of the drug;

[0063] Step S3: Obtain drug chemical structure information from the DrugBank database and extract the two-dimensional chemical space structural characteristics of the drug;

[0064] Step S4: extracting the chemical space structural features of the drug based on edge coding; extracting the chemical space structural features of the drug based on centrality coding; extracting the chemical space structural features of the drug based on space coding;

[0065] Step S5: Reduce the dimension of chemical space structural features based on lightweight self-attention mechanism;

[0066] Step S6: The one-dimensional similarity features and the two-dimensional chemical space structure features are fused through a multi-head cross-attention mechanism to generate multi-dimensional fusion features and predict drug interactions.

[0067] Furthermore, in step S1, the drug attribute information includes: drug chemical structure characteristics (548×881), drug target characteristics (548×780), drug transporter characteristics (548×78), drug enzyme characteristics (548×129), drug indication characteristics (548×4897), and drug pathway characteristics (548×253); drug side effect information: SIDER side effect characteristics (548×4897) and OFFSIDES side effect characteristics (548×9496).

[0068] The dataset constructed in this example contains 548 drugs. The standardized drug names are obtained from DrugBank. Drug interaction information is obtained from the TWOSIDES dataset based on the standardized names. Drug attribute information is obtained from DrugBank, PubChem, and KEGG. Drug side effect information is obtained from the SIDER side effect dataset and the OFFSIDER side effect dataset.

[0069] Among them, drug interaction information of 548 drugs was obtained from the TWOSIDES dataset, and a matrix with drugs as rows and drugs as columns was constructed to obtain a 548×548 drug interaction matrix. The drug interaction matrix contained 48,584 pairs of drug interactions. The chemical structure information of 881 drugs was obtained from PubChem, 780 drug target information, 78 drug transporter information, 129 drug enzyme information, and 4,897 drug indication information were obtained from DrugBank, 253 drug pathway information was obtained from KEGG, 4,897 drug side effect information was obtained from the SIDER side effect dataset, and 9,496 drug Off side effect information was obtained from the OFFSIDES side effect dataset.

[0070] For drug attribute information, taking the drug chemical structure characteristics as an example, there are 881 chemical substructures in the chemical structure characteristics. A matrix with drugs as rows and chemical substructures as columns is constructed to obtain a 548×881 chemical structure feature matrix. Repeat the above operation to obtain the remaining five drug attribute matrices, namely: a 548×780 drug target feature matrix, a 548×78 drug transporter feature matrix, a 548×129 drug enzyme feature matrix, a 548×4897 drug indication feature matrix, and a 548×253 drug pathway feature matrix.

[0071] For drug side effect information, we used the SIDER side effect feature dataset as an example. The SIDER side effect dataset contains 4897 side effects. We constructed a matrix with drugs as rows and side effects as columns, resulting in a 548×4897 SIDER side effect feature matrix. Repeating the above steps yielded a 548×9496 OFFSIDES side effect feature matrix.

[0072] Furthermore, in step S2, eight drug feature similarity matrices are generated by Jaccard similarity calculation: drug chemical structure similarity matrix, drug target similarity matrix, drug transporter similarity matrix, drug enzyme similarity matrix, drug indication similarity matrix, drug pathway similarity matrix, SIDER side effect similarity matrix, and OFFSIDES side effect similarity matrix. The Jaccard calculation formula is:

[0073]

[0074] Among them, P and Q are the feature sets of two drugs in the same drug feature matrix, respectively. The 8 drug feature similarity matrices are converted into column feature vectors, and the 8 column feature vectors are spliced ​​column by column using the feature splicing method to obtain the 300304×8-dimensional drug one-dimensional similarity feature D S ;

[0075] Furthermore, in step S3, the chemical spatial structure information of the drug is obtained in the following way: 548 SDF files are downloaded from the DrugBank database, and the Chem.MolFromMolFile() function of the RDKit library is used to read the SDF file to generate the Mol object of RDKit; each Mol object is traversed, and the atom and bond information is obtained through the mol.GetAtoms() and mol.GetBonds() functions. This information includes the type, charge, position of the atom, and the type and length of the bond, etc. In this way, we can comprehensively and accurately extract the chemical spatial structure information of the drug, providing a solid foundation for further chemical structure analysis and similarity calculation.

[0076] Furthermore, in step S4, edge coding is implemented in the following way: the atomic feature vector consists of 9 elements, including the number of atoms, chirality information, degree, formal charge, number of connected hydrogen atoms, number of free radical electrons, hybridization state, whether it participates in aromatic bonds, and whether it is located in a ring structure; the edge feature vector represents the bond type, stereochemical bond, and conjugation of adjacent atomic pairs. The edge coding is composed of the feature vector x of atom i. i and edge (i,j) constitute;

[0077] Formula (2) represents the characteristic vector x of atom i in the drug i :

[0078] x i =[α1,α2,α3,α4,α5,α6,α7,α8,α9] (2)

[0079] Among them: eigenvector x i It is composed of 9 elements α1α2α3,α4,α5,α6,α7,α8,α9, which respectively represent the number of the atom in the drug, chirality information (unspecified type, r-type or s-type), degree, formal charge, number of connected hydrogen atoms, number of free radical electrons, hybridization state, whether it participates in aromatic bonds and whether it is located in a ring structure;

[0080] Formula (3) represents the edge e of the bonded pair of adjacent atoms i and j. (i,j) :

[0081] e (i,j) =[β1,β2,β3] (3)

[0082] Among them: β1 represents the type of bond, β2 represents the stereochemical bond, and β3 represents whether the bond is conjugated;

[0083] The edge encoding consists of the eigenvector of atom i and the edge, as shown in formula (2), where each element corresponds to a specific property of the atom in the drug, and the edge eigenvector, as shown in formula (3), further enriches the information content of the encoding by representing the type of bond, stereochemical bond and conjugation.

[0084] Furthermore, in step S4, centrality coding is implemented in the following way: Based on the atomic feature vector x in formula (2) i , combined with the centrality characteristics of drug atom i, generate the centrality code of drug atom i and obtain the centrality code:

[0085]

[0086] Where: W X x i represents the atomic characteristics of drug atom i; O i represents the out-degree centrality vector of drug atom i, I i represents the in-degree centrality vector of drug atom i, W out O i +W in I i is the central feature of drug atom i, W X 、W out and W in It is a learnable shared weight matrix that shares the weights of the embedding layer and attention block between drugs through the weight-sharing Siamese architecture to learn the characteristics of atoms during the parameter update process.

[0087] Furthermore, in step S4, based on the edge feature vector e in formula (3) (i,j) To obtain the spatial encoding, first find the shortest path P from node i to node j, and then use the shortest path distance to represent the relationship between the positions of the two nodes in the spatial structure. If the node pair i, j is connected (can be non-adjacent), the shortest path distance is used as the edge e (i,j) The spatial position s (i,j) , for a node pair (i, j), if the node pair i, j is not connected, its edge e (i,j) The spatial position s (i,j) Set to -1, and after obtaining the representation of the edge and spatial structure, the spatial encoding of the drug is expressed by formula (5)

[0088]

[0089] in: It is the side of the drug (i,j) The embedding is done by taking the edge feature P in the shortest path of the atom pair (i, j) in the drug l and the shared weight matrix Wedge The dot product is averaged to obtain, l is the lth edge between the atom pair (i, j) in the drug, and the total number of edges is k. (i,j) It is edge e (i,j) The spatial position of W edge and W spat They are two learnable shared weight moments. Edge encoding characterizes the drug structure through atomic features (such as charge, hybridization state) and bond properties (type, conjugation), centrality encoding quantifies the topological importance of nodes in molecules, and spatial encoding integrates atomic three-dimensional coordinate information. The three work together to improve the expression accuracy and predictive efficiency of chemical space features.

[0090] Furthermore, the lightweight self-attention mechanism in step S5 is a stacked encoder based on sparse self-attention, which includes three probabilistic sparse self-attention blocks. The first two self-attention blocks are followed by a convolution module and a maximum pooling module for distillation. The sparse self-attention formula is:

[0091]

[0092] in: is a sparse matrix of the same size as the query, containing only the top-u queries, K is the key vector, V is the numeric vector, is the dimension scaling factor of the key vector;

[0093] After the features obtained in step S4 are input into the encoder, the probabilistic sparse self-attention calculation formula is:

[0094]

[0095] in: W K and W V is a learnable shared weight matrix;

[0096] After the first two attention blocks, distillation is performed to prioritize the more important mappings and generate a more focused feature at the next level. The dimensionality reduction feature calculation formula is:

[0097] X j+1 =MaxPool(ELU(Conv1d([X j ] op )))(8)

[0098] Where: [X j ] op Represents the ProbSparse self-attention mechanism input to the j-th attention block, Conv1d() represents the one-dimensional convolution layer, ELU() is the activation function, and MaxPool() is the maximum pooling layer;

[0099] The drug embedding vector is processed by the encoder to generate the corresponding low-dimensional feature vector, and the low-dimensional feature vectors of the drug are spliced ​​to obtain the chemical space structure feature D C .

[0100] Furthermore, in step S6, a multi-head cross attention mechanism is used to realize the drug similarity feature D S and chemical spatial structural characteristics D C The fusion of chemical space structure characteristics D C Projection into two independent matrices V C , K C On the above, we get a set of values ​​and keys, and use the similarity feature D S Projected into another separate matrix Q S middle;

[0101] V C =D C W V ,K C =D C W K ,Q S =D S W Q (9)

[0102] Where: W V 、W K 、W Q is the weight matrix.

[0103] Furthermore, the correlation matrix and vector V C Multiply to obtain vector Z, and use cross-feature correlation to complement chemical structure features;

[0104]

[0105] in: Used to represent D C With D S The correlation between Q S is the query vector, K C is the key vector, V C is a numeric vector, is the dimension scaling factor of the key vector, which is transformed into Z by reprojecting the vector Z into the original space through nonlinear transformation C , and perform residual connection through formula (11), where W O is the output weight matrix before the feedforward neural network layer (FFN).

[0106] D′=α·Z C W O +β·D C (11)

[0107] Finally, a feedforward neural network (FFN) is applied to further refine the global information, thereby improving the robustness and accuracy of the model, and outputting the fused feature D″ through formula (12).

[0108] D″=γ·D′+δ·FFN(D′) (12)

[0109] A learnable coefficient is applied to each residual connection branch in Equations (11) and (12) to adaptively learn data from different branches to achieve performance improvement, where α, β, γ, δ are learnable parameters that are initialized to 1 during training.

[0110] The fusion features and drug interaction matrix are input into a multi-layer perceptron with 3 layers to predict drug interactions, and the drug interaction prediction result Y is calculated by formula (13): pred .

[0111] Y pred =MLP(D″,θ MLP ) (13)

[0112] Where: θ MLP Refers to the training parameters of MLP.

[0113] Working principle: By constructing a multidimensional feature fusion framework, the dual technical bottlenecks of incomplete multidimensional feature representation and insufficient feature fusion in existing drug interaction prediction methods are systematically solved. Specifically, drug properties (chemical structure, target, enzyme, indication, etc.), side effects (SIDER, OFFSIDES) and chemical space structure information are first integrated from multiple source databases such as DrugBank, PubChem, and KEGG to construct a multidimensional feature covering drug molecular properties, pharmacological properties and drug chemical two-dimensional structure. Subsequently, the Jaccard similarity algorithm is used to calculate the similarity between drug properties and side effect features to generate a one-dimensional similarity feature of the drug. Simultaneously, the chemical space structure of the drug is extracted based on the RDKit tool, and the chemical space structure feature of the drug is constructed through three complementary methods: edge encoding, centrality encoding and spatial encoding. To reduce the feature dimension and retain key information, a lightweight self-attention mechanism is introduced to achieve efficient dimensionality reduction of chemical space features. On this basis, a multi-head cross-attention mechanism is used to realize the fusion of one-dimensional drug similarity features and two-dimensional chemical space structure features. Finally, the fused multidimensional features are input into a three-layer MLP to complete drug interaction prediction. This method not only overcomes the insufficient representation ability of traditional methods due to the single feature dimension, but also solves the information loss problem caused by insufficient fusion of multidimensional features, thereby significantly improving the accuracy and generalization ability of drug interaction prediction.

[0114] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

Claims

1. A drug interaction prediction method based on multidimensional drug feature fusion, characterized by: The following steps are involved: Step S1: Acquire multi-source characteristic information of a drug, wherein the multi-source characteristic information includes drug property information, drug side effect information, and drug chemical structure diagram information; Step S2: Calculate the drug feature similarity matrix based on the Jaccard similarity algorithm to generate a one-dimensional similarity feature of the drug; Step S3: Obtain drug chemical structure information from the DrugBank database and extract the two-dimensional chemical space structural characteristics of the drug; Step S4: extracting the chemical space structural features of the drug based on edge coding; extracting the chemical space structural features of the drug based on centrality coding; Extract the chemical spatial structural characteristics of drugs based on spatial coding; Step S5: Reduce the dimension of chemical space structural features based on lightweight self-attention mechanism; Step S6: The one-dimensional similarity features and the two-dimensional chemical space structure features are fused through a multi-head cross-attention mechanism to generate multi-dimensional fusion features and predict drug interactions.

2. The drug interaction prediction method based on multidimensional drug feature fusion according to claim 1, characterized in that: In step S1, the drug attribute information includes: drug chemical structure characteristics (548×881), drug target characteristics (548×780), drug transporter characteristics (548×78), drug enzyme characteristics (548×129), drug indication characteristics (548×4897), and drug pathway characteristics (548×253); drug side effect information: SIDER side effect characteristics (548×4897) and OFFSIDES side effect characteristics (548×9496).

3. The method for predicting drug interactions based on multidimensional drug feature fusion according to claim 1, characterized in that: In step S2, eight drug feature similarity matrices are generated by Jaccard similarity calculation, including: drug chemical structure similarity matrix, drug target similarity matrix, drug transporter similarity matrix, drug enzyme similarity matrix, drug indication similarity matrix, drug pathway similarity matrix, SIDER side effect similarity matrix and OFFSIDES side effect similarity matrix, wherein the Jaccard calculation formula is: Among them, P and Q are the feature sets of two drugs in the same drug feature matrix, respectively. The 8 drug feature similarity matrices are converted into column feature vectors, and the 8 column feature vectors are spliced ​​column by column using the feature splicing method to obtain the 300304×8-dimensional drug one-dimensional similarity feature D S .

4. The method for predicting drug interactions based on multidimensional drug feature fusion according to claim 1, characterized in that: In step S3, the chemical spatial structure information of the drug is obtained by downloading 548 SDF files from the DrugBank database, reading the SDF files using the Chem.MolFromMolFile() function of the RDKit library, and generating the Mol object of the RDKit; traversing each Mol object and obtaining the atom and bond information using the mol.GetAtoms() and mol.GetBonds() functions.

5. The method for predicting drug interactions based on multidimensional drug feature fusion according to claim 1, characterized in that: In step S4, edge coding is implemented in the following way: the atomic feature vector consists of 9 elements, including the number of atoms, chirality information, degree, formal charge, number of connected hydrogen atoms, number of free radical electrons, hybridization state, whether it participates in aromatic bonds and whether it is located in a ring structure; the edge feature vector represents the bond type, stereochemical bond and conjugation of adjacent atomic pairs, and the edge coding is composed of the feature vector x of atom i i and edge (i,j) constitute. Formula (2) represents the characteristic vector x of atom i in the drug i : x i =[α1,α2,α3,α4,α5,α6,α7,α8,α9] (2) Among them: eigenvector x i It is composed of 9 elements α1α2α3,α4,α5,α6,α7,α8,α9, which respectively represent the number of the atom in the drug, chirality information (unspecified type, r-type or s-type), degree, formal charge, number of connected hydrogen atoms, number of free radical electrons, hybridization state, whether it participates in aromatic bonds and whether it is located in a ring structure; Formula (3) represents the edge e of the bonded pair of adjacent atoms i and j. (i,j) : e (i,j) =[β1,β2,β3] (3) Among them: β1 represents the type of bond, β2 represents the stereochemistry of the bond, and β3 represents whether the bond is conjugated.

6. The method for predicting drug interactions based on multidimensional drug feature fusion according to claim 1, characterized in that: In step S4, the centrality coding is implemented in the following way: based on the atomic feature vector x in formula (2) i , combined with the centrality characteristics of drug atom i, generate the centrality code of drug atom i and obtain the centrality code: Where: W X x i represents the atomic characteristics of drug atom i; O i represents the out-degree centrality vector of drug atom i, I i represents the in-degree centrality vector of drug atom i, W out O i +W in I i is the central feature of drug atom i, W X 、W out and W in It is a learnable shared weight matrix that shares the weights of the embedding layer and attention block between drugs through the weight-sharing Siamese architecture to learn the characteristics of atoms during the parameter update process.

7. The method for predicting drug interactions based on multidimensional drug feature fusion according to claim 1, characterized in that: In step S4, based on the edge feature vector e in formula (3) (i,j) To obtain the spatial encoding, first find the shortest path P from node i to node j, and then use the shortest path distance to represent the relationship between the positions of the two nodes in the spatial structure. If the node pair i, j is connected (can be non-adjacent), the shortest path distance is used as the edge e (i,j) The spatial position s (i,j) , for a node pair (i, j), if the node pair i, j is not connected, its edge e (i,j) The spatial position s (i,j) Set to -1, and after obtaining the representation of the edge and spatial structure, the spatial encoding of the drug is expressed by formula (5) in: It is the side of the drug (i,j) The embedding is done by taking the edge feature P in the shortest path of the atom pair (i, j) in the drug l and the shared weight matrix W edge The dot product is averaged to obtain, l is the lth edge between the atom pair (i, j) in the drug, and the total number of edges is k. (i,j) It is edge e (i,j) The spatial position of W edge and W spat are two learnable shared weight moments.

8. The method for predicting drug interactions based on multidimensional drug feature fusion according to claim 1, characterized in that: The lightweight self-attention mechanism in step S5 is a stacked encoder based on sparse self-attention, which includes three probabilistic sparse self-attention blocks. The first two self-attention blocks are followed by a convolution module and a maximum pooling module for distillation. The sparse self-attention formula is: in: is a sparse matrix of the same size as the query, containing only the top-u queries, K is the key vector, V is the numeric vector, is the dimension scaling factor of the key vector; After the features obtained in step S4 are input into the encoder, the probabilistic sparse self-attention calculation formula is: in: W K and W V is a learnable shared weight matrix. After the first two attention blocks, distillation is performed to prioritize the more important mappings and generate a more focused feature at the next level. The dimensionality reduction feature calculation formula is: X j+1 =MaxPool(ELU(Conv1d([X j ] op )))(8) Where: [X j ] op Represents the ProbSparse self-attention mechanism input to the j-th attention block, Conv1d() represents the one-dimensional convolution layer, ELU() is the activation function, and MaxPool() is the maximum pooling layer; The drug embedding vector is processed by the encoder to generate the corresponding low-dimensional feature vector, and the low-dimensional feature vectors of the drug are spliced ​​to obtain the chemical space structure feature D C .

9. The method for predicting drug interactions based on multidimensional drug feature fusion according to claim 1, characterized in that: In step S6, a multi-head cross attention mechanism is used to realize the drug similarity feature D S and chemical spatial structural characteristics D C The fusion of chemical space structure characteristics D C Projection into two independent matrices V C , K C On the above, we get a set of values ​​and keys, and use the similarity feature D S Projected into another separate matrix Q S middle, V C =D C W V ,K C =D C W K ,Q S =D S W Q (9) Where: W V 、W K 、W Q is the weight matrix.

10. The method for predicting drug interactions based on multi-dimensional drug feature fusion according to claim 9, characterized in that: The correlation matrix and vector V C Multiply to obtain vector Z, and use cross-feature correlation to complement chemical structure features; in: Used to represent D C With D S The correlation between S is the query vector, K C is the key vector, V C is a numeric vector, is the dimension scaling factor of the key vector, which is transformed into Z by reprojecting the vector Z into the original space through nonlinear transformation C , and perform residual connection through formula (11), where W O is the output weight matrix before the feedforward neural network layer (FFN); D′=α·Z C W O +β·D C (11) Finally, a feedforward neural network (FFN) is applied to further refine the global information, thereby improving the robustness and accuracy of the model, and outputting the fusion feature D″ through formula (12); D″=γ·D′+δ·FFN(D′) (12) A learnable coefficient is applied to each residual connection branch in Equations (11) and (12) to adaptively learn data from different branches to achieve performance improvement, where α, β, γ, δ are learnable parameters that are initialized to 1 during training; The fusion features and drug interaction matrix are input into a multi-layer perceptron with 3 layers to predict drug interactions, and the drug interaction prediction result Y is calculated by formula (13): pred ; Y pred =MLP(D″,θ MLP ) (13) Where: θ MLP Refers to the training parameters of MLP.