A Drug Repositioning Method and System Based on Heterogeneous Bioinformatics Networks
By using heterogeneous biological information network and graph representation learning method in the drug relocation model, the graph representation feature matrix of drugs and diseases is generated, and the problem of insufficient accuracy of drug and disease association relationship in the prior art is solved, and more efficient drug relocation is achieved.
Patent Information
- Application Number
- CN202310994556.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-08
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2043-08-08
AI Technical Summary
Existing drug relocation models lack the ability to properly handle multiple relationships in bioisomeric bioinformatic networks, resulting in poor accuracy of predicted drug and disease associations.
A graph representation learning model based on heterogeneous biological information network is adopted. By constructing a heterogeneous biological information network, the biological attribute matrix of drugs, proteins and diseases is extracted, and a multi-level sub-graph represents learning strategies, a graph representation feature matrix of drugs and diseases is generated, and the drug relocation model is finally trained using a random forest classifier.
By learning biomolecular characteristics more comprehensively, improving the accuracy of drug relocation, the problem of insufficient accuracy in drug relocation in the prior art in biological networks is solved.
Smart Images

Figure CN117198383B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer data processing, and particularly relates to a drug repositioning method and system based on a heterogeneous biological information network. Background Art
[0002] Currently, drug repositioning is a promising development strategy for discovering potential candidate drugs for diseases. However, existing drug repositioning models lack the ability to predict the association between potential drugs and diseases by considering the mechanism of action between biomolecules. For example, valproic acid can affect the life cycle of breast cancer cells by acting on target proteins. If only relying on the bipartite graph of drugs and diseases, this biological cell mechanism is difficult to be discovered.
[0003] Although some drug repositioning methods considering heterogeneous networks have been proposed, such as:
[0004] The Chinese patent application with publication number CN107545151A discloses a drug repositioning method based on low-rank matrix completion. This method uses known disease data, drug data, and disease-drug association data to construct a disease-drug heterogeneous network; determines the optimal rank of the filling matrix based on a verification method; fills the adjacency matrix of the drug-disease heterogeneous network using the singular value threshold algorithm based on the selected optimal rank; predicts potential new drug-disease associations based on the filled matrix.
[0005] The Chinese patent application with publication number CN114613452A discloses a drug repositioning method and system based on a drug classification graph neural network. This method includes: obtaining drug similarity data and disease similarity data, and respectively constructing a drug attribute heterogeneous edge network G1 and a disease attribute network G2; respectively extracting features from the drug attribute edge heterogeneous network G1 and the disease attribute network G2 to obtain a drug attribute matrix A and a disease attribute matrix B; constructing a drug-disease association network G3 according to the drug attribute matrix A and the disease attribute matrix B; performing message passing and in-domain alignment based on the drug-disease association network G3 to obtain a drug embedding matrix U and a disease embedding matrix V; training an edge confidence MLP network based on the drug embedding matrix U and the disease embedding matrix V to obtain a trained MLP network, and using the trained MLP network for drug repositioning.
[0006] However, the predicted drug-disease association relationships they obtained show poor performance in terms of accuracy. Summary of the Invention
[0007] The object of the present invention is to overcome the defects that the prior art lacks the ability to properly handle various relationships in a heterogeneous biological information network and the predicted drug-disease association relationships show poor performance in terms of accuracy.
[0008] To achieve the above object, the present invention proposes a drug repositioning method based on a heterogeneous biological information network, and the method includes:
[0009] Input the disease and candidate drug names into the trained drug repositioning model to obtain the prediction scores of the candidate drugs for the disease, and complete the drug repositioning according to the prediction scores; wherein,
[0010] The drug repositioning model is a graph representation learning model, and the process of establishing the training data set during its training includes:
[0011] Construct a heterogeneous biological information network based on the existing known association relationship data of drugs and diseases, the association relationship data of drugs and proteins, the association relationship data of proteins and diseases, the chemical structure information of drugs, the biological semantic information of diseases, and the base sequence information of proteins;
[0012] Perform feature learning on the heterogeneous biological information network to respectively obtain the biological attribute matrices of drugs, proteins, and diseases;
[0013] Extract the multi-level subgraphs of drugs and diseases in the heterogeneous biological information network;
[0014] According to the biological attribute matrices and multi-level subgraphs of drugs, proteins, and diseases, obtain the graph representation feature matrix of drugs and diseases, that is, the training data set.
[0015] As an improvement of the above technical solution, the performing feature learning on the heterogeneous biological information network to respectively obtain the biological attribute matrices of drugs, proteins, and diseases specifically includes:
[0016] Apply the chemical structure similarity theory to calculate the chemical structure information of drugs to obtain the biological attribute matrix of drugs; apply the biological semantic similarity theory to calculate the biological semantic information of diseases to obtain the biological attribute matrix of diseases; apply the biological sequence statistics theory to calculate the base sequence information of proteins to obtain the biological attribute matrix of proteins.
[0017] As an improvement of the above technical solution, the applying the chemical structure similarity theory to calculate the chemical structure information of drugs in the heterogeneous biological information network to obtain the biological attribute matrix of drugs specifically is:
[0018] Apply the chemical structure similarity theory to calculate the chemical structure information of drugs using the RDKit open source toolkit to obtain the biological attribute matrix of drugs.
[0019] As one of the improvements to the above technical solution, obtaining the graph representation feature matrix of drugs and diseases specifically includes: calculating the graph representation features of drugs and diseases by applying a multi-level subgraph representation learning strategy to obtain the graph representation feature matrix of drugs and diseases.
[0020] As one of the improvements to the above technical solution, the heterogeneous biological information network HBIN is represented as:
[0021] HBIN = {V, E, C}
[0022] Among them, V represents the set composed of all nodes of the heterogeneous biological information network, V = {V DR , V DI , V PR}; V DR represents the drug abstracted as a node in the network, represents the i-th drug, and N represents the number of drugs; V DI represents the disease abstracted as a node in the network represents the j-th disease, and M represents the number of diseases; V PR represents the protein molecule abstracted as a node in the network represents the l-th protein molecule, and P represents the number of protein molecules; the meaning of abstracting drugs as a network, diseases as a network, and protein molecules as a network is to mathematize entities such as drugs, diseases, and proteins into nodes V in HBIN for better calculation;
[0023] E represents the set composed of all edges of the heterogeneous biological information network, E = {E DD , E DP , E PD}; E DD represents abstracting the association relationship between drugs and diseases as an edge in the network; E DP represents abstracting the association relationship between drugs and protein molecules as an edge in the network, and E PD abstracts the association relationship between protein molecules and diseases as an edge in the network;
[0024] C represents the set of node attributes of the heterogeneous biological information network, C = {C DR , C DI , C PR}; C DR represents abstracting drug biological knowledge as the node biological attribute in the network, and C DI represents abstracting the biological knowledge of diseases as the node biological attribute in the network; C PR represents abstracting the biological knowledge of proteins as the node biological attribute in the network.
[0025] As one of the improvements to the above technical solution, calculating the biological semantic information of diseases in the heterogeneous biological information network by applying the biological semantic similarity theory to obtain the biological attribute matrix of diseases, specifically including:
[0026] Step A1: Define a directed acyclic graph
[0027]
[0028] where represents all ancestor nodes of the disease and represents the set of all edges of the disease ;
[0029] Calculate the contribution of the disease to the disease
[0030]
[0031] where γ represents the semantic contribution factor; children of represents the set of child nodes of the disease node ;
[0032] Step A2: Obtain the semantic value of the disease by calculating the contributions of all ancestor nodes in
[0033]
[0034] Step A3: Apply the Jaccard correlation coefficient formula to calculate the semantic similarity between the disease and
[0035]
[0036] Step A4: Obtain the biological attribute matrix S DI of the diseases:
[0037]
[0038] where
[0039]
[0040] where M is the number of diseases.
[0041] As one of the improvements of the above technical solution, the graph representation features of drugs and diseases are calculated by using a multi-level subgraph representation learning strategy, and a graph representation feature matrix of drugs and diseases is obtained, specifically including:
[0042] Step B1: Define three multi-level subgraph patterns MP based on the interaction mechanism between biomolecules:
[0043] MP = {MP1, MP2, MP3}
[0044] where MP1: {drug → protein → disease}; MP2: {drug → protein → drug → disease}; MP3: {drug → protein → disease → protein → disease};
[0045] Define that each subgraph has G paths. For any path P r is composed of L nodes, that is, v1 → … → v i → v i+1 → v L ; where the path pattern r ∈ MP, the node v ∈ V, and V represents the set composed of all nodes of the heterogeneous bioinformatics network; the transfer probability from the i-th node v i to the i + 1-th node v i+1 is prob(v i , v i+1 ):
[0046]
[0047] where t ∈ {drug, protein, disease}, and t i is the type of the i-th node, representing the type set of drugs, proteins, and diseases; Φ(t i-1 ) represents the type of the next node when the current node type is t i-1 ; N(v i , Φ(t i-1 )) represents the set of neighbor node types of node v i being t i-1 ; E represents the set composed of all edges of the heterogeneous bioinformatics network;
[0048] Step B2: Obtain a multi-level subgraph G P from the paths based on the multi-level subgraph pattern obtained in Step B1:
[0049] G P = [P r r ∈ MP
[0050] Step B3: Obtain a graph representation feature matrix of drugs and diseases from the multi-level subgraph G P Use a graph neural network to learn the feature representation matrix X of the nodes in each level of subgraphs:
[0051] X k = σ(D α-1 AD -α X k-1 W k-1 )
[0052] where X k represents the representation of the nodes when the number of neural network layers is k, D is the diagonal node degree matrix of G P , α ∈ [0, 1] is the convolution coefficient, is the trainable weight matrix of the (k - 1)th layer, σ(·) is the activation function. When k = 0, X 0 is the biological attribute matrix C, and A is an |V|×|V| adjacency matrix of G P ;
[0053] According to X k , three |V|×d matrices can be formed to represent different representations of all nodes on G P , where d represents; Define a feature tensor to collect the above three different node representation matrices X k ;
[0054] Step B4: Use the attention mechanism to obtain the graph representation feature matrix H of the drug and the disease. Specifically, according to all subgraph-level representations X of each node v, construct the graph representation feature matrix of the drug and the disease
[0055]
[0056] e p = q T · t P
[0057]
[0058]
[0059] where e p represents the graph representation feature of node v in all subgraph patterns, e Pr represents the graph representation feature of node v when the subgraph pattern is r, T represents the transpose, and t P represents the normalized representation of the subgraph feature of node v, v ∈ V, represents the transformed representations of all nodes in the subgraph G P , and are trainable parameters; is the parameterized attention weight, βP For path instance P r Relative to G P Relative importance;
[0060] According to all sub - graph - level representations X of each node v, construct the graph - representation feature matrix of drugs and diseases
[0061] and the feature - representation matrix X of nodes in each - level sub - graph, as well as all sub - graph - level representations X in the following text. Here, X all refers to the same meaning.
[0062] As one of the improvements of the above - mentioned technical solution, when training a drug re - positioning model using the graph - representation feature matrix of drugs and diseases, apply a random - forest classifier to train and obtain the drug re - positioning model.
[0063] The present invention also proposes a drug re - positioning system based on a heterogeneous biological information network, constructed based on any one of the above - mentioned methods. The system includes:
[0064] A model training module, used to construct the biological association network and biological knowledge of drugs, proteins, and diseases into a heterogeneous biological information network, extract the multi - level sub - graphs of drugs and diseases in the heterogeneous biological information network to obtain the graph - representation feature matrix of drugs and diseases, and train a drug re - positioning model; and
[0065] A drug discovery module, used to input the relationship pair of a disease and candidate drugs into the trained drug re - positioning model to obtain the prediction scores of the candidate drugs for the disease.
[0066] As one of the improvements of the above - mentioned technical solution, the model training module includes:
[0067] A network construction sub - module, used to construct the biological association network and biological knowledge of drugs, proteins, and diseases into a heterogeneous biological information network, and transfer the constructed heterogeneous biological information network to the feature learning sub - module;
[0068] A feature learning sub - module, used to perform feature learning on the heterogeneous biological information network to respectively obtain the biological - attribute matrices of drugs, proteins, and diseases, and extract the multi - level sub - graphs of drugs and diseases in the heterogeneous biological information network. According to the biological - attribute matrices and multi - level sub - graphs of drugs, proteins, and diseases, obtain the graph - representation feature matrix of drugs and diseases, and transfer the graph - representation feature matrix of drugs and diseases to the model training sub - module; and
[0069] A model training sub - module, used to train a drug re - positioning model using the graph - representation features of drugs and diseases to obtain a trained drug re - positioning model.
[0070] Compared with the prior art, the advantages of the present invention are as follows:
[0071] 1. A drug repositioning method and system based on multi-level subgraph representation learning and heterogeneous biological information network provided by the present invention utilize the semantic association information of the heterogeneous biological information network and the biological function information of biomolecules to complete the drug repositioning task during the algorithm design process, improve the accuracy of drug repositioning by more comprehensively learning biomolecule features, and solve the defects of the prior art in drug repositioning in biological networks;
[0072] 2. The present invention takes into account the semantic high-order connection patterns presented in the heterogeneous biological information network; defines subgraph representation learning to enhance the model's understanding of the interactions between drugs and targets, and how they affect the occurrence and development of diseases, which may have been overlooked in previous drug repositioning models. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1 The figure shows the architecture diagram of a drug repositioning system based on a graph representation learning framework. DETAILED DESCRIPTION OF THE INVENTION
[0074] The technical solutions of the present invention will be described in detail below with reference to the accompanying drawings.
[0075] The present invention provides a drug repositioning method based on multi-level subgraph representation learning and heterogeneous biological information network, which is carried out according to the following steps:
[0076] Step 1: Collect the association relationship data of known drugs and diseases, the association relationship data of drugs and proteins, the association relationship data of proteins and diseases, the chemical structure information of drugs, the biological semantic information of diseases, and the amino acid sequence information of proteins;
[0077] Step 2: Based on the association relationship data of drugs and diseases, the association relationship data of drugs and proteins, the association relationship data of proteins and diseases, and the biological knowledge of drugs, proteins, and diseases in Step 1, construct a heterogeneous biological information network;
[0078] Define the heterogeneous biological information network as HBIN = {V, E, C};
[0079] wherein, V represents the set composed of all nodes of the heterogeneous biological information network, V = {V DR , V DI , V PR}; V DR represents the drug abstracted as a node in the network, represents the i-th drug, and N represents the number of drugs; V DI represents the disease abstracted as a node in the network represents the j-th disease, M represents the number of diseases; V PR represents the protein molecule abstracted as a node in the network represents the l-th protein molecule, P represents the number of protein molecules;
[0080] E represents the set of all edges in the heterogeneous biological information network, E = {E DD , E DP , E PD}; E DD represents the association relationship between drugs and diseases abstracted as an edge in the network; E DP represents the association relationship between drugs and protein molecules abstracted as an edge in the network, E PD abstracts the association relationship between protein molecules and diseases as an edge in the network;
[0081] C represents the set of node attributes of the heterogeneous biological information network, C = {C DR , C DI , C PR}; C DR represents the biological knowledge of drugs abstracted as the biological attributes of nodes in the network, C DI represents the biological knowledge of diseases abstracted as the biological attributes of nodes in the network; C PR represents the biological knowledge of proteins abstracted as the biological attributes of nodes in the network.
[0082] Step 3: Apply the chemical structure similarity theory to calculate the chemical structure information of the drugs in Step 1 using the RDKit open-source toolkit to obtain the biological attribute matrix of the drugs;
[0083] Obtain the biological attribute matrix of the drugs Calculate the biological attribute matrix of the drugs according to the chemical structure information of the drugs;
[0084] The specific method for calculating the biological attribute matrix of the drugs is as follows:
[0085] Apply the RDKit software package to calculate the chemical structure information of the drugs to obtain the biological attribute matrix of the drugs, run the Python code:
[0086] from rdkit.Chem import AllChem
[0087] S DR = AllChem.GetMorganFingerprintAsBitVect(A DR , 2, nBits = 1024).
[0088] Step 4: Apply the biological semantic similarity theory to calculate the biological semantic information of the diseases in Step 1 using the Jaccard correlation coefficient formula to obtain the biological attribute matrix of the diseases;
[0089] Obtain the biological attribute matrix of the diseases Calculate the biological attribute matrix of the diseases according to the disease node attributes;
[0090] The specific method for calculating the biological attribute matrix of the diseases is as follows:
[0091] Step 4-1: Define a directed acyclic graph Apply the Jaccard correlation coefficient formula to calculate the disease semantic similarity, Denote all ancestor nodes of the disease as, Denote the set of all edges of the disease as, the contribution of the disease to the disease is
[0092]
[0093] where γ represents the semantic contribution factor; childern of Denote the set of child nodes of the disease node as.
[0094] Step 4-2: By calculating the contributions of all ancestor nodes in , obtain the semantic value of the disease as:
[0095]
[0096] Step 4-3: Combine formulas (1-1) and (1-2), and The semantic similarity between is:
[0097]
[0098] Step 4-4: Obtain the biological attribute matrix S of the diseases according to formula (1-3) DI :
[0099]
[0100] where,
[0101]
[0102] where M is the number of diseases.
[0103] Step 5: Based on the heterogeneous biological information network in Step 2, apply the interaction mechanism between biomolecules to construct the multi-level subgraph representation features of drugs and diseases, and obtain the graph representation feature matrix of drugs and diseases;
[0104] Specific methods for calculating the graph representation feature matrix of drugs and diseases are as follows:
[0105] Step 5-1: Define three multi-level subgraph patterns based on the interaction mechanism between biomolecules
[0106] MP = {MP1, MP2, MP3}; where MP1: {drug → protein → disease};
[0107] MP2: {drug → protein → drug → disease};
[0108] MP3: {drug → protein → disease → protein → disease};
[0109] Define that each subgraph has G paths. For any path P r (r ∈ MP) is composed of L nodes, that is, v1 → … → v i → v i+1 → v L (v ∈ V); V represents the set composed of all nodes in the heterogeneous biological information network; node v i to v i+1 The transition probability is prob(v i , v i+1 ):
[0110]
[0111] In the above formula, t ∈ {drug, protein, disease}, v i represents the i-th node; t i is the type of the i-th node, representing the set of types of drugs, proteins, and diseases; Φ(t i-1 ) represents the type of the next node when the current node type is t i-1 ; N(v i , Φ(t i-1 )) represents the set of neighbor node types of node v i whose type is t i-1 ; E represents the set composed of all edges in the heterogeneous biological information network;
[0112] Step 5-2: Obtain the multi-level subgraph G P :
[0113] G P = [P r r = MP
[0114] Step 5-3: From the multi-level subgraph G P Obtain the graph representation feature matrix of drugs and diseases Use a specific graph neural network to learn the feature representation matrix X of the nodes in each level of the subgraph:
[0115] X k = σ(D α-1 AD -α X k-1 W k-1 )
[0116] where, in the formula X k represents the representation of the nodes when the number of neural network layers is k. D is the diagonal node degree matrix of G P , α ∈ [0, 1] is the convolution coefficient, is the trainable weight matrix of the (k - 1)th layer, and σ(·) is the activation function. When k = 0, X 0 is the biological attribute matrix C, and A is an |V|×|V| adjacency matrix of G P .
[0117] According to X k Three |V|×d matrices can be formed to represent different representations of all nodes on G P . Define a feature tensor to collect the above three different node representation matrices X k .
[0118] Step B4: Use the attention mechanism to obtain the graph representation feature matrix H of drugs and diseases:
[0119] e p = q T · t P
[0120]
[0121]
[0122]
[0123] where, v ∈ V, represents the transformed representation of all nodes in the subgraph G P , and are trainable parameters. is the parameterized attention weight, and β P is the relative importance of the path instance P r relative to G P .
[0124] Construct a graph representation feature matrix of drugs and diseases based on all subgraph-level representations X of each node v.
[0125] Step 6: Based on the graph representation feature matrix of drugs and diseases in Step 5, apply a random forest classifier to train a drug repositioning model.
[0126] Step 6-1: Input the parameter t of the random forest classifier.
[0127] Step 6-2: According to the graph representation feature matrix H of drugs and diseases, apply a random forest classifier to train the input feature matrix H to obtain a drug discovery model.
[0128] Step 7: According to the drug repositioning model obtained in Step 6, obtain the required candidate drugs.
[0129] Construct the association relationship between each disease and candidate drug into a test input feature matrix Y, and use the above-obtained drug discovery model to test the input feature matrix Y to obtain the prediction score of the candidate drug for the disease. The higher the prediction score of the candidate drug, the more suitable the candidate drug is for the input disease.
[0130] The drug repositioning model is a graph neural network model.
[0131] As Figure 1 shown, the present invention also provides a drug repositioning system based on a graph representation learning framework. The system includes a drug discovery module, a model training module, and a model result display module.
[0132] The model training module is connected to the drug discovery module, and the drug discovery module is connected to the result display module.
[0133] The model training module is used to construct a heterogeneous biological information network from the biological association network and biological knowledge of drugs, proteins, and diseases, extract multi-level subgraphs of drugs and diseases in the heterogeneous biological information network to obtain a graph representation feature matrix of drugs and diseases, and train a drug repositioning model.
[0134] The model training module includes a network construction sub-module, a feature learning sub-module, and a model training sub-module.
[0135] The network construction sub-module is used to construct a heterogeneous biological information network from the biological association network and biological knowledge of drugs, proteins, and diseases, and transfer the constructed heterogeneous biological information network to the feature learning sub-module.
[0136] The feature learning sub-module is used to extract multi-level subgraphs of drugs and diseases in the heterogeneous biological information network to obtain a graph representation feature matrix of drugs and diseases, and transfer these matrices to the model training sub-module.
[0137] The model training sub-module is used to train a drug repositioning model based on the input model parameters and the results of the feature learning module, and transfer the model to the drug discovery module.
[0138] The drug discovery module is used to generate prediction results based on the input prediction data of the drug repositioning model, and transfer these results to the result display module.
[0139] The network construction sub-module constructs a heterogeneous biological information network containing drugs, proteins, and diseases by introducing proteins as mediators to enhance the connectivity between drugs and diseases.
[0140] The feature learning sub-module more comprehensively considers the features of drugs and diseases, and trains the drug repositioning model by calculating the biological function similarity matrix and the graph representation feature matrix of drugs and diseases, and can more accurately give candidate drugs for a given disease.
[0141] Network construction sub-module:
[0142] First step, construct a heterogeneous biological information network: Define the heterogeneous biological information network as HBIN = {V, E, C}; abstract drugs, diseases, and protein molecules in the biological network as nodes in the network and Abstract the biological knowledge of drugs and diseases as the node biological attributes A DR and A DI , and abstract the association relationships between drugs, diseases, and protein molecules as the edges E DD , E DP and E PD ;
[0143] Second step, store the heterogeneous network information. Organize the set V = {V DR , V DI , V PR} composed of all nodes in the heterogeneous biological information network, the set E = {E DD , E DP , E PD} composed of all edges, and the set C = {C DR , C DI , C PR} of node attributes, and store them;
[0144] Feature learning sub-module:
[0145] First step, obtain the biological attribute matrix of drugs Calculate the biological attribute matrix of drugs according to the chemical structure information of drugs obtained in the network construction module;
[0146] The specific method for calculating the biological attribute matrix of drugs is as follows:
[0147] 2-1 Apply the RDKit software package to calculate the chemical structure information of drugs and obtain the biological attribute matrix of drugs. Run the Python code:
[0148] from rdkit.Chem import AllChem
[0149] S DR = AllChem.GetMorganFingerprintAsBitVect(A DR , 2, nBits = 1024);
[0150] The second step is to obtain the biological attribute matrix of diseases Calculate the biological attribute matrix of diseases according to the disease node attributes obtained in the network construction module;
[0151] The specific method for calculating the biological attribute matrix of diseases is as follows:
[0152] 2-2 Define a directed acyclic graph Apply the Jaccard correlation coefficient formula to calculate the disease semantic similarity, denote all ancestor nodes of disease , denote all edge sets of disease , the contribution of disease to disease is
[0153]
[0154] where γ represents the semantic contribution factor; childern of denote the set of child nodes of disease node .
[0155] 2-3 By calculating the contributions of all ancestor nodes in , the semantic value of disease is:
[0156]
[0157] 2-4 Combine formulas (1-1) and (1-2), and The semantic similarity between
[0158]
[0159] 2-5 Obtain the biological attribute matrix S of the disease according to formula (1-3). DI :
[0160]
[0161] In the above formula,
[0162]
[0163] Step 3: Obtain the graph representation feature matrix H of drugs and diseases. According to the heterogeneous biological information network data obtained in the network construction module, apply the interaction mechanism between biomolecules to construct the multi-level sub-graph representation features of drugs and diseases, and obtain the graph representation feature matrix of drugs and diseases;
[0164] The specific method for calculating the graph representation feature matrix of drugs and diseases is as follows:
[0165] 2-6 Define three multi-level sub-graph patterns MP = {MP1, MP2, MP3} based on the interaction mechanism between biomolecules; where MP1: {drug → protein → disease}; MP2: {drug → protein → drug → disease}; MP3: {drug → protein → disease → protein → disease};
[0166] Define that each sub-graph has G paths. For any path P r (r ∈ MP) is composed of L nodes, that is, v1 → … → v i → v i+1 → v L (v ∈ V); V represents the set composed of all nodes of the heterogeneous biological information network; node v i to v i+1 The transition probability is prob(v i , v i+1 ):
[0167]
[0168] In the above formula, t ∈ {drug, protein, disease}, v i represents the i-th node; t i is the type of the i-th node, representing the type set of drugs, proteins, and diseases; Φ(t i-1 ) represents the type of the next node when the current node type is t i-1 ; N(v i , Φ(t i-1 )) represents the set of neighbor node types of node v i being t i-1 ; E represents the set composed of all edges of the heterogeneous biological information network;
[0169] 2-7 The path based on the multi-level subgraph pattern is obtained from the above formula to get the multi-level subgraph G P :
[0170] G P =[P r r = MP
[0171] 2-8 From the multi-level subgraph G P Get the graph representation feature matrix of drugs and diseases Adopt a specific graph neural network to learn the feature representation matrix X of the nodes in each level of the subgraph:
[0172] X k =σ(D α-1 AD -α X k-1 W k-1 )
[0173] Among them, in the formula X k represents the representation of the node when the number of neural network layers is k. D is the diagonal node degree matrix of G P , α ∈ [0,1] is the convolution coefficient, is the trainable weight matrix of the k-1 layer, and σ(·) is the activation function. When k = 0, X 0 is the biological attribute matrix C, and A is an |V|×|V| adjacency matrix of G P .
[0174] According to X k Three |V|×d matrices can be formed to represent different representations of all nodes on G P Define a feature tensor to collect the above three different node representation matrices X k .
[0175] 2-9 Adopt the attention mechanism to get the graph representation feature matrix H of drugs and diseases:
[0176] e p =q T ·t P
[0177]
[0178]
[0179]
[0180] Among them, v ∈ V, represents the transformed representation of all nodes in the subgraph G P , and are trainable parameters. is the parameterized attention weight, β P is the path instance P r relative to G P relative importance.
[0181] Construct a graph representation feature matrix of drugs and diseases based on all subgraph-level representations X of each node v
[0182] Model training sub-module:
[0183] In the first step, input the parameter t of the random forest classifier;
[0184] In the second step, according to the graph representation feature matrix H of drugs and diseases, apply the random forest classifier to train the input feature matrix H to obtain a drug discovery model.
[0185] Use the drug discovery module to predict the association relationship between each disease and candidate drug. Specifically, construct the association relationship between each disease and candidate drug as a test input feature matrix Y, and use the obtained drug discovery model above to test the input feature matrix Y to obtain the prediction scores of the candidate drugs for the disease for use by the result display module;
[0186] Result display module:
[0187] According to the results obtained by the drug discovery module, this module takes the association relationship between each disease and candidate drug as a row, where the disease, drug, and prediction score are elements in the row, and processes all prediction results into a text file for output display.
[0188] The present invention can also provide a computer device, including: at least one processor, a memory, at least one network interface, and a user interface. Each component in the device is coupled together through a bus system. It can be understood that the bus system is used to realize the connection and communication between these components. In addition to the data bus, the bus system also includes a power bus, a control bus, and a status signal bus.
[0189] Among them, the user interface can include a display, a keyboard, or a pointing device. For example, a mouse, a trackball, a touchpad, or a touch screen, etc.
[0190] It can be understood that the memory in the disclosed embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DRRAM). The memories described herein are intended to include but not be limited to these and any other suitable types of memories.
[0191] In some embodiments, the memory stores the following elements, executable modules, or data structures, or subsets or supersets thereof: an operating system and application programs.
[0192] Among them, the operating system includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., and is used to implement various basic services and process hardware-based tasks. The application programs include various application programs, such as a media player and a browser, etc., and are used to implement various application services. The program for implementing the method of the disclosed embodiments of the present application can be included in the application programs.
[0193] In the above embodiments, by calling the programs or instructions stored in the memory, specifically, the programs or instructions stored in the application programs, the processor is configured to:
[0194] Execute the steps of the above method.
[0195] The above method can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with the ability to process signals. During implementation, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute each of the above-disclosed methods, steps, and logic block diagrams. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Combining the steps of the above-disclosed method can be directly embodied as being completed by the execution of a hardware decoding processor, or by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0196] It can be understood that these embodiments described in the present invention can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application, or a combination thereof.
[0197] For software implementation, the techniques of the present invention can be implemented by executing the functional modules of the present invention (such as procedures, functions, etc.). The software code can be stored in the memory and executed by the processor. The memory can be implemented inside or outside the processor.
[0198] The present invention can also provide a non-volatile storage medium for storing a computer program. When the computer program is executed by a processor, each step in the above method embodiments can be implemented.
[0199] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that any modification or equivalent replacement of the technical solutions of the present invention does not depart from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A drug repositioning method based on a heterogeneous biological information network, the method comprising: Input the disease and candidate drug names into the trained drug repositioning model to obtain the predicted scores of the candidate drugs for the disease, and complete drug repositioning according to the predicted scores; among them, The drug repositioning model is a graph representation learning model, and the process of establishing a training data set during its training includes: Construct a heterogeneous biological information network based on the existing known association relationship data between drugs and diseases, the association relationship data between drugs and proteins, the association relationship data between proteins and diseases, the chemical structure information of drugs, the biological semantic information of diseases, and the base sequence information of proteins; Perform feature learning on the heterogeneous biological information network to obtain the biological attribute matrices of drugs, proteins, and diseases respectively; Extract the multi-level subgraphs of drugs and diseases in the heterogeneous biological information network; According to the biological attribute matrices and multi-level subgraphs of drugs, proteins, and diseases, obtain the graph representation feature matrix of drugs and diseases, that is, the training data set; The obtaining of the graph representation feature matrix of drugs and diseases is as follows: Calculate the graph representation features of drugs and diseases using the multi-level subgraph representation learning strategy to obtain the graph representation feature matrix of drugs and diseases, specifically including: Step B1: Define three multi-level subgraph patterns MP based on the interaction mechanism between biomolecules: MP = {MP1, MP2, MP3} Among them, MP1: {drug → protein → disease}; MP2: {drug → protein → drug → disease}; MP3: {drug → protein → disease → protein → disease}; Define that each subgraph has G paths. For any path P r is composed of L nodes, that is, v1 → … → v i → v i+1 → v L ; where the path pattern r ∈ MP, the node v ∈ V, and V represents the set composed of all nodes in the heterogeneous bioinformatics network; the i-th node v i to the (i + 1)-th node v i+1 has a transition probability of prob(v i , v i+1 ): where \(t\in\{\text{drug},\text{protein},\text{disease}\}\), \(t\) i is the type of the \(i\)-th node, representing the set of types of drugs, proteins, and diseases; \(\varPhi(t\) i-1 ) represents the type of the next node when the current node type is \(t\) i-1 ; \(N(v\) i ,\varPhi(t\) i-1 )) represents the set of neighbor node types of node \(v\) i with type \(t\) i-1 ; \(E\) represents the set of all edges in the heterogeneous biological information network; Step B2: Obtain the multi-level subgraph G from the path based on the multi-level subgraph pattern obtained according to Step B1 P : G P = [P r r ∈ MP Step B3: Obtain the graph representation feature matrix of drugs and diseases from the multi-level subgraph G P to obtain the graph representation feature matrix of drugs and diseases Use a graph neural network to learn the feature representation matrix X of the nodes in each level of the subgraph: X k = σ(D α-1 AD -α X k-1 W k-1 ) Among them, X k represents the representation of a node when the number of neural network layers is k, D is the P diagonal node degree matrix of G, α ∈ [0, 1] is the convolution coefficient, is the trainable weight matrix of the (k - 1)-th layer, σ(·) is the activation function, when k = 0, X 0 is the biological attribute matrix C, and A is a P |V|×|V| adjacency matrix of G; According to X k Three |V|×d matrices can be formed to represent different representations of all nodes in G P ; Define a feature tensor to collect the above three different representation matrices X of nodes k ; Step B4: Use the attention mechanism to obtain the graph representation feature matrix H of the drug and the disease. Specifically, based on all subgraph-level representations X of each node v, construct the graph representation feature matrix of the drug and the disease e p = q T · t P Among them, e p represents the graph representation feature of node v in all subgraph patterns, represents the graph representation feature of node v when the subgraph pattern is r, T represents transpose, t P represents the representation after normalization of the subgraph feature of node v, v ∈ V, represents the transformed representation of all nodes in subgraph G P ; and are trainable parameters; is the parameterized attention weight, β P is the relative importance of path instance P r relative to G P ; Construct a graph representation feature matrix of drugs and diseases based on all subgraph-level representations X of each node v 2. The drug repositioning method based on a heterogeneous biological information network according to claim 1, wherein The performing of feature learning on the heterogeneous biological information network to obtain the biological attribute matrices of drugs, proteins, and diseases respectively specifically includes: Apply the chemical structure similarity theory to calculate the chemical structure information of drugs to obtain the biological attribute matrix of drugs; apply the biological semantic similarity theory to calculate the biological semantic information of diseases to obtain the biological attribute matrix of diseases; apply the biological sequence statistics theory to calculate the base sequence information of proteins to obtain the biological attribute matrix of proteins.
3. The drug repositioning method based on a heterogeneous biological information network according to claim 2, wherein The applying of the chemical structure similarity theory to calculate the chemical structure information of drugs in the heterogeneous biological information network to obtain the biological attribute matrix of drugs is specifically: Apply the chemical structure similarity theory to calculate the chemical structure information of drugs using the RDKit open source toolkit to obtain the biological attribute matrix of drugs.
4. The drug repositioning method based on a heterogeneous biological information network according to any one of claims 1-3, wherein The heterogeneous biological information network HBIN is represented as: HBIN = {V, E, C} Among them, V represents the set composed of all nodes in the heterogeneous biological information network, V = {V DR , V DI , V PR}; V DR represents the drug abstracted as a node in the network, represents the i-th drug, N represents the number of drugs; V DI represents the disease abstracted as a node in the network represents the j-th disease, M represents the number of diseases; V PR represents the protein molecule abstracted as a node in the network represents the l-th protein molecule, P represents the number of protein molecules; Let \(E\) denote the set composed of all edges in the heterogeneous biological information network, \(E = \{E DD , E DP , E PD \}\); \(E DD denotes the association relationship between drugs and diseases abstracted as an edge in the network; \(E DP denotes the association relationship between drugs and protein molecules abstracted as an edge in the network, and \(E PD abstracts the association relationship between protein molecules and diseases as an edge in the network; C represents the set of heterogeneous bioinformatics network node attributes, C = {C DR , C DI , C PR}; C DR represents the abstraction of drug biological knowledge into node biological attributes in the network, C DI represents the abstraction of disease biological knowledge into node biological attributes in the network; C PR represents the abstraction of protein biological knowledge into node biological attributes in the network.
5. The drug repositioning method based on a heterogeneous biological information network according to claim 4, characterized in that, The performing of feature learning on the heterogeneous biological information network to obtain the biological attribute matrix of diseases specifically includes: Step A1: Define a directed acyclic graph Among them, represents all ancestor nodes of the disease , and represents the set of all edges of the disease ; Calculate the disease For the disease Contribution Among them, γ represents the semantic contribution factor; childrenof represents the set of child nodes of the disease node ; Step A2: By calculating the contributions of all ancestor nodes in to obtain the semantic value of the disease Step A3: Calculate the semantic similarity between diseases using the Jaccard correlation coefficient formula and using the Jaccard correlation coefficient formula Step A4: Obtain the biological attribute matrix S of the disease DI : Among them, where M is the number of diseases.
6. The drug repositioning method based on a heterogeneous biological information network according to any one of claims 1 - 3, characterized in that, When using the graph representation feature matrix of drugs and diseases to train a drug repositioning model, apply a random forest classifier to train and obtain the drug repositioning model.
7. A drug repositioning system based on a heterogeneous biological information network, constructed based on the method according to any one of claims 1 - 6, characterized in that, The system includes: A model training module, which is used to construct a heterogeneous biological information network from the biological association network and biological knowledge of drugs, proteins, and diseases, extract the multi-level subgraphs of drugs and diseases in the heterogeneous biological information network, obtain the graph representation feature matrix of drugs and diseases, and train a drug repositioning model; and A drug discovery module for inputting the relationship pairs of diseases and candidate drugs into a trained drug repositioning model to obtain the prediction scores of the candidate drugs for the diseases; The model training module includes: A network construction sub-module for constructing a heterogeneous biological information network from the biological association networks and biological knowledge of drugs, proteins, and diseases, and transmitting the constructed heterogeneous biological information network to the feature learning sub-module; A feature learning sub-module for performing feature learning on the heterogeneous biological information network to respectively obtain the biological attribute matrices of drugs, proteins, and diseases, and for extracting the multi-level subgraphs of drugs and diseases in the heterogeneous biological information network, obtaining the graph representation feature matrices of drugs and diseases based on the biological attribute matrices and multi-level subgraphs of drugs, proteins, and diseases, and transmitting the graph representation feature matrices of drugs and diseases to the model training sub-module; and A model training sub-module for training a drug repositioning model using the graph representation features of drugs and diseases to obtain a trained drug repositioning model.
Citation Information
Patent Citations
Drug repositioning method based on low-rank matrix completion
CN107545151A
Drug relocation method and system based on biological knowledge and network topology structure
CN114171113A
Drug relocation method and system based on drug classification map neural network
CN114613452A