A drug repositioning method based on topological information and biochemical information

By building a heterogeneous network, comprehensively utilizing topological information and biochemical information, screening out reasonable negative samples, and using integrated learning classifiers for drug relocation, the problems of insufficient utilization of heterogeneous network information and unreliable negative samples in the existing technology are solved, and the prediction accuracy is improved.

CN119339782BActive Publication Date: 2025-07-22XINYU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411337472.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-25
Publication Date
2025-07-22
Estimated Expiration
2044-09-25

AI Technical Summary

Technical Problem

Existing drug relocation technologies are insufficiently utilized by heterogeneous network information and lack reliable negative samples, resulting in low prediction accuracy.

Method used

By building a heterogeneous network, comprehensively utilize topological information and biochemical information, more reasonable negative samples are selected and predictions are used using an integrated learning classifier.

Benefits of technology

It improves the prediction accuracy of drug relocation methods and improves the comprehensive consideration of the possibility of drug and disease interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119339782B_ABST
    Figure CN119339782B_ABST
Patent Text Reader

Abstract

The present application relates to a drug repositioning method based on topological information and biochemical information, which comprises the following steps: obtaining a data set to construct a heterogeneous network; calculating and obtaining a comprehensive topological information matrix and a biochemical affinity matrix of drugs, a comprehensive topological information matrix and a biochemical affinity matrix of diseases, a drug feature matrix, and a disease feature matrix; expanding the meta-paths in the heterogeneous network and screening out more reasonable negative samples; using an ensemble learning classifier for prediction, and outputting the probability of a treatment relationship between a drug-disease pair as the prediction result. The present invention overcomes the limitation of the existing drug repositioning technology in insufficient utilization of heterogeneous network information, uses biochemical information from multiple perspectives and optimizes it, so as to obtain a more comprehensive feature information representation; at the same time, in view of the fact that there is a lack of drug-disease deterministic negative samples, it fully considers various interaction possibilities between drugs and diseases, screens out reliable negative samples, and improves the prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of bioinformatics, and particularly relates to a drug repositioning method based on topological information and biochemical information. Background Art

[0002] Discovering new drugs is a complex task that requires knowledge in many biological and chemical fields, as well as a large amount of time and money, and is accompanied by a high failure rate. At the same time, the utilization value of already marketed drugs has far from been fully exploited, resulting in a great waste of existing drug resources. Therefore, drug repositioning methods are of great significance for the application and development of drugs.

[0003] Yajie Meng et al. disclosed a drug repositioning method DRAGNN using a weighted local information augmented graph neural network in the paper "Drug repositioning based on weighted local information augmented graph neural network" published in 2024. This method integrates relevant data of drugs and diseases for drug repositioning. By integrating a drug-drug similarity network, a disease-disease similarity network, and a drug-disease treatment relationship network, a heterogeneous network was constructed; heterogeneous information and neighborhood homogeneous information of drugs and diseases were collected and integrated, and an MLP was used for the final prediction. However, the integration of drug and disease information and the utilization of heterogeneous network information by this method are not sufficient enough, so it has certain limitations. At the same time, when training a model using a supervised learning method, both positive samples and negative samples are required. However, drug-disease association data is similar to other biological data and lacks experimentally verified negative samples. Currently, most methods randomly select some from unknown drug-disease pairs as negative samples. However, such negative samples are unreliable, and there may be positive samples that have not been discovered yet. Summary of the Invention

[0004] The purpose of the present invention is to provide a drug repositioning method based on topological information and biochemical information, which overcomes the limitation that the existing drug repositioning technology does not utilize heterogeneous network information sufficiently, can more fully integrate and utilize heterogeneous network information, uses biochemical information from multiple perspectives and optimizes it, so as to obtain a more comprehensive feature information representation; at the same time, in view of the fact that there is a lack of drug-disease deterministic negative samples, various interaction possibilities between drugs and diseases are fully considered, reliable negative samples can be screened out, the defect of randomly selecting negative samples by general methods is overcome, and the prediction accuracy is improved.

[0005] The technical solution adopted by the present invention is: a drug repositioning method based on topological information and biochemical information, including the following steps:

[0006] S1: Obtain a dataset to construct a heterogeneous network; the dataset includes a drug-disease treatment relationship dataset, a drug-protein interaction relationship dataset, and a disease-protein interaction relationship dataset; obtain 5 similarity matrices of drugs and 2 similarity matrices of diseases according to the dataset, the 5 similarity matrices of drugs include a chemical structure similarity matrix, an ATC code similarity matrix, a side effect similarity matrix, a drug-drug interaction similarity matrix, and a target spectrum similarity matrix, and the 2 similarity matrices of diseases include a disease phenotype similarity matrix and a disease ontology similarity matrix;

[0007] S2: Use the number of one-step neighbors of drugs or diseases as the measurement basis for the corresponding node topology information, calculate the number of common one-step heterogeneous neighbors between drugs and other drugs and between diseases and other diseases, construct a drug-drug common one-step neighbor matrix and a disease-disease common one-step neighbor matrix, and use cosine similarity to calculate the similarity between two nodes, construct a common one-step neighbor similarity matrix between drug nodes and a common one-step neighbor similarity matrix between disease nodes, and obtain a comprehensive topology information matrix of drugs and a comprehensive topology information matrix of diseases;

[0008] S3: Calculate the average values of the 5 similarity matrices of drugs and the 2 similarity matrices of diseases from multiple biochemical perspectives in step S1 respectively to obtain an integrated drug similarity matrix and an integrated disease similarity matrix, and adjust the integrated drug similarity matrix and the integrated disease similarity matrix according to the pruning algorithm to obtain a biochemical proximity matrix of drugs and a biochemical proximity matrix of diseases;

[0009] S4: Extract and integrate the topological information and biochemical information of drugs and diseases, remove the noise of the data, and obtain low-dimensional feature representations of drugs and diseases to obtain a drug feature matrix and a disease feature matrix;

[0010] S5: Expand the "protein" in the meta-path in the heterogeneous network to "protein → drug → protein" or "protein → disease → protein", and screen out more reasonable negative samples, where the negative samples are drug-disease pairs without a treatment relationship;

[0011] S6: After concatenating the drug features and disease features obtained in S4, use the feature vector of the drug-disease pair as the input, and use the integrated learning classifier XGBoost for prediction, and output the probability that the drug-disease pair has a treatment relationship as the prediction result.

[0012] Furthermore, the similarity value range in the 5 similarity matrices of drugs and the 2 similarity matrices of diseases is [0, 1], and the higher the similarity, the more similar the two drugs or diseases are.

[0013] Furthermore, the specific steps of step S2 are as follows:

[0014] S201: Write the drug-disease treatment relationship dataset, drug-protein interaction relationship dataset, and disease-protein interaction relationship dataset in matrix form, denoted as drug-disease treatment relationship matrix A ds , drug-protein interaction relationship matrix A dp and disease-protein interaction relationship matrix A sp ;

[0015] S202: Calculate the drug-drug common one-step neighbor matrix N ds and disease-disease common one-step neighbor matrix N dp according to the drug-disease treatment relationship matrix A sp , drug-protein interaction relationship matrix A dd and disease-protein interaction relationship matrix A ss . The specific calculation formula is:

[0016] N dd =A dp ×(A dp ) T +A ds ×(A ds ) T ;

[0017] N ss =A sp ×(A sp ) T +(A ds ) T ×A ds ;

[0018] Among them, (A dp ) T represents the transpose matrix of the drug-protein interaction relationship matrix A dp , (A ds ) T represents the transpose matrix of the drug-disease treatment relationship matrix A ds , and (A sp ) T represents the transpose matrix of the disease-protein interaction relationship matrix A sp ;

[0019] Use the data in the i-th row of the drug-drug common one-step neighbor matrix N dd as the common one-step neighbor information of the i-th drug node and other drug nodes, and the data in the j-th row of the disease-disease common one-step neighbor matrix N ss as the common one-step neighbor information of the j-th disease node and other disease nodes;

[0020] S203: Calculate the common one-step neighbor similarity matrix S between drug nodes using cosine similarity d and the common one-step neighbor similarity matrix S between disease nodes s , and the specific calculation formula is as follows:

[0021]

[0022] where m represents the total number of drug nodes, and n represents the total number of disease nodes; (x1, x2,..., x m ) represents the common one-step neighbor information of drug x, (y1, y2,..., y m ) represents the common one-step neighbor information of drug y, x i represents the common one-step neighbor information between drug x and the i-th drug, y i represents the common one-step neighbor information between drug y and the i-th drug; (a1, a2,..., a n ) represents the common one-step neighbor information of disease a, (b1, b2,..., b n ) represents the common one-step neighbor information of disease b, a j represents the common one-step neighbor information between disease a and the j-th disease, b j represents the common one-step neighbor information between disease b and the j-th disease;

[0023] S204: Concatenate the common one-step neighbor similarity matrix S between drug nodes d with the drug-disease treatment relationship matrix A ds and the drug-protein interaction relationship matrix A dp to obtain the comprehensive topological information matrix T of drugs d ; at the same time, concatenate the common one-step neighbor similarity matrix S between disease nodes s with the transposed matrix of the drug-disease treatment relationship matrix A ds (A ds ) T and the disease-protein interaction relationship matrix A sp to obtain the comprehensive topological information matrix T of diseases s , and the specific calculation formula is as follows:

[0024] T d = [S d , A ds , A dp ;

[0025] T s = [S s , (A ds ) T , Asp .

[0026] Further, the specific steps of step S3 are as follows:

[0027] S301: Calculate the average values of the drug similarity and disease similarity from multiple biochemical perspectives respectively to obtain the integrated drug similarity matrix Sim d and disease similarity matrix Sim s . The specific calculation formula is:

[0028]

[0029] where ChemSim d is the chemical structure similarity matrix, ATCSim d is the ATC code similarity matrix, SESim d is the side effect similarity matrix, DrDrSim d is the drug-drug interaction similarity matrix, TargetSim d is the target spectrum similarity matrix; PhenomeSim s is the disease phenotype similarity matrix, MeshSim s is the disease ontology similarity matrix;

[0030] S302: Adjust the integrated drug similarity matrix and disease similarity matrix according to the pruning algorithm, that is, adjust the first k similarity values in descending order of similarity values in each row of the drug similarity matrix and disease similarity matrix, and set the remaining values to zero, so as to obtain the biochemical affinity matrix B d of drugs and the biochemical affinity matrix B s of diseases. The specific calculation formula is:

[0031] B d = pruningSim(Sim d , k, η);

[0032] B s = pruningSim(Sim s , k, η);

[0033] where pruningSim(·) represents the pruning operation, k represents the number of adjustments, and η represents the decay factor;

[0034] The steps of the pruning algorithm are:

[0035] Initialize the biochemical affinity matrix B d of drugs as a 0 matrix with m rows and m columns, and initialize the biochemical affinity matrix B sInitialize it as an n×n zero matrix; for each row of the drug similarity matrix Sim d , store the column subscripts corresponding to the top k numerical values in descending order of similarity values in each row of the matrix indexes1. For each row of the disease similarity matrix Sim s , store the column subscripts corresponding to the top k numerical values in descending order of similarity values in each row of the matrix indexes2; according to the column subscripts recorded in each row and column of the indexes1 matrix, attenuate the element values at the corresponding positions in the drug similarity matrix Sim d , and fill the attenuated numerical values into the corresponding positions of the drug biochemical affinity matrix B d to obtain the drug biochemical affinity matrix B d ; according to the column subscripts recorded in each row and column of the indexes2 matrix, attenuate the element values at the corresponding positions in the disease similarity matrix Sim s , and fill the attenuated numerical values into the corresponding positions of the disease biochemical affinity matrix B s to obtain the disease biochemical affinity matrix B s , and the specific calculation formula is:

[0036] B d (M, indexes1(M, K)) = η K-1 ×Sim d (M, indexes1(M, K));

[0037] B s (N, indexes2(N, K)) = η K-1 ×Sim s (N, indexes2(N, K));

[0038] where M = 1, 2, …, m; N = 1, 2, …, n; K = 1, 2, …, k.

[0039] Furthermore, the specific steps of step S4 are: perform singular value decomposition on the comprehensive topological information matrix T d of drugs, the drug biochemical affinity matrix B d , the comprehensive topological information matrix T s of diseases, and the disease biochemical affinity matrix B d respectively, and integrate the matrices obtained after decomposition to obtain the drug feature matrix F d and the disease feature matrix F s , and the specific calculation formula is:

[0040] F d = [U Td , U Bd ;

[0041] F s = [U Ts , U Bs ;

[0042] Among them, U Td is the eigenmatrix after the singular value decomposition of the comprehensive topological information matrix T d of the drug, U Bd is the eigenmatrix after the singular value decomposition of the biochemical affinity matrix B d of the drug, U Ts is the eigenmatrix after the singular value decomposition of the comprehensive topological information matrix T s of the disease, U Bs is the eigenmatrix after the singular value decomposition of the biochemical affinity matrix B s of the disease.

[0043] Furthermore, the specific steps of the step S5 are as follows:

[0044] S501: Denote the "drug → disease" path as the meta-path M1, the "drug → protein → disease" path as the meta-path M2, the "drug → protein → drug → disease" path as the meta-path M31, the "drug → disease → drug → disease" path as the meta-path M32, and the "drug → disease → protein → disease" path as the meta-path M33;

[0045] S502: Expand the "protein" in the meta-path M2 into "protein → drug → protein" to obtain the meta-path M2P1; expand the "protein" in the meta-path M2 into "protein → disease → protein" to obtain the meta-path M2P2; expand the "protein" in the meta-path M31 into "protein → drug → protein" to obtain the meta-path M31P1; expand the "protein" in the meta-path M31 into "protein → disease → protein" to obtain the meta-path M31P2; expand the "protein" in the meta-path M33 into "protein → drug → protein" to obtain the meta-path M33P1; expand the "protein" in the meta-path M33 into "protein → disease → protein" to obtain the meta-path M33P2;

[0046] S503: The meta-path M1, the meta-path M2, the meta-path M31, the meta-path M32, the meta-path M33, the meta-path M2P1, the meta-path M2P2, the meta-path M31P1, the meta-path M31P2, the meta-path M33P1, and the meta-path M33P2 together constitute the meta-path matrix M of the heterogeneous network, and the specific calculation formula is:

[0047] M = M1 + M2 + M31 + M32 + M33 + M2P1 + M2P2 + M31P1

[0048] +M31P2 + M33P1 + M33P2;

[0049] M1 = A ds ;

[0050] M2 = A dp ×(A sp ) T ;

[0051] M31 = A dp ×(A dp ) T ×A ds ;

[0052] M32 = A ds ×(A ds ) T ×A ds ;

[0053] M33 = A ds ×A sp ×(A sp ) T ;

[0054] M2P1 = A dp ×A pdp ×(A sp ) T ;

[0055] M2P2 = A dp ×A pSp ×(A sp ) T ;

[0056] M31P1 = A dp ×A pdp ×(A dp ) T ×A ds ;

[0057] M31P2 = A dp ×A psp ×(A dp ) T ×A ds ;

[0058] M33P1 = A ds ×A sp ×A pdp ×(A sp ) T ;

[0059] M33P2 = A ds ×A sp ×A psp ×(A sp) T ;

[0060] A pdp =(A dp ) T ×A dp ;

[0061] A psp =(A sp ) T ×A sp ;

[0062] Wherein, A pdp represents the "protein → drug → protein" expansion mode, and A psp represents the "protein → disease → protein" expansion mode;

[0063] S504: Determine whether each element value in the meta-path matrix M is 0, and record the drug-disease combination mapped by the subscript combination corresponding to the element with the element value of 0 as a negative sample.

[0064] Furthermore, the specific steps of step S6 are as follows: Regarding the drug-disease association prediction as a binary classification problem, where the positive sample is the drug-disease pair with a treatment relationship, and the negative sample is the drug-disease pair without a treatment relationship; After concatenating the drug features and disease features obtained in S4, the feature vector of the drug-disease pair is used as the input, and an initial decision tree is created through the ensemble learning classifier XGBoost, and then new regression trees are iteratively constructed to reduce the calculation residuals. The scores corresponding to the leaf nodes of each feature vector on each tree are added up as the prediction result corresponding to the feature vector, that is, the probability that the drug-disease pair has a treatment relationship.

[0065] The beneficial effects of the present invention are as follows:

[0066] (1) When extracting drug and disease information, the present invention comprehensively considers topological information and biochemical information, makes more reasonable use of the heterogeneous network, extracts more sufficient and effective heterogeneous network information, uses biochemical information from multiple perspectives and optimizes it, so as to obtain a more comprehensive feature information representation and improve the performance of the repositioning method described in the present invention;

[0067] (2) Compared with the method of randomly selecting negative samples from unknown samples adopted in the prior art, the present invention adopts the method of expanding "protein" in the meta-path to "protein → drug → protein" or "protein → disease → protein", comprehensively considering the 11 meta-path situations starting from drugs and ending with diseases, fully considering various interaction possibilities between drugs and diseases, so as to screen out more reliable negative samples and further improve the accuracy of the prediction result. Description of the Drawings

[0068] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0069] Figure 1 It is the method flowchart of the embodiment of the present invention. Detailed implementation manners

[0070] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0071] As Figure 1 shown, the embodiment of the present invention proposes a drug repositioning method based on topological information and biochemical information, including the following steps:

[0072] S1: Obtain a data set to construct a heterogeneous network; the data set includes a drug-disease treatment relationship data set, a drug-protein interaction relationship data set, and a disease-protein interaction relationship data set; obtain 5 similarity matrices of drugs and 2 similarity matrices of diseases according to the data set. The 5 similarity matrices of drugs include a chemical structure similarity matrix, an ATC code similarity matrix, a side effect similarity matrix, a drug-drug interaction similarity matrix, and a target spectrum similarity matrix. The 2 similarity matrices of diseases include a disease phenotype similarity matrix and a disease ontology similarity matrix.

[0073] In the embodiments of the present invention, the similarity values in the five similarity matrices of drugs and the two similarity matrices of diseases range from [0, 1]. The higher the similarity, the more similar two drugs or diseases are. The drug-disease association dataset includes the Fdataset dataset and the Cdataset dataset. Among them, the Fdataset dataset can be directly obtained from the paper "PREDICT: a method for inferring novel drug indications with application to personalized medicine" published by Assaf Gottlieb et al. in 2011; the Cdataset dataset can be directly obtained from the paper "Drug repositioning based on comprehensive similarity measures and Bi-Random Walk algorithm" published by Huimin Luo et al. in 2016. The Fdataset dataset contains 593 drugs, 313 diseases, 1753 proteins, 1933 verified drug-disease associations, 3184 drug-protein interaction relationships, and 1066 disease-protein interaction relationships. The Cdataset dataset includes 663 drugs, 409 diseases, 1805 proteins, 2532 verified drug-disease associations, 3251 drug-protein interaction relationships, and 1280 disease-protein interaction relationships.

[0074] The five similarity matrices of drugs and the two similarity matrices of diseases can be directly obtained from the paper "Computational drug repositioning based on multi-similarities bilinear matrix factorization" published by Mengyun Yang et al. in 2021, or can be calculated or obtained respectively using the following methods:

[0075] (1) Chemical structure similarity of drugs

[0076] First, download the SMILES string information of drugs from the DrugBank database, and then use an open-source tool CDK to convert the SMILES strings of drugs into molecular fingerprint information. Based on the molecular fingerprint information of two drugs, use the Tanimoto metric method, that is, the method of dividing the intersection of the two molecular fingerprint information by the union to calculate the similarity between the two molecular fingerprints, and obtain the chemical structure similarity of drug-drug.

[0077] (2) ATC code similarity of drugs

[0078] First, download the ATC codes of drugs from the DrugBank database, and then calculate the ATC code similarity of drug-drug according to the semantic similarity algorithm introduced in the paper "Using Information Content to Evaluate Semantic Similarity in a Taxonomy" published by Philip Resnik in 1995.

[0079] (3) Side effect similarity of drugs

[0080] First, download the side effect information of drugs from the SIDER database, with a total of 5,868 side effects. Then, represent each drug as a 5,868-bit binary string. If the drug has a certain side effect, set the value at the corresponding position in the binary string to 1, otherwise set it to 0. Finally, calculate the Jaccard similarity between the binary strings to obtain the side effect similarity of drug-drug.

[0081] (4) Drug-drug interaction similarity

[0082] First, download the drug-drug interaction data from the DrugBank database, and then represent each drug as a binary interaction expression profile of the drug. In the binary interaction expression profile of the drug, for the drugs that interact with this drug, mark the corresponding positions as 1, otherwise mark them as 0. Finally, calculate the Jaccard similarity between the binary strings to obtain the drug-drug interaction similarity.

[0083] (5) Target spectrum similarity of drugs

[0084] First, download the drug-target interaction data from the DrugBank database, and then represent each drug as a binary interaction expression profile of the target. In the binary interaction expression profile of the target, for the targets that interact with this drug, mark the corresponding positions as 1, otherwise mark them as 0. Finally, calculate the Jaccard similarity between the binary strings to obtain the target spectrum similarity of drug-drug.

[0085] (6) Phenotypic similarity of diseases

[0086] Download the phenotypic similarity data of diseases from the MimMiner database recorded in the paper "A text-mining analysis of the human phenome" published by Marc A van Driel et al. in 2006. This database provides the phenotypic similarities of 5,080 diseases in the Online Mendelian Inheritance in Man (OMIM) database. Extract the similarities of the diseases required in the embodiments of the present invention to form a disease phenotypic similarity matrix.

[0087] (7) Ontological similarity of diseases

[0088] First, download the MeSH descriptors of diseases from the NIH website. According to the MeSH descriptors, represent the diseases as a directed acyclic graph, and then calculate the ontological semantic similarity of the diseases according to the method introduced in the paper "Inferring the human microRNA functional similarity and functional network based on microRNA-associated diseases" published by Dong Wang et al. in 2010.

[0089] S2: If some proteins or some diseases are shared between drugs, then these drugs may be more similar. Similarly, if some proteins or some drugs are shared between diseases, then these diseases may be more similar. Therefore, in the embodiments of the present invention, the number of one-step neighbors of drugs or diseases is used as the measurement basis for the topological information of the corresponding nodes, calculate the number of common one-step heterogeneous neighbors (proteins, diseases) between drugs and other drugs, and the number of common one-step heterogeneous neighbors (proteins, drugs) between diseases and other diseases, construct a drug-drug common one-step neighbor matrix and a disease-disease common one-step neighbor matrix, and use cosine similarity to calculate the similarity between two nodes, construct a common one-step neighbor similarity matrix between drug nodes and a common one-step neighbor similarity matrix between disease nodes, and obtain a comprehensive topological information matrix of drugs and a comprehensive topological information matrix of diseases; the specific steps are as follows:

[0090] S201: Write the drug-disease treatment relationship dataset, the drug-protein interaction relationship dataset, and the disease-protein interaction relationship dataset in matrix form, and denote them as the drug-disease treatment relationship matrix A ds 、the drug-protein interaction relationship matrix A dp and the disease-protein interaction relationship matrix A sp . Among them. The drug-disease treatment relationship matrix A dsIf the data in the i-th row and j-th column of the matrix is 1, it indicates that the i-th drug has a treatment relationship with the j-th disease; if the data in the i-th row and j-th column is 0, it indicates that the i-th drug has no treatment relationship with the j-th disease; drug-protein interaction relationship matrix A dp If the data in the i-th row and j-th column of the matrix is 1, it indicates that the i-th drug has an interaction relationship with the j-th protein; if the data in the i-th row and j-th column is 0, it indicates that the i-th drug has no interaction relationship with the j-th protein; if the data in the i-th row and j-th column of the disease-protein interaction relationship matrix is 1, it indicates that the i-th disease has an interaction relationship with the j-th protein; if the data in the i-th row and j-th column is 0, it indicates that the i-th disease has no interaction relationship with the j-th protein.

[0091] S202: According to the drug-disease treatment relationship matrix A ds , drug-protein interaction relationship matrix A dp and disease-protein interaction relationship matrix A sp Calculate the drug-drug common one-step neighbor matrix N dd and the disease-disease common one-step neighbor matrix N ss , and the specific calculation formula is:

[0092] N dd = A dp × (A dp ) T + A ds × (A ds ) T ;

[0093] N ss = A sp × (A sp ) T + (A ds ) T × A ds ;

[0094] Among them, (A dp ) T represents the transpose matrix of the drug-protein interaction relationship matrix A dp , (A ds ) T represents the transpose matrix of the drug-disease treatment relationship matrix A ds , (A sp ) T represents the transpose matrix of the disease-protein interaction relationship matrix A sp .

[0095] Using the data in the i-th row of the drug-drug common one-step neighbor matrix N dd as the common one-step neighbor information of the i-th drug node and other drug nodes, and the disease-disease common one-step neighbor matrix N ssThe j-th row data in it is used as the common one-step neighbor information between the j-th disease node and other disease nodes.

[0096] S203: Calculate the common one-step neighbor similarity matrix S between drug nodes using cosine similarity d and the common one-step neighbor similarity matrix S between disease nodes s , and the specific calculation formula is:

[0097]

[0098] where m represents the total number of drug nodes, and n represents the total number of disease nodes; (x1, x2,..., x m ) represents the common one-step neighbor information of drug x, (y1, y2,..., y m ) represents the common one-step neighbor information of drug y, x i represents the common one-step neighbor information between drug x and the i-th drug, y i represents the common one-step neighbor information between drug y and the i-th drug; (a1, a2,..., a n ) represents the common one-step neighbor information of disease a, (b1, b2,..., b n ) represents the common one-step neighbor information of disease b, a j represents the common one-step neighbor information between disease a and the j-th disease, b j represents the common one-step neighbor information between disease b and the j-th disease. In the embodiments of the present invention, the Fdataset dataset contains 593 drugs and 313 diseases, and the Cdataset dataset includes 663 drugs and 409 diseases.

[0099] S204: Concatenate the common one-step neighbor similarity matrix S between drug nodes d with the drug-disease treatment relationship matrix A ds and the drug-protein interaction relationship matrix A dp to obtain the comprehensive topological information matrix T of drugs d ; at the same time, concatenate the common one-step neighbor similarity matrix S between disease nodes s with the transposed matrix of the drug-disease treatment relationship matrix A ds (A ds ) T and the disease-protein interaction relationship matrix A sp to obtain the comprehensive topological information matrix T of diseases s , and the specific calculation formula is:

[0100] T d = [S d , A ds , A dp;

[0101] T s = s ,(A ds ) T ,A sp .

[0102] S3: Calculate the average values of the similarity matrices of 5 drugs and the similarity matrices of 2 diseases from multiple biochemical perspectives in step S1 respectively to obtain the integrated drug similarity matrix and disease similarity matrix, and adjust the integrated drug similarity matrix and disease similarity matrix according to the pruning algorithm to obtain the biochemical proximity matrix of drugs and the biochemical proximity matrix of diseases; the specific steps are as follows:

[0103] S301: Calculate the average values of the drug similarities and disease similarities from multiple biochemical perspectives respectively to obtain the integrated drug similarity matrix Sim d and disease similarity matrix Sim s , and the specific calculation formula is:

[0104]

[0105] where ChemSim d is the chemical structure similarity matrix, ATCSim d is the ATC code similarity matrix, SESim d is the side effect similarity matrix, DrDrSim d is the drug-drug interaction similarity matrix, TargetSim d is the target spectrum similarity matrix; PhenomeSim s is the disease phenotype similarity matrix, MeshSim s is the disease ontology similarity matrix.

[0106] S302: Adjust the integrated drug similarity matrix and disease similarity matrix according to the pruning algorithm, that is, adjust the top k similarity values in descending order of the similarity values in each row of the drug similarity matrix and disease similarity matrix, and set the remaining values to zero, so as to obtain the biochemical proximity matrix B d of drugs and the biochemical proximity matrix B s of diseases, and the specific calculation formula is:

[0107] B d = pruningSim(Sim d ,k,η);

[0108] B s = pruningSim(Sim s ,k,η);

[0109] Among them, pruningSim(·) represents the pruning operation, k represents the number of adjustments, and η represents the decay factor. In the embodiment of the present invention, the value of k is 35, and the value of the decay factor η is 0.65.

[0110] The steps of the pruning algorithm are as follows:

[0111] Initialize the biochemical affinity matrix B of the drug d as a 0 matrix with m rows and m columns, and initialize the biochemical affinity matrix B of the disease s as a 0 matrix with n rows and n columns; store the column subscripts corresponding to the top k numerical values in descending order of the similarity values in each row of the drug similarity matrix Sim d in each row of the matrix indexes1, and store the column subscripts corresponding to the top k numerical values in descending order of the similarity values in each row of the disease similarity matrix Sim s in each row of the matrix indexes2; according to the column subscripts recorded in each row and each column of the indexes1 matrix, decay the element values at the corresponding positions in the drug similarity matrix Sim d and fill the decayed numerical values into the corresponding positions of the biochemical affinity matrix B of the drug d to obtain the biochemical affinity matrix B of the drug d ; according to the column subscripts recorded in each row and each column of the indexes2 matrix, decay the element values at the corresponding positions in the disease similarity matrix Sim s and fill the decayed numerical values into the corresponding positions of the biochemical affinity matrix B of the disease s to obtain the biochemical affinity matrix B of the disease s , and the specific calculation formula is:

[0112] B d (M, indexes1(M, K)) = η K-1 × Sim d (M, indexes1(M, K));

[0113] B s (N, indexes2(N, K)) = η K-1 × Sim s (N, indexes2(N, K));

[0114] Among them, M = 1, 2,..., m; N = 1, 2,..., n; K = 1, 2,..., k.

[0115] S4: Extract and integrate the topological and biochemical information of drugs and diseases, remove the noise in the data, and obtain the low-dimensional feature representations of drugs and diseases, resulting in a drug feature matrix and a disease feature matrix. Since any matrix F can be approximately decomposed into the product of three matrices U, Σ, and V through singular value decomposition, i.e.:

[0116] F≈UΣV T ;

[0117] where U and V are orthogonal matrices, and Σ is a non-negative diagonal matrix. Therefore, the singular value decomposition can be performed on the drug comprehensive topological information matrix T d , the drug biochemical proximity matrix B d , the disease comprehensive topological information matrix T s , and the disease biochemical proximity matrix B d respectively to obtain the drug feature matrix F d and the disease feature matrix F s . The specific calculation formula is:

[0118] F d =[U Td ,U Bd ;

[0119] F s =[U Ts ,U Bs ;

[0120] where U Td is the feature matrix obtained after the singular value decomposition of the drug comprehensive topological information matrix T d , U Bd is the feature matrix obtained after the singular value decomposition of the drug biochemical proximity matrix B d , U Ts is the feature matrix obtained after the singular value decomposition of the disease comprehensive topological information matrix T s , and U Bs is the feature matrix obtained after the singular value decomposition of the disease biochemical proximity matrix B s .

[0121] S5: In the heterogeneous network G=(V, E), it includes "targeted / being targeted" links between drugs and proteins, "triggered / being triggered" links between proteins and diseases, and "treated / being treated" links between drugs and diseases. Among them, V is the set of all drug, protein, and disease nodes in the heterogeneous network, and E is the set of all heterogeneous links in the heterogeneous network. Expand the "protein" in the meta-path in the heterogeneous network to "protein→drug→protein" or "protein→disease→protein" to screen out more reasonable negative samples, where the negative samples are drug-disease pairs without a treatment relationship. The specific method is:

[0122] S501: Denote the "drug→disease" path as meta-path M1, the "drug→protein→disease" path as meta-path M2, the "drug→protein→drug→disease" path as meta-path M31, the "drug→disease→drug→disease" path as meta-path M32, and the "drug→disease→protein→disease" path as meta-path M33.

[0123] S502: Expand the "protein" in meta-path M2 to "protein→drug→protein" to obtain meta-path M2P1; expand the "protein" in meta-path M2 to "protein→disease→protein" to obtain meta-path M2P2; expand the "protein" in meta-path M31 to "protein→drug→protein" to obtain meta-path M31P1; expand the "protein" in meta-path M31 to "protein→disease→protein" to obtain meta-path M31P2; expand the "protein" in meta-path M33 to "protein→drug→protein" to obtain meta-path M33P1; expand the "protein" in meta-path M33 to "protein→disease→protein" to obtain meta-path M33P2.

[0124] S503: The meta-path matrix M of the heterogeneous network is jointly composed of meta-path M1, meta-path M2, meta-path M31, meta-path M32, meta-path M33, meta-path M2P1, meta-path M2P2, meta-path M31P1, meta-path M31P2, meta-path M33P1, and meta-path M33P2. The specific calculation formula is:

[0125] M = M1 + M2 + M31 + M32 + M33 + M2P1 + M2P2 + M31P1

[0126] + M31P2 + M33P1 + M33P2;

[0127] M1 = A ds ;

[0128] M2 = A dp ×(A sp ) T ;

[0129] M31 = A dp ×(A dp ) T ×A ds ;

[0130] M32 = A ds ×(A ds ) T ×A ds ;

[0131] M33 = A ds ×A sp×(A sp ) T ;

[0132] M2P1 = A dp ×A pdp ×(A sp ) T ;

[0133] M2P2 = A dp ×A pSp ×(A sp ) T ;

[0134] M31P1 = A dp ×A pdp ×(A dp ) T ×A ds ;

[0135] M31P2 = A dp ×A psp ×(A dp ) T ×A ds ;

[0136] M33P1 = A ds ×A sp ×A pdp ×(A sp ) T ;

[0137] M33P2 = A ds ×A sp ×A psp ×(A sp ) T ;

[0138] A pdp =(A dp ) T ×A dp ;

[0139] A psp =(A sp ) T ×A sp ;

[0140] Among them, A pdp represents the "protein → drug → protein" expansion mode, and A psp represents the "protein → disease → protein" expansion mode.

[0141] S504: Determine whether each element value in the meta-path matrix M is 0, and record the drug-disease combination mapped by the subscript combination corresponding to the element with the element value of 0 as a negative sample.

[0142] Since the meta-path matrix M comprehensively considers the situations of 11 meta-paths starting from drugs and ending at diseases, if the value of a certain element in M is still equal to 0, the possibility of a treatment relationship between the drug and the disease mapped by its row and column subscripts is relatively low. This drug-disease combination can be regarded as a more reliable negative sample. Compared with the method of randomly selecting an unknown drug-disease combination as a negative sample in existing similar research methods, the negative samples screened in the embodiments of the present invention are more reliable.

[0143] S6: Use the integrated learning classifier XGBoost to predict the drug feature matrix and the disease feature matrix, and output the probability of a treatment relationship between the drug-disease pair as the prediction result. The specific steps are as follows: Regard the drug-disease association prediction as a binary classification problem. Among them, the positive sample is the drug-disease pair with a treatment relationship, and the negative sample is the drug-disease pair without a treatment relationship. After concatenating the drug features and disease features obtained in S4, the feature vector of the drug-disease pair is used as the input. An initial decision tree is created through the integrated learning classifier XGBoost, and then new regression trees are iteratively constructed to reduce the calculation residuals. The scores corresponding to the leaf nodes of each feature vector on each tree are added together as the prediction result corresponding to the feature vector, that is, the probability of a treatment relationship between the drug-disease pair. In the embodiments of the present invention, in order to prevent overfitting in the training stage, in the embodiments of the present invention, the number of decision trees is set to 500, and the maximum depth of the tree is set to 6.

[0144] In terms of experiments, the embodiments of the present invention compare the prediction performance of the method described in the embodiments of the present invention (denoted as TBDR in Tables 1-3) with other classical recommendation system methods in the datasets Fdataset and Cdataset based on the 5-fold cross-validation method. The negative samples in the embodiments of the present invention are randomly selected with the same number as the positive samples after being screened by the method described in the embodiments of the present invention. To ensure fairness, when performing 5-fold cross-validation, the training set and test set used by each method each time are the same. The average value of the area under the ROC curve AUC obtained 5 times is used as the final AUC result score, and the average value of the area under the PR curve AUPR is used as the final AUPR result score. The experimental results shown in Table 1 can be obtained:

[0145] Table 1 AUC results and AUPR results of 5-fold cross-validation experiments of various methods on different datasets

[0146]

[0147] As can be seen from Table 1, since the method described in the embodiments of the present invention comprehensively considers topological information and biochemical information when extracting drug and disease information, makes more reasonable use of the heterogeneous network, and optimizes the biochemical information to a certain extent, so as to obtain a more comprehensive feature information representation. Therefore, the performance of the method described in the embodiments of the present invention is better than other methods. The AUC result score and AUPR result score of the method described in the embodiments of the present invention in the Fdataset dataset are 0.952 and 0.965 respectively, and the AUC result score and AUPR result score obtained in the Cdataset dataset are 0.963 and 0.973 respectively, both of which are better than other methods.

[0148] In addition, the method described in the embodiments of the present invention was also used to predict candidate drugs for treating breast cancer and small cell lung cancer during the experiment. When predicting candidate drugs for diseases, for the Cdataset dataset, all known drug-disease association datasets can be used as the training set, and for the remaining candidate sets, after the method described in the embodiments of the present invention calculates the prediction probabilities of the treatment relationships of all drug-disease pairs in the candidate set, they are sorted, and the prediction results shown in Tables 2 and 3 can be obtained. Among the top 10 drugs predicted by the method described in the embodiments of the present invention to potentially treat breast cancer, 8 drugs have been verified by various evidences from authoritative sources and clinical trials; among the top 10 drugs predicted to potentially treat small cell lung cancer, 7 drugs have been verified by various evidences from authoritative sources and clinical trials.

[0149] Table 2 Top ten drugs predicted to potentially treat breast cancer

[0150] Serial number DrugBank IDs Predicted candidate drugs Literature verification 1 DB00255 Diethylstilbestrol Not yet verified 2 DB00635 Prednisone √ 3 DB01204 Mitoxantrone √ 4 DB00860 Dehydrocortisol √ 5 DB00755 Tretinoin √ 6 DB01196 Estramustine √ 7 DB00305 Mitomycin √ 8 DB00773 Etoposide √ 9 DB00531 Cyclophosphamide √ 10 DB00687 Fludrocortisone Not yet verified

[0151] Table 3 Top ten drugs predicted to potentially treat small cell lung cancer

[0152]

[0153]

[0154] In summary, in the embodiments of the present invention, when extracting drug and disease information, topological information and biochemical information are comprehensively considered, the utilization of the heterogeneous network is more reasonable, the extraction of heterogeneous network information is more sufficient and effective, biochemical information from multiple perspectives is used and optimized, so as to obtain a more comprehensive feature information representation and improve the performance of the repositioning method described in the present invention; compared with the method of randomly selecting negative samples from unknown samples adopted in the prior art, in the embodiments of the present invention, the method of expanding the "protein" in the metapath to "protein → drug → protein" or "protein → disease → protein" is adopted, which more comprehensively considers the "targeted / being targeted" link relationship between drugs and proteins and the "causing / being caused" link relationship between proteins and diseases in the heterogeneous network. Finally, considering the situations of 11 metapaths starting from drugs and ending at diseases, more reasonable negative samples can be screened out, thereby further improving the accuracy of the prediction results.

[0155] The foregoing are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A drug repositioning method based on topological information and biochemical information, characterized in that, It includes the following steps: S1: Obtain a dataset to construct a heterogeneous network; the dataset includes a drug-disease treatment relationship dataset, a drug-protein interaction relationship dataset, and a disease-protein interaction relationship dataset; Obtain 5 similarity matrices of drugs and 2 similarity matrices of diseases according to the dataset. The 5 similarity matrices of drugs include a chemical structure similarity matrix, an ATC code similarity matrix, a side effect similarity matrix, a drug-drug interaction similarity matrix, and a target spectrum similarity matrix. The 2 similarity matrices of diseases include a disease phenotype similarity matrix and a disease ontology similarity matrix; S2: Use the number of one-step neighbors of drugs or diseases as the measurement basis for the topological information of the corresponding nodes, calculate the number of common one-step heterogeneous neighbors between drugs and other drugs and between diseases and other diseases, construct a drug-drug common one-step neighbor matrix and a disease-disease common one-step neighbor matrix, and use cosine similarity to calculate the similarity between two nodes, construct a common one-step neighbor similarity matrix between drug nodes and a common one-step neighbor similarity matrix between disease nodes, and obtain a comprehensive topological information matrix of drugs and a comprehensive topological information matrix of diseases; S3: Calculate the average values of the 5 similarity matrices of drugs and the 2 similarity matrices of diseases from multiple biochemical perspectives in step S1 respectively to obtain an integrated drug similarity matrix and an integrated disease similarity matrix, and adjust the integrated drug similarity matrix and the integrated disease similarity matrix according to the pruning algorithm to obtain a biochemical affinity matrix of drugs and a biochemical affinity matrix of diseases; S4: Extract and integrate the topological information and biochemical information of drugs and diseases, remove the noise of the data, and obtain low-dimensional feature representations of drugs and diseases to obtain a drug feature matrix and a disease feature matrix; The specific steps are as follows: separately perform singular value decomposition on the comprehensive topological information matrix T of the drug d , the biochemical affinity matrix B of the drug d , the comprehensive topological information matrix T of the disease s , and the biochemical affinity matrix B of the disease d , and integrate the matrices obtained after decomposition to obtain the drug feature matrix F d and the disease feature matrix F s . The specific calculation formula is as follows: F d = [U Td , U Bd ; F s = [U Ts , U Bs ; Among them, U Td is the comprehensive topological information matrix T d of the drug after singular value decomposition, and U Bd is the biochemical affinity matrix B d of the drug after singular value decomposition, and U Ts is the comprehensive topological information matrix T s of the disease after singular value decomposition, and U Bs is the biochemical affinity matrix B s of the disease after singular value decomposition; S5: Expand the "protein" in the meta-path in the heterogeneous network to "protein → drug → protein" or "protein → disease → protein", and screen out more reasonable negative samples. The negative samples are drug-disease pairs without a treatment relationship; S6: After concatenating the drug features and disease features obtained in S4, use the feature vector of the drug-disease pair as the input, and use the integrated learning classifier XGBoost for prediction, and output the probability that the drug-disease pair has a treatment relationship as the prediction result.

2. The drug repositioning method based on topological information and biochemical information according to claim 1, wherein The similarity value range in the 5 similarity matrices of drugs and the 2 similarity matrices of diseases is [0,1]. The higher the similarity, the more similar the two drugs or diseases are.

3. A drug repositioning method based on topological information and biochemical information according to claim 1, characterized in that, The specific steps of step S2 are: S201: Write the drug-disease treatment relationship dataset, drug-protein interaction relationship dataset, and disease-protein interaction relationship dataset in matrix form, denoted as drug-disease treatment relationship matrix A ds , drug-protein interaction relationship matrix A dp , and disease-protein interaction relationship matrix A sp ; S202: According to the drug-disease treatment relationship matrix A ds , the drug-protein interaction relationship matrix A dp and the disease-protein interaction relationship matrix A sp Calculate the drug-drug common one-step neighbor matrix N dd and the disease-disease common one-step neighbor matrix N ss , and the specific calculation formula is: N dd = A dp × (A dp ) T + A ds × (A ds ) T ; N ss = A sp ×(A sp ) T +(A ds ) T ×A ds ; Among them, (A dp ) T represents the transpose matrix of the drug-protein interaction relationship matrix A dp , (A ds ) T represents the transpose matrix of the drug-disease treatment relationship matrix A ds , (A sp ) T represents the transpose matrix of the disease-protein interaction relationship matrix A sp ; Using the data in the i-th row of the drug-drug common one-step neighbor matrix N dd as the common one-step neighbor information between the i-th drug node and other drug nodes, and using the data in the j-th row of the disease-disease common one-step neighbor matrix N ss as the common one-step neighbor information between the j-th disease node and other disease nodes; S203: Calculate the common one-step neighbor similarity matrix S between drug nodes using cosine similarity d and the common one-step neighbor similarity matrix S between disease nodes s , and the specific calculation formula is as follows: Among them, m represents the total number of drug nodes, and n represents the total number of disease nodes; (x1, x2, …, x m ) represents the common one-step neighbor information of drug x, (y1, y2, …, y m ) represents the common one-step neighbor information of drug y, x i represents the common one-step neighbor information between drug x and the i-th drug, y i represents the common one-step neighbor information between drug y and the i-th drug; (a1, a2, …, a n ) represents the common one-step neighbor information of disease a, (b1, b2, …, b n ) represents the common one-step neighbor information of disease b, a j represents the common one-step neighbor information between disease a and the j-th disease, b j represents the common one-step neighbor information between disease b and the j-th disease; S204: Concatenate the common one-step neighbor similarity matrix S among drug nodes d with the drug-disease treatment relationship matrix A ds and the drug-protein interaction relationship matrix A dp to obtain the comprehensive topological information matrix T of drugs d ; meanwhile, concatenate the common one-step neighbor similarity matrix S among disease nodes s with the transposed matrix of the drug-disease treatment relationship matrix A ds (A ds ) T and the disease-protein interaction relationship matrix A sp to obtain the comprehensive topological information matrix T of diseases s . The specific calculation formula is as follows: T d = [S d , A ds , A dp ; T s = [S s , (A ds ) T , A sp .

4. A drug repositioning method based on topological information and biochemical information according to claim 3, characterized in that The specific steps of step S3 are: S301: Calculate the average values of drug similarity and disease similarity from multiple biochemical perspectives respectively to obtain the integrated drug similarity matrix Sim d and the disease similarity matrix Sim s , and the specific calculation formula is as follows: Among them, ChemSim d is the chemical structure similarity matrix, ATCSim d is the ATC code similarity matrix, SESim d is the side effect similarity matrix, DrDrSim d is the drug-drug interaction similarity matrix, TargetSim d is the target spectrum similarity matrix; PhenomeSim s is the disease phenotype similarity matrix, MeshSim s is the disease ontology similarity matrix; S302: Adjust the integrated drug similarity matrix and disease similarity matrix according to the pruning algorithm, that is, adjust the top k similarity values in descending order of similarity values in each row of the drug similarity matrix and the disease similarity matrix, and set the remaining values to zero, so as to obtain the biochemical affinity matrix B of drugs d and the biochemical affinity matrix B of diseases s , and the specific calculation formula is: B d = pruningSim(Sim d , k, η); B s = pruningSim(Sim s , k, η); Among them, pruningSim(·) represents the pruning operation, k represents the number of adjustments, and η represents the decay factor; The steps of the pruning algorithm are: Initialize the biochemical affinity matrix B of drugs d as a 0 matrix with m rows and m columns, and initialize the biochemical affinity matrix B of diseases s as a 0 matrix with n rows and n columns; store the column subscripts corresponding to the first k numerically descending similarity values in each row of the drug similarity matrix Sim d in each row of the matrix indexes1, and store the column subscripts corresponding to the first k numerically descending similarity values in each row of the disease similarity matrix Sim s in each row of the matrix indexes2; according to the column subscripts recorded in each row and column of the indexes1 matrix, attenuate the element values at the corresponding positions in the drug similarity matrix Sim d and fill the attenuated values into the corresponding positions of the biochemical affinity matrix B of drugs d to obtain the biochemical affinity matrix B of drugs d ; according to the column subscripts recorded in each row and column of the indexes2 matrix, attenuate the element values at the corresponding positions in the disease similarity matrix Sim s and fill the attenuated values into the corresponding positions of the biochemical affinity matrix B of diseases s to obtain the biochemical affinity matrix B of diseases s , and the specific calculation formula is: B d (M, indexes1(M, K)) = η K-1 × Sim d (M, indexes1(M, K)); B s (N, indexes2(N, K)) = η K-1 × Sim s (N, indexes2(N, K)); Among them, M = 1, 2, …, m; N = 1, 2, …, n; K = 1, 2, …, k.

5. The drug repositioning method based on topological information and biochemical information according to claim 4, wherein The specific steps of step S5 are: S501: Denote the "drug→disease" path as meta-path M1, the "drug→protein→disease" path as meta-path M2, the "drug→protein→drug→disease" path as meta-path M31, the "drug→disease→drug→disease" path as meta-path M32, and the "drug→disease→protein→disease" path as meta-path M33; S502: Expand the "protein" in meta-path M2 to "protein→drug→protein" to obtain meta-path M2P1; expand the "protein" in meta-path M2 to "protein→disease→protein" to obtain meta-path M2P2; expand the "protein" in meta-path M31 to "protein→drug→protein" to obtain meta-path M31P1; expand the "protein" in meta-path M31 to "protein→disease→protein" to obtain meta-path M31P2; expand the "protein" in meta-path M33 to "protein→drug→protein" to obtain meta-path M33P1; expand the "protein" in meta-path M33 to "protein→disease→protein" to obtain meta-path M33P2; S503: The meta-path M1, meta-path M2, meta-path M31, meta-path M32, meta-path M33, meta-path M2P1, meta-path M2P2, meta-path M31P1, meta-path M31P2, meta-path M33P1, and meta-path M33P2 together constitute the meta-path matrix M of the heterogeneous network. The specific calculation formula is: M = M1 + M2 + M31 + M32 + M33 + M2P1 + M2P2 + M31P1 + M31P2 + M33P1 + M33P2; M1 = A ds ; M2 = A dp ×(A sp ) T ; M31 = A dp × (A dp ) T × A ds ; M32 = A ds × (A ds ) T × A ds ; M33 = A ds × A sp × (A sp ) T ; M2P1 = A dp × A pdp × (A sp ) T ; M2P2 = A dp × A pSp × (A sp ) T ; M31P1 = A dp × A pdp × (A dp ) T × A ds ; M31P2 = A dp × A psp × (A dp ) T × A ds ; M33P1 = A ds × A sp × A pdp × (A sp ) T ; M33P2 = A ds × A sp × A psp × (A sp ) T ; A pdp =(A dp ) T ×A dp ; A psp =(A sp ) T ×A sp ; Among them, A pdp represents the "protein → drug → protein" expansion mode, and A psp represents the "protein → disease → protein" expansion mode; S504: Determine whether each element value in the meta-path matrix M is 0, and denote the drug-disease combination mapped by the subscript combination corresponding to the element with the element value of 0 as the negative sample.

6. The drug repositioning method based on topological information and biochemical information according to claim 5, wherein The specific steps of step S6 are as follows: Consider the drug-disease association prediction as a binary classification problem, where the positive sample is the drug-disease pair with a treatment relationship, and the negative sample is the drug-disease pair without a treatment relationship; After concatenating the drug features and disease features obtained in S4, use the feature vector of the drug-disease pair as the input, create an initial decision tree through the ensemble learning classifier XGBoost, and then iteratively construct new regression trees to reduce the calculation residuals. Add the scores corresponding to the leaf nodes of each feature vector on each tree as the prediction result corresponding to the feature vector, that is, the probability that the drug-disease pair has a treatment relationship.

Citation Information

Patent Citations

  • Drug target prediction method based on multiple similarity network walk

    CN108520166A

  • Drug target interaction relationship prediction method based on collaborative matrix decomposition

    CN110957002A