Drug screening method based on drug relocation mapping knowledge domain
By employing the OpenKE framework, supplementing drug information, and constructing novel triples within the DRKG knowledge graph, combined with the Ensemble drug recommendation method, the shortcomings of the DRKG knowledge graph in information updates and the defects of the TransE algorithm are addressed, thereby improving the accuracy and stability of drug recommendations.
Patent Information
- Application Number
- CN202511185927.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-12-05
AI Technical Summary
Existing DRKG knowledge graphs suffer from insufficient information updates and the TransE algorithm's inadequate ability to model complex relationships, resulting in poor predictive performance of complex drug recommendation models and making it difficult to modify the model structure or embed innovative algorithms.
We replace DGL-KE with the OpenKE representation learning framework, supplement drug-related node information, add drug structure, target gene and side effect information, construct drug structure similarity triples, and integrate multiple drug recommendation algorithms through the ensemble method to improve drug recommendation performance.
It improves the general link prediction performance and proprietary drug recommendation performance of knowledge graphs, enabling high-throughput drug screening and intelligent drug development.
Smart Images

Figure BDA0005562106110000081 
Figure BDA0005562106110000082 
Figure BDA0005562106110000083
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of drug screening, and particularly relates to a drug screening method based on a drug repositioning knowledge graph. BACKGROUND
[0002] The DRKG knowledge graph is a comprehensive biomedical knowledge graph that can realize drug repositioning. The knowledge graph integrates a series of entities such as genes, pathways, molecular functions, biological processes, cell components, diseases, symptoms, compounds and their related attributes, and the mutual relationships between the entities, to construct a connection graph with drugs and diseases as the core. The knowledge graph has been used for new crown drug research and development, and has screened out a variety of clinically proven anti-new crown virus repositioning drugs.
[0003] However, the DRKG knowledge graph still has some deficiencies. First, the knowledge graph was established in 2020, and the disease and drug related information contained therein has been lacking in real-time updates for a long time. On the other hand, the knowledge graph is modeled based on the DGL-KE representation learning framework using the TransE embedding algorithm to realize the relationship of triplets. The TransE algorithm has the defect of oversimplifying the interaction logic between entities and relationships, which leads to insufficient modeling ability of complex relationships of the algorithm, and may seriously affect the prediction performance of the drug recommendation complex model. Although DGL-KE has a great advantage in training efficiency, the price is that it is difficult to modify the model structure or embed innovative algorithms. SUMMARY
[0004] The purpose of the present application is to provide a drug screening method based on a drug repositioning knowledge graph.
[0005] A drug screening method based on a drug repositioning knowledge graph, according to the following steps:
[0006] (1) DRKG knowledge graph representation learning framework optimization and drug node information supplement;
[0007] (2) Drug structure similarity information triplet establishment;
[0008] (3) Drug target gene information triplet expansion;
[0009] (4) Drug side effect information triplet expansion;
[0010] (5) Establishment of Ensemble drug recommendation method based on multiple benchmark algorithms;
[0011] (6) Calculation of the degree of overlap of disease recommended drugs based on the "disease-drug" and "gene-drug" relationships;
[0012] (7) Knowledge graph performance evaluation based on link prediction task;
[0013] (8) Calculation of the proportion of known drugs in the specific disease drug recommendation result and the average proportion of known drugs in all disease recommended drugs.
[0014] The framework optimization includes converting the DRKG dataset into a data format supported by OpenKE.
[0015] The triple establishment includes: extracting all drug entities in the knowledge graph, obtaining the SMILES representation of the corresponding drug from DrugBank through the unique ID of the drug entity; obtaining the vector representation of the drug structure information based on the SMILES of the drug by using ECFP and PCP methods, thereby realizing the feature extraction of the drug molecular structure; and performing drug structure similarity analysis based on the PCP and ECFP vector representations of the drug structure information by using cosine similarity.
[0016] The drug target gene information triple expansion includes: obtaining high-confidence drug target information directly derived from experimental verification results from D1 and D2 databases of DSigDB, and constructing triples from the D1 and D2 databases respectively through data analysis, entry extraction, entity conversion and deduplication.
[0017] The drug side effect information triple expansion includes: adding the drug side effect triples extracted from T-ARDIS to the knowledge graph.
[0018] The Ensemble drug recommendation method establishment of step (5) includes: obtaining drug recommendation lists and drug ranking information of different benchmark algorithms; then statistically analyzing the recommendation ranking of each drug, and weighting and averaging the ranking of each drug, that is, selecting the top 3 in the ranking of the benchmark algorithm to take the average value; finally, sorting according to the comprehensive score, and selecting the TOP-100 drugs as the final recommended drug list.
[0019] The drug coincidence degree calculation of step (6) includes: using all drugs in DrugBank as candidate drugs, first performing drug recommendation based on the disease ID and the treatment relationship in the knowledge graph using TransE and other algorithms based on the "disease-drug" mode to obtain a recommended drug list; secondly, obtaining a list of drugs with known treatment relationships, and calculating the coincidence degree between the recommended result and the drugs with treatment relationships; finally, considering the inhibition and target relationship, performing drug recommendation based on the "gene-drug" mode using TransE and other algorithms and calculating the coincidence degree.
[0020] The knowledge graph performance evaluation of step (7) comprises the following steps: firstly, dividing the knowledge graph data into a training set, a validation set and a test set; then generating negative samples for each test positive sample by replacing the head or tail entity; then calculating the prediction score of all samples by using the trained embedding model; and finally evaluating the link prediction ability of the knowledge graph by hit@k, MR and MRR indexes.
[0021] The calculation of the average proportion comprises the following steps: firstly, obtaining the candidate drug recommendation result of a specific disease; secondly, obtaining a list of drugs known to have a therapeutic relationship for the specific disease according to the therapeutic relationship in the DRKG data set; then calculating the proportion of the candidate drug in the known drugs; and finally, performing candidate drug recommendation for all diseases in the DRKG data set, obtaining the proportion information of each disease, and then calculating the average value to obtain the average proportion of known drugs in all recommended drugs.
[0022] The beneficial effects of the present application are as follows: the present application first uses knowledge graph technology to realize high-throughput drug screening based on intelligent drug repositioning. The main strategies for developing drugs based on intelligent drugs are as follows: one is to develop new drugs according to known targets. Two is to establish a high-throughput drug screening strategy for drug screening. Three is to use artificial intelligence technology or bioinformatics analysis based on disease-related data sets for drug efficacy prediction. We attempt to replace the original DGL-KE with the open-source and flexible editable OpenKE representation learning framework; based on this, we improve the original triple of the DRKG drug-related node information based on the DrugBank database; and by adding new triple information, we add drug target gene information and drug side effect information to the existing knowledge graph. In addition, based on the logic behind the "structure determines function", we also innovatively integrate drug structure information into the drug recommendation process by incorporating structural similarity analysis. The optimization of the learning framework and the supplement of drug-related information both effectively improve the general link prediction performance and the specific drug recommendation performance of the DRKG knowledge graph. In view of the design defects of the TransE embedding algorithm, we realize the original "disease-drug" and "gene-drug" two drug recommendation methods based on a variety of embedding algorithms, and on this basis, integrate all implementation forms of the above two recommendation methods, and establish a third Ensemble recommendation method using ensemble learning strategy. The ensemble algorithm effectively realizes the complementary advantages of the two drug recommendation methods and a variety of embedding algorithms, further effectively improving the drug recommendation performance of the DRKG knowledge graph. In summary, we effectively realize the performance improvement of the DRKG knowledge graph through knowledge graph expansion and algorithm optimization. DETAILED DESCRIPTION
[0023] To facilitate understanding of the present invention, a more comprehensive description will be given below. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure of the present invention.
[0024] Example 1
[0025] I. Experimental Methods
[0026] 1. Optimization of the DRKG Knowledge Graph Representation Learning Framework: Download and configure the OpenKE-PyTorch version from the OpenKE website (http: / / openke.thunlp.org / ) according to the instructions. Write an automated Python script to convert the DRKG dataset to an OpenKE-supported format, following the standard three-file structure of OpenKE (train2id.txt, entity2id.txt, and relation2id.txt). Run and test the knowledge graph representation learning algorithm according to the steps in the "How to Train" and "How to Test" links on the OpenKE website documentation page (http: / / 139.129.163.161 / static / index.html).
[0027] 2. Supplementing Drug-Related Node Information: Apply for and download the relevant drug database from the DrugBank website. Use a Python program to convert the original XML file into a JSON file, extract and match fields, and extract the required drug information (such as name, average-mass, physical state, etc.) and store it as a CSV file. The specific information is shown in Table 4.
[0028] 3. New triplets of drug structure similarity information: All drug entities in the knowledge graph were extracted, and the SMILES representation of the corresponding drug was obtained from DrugBank through the unique ID of the drug entity. Two methods, ECFP (Extended Connectivity Fingerprints) and PCP (Physicochemical Properties), were used to obtain vector representations of drug structure information based on the SMILES of the drug, thereby achieving feature extraction of drug molecular structure. Cosine Similarity was used to analyze drug structure similarity based on the PCP and ECFP vector representations of drug structure information. Through the two vector representation methods of ECFP and PCP, 21075 (DRUGBANK::ECFP::Compound:Compound) and 38356 (DRUGBANK::PCP::Compound:Compound) structure similarity triplets were constructed and included in the knowledge graph (Table 5).
[0029] 4. Expansion of drug target gene information triplets: High-confidence drug target information directly derived from experimental verification results was obtained from the D1 and D2 databases of DSigDB (http: / / DSigDB.tanlab.org / DSigDBv1.0 / ). Through data analysis, entry extraction, entity conversion, and deduplication, 10245 and 2661 triplets were finally constructed from the D1 and D2 databases, respectively. New triplets were established through the DSig::INHIBIT:Compound:Gene and DSig::INHIBITOR:Compound:Gene relationships (Table 5), and were included in the knowledge graph.
[0030] 5. Expansion of drug side effect information triplets: T-ARDIS (http: / / www.bioinsilico.org / T-ARDIS / ) extracted 39640 drug side effect triplets and added them to the knowledge graph. New triplets were established through the Tard::DrugSe:Compound:Se relationship and were included in the knowledge graph (Table 5).
[0031] 6. Ensemble drug recommendation method based on multiple benchmark algorithms: First, obtain the drug recommendation list and drug ranking information of different benchmark algorithms (a total of 8, including TransE, TransD, TransH and RotateE based on "disease-drug" recommendation method, and the same 4 benchmark algorithms based on "gene-drug" recommendation method); Then, the recommendation ranking of each drug is counted, and the ranking of each drug is weighted and averaged, that is, the top 3 rankings of the 8 benchmark algorithms are selected to take the average value; Finally, according to the comprehensive score, the TOP-100 drugs are selected as the final recommended drug list.
[0032] 7. Calculation of PTSD disease recommended drug coincidence degree based on "disease-drug" and "gene-drug" relationships: Use all drugs (24313) in DrugBank as candidate drugs, first based on PTSD disease ID and knowledge graph treatment relationship (such as 'Hetionet::CtD::Compound:Disease', 'GNBR::T::Compound:Disease'), use TransE algorithm based on "disease-drug" method to recommend drugs, and obtain the recommended drug list. Secondly, obtain the drug list known to have a treatment relationship for PTSD disease, and calculate the coincidence degree between the recommended results and the drugs known to have a treatment relationship for PTSD. Finally, considering the inhibition and target relationship (such as 'GNBR::N::Compound:Gene' and 'DRUGBANK::target::Compound:Gene'), use TransE algorithm based on "gene-drug" method to recommend drugs and calculate the coincidence degree.
[0033] 8. Knowledge graph performance evaluation based on link prediction task: First, divide the knowledge graph data into training set, validation set and test set; Then generate negative samples for each test positive sample by replacing the head or tail entity; Then use the trained embedding model (such as TransE) to calculate the prediction score of all samples; Finally, evaluate the link prediction ability of the knowledge graph through hit@k, MR (Mean Rank) and MRR (Mean Reciprocal Ranking) and other indicators.
[0034] 9. Calculation of the proportion of known drugs in the specific disease drug recommendation results and the average proportion of known drugs in all disease recommended drugs: First, the candidate drug recommendation results for a specific disease are obtained, and then the list of known drugs with therapeutic relationships for the specific disease is obtained according to the therapeutic relationships in the DRKG dataset. Then, the proportion of candidate drugs in known drugs is calculated. Finally, candidate drug recommendations are made for all diseases in the DRKG dataset. After obtaining the proportion information for each disease, the average value is calculated to obtain the average proportion of known drugs in all disease recommended drugs.
[0035] II. Experimental results
[0036] Optimization of DRKG knowledge graph representation learning framework: DRKG knowledge graph is a comprehensive biomedical knowledge graph that can realize drug repositioning. Based on the DGL-KE representation learning framework, the TransE embedding algorithm is used to model the relationship of triples, supporting two drug screening modes of "gene-drug" and "disease-drug". Although DGL-KE has great advantages in training efficiency, it is difficult to modify the model structure or embed innovative algorithms, and the performance of the embedded algorithm needs to be further tested. To optimize the performance of the DRKG knowledge graph, we try to replace the original DGL-KE with the open-source and flexible OpenKE representation learning framework, laying the foundation for subsequent optimization. To achieve comparability, we use the TransE embedding algorithm to evaluate the prediction accuracy of different types of relationships or triples through the link prediction task (see Appendix for specific evaluation indicators). The results show that the DGL-KE framework is superior to the OpenKE framework in the comprehensive evaluation of prediction accuracy for different types of relationships or triples (Table 1). However, when performing drug repositioning, our focus is on the accuracy of the recommendation results of the "treatment" relationship under different frameworks. Therefore, we compare the proportion of known drugs in the recommendation results of the PTSD disease in the knowledge graph to evaluate the drug repositioning recommendation effect of OpenKE and DGL-KE respectively. The results show that the OpenKE framework is significantly better than the DGL-KE framework in the accuracy of drug recommendation based on two types of relationships (Table 2). We further compared the overlap of PTSD disease recommended drugs obtained by DGL-KE and OpenKE frameworks based on two types of relationships. The results show that the overlap of known drugs recommended by the two recommendation methods is higher after the OpenKE framework is recommended (Table 3). This result further suggests that the OpenKE framework has higher stability and reliability in predicting drug repositioning related relationships than the DGL-KE framework. In summary, the OpenKE framework is more suitable for PTSD drug repositioning related relationship prediction and recommendation than the DGL-KE framework.
[0037] Table 1D Comparison table of evaluation indexes of DGL-KE and OpenKE algorithm
[0038]
[0039] Table 2 Known drug proportion in PTSD drug recommendation results
[0040]
[0041] Table 3 Comparison of drug recommendation coincidence degree based on "disease-drug" relationship and "gene-drug" relationship
[0042]
[0043] DRKG knowledge graph drug repositioning information expansion: further expand and optimize the DRKG knowledge graph by supplementing drug repositioning related information. First, supplement drug-related node information based on the original triplets. We obtained the following drug-related information from the DrugBank website (https: / / www.drugbank.com / ) (Table 4). Further, by adding new triplet information, drug structure information, drug target gene information, and drug side effect information are added to the existing knowledge graph. Drug structure information is a key information for drug activity prediction, but it is not included in the DRKG database. We calculated the structural similarity between drugs based on the SMILES representation of drug structure in DrugBank, thereby constructing new triplet information of drug structure similarity (Table 5). In addition, more drug target gene information and drug side effect information were obtained from DSigDB (http: / / DSigDB.tanlab.org / DSigDBv1.0 / ) and T-ARDIS (http: / / www.bioinsilico.org / T-ARDIS / ) respectively, and new triplets were constructed to include this information (Table 5). We compared the performance of the knowledge graph before and after expansion based on the TransE algorithm under the OpenKE framework. The results of the link prediction task showed that the link prediction performance improved significantly after the expansion of the triplets (Table 6). Drug recommendations were made for all diseases in the knowledge graph, and the average proportion of known drugs in all disease recommended drugs was calculated. The results showed that the new triplet information effectively improved the accuracy and relevance of drug recommendation results (Table 7). Further, PTSD and Anxiety diseases were used as representative brain diseases for drug recommendation. The results of the analysis showed that for PTSD and Anxiety diseases, the proportion of known drugs in the drug recommendation results based on “disease-drug” and “gene-drug” relationships increased significantly (Tables 8 and 9), proving that the new triplet information effectively improved the drug recommendation ability of the knowledge graph. The above results show that the expansion of DRKG knowledge graph drug repositioning information further improves the drug repositioning performance of the knowledge graph.
[0044] Table 4 Drug information supplement table
[0045]
[0046]
[0047] Table 5 Extended triplets table
[0048]
[0049] Table 6 TransE evaluation indicators before and after expansion
[0050]
[0051] Table 7 Comparison of the average proportion of known drugs in the drug recommendation results of all diseases before and after expansion
[0052]
[0053] Table 8 Comparison of the proportion of known drugs in the drug recommendation results of PTSD before and after expansion
[0054]
[0055] Table 9 Comparison of the proportion of known drugs in the drug recommendation results of Anxiety before and after expansion
[0056]
[0057] Ensemble drug recommendation method based on multiple benchmark algorithms: Finally, we integrated the existing "disease-drug" and "gene-drug" two drug recommendation methods, and established a weighted recommendation method Ensemble based on multiple benchmark algorithms (TransE, TransD, TransH and RotatE) using ensemble learning strategy. The candidate drug set was set to the 8104 drugs approved by FDA, and the average proportion of known drugs in the recommended drugs of PTSD was calculated to compare the recommendation performance of the weighted ensemble recommendation method and the existing "disease-drug" and "gene-drug" two methods. Table 10 shows that the Ensemble algorithm achieves the maximum value on the Top-10 / 50 / 100 indicators, and the second largest value on the remaining one indicator. In summary, based on the use of the OpenKE framework, we expanded the DRKG knowledge graph information and established an Ensemble drug recommendation method using ensemble learning strategy, which effectively improved the recommendation performance of the knowledge graph drug repositioning.
[0058] Table 10 Comparison of the proportion of known drugs in the drug recommendation results of PTSD before and after expansion
[0059]
[0060] The above-described embodiments only express several embodiments of the present application, and the description is more specific and detailed, but it should not be understood as limiting the scope of the patent. It should be noted that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.
Claims
1. A drug screening method based on a drug repositioning knowledge graph, characterized in that, The following steps are taken: (1) DRKG knowledge graph representation learning framework optimization and drug node information supplement; (2) Drug structure similarity information triplets are established; (3) Drug target gene information triplets are expanded; (4) Drug side effect information triplets are expanded; (5) Ensemble drug recommendation method based on multiple benchmark algorithms is established; (6) The calculation of the overlap of recommended drugs based on the "disease-drug" and "gene-drug" relationships; (7) Knowledge graph performance evaluation based on link prediction task; (8) The calculation of the proportion of known drugs in the drug recommendation results for a specific disease and the average proportion of known drugs in the recommended drugs for all diseases. 2.The method of claim 1, wherein, The framework optimization includes converting the DRKG dataset into a data format supported by OpenKE. 3.The method of claim 1, wherein, The triplet establishment includes: extracting all drug entities in the knowledge graph, obtaining the SMILES representation of the corresponding drug from DrugBank through the unique ID of the drug entity; using ECFP and PCP methods to obtain vector representations of drug structure information based on the SMILES of the drug, thereby realizing feature extraction of drug molecular structure; using cosine similarity to analyze drug structure similarity based on PCP and ECFP vector representations of drug structure information. 4.The method of claim 1, wherein, The drug target gene information triplet expansion includes: obtaining high-confidence drug target information directly from experimental verification results from D1 and D2 databases of DSigDB, constructing triplets from D1 and D2 databases through data analysis, entry extraction, entity conversion, and deduplication. 5.The method of claim 1, wherein, The drug side effect information triplet expansion includes: adding drug side effect triplets extracted from T-ARDIS to the knowledge graph. 6.The method of claim 1, wherein, The establishment of the Ensemble drug recommendation method in step (5) includes: obtaining drug recommendation lists and drug ranking information of different benchmark algorithms; then, the recommendation ranking of each drug is counted, and the ranking of each drug is weighted and averaged, that is, the top 3 rankings in the benchmark algorithm are selected and averaged; finally, the drugs are sorted according to the comprehensive score, and the top 100 drugs are selected as the final recommended drug list.
7. The drug screening method based on drug relocation knowledge graph according to claim 1, characterized in that, The calculation of the drug overlap in step (6) includes: using all drugs in DrugBank as candidate drugs, first using TransE and other algorithms to recommend drugs based on the "disease-drug" relationship based on disease ID and treatment relationship in the knowledge graph, obtaining the recommended drug list; secondly, obtaining the list of drugs known to have a treatment relationship for the disease, and calculating the overlap between the recommended results and the drugs known to have a treatment relationship; finally, considering the inhibition and target relationship, using TransE and other algorithms to recommend drugs based on the "gene-drug" relationship and calculating the overlap. 8.The method of claim 1, wherein, The knowledge graph performance evaluation of step (7) comprises the following steps: firstly, dividing the knowledge graph data into a training set, a validation set and a test set; then generating negative samples for each test positive sample by replacing the head or tail entity; then calculating the prediction score of all samples by using the trained embedding model; finally, evaluating the link prediction ability of the knowledge graph by hit@k, MR and MRR indexes. 9.The method of claim 1, wherein, The calculation of the average proportion comprises the following steps: firstly, obtaining the candidate drug recommendation result of a specific disease; secondly, obtaining the list of drugs known to have a therapeutic relationship with the specific disease according to the therapeutic relationship in the DRKG dataset; then calculating the proportion of the candidate drug in the known drugs; finally, performing candidate drug recommendation for all diseases in the DRKG dataset, obtaining the proportion information of each disease, and then calculating the average value to obtain the average proportion of the known drugs in the recommended drugs of all diseases.