A method for predicting antidepressant drugs

By integrating biomedical knowledge graphs and transcriptome data, a neural network model was constructed to screen antidepressants, solving the problems of high cost and low efficiency in existing technologies. This resulted in efficient and accurate drug screening and has broad application prospects.

CN119517448BActive Publication Date: 2025-10-31KUNMING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411430075.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-14
Publication Date
2025-10-31
Estimated Expiration
2044-10-14

AI Technical Summary

Technical Problem

Existing technologies for developing antidepressant drugs are costly, inefficient, and lack an understanding of the complex network between drugs, targets, and depression, making cost-effective treatments difficult to achieve.

Method used

By integrating biomedical knowledge graphs and transcriptomic data, a neural network model is constructed. The trained neural network model is used to screen for potential antidepressants. By combining the triple information in the knowledge graph and differentially expressed genes in the transcriptome, the relationship between drugs and depression is predicted.

Benefits of technology

This improved the efficiency and accuracy of antidepressant screening, reduced costs, and provided new strategies for drug screening for other diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119517448B_ABST
    Figure CN119517448B_ABST
Patent Text Reader

Abstract

This invention provides a method for predicting antidepressant drugs. Based on a combination of knowledge graph, transcriptomics, and machine learning technologies, it screens potential drugs for depression. Specifically, it integrates disease-related node files, relation files, and transcriptomic data to construct information in triplet form, which is then imported into a model and converted into digital tensor information for scoring. The training dataset includes antidepressant tags from ClinicTrials and DrugBank, as well as other non-antidepressant drugs, combined with differential gene expression data obtained using RNA-seq technology. Multiple scoring models, such as TransE, TransR, DistMult, and ComplEx, are used for training and evaluation to screen potential antidepressant drugs. This method has the advantages of low cost, high efficiency, and wide applicability, significantly improving the accuracy and efficiency of antidepressant drug screening, providing a new drug discovery pathway for the treatment of depression, and helping to accelerate the research and development process of antidepressant drugs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of drug screening technology for depression, specifically to a method for screening potential drugs for depression based on a combination of knowledge graphs, transcriptomics, and machine learning. Background Technology

[0002] Depression is a serious mental illness affecting approximately 16% of the global population. It is projected to become a leading cause of disability by the end of 2030. Depression, also known as depressive disorder, is a mental disorder with a high incidence, high clinical cure rate but low treatment acceptance rate and high relapse rate. Its main characteristic is significant and persistent low mood. Patients may exhibit self-harm, suicidal behavior, and even psychotic symptoms such as delusions or hallucinations. Clinically, depressive disorders can be classified as mild, moderate, and severe based on the number, type, and severity of symptoms. In addition to low mood, loss of interest, and lack of energy, there are some early symptoms such as slowed reaction time, slowed thinking, and memory loss; these symptoms vary among individuals. The causes of depression are related to chronic stress, gut microbiota, genetic factors, and substance abuse. Although the specific pathogenesis of depression is not yet fully understood, it is generally believed to be closely related to a decrease in monoamine neurotransmitter levels, and most antidepressants are currently developed based on this theory.

[0003] RNA sequencing (RNA-seq), a high-throughput gene expression analysis technique, has been widely used in molecular biology research since its inception, contributing to the understanding of gene function and regulatory mechanisms. RNA-seq is commonly used to analyze differential gene expression (DGE), and its standard workflow generally consists of three steps: First, a wet assay is performed to extract RNA and enrich mRNA or remove rRNA, followed by cDNA synthesis and construction of a sequencing library; next, sequencing is performed on a high-throughput sequencing platform, typically with a sequencing depth of 10-30 million reads per sample; finally, data analysis is performed, including aligning reads to a reference genome or transcriptome, calculating the number of read-to-transcriptome alignments, and performing normalization and statistical differential analysis between samples. The application of RNA-seq helps to reveal the complexity of mRNA splicing and the regulatory mechanisms of non-coding RNAs and enhancer RNAs on gene expression.

[0004] The rich biomedical knowledge gained from biological experiments and clinical practice is a valuable resource in biomedicine. Knowledge graphs (KGs), as an effective knowledge management tool, support the integration of knowledge in the biomedical and life sciences fields. A KG is a multi-relationship graph or network used to integrate, coordinate, and store biomedical knowledge collected from multiple knowledge sources. A KG consists of nodes representing biomedical entities (such as diseases, drugs, genes, and biological processes), connected by a series of edges representing the relationships between them (e.g., drug-treatment-disease, disease-association-gene and drug interaction-drug relationship, etc.).

[0005] Statistics show that pharmaceutical companies invest heavily in developing new chemical entities for FDA approval, with an average R&D cost of $2.6 billion per drug. Utilizing existing drugs to discover new indications is a cost-effective strategy that can help develop disease prevention and treatment plans. Recent research indicates that network-based approaches can effectively screen FDA-approved drugs with favorable pharmacokinetic / pharmacodynamic properties, safety, and tolerability for potential new indications by leveraging the relationship between drug targets and diseases. However, developing cost-effective treatments for MDD remains challenging due to a lack of understanding of the complex networks between drugs, targets, depression (MDD), and the disease. Therefore, extracting relevant information from large-scale structured medical databases is a daunting task. Summary of the Invention

[0006] The objective of this invention is to provide a screening method for drugs with potential antidepressant effects based on knowledge graphs combined with transcriptomics, aiming to find new antidepressant drugs and improve the efficiency of existing drug screening technologies.

[0007] To achieve the above objectives, the technical solution of the present invention is as follows:

[0008] A method for predicting antidepressant medication includes:

[0009] Acquire existing knowledge graphs in the biomedical and life sciences fields, information on drugs to be predicted, and information on depression; among which, information on depression includes differentially expressed genes associated with depression.

[0010] The collected information on drugs to be predicted and information on depression are integrated into triples according to nodes and relationships, and combined with triples from existing knowledge graphs in the biomedical and life science fields to construct a dataset.

[0011] The known triples in the dataset are input into the neural network model. The model is trained by minimizing the error between the predicted score output by the neural network model and the actual relationship of the input triples.

[0012] The drug to be predicted is paired with depression as nodes to form a triplet with unknown relationship, and then input into a trained neural network model. Based on the treatment relationship scores of the predicted drug and depression output by the trained neural network model, potential antidepressants are screened.

[0013] Furthermore, the information related to the drug to be predicted includes the disease, anatomy, pathway, drug, symptoms, dietary supplement ingredients, dietary supplement products, treatment category, genes, and side effects related to the treatment of the drug to be predicted.

[0014] Furthermore, the node pairs of the triples include: disease-disease, drug-gene, drug-disease, and gene-gene; among them, differentially expressed genes associated with depression constitute the depression-gene triple as potential therapeutic targets.

[0015] Furthermore, the neural network model is one or more of TransE, TransR, DistMult, and ComplEx.

[0016] Furthermore, the error between the predicted score output by the minimized neural network model and the actual relationship of the input triplet is specifically as follows:

[0017] Assuming that the closer the predicted score is to 1, the stronger the association between the node pairs, then we should maximize the score of the correct triples while minimizing the score of the incorrect triples.

[0018] Furthermore, potential antidepressants are screened based on the relationship scores between predicted drugs and the treatment of depression, as output by the trained neural network model. Specifically:

[0019] The predicted drug-to-depression treatment relationship scores output by the trained neural network model are sorted from highest to lowest, and drugs ranked higher than known antidepressants are selected as potential antidepressants.

[0020] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aforementioned method for predicting antidepressant drugs.

[0021] A storage medium containing computer-executable instructions that, when executed by a computer processor, implement the aforementioned method for predicting antidepressant drugs.

[0022] A computer program product includes a computer program / instructions that, when executed by a processor, implement the steps of the antidepressant drug prediction method.

[0023] This invention can effectively screen for new drugs with antidepressant potential, while improving screening efficiency and accuracy. This method is not only low-cost but also has broad application prospects, and can be used to screen potential drugs and related targets and mechanisms for other diseases. Attached Figure Description

[0024] Figure 1 Here is a flowchart of an antidepressant drug prediction method provided by the present invention;

[0025] Figure 2 In this embodiment of the invention, existing knowledge graph iBKH data is reorganized into triple information and input into the model for training. The optimal triple type is selected for input into the model for training, resulting in a trained model. Subsequently, the optimal disease-drug relationship type is selected, and clinical drugs are scored using AUC scores from different models TransE, TransR, DistMult, and ComplEx, as well as the Ensembl scoring results after integrating the four models. Specifically, the Ensembl scoring after integrating the four models involves summing the rankings of each drug in the results of the four training models, normalizing the data, scaling the data to 0 to 1, re-ranking the drugs, and selecting drugs ranked before known antidepressants as potential antidepressants.

[0026] Figure 3 In this embodiment of the invention, based on existing knowledge graphs iBKH and PrimeKG, data is processed and integrated to reorganize into triple information. The best triple type is selected and input into the model for training to obtain a trained model. Then, the best disease-drug relationship type is selected, and the AUC scores of clinical drugs are obtained through different models TransE, TransR, DistMult and ComplEx. The Ensembl scoring results after the four models are integrated are also shown in the figure.

[0027] Figure 4 In this embodiment of the invention, differentially expressed genes obtained from iBKH and primeKG, along with differential analysis of different transcriptome data, are processed and integrated to form triplet information. The optimal triplet type is selected and input into the model for training to obtain a trained model. Subsequently, the optimal disease-drug relationship type is selected, and the AUC scores of clinical drugs are obtained through different models TransE, TransR, DistMult, and ComplEx. The Ensembl scoring results after the integration of the four models are also presented. Detailed Implementation

[0028] To better illustrate the purpose, technical solution, and advantages of the present invention, the present invention will be further described below in conjunction with specific embodiments.

[0029] This invention provides a method for predicting antidepressant drugs. By integrating information on MDD-related drugs, targets, and pathways from existing knowledge graphs (KGs) and combining it with depression-related transcriptomic data, a new knowledge graph is constructed and a neural network model is trained. This trained neural network model is then used as a scoring model, and potential antidepressant drugs are screened based on the output of the scoring model. Figure 1 As shown, it includes the following steps:

[0030] (1) Construct a training dataset, which contains various triples that form a new knowledge graph. This dataset is obtained through disease node files, disease-related other factor relationship type files, disease-related transcriptome data, and antidepressant clinical drugs and antidepressant tags. Specifically, this includes the following sub-steps:

[0031] (1.1) Obtain existing knowledge graphs in the biomedical and life science fields, information on drugs to be predicted, and information on depression;

[0032] The existing knowledge graphs in the biomedical and life sciences fields refer to existing knowledge graphs related to biomedicine and life sciences. These generally include disease node files and relationship type files for other disease-related factors, which can be downloaded from various databases. This invention downloaded 10 types of node files and 18 types of relationship files from iBKH (https: / / github.com / wcm-wanglab / iBKH / ), and 10 types of node files and 26 types of relationship files from PrimeKG (https: / / github.com / mims-harvard / PrimeKG / ). Based on the node and relationship files from iBKH, the two datasets were integrated and processed, ultimately yielding 366,770 nodes, encompassing Disease, Anatomy, Pathway, Drug, Symptom, Dietary Supplement Ingredient (DSI), Dietary Supplement Product (DSP), and Therapeutic Category. Ten types of node information, including Class, Gene, and Side-Effect, were obtained, yielding 50,312,276 relationship types. These included 18 node pairs such as DSI-Symptom, DSI-Disease, DSI-Drug, DSI-Anatomy, DSP-DSI, DSI-TC, Disease-Pathway, Drug-Pathway, Drug-Side Effect, Disease-Gene, Disease-Symptom, Anatomy-Gene, Gene-Pathway, Disease-Disease, Drug-Drug, Drug-Gene, Drug-Disease, and Gene-Gene. The results are shown in Table 1.

[0033] The depression-related information in this invention mainly includes differentially expressed genes associated with depression. In this embodiment, six datasets related to depression, namely GSE45468, GSE45642, GSE66277, GSE87610, GSE102556 and GSE185855, were downloaded from the GEO database. The transcriptome data were analyzed for differential expression using the limma and DESeq2 packages in R language to obtain differentially expressed genes (DEGs, with a screening threshold of p<0.05; |log2FC|>=1.2). The results of differential expression analysis for each transcriptome are shown in Table 2.

[0034] The drug-related information to be predicted in this invention mainly refers to whether it is an antidepressant. In this embodiment of the invention, known approved drugs and clinical trial drugs were downloaded from the DrugBank (https: / / go.drugbank.com / drugs / ) and ClinicalTrials (https: / / clinicaltrials.gov / ) databases. Drugs related to depression were set as antidepressants and labeled with the corresponding tag 1, while other drugs were set as non-antidepressants and labeled with the corresponding tag 0. The drug results for each stage are shown in Table 3.

[0035] Table 1: Node and Relationship Information of the Training Dataset

[0036] Entity type quantity disease 37,997 Anatomy 23,186 path 31,658 drug 37,628 symptom 1,361 Dietary supplement ingredients 4,101 Dietary supplement products 137,568 Treatment Category 605 Gene 88,415 side effect 4,251 total 366,770

[0037] Relationship type quantity DSI-Symptom 2,093 DSI-Disease 5,134 DSI-Drug 3,057 DSI-Anatomy 4,334 DSP-DSI 689,297 DSI-TC 5,430 Disease-Pathway 1,941 Drug-Pathway 3,231 Drug-Side Effect 163,206 Disease-Gene 27,538,774 Disease-Symptom 3,357 Anatomy-Gene 12,171,021 Gene-Pathway 152,243 Disease-Disease 18,904 Drug-Drug 2,684,682 Drug-Gene 1,275,949 Drug-Disease 2,727,267 Gene-Gene 802,707 Total 50,312,276

[0038] Table 2: Results of Transcriptional Differential Analysis

[0039]

[0040]

[0041] Table 3: Drug Clinical Information

[0042] Clinical stage Depression-related medications Phase I clinical trial 29 Phase II clinical trials 72 Phase III clinical trials 27 Phase IV clinical trials 40 FDA approved for market launch 65

[0043] MDD-related drugs represent the number of antidepressants in different stages of drug use, and are also the number of antidepressant labels.

[0044] (1.2) The collected information on drugs to be predicted and information on depression are integrated into triples according to nodes and relationships, and then combined with the triples of existing knowledge graphs in the fields of biomedicine and life sciences to construct a dataset.

[0045] Specifically, this embodiment uses data collected by iBKH as a basis to select different relationship types and reorganize them into new triplet information. There are three main selection methods: The first method selects 11 relationship types, which, compared to the remaining eight, have direct and indirect relationships with diseases and drugs: Disease-Pathway, Drug-Pathway, Drug-SideEffect, Disease-Gene, Disease-Symptom, Gene-Pathway, Disease-Disease, Drug-Drug, Drug-Gene, Drug-Disease, and Gene-Gene; The second method selects five relationships: Disease-Disease, Drug-Drug, Drug-Gene, Drug-Disease, and Gene-Gene, which further reduces the influence of other factors and only considers the influence of drugs and diseases; The third method selects Disease-Disease, Drug-Gene, Drug-Disease, and Gene-Gene, excluding the influence of Drug-Drug from the second method. Based on the iBKH dataset, PrimeKG data was added, and then the data was reorganized into new triplet information following the selection method in step one. Based on the iBKH and PrimeKG datasets, differentially expressed genes obtained from differential analysis of six datasets were added to the Disease-Gene relationship type. These differentially expressed genes were added to the Disease-Gene relationship as potential therapeutic targets, and then the data was reorganized into new triplet information following the selection method in step one, selecting different relationship types.

[0046] (2) The scoring model is trained by using the integrated triplet information as input to a neural network model for feature learning. Nodes and relations are converted into tensors, and the model is trained with the goal of minimizing the error between the predicted score and the actual relation. The trained neural network model is then used as the scoring model. The scoring model in this embodiment of the invention uses four models from the Python package DGL-KE: TransE, TransR, DistMult, and ComplEx. The specific description is as follows:

[0047] The scoring model based on TransE: TransE is a machine learning model for Knowledge Graph Embedding (KGE). Its core idea is to embed entities and relations in a knowledge graph into a low-dimensional vector space, allowing semantic relationships between entities to be captured through vector operations. The model aims to learn a function f that maps head entity h, relation r, and tail entity t to a score representing the probability that the triple (h, r, t) is true in the knowledge graph. TransE models relation r as a vector such that for a correct triple (h, r, t), the embedding vector of head entity h plus the relation vector r should be as close as possible to the embedding vector of tail entity t. It is assumed that the closer the score is to 1, the stronger the association between h and t.

[0048] TransR is an improvement on the TransE model, designed to address some of TransE's limitations in handling complex relationships, particularly many-to-many and many-to-one relationships. It retains some fundamental concepts from TransE, including entities, relations, and triples. However, TransR introduces a relation matrix to enhance the model's expressive power. The core idea is to combine the relation matrix with entity vectors to capture more complex relational patterns. In TransR, each relation r has a corresponding relation matrix Mr, which is used to adjust the interactions between entity vectors. It is assumed that the closer the score is to 1, the stronger the association between h and t.

[0049] DistMult is a machine learning model for predicting entity relationships in knowledge graphs. Based on tensor decomposition, it aims to capture complex relationships between entities. Its core idea is to map entities and relationships to a low-dimensional continuous vector space. Specifically, the DistMult model is trained using triples (head entity, relationship, tail entity), where each entity and relationship is represented as a vector. The training objective is to maximize the score of correct triples while minimizing the score of incorrect triples. Using a multinomial function as the distance metric effectively captures the semantic associations between entities, improving the accuracy of entity relationship prediction. Specifically, for a triple, the DistMult model performs a dot product operation on the vectors of the head entity and relationship, and then performs a dot product operation on the vector of the tail entity to obtain a score value representing the confidence of the triple. It is assumed that the closer the score is to 1, the stronger the association between h and t.

[0050] ComplEX extends DistMult by introducing complex-valued embeddings. Its core feature is that entity and relation embeddings h, r, t no longer exist in the real space but in the complex space. It uses complex vectors to represent entities and relations in the knowledge graph, thus enabling more flexible capture of complex interactions between entities, especially symmetric and asymmetric relations. It is assumed that the closer the score is to 1, the stronger the association between h and t.

[0051] The formulas for the four scoring models, TransE, TransR, DistMult, and ComplEx, are as follows:

[0052] TransE=-||h+rt||

[0053]

[0054] DisMult=h T diag(r)t

[0055] ComplEx=h T Re(diag(r)t)

[0056] In the formula, M r It is the mapping matrix of relation r, used to project entities from entity space to relation space. diag(*) means converting relation vector r into a diagonal matrix, and Re(*) means taking the real part.

[0057] (3) Screen potential antidepressants based on the relationship scores between the predicted drugs and the treatment of depression output by the trained neural network model.

[0058] In one specific implementation, the predicted drug-to-depression treatment relationship scores output by the trained neural network model are sorted from highest to lowest, and drugs ranked higher than known antidepressants are selected as potential antidepressants. Furthermore, the rankings of each drug across the four trained models can be summed, normalized, and scaled to 0 to 1. The drugs are then re-ranked, and those ranked higher than known antidepressants are selected as potential antidepressants.

[0059] In this embodiment, clinical drug information is downloaded from the DrugBank and ClinicaTrials databases. There are 4110 FDA-approved drugs, of which 65 have antidepressant tags; 4113 drugs from Phase IV clinical trials to market approval, of which 105 have antidepressant tags; 4120 drugs from Phase III clinical trials to market approval, of which 132 have antidepressant tags; 4157 drugs from Phase II clinical trials to market approval, of which 204 have antidepressant tags; and 4176 drugs from Phase I clinical trials to market approval, of which 233 have antidepressant tags. The clinical drug and depression triad information is input into a trained scoring model. Based on the drug ranking results of the scoring model, potential antidepressant drugs are screened. Among them, the optimal scoring model built based on the existing iBKH knowledge graph uses the triplet information of three types of nodes: Disease-Gene, Drug-Gene, Drug-Disease, and Gene-Gene. The model was trained based on these three types of relationships between nodes. Subsequently, the scoring model selected six relationship types between diseases and drugs: treatment, mitigation, effect, association, associative relationship, and semantic relationship. The final AUC score was around 0.7. Figure 2 As shown;

[0060] The optimal scoring model, built by incorporating PrimeKG knowledge graph information into the existing iBKH knowledge graph, uses three types of triplet information—Disease-Gene, Drug-Gene, Drug-Disease, and Gene-Gene—as the corresponding relationship types between nodes. The model is then trained based on these relationships, selecting six types of relationships between diseases and drugs: treatment, mitigation, effect, association, associative relationship, and semantic relationship. The highest AUC score achieved is 0.75, indicating a slight improvement in model performance. Figure 3 As shown;

[0061] In existing knowledge graphs, diseases and genes have the following relationships: association, downregulation, upregulation, inferred relationship, inappropriate regulation, disease-related, pathogenic mutation, polymorphic alteration, risk role in pathogenesis, potential therapeutic effect, biomarker (diagnosis), progression promotion, drug target, overexpression in disease, and mutations affecting disease course. Differential analysis of MDD patients and normal samples revealed differentially expressed genes closely related to the pathogenesis of MDD, aiding in the discovery of MDD-related drugs. For potentially related genes (differentially expressed genes) identified by the transcriptome, this invention adds a '1' to the association relationship to indicate that the gene is associated with the disease; other non-differentially expressed genes are not added. Other relationships remain unchanged. Subsequently, after incorporating information on differentially expressed genes from different transcriptomes into the iBKH and integrated knowledge graph, model training results show varying degrees of improvement in model performance, such as... Figure 4As shown. However, the triplet information and disease-drug relationship types in the optimal scoring model of the integrated knowledge graph differ after incorporating corresponding differentially expressed transcriptome genes. For example, with GSE53987, the corresponding triplet information is Disease-Gene, Drug-Gene, Drug-Disease, and Gene-Gene. The disease-drug relationship types selected by the scoring model are treatment and palliative, with a final AUC score as high as 0.76. With GSE185885, the corresponding triplet information is Disease-Gene, Drug-Gene, Drug-Disease, Drug-Drug, and Gene-Gene. The disease-drug relationship types selected by the scoring model are Treats and Palliates, with a final AUC score as high as 0.76. With GSE66277, the corresponding triplet information is Disease-Pathway, Drug-Pathway, and Drug-Side. Eleven disease-drug relationship types were identified: Effect, Disease-Gene, Disease-Symptom, Gene-Pathway, Disease-Disease, Drug-Drug, Drug-Gene, Drug-Disease, and Gene-Gene. The scoring model selected two relationship types based on the disease-drug relationship: Treats and Palliates. The highest AUC score reached 0.83. Adding GSE87610 resulted in five triplet types: Disease-Gene, Drug-Gene, Drug-Disease, Drug-Drug, and Gene-Gene. The scoring model selected two relationship types based on the disease-drug relationship: Treats and Palliates. The highest AUC score reached 0.78. Adding GSE45642 resulted in three triplet types: Disease-Pathway, Drug-Pathway, Drug-Side... There are 11 types of relationships between disease and drug, including Effect, Disease-Gene, Disease-Symptom, Gene-Pathway, Disease-Disease, Drug-Drug, Drug-Gene, Drug-Disease, and Gene-Gene. The relationship type between the disease and drug selected according to the scoring model is either treatment or mitigation. The highest AUC score can reach 0.85; Adding GSE45468, the corresponding triplet information includes five types: Disease-Gene, Drug-Gene, Drug-Disease, Drug-Drug, and Gene-Gene. Based on the scoring model, the relationship between disease and drug is selected as either treatment or mitigation. The highest AUC score reached 0.85, and the final integrated model also achieved 0.81. The drugs selected based on the scoring results are shown in Table 4.

[0062] Table 4: Top 20 drugs predicted by Ensembl AUC score of 0.81

[0063]

[0064]

[0065] Where candidate represents a potential antidepressant, and MDD represents a known marketed antidepressant.

[0066] An electronic device provided in this invention includes one or more processors for implementing an antidepressant drug prediction method as described in the above embodiments.

[0067] The electronic device of the present invention can be applied to any device with data processing capabilities, such as a computer or other similar device.

[0068] The device embodiment can be implemented through software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of any data processing device in which it resides reading the corresponding computer program instructions from non-volatile memory into memory for execution. From a hardware perspective, in addition to the processor, memory, network interface, and non-volatile memory, the data processing device in which the device in the embodiment resides may also include other hardware depending on the actual function of that data processing device, which will not be elaborated further.

[0069] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0070] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the present invention according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0071] This invention also provides a computer-readable storage medium storing a program thereon, which, when executed by a processor, implements an antidepressant drug prediction method as described in the above embodiments.

[0072] The computer-readable storage medium can be an internal storage unit of any data processing device described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be any data processing device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., mounted on the device. Furthermore, the computer-readable storage medium can include both internal storage units of any data processing device and external storage devices. The computer-readable storage medium is used to store the computer program and other programs and data required by the data processing device, and can also be used to temporarily store data that has been output or will be output.

[0073] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. A method for predicting antidepressant medication, characterized in that, include: Acquire existing knowledge graphs in the biomedical and life sciences fields, information on drugs to be predicted, and information on depression; The depression-related information includes differentially expressed genes associated with depression. These differentially expressed genes were obtained by downloading six depression-related datasets (GSE45468, GSE45642, GSE66277, GSE87610, GSE102556, and GSE185855) from the GEO database and performing differential analysis on the transcriptome data using the R language packages limma and DESeq2. The collected information on drugs to be predicted and information on depression are integrated into triples according to nodes and relationships, and combined with triples from existing knowledge graphs in the biomedical and life science fields to construct a dataset. The node pairs of the triples include: disease-disease, drug-gene, drug-disease, and gene-gene. Among them, differentially expressed genes related to depression serve as potential therapeutic targets and constitute the depression-gene triple. The known triples in the dataset are input into a neural network model. The model is trained by minimizing the error between the predicted score output by the neural network model and the actual relationship of the input triples. The neural network model is TransE, TransR, DistMult, or ComplEx. The drug to be predicted is paired with depression as nodes to form triplets representing unknown relationships, which are then input into a trained neural network model. Based on the treatment relationship scores between the predicted drug and depression output by the trained neural network model, potential antidepressants are screened. Specifically: The rankings of the drugs to be predicted in the four training models are summed, then normalized, and the data is scaled to 0 to 1. The drugs to be predicted are then re-ranked, and the drugs ranked before the known antidepressants are selected as potential antidepressants.

2. The method according to claim 1, characterized in that, The error between the predicted score output by the minimized neural network model and the actual relationship of the input triples is specifically as follows: Assuming that the closer the predicted score is to 1, the stronger the association between the node pairs, then we should maximize the score of the correct triples while minimizing the score of the incorrect triples.

3. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements an antidepressant drug prediction method as described in any one of claims 1-2.

4. A storage medium containing computer-executable instructions that, when executed by a computer processor, implement an antidepressant drug prediction method as described in any one of claims 1-2.

5. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the antidepressant drug prediction method according to any one of claims 1-2.

Citation Information

Patent Citations

  • Novel active compound calculation and screening method for various viruses

    CN116469485A

  • Drug curative effect prediction method based on artificial intelligence and related equipment

    CN116524995A