Graph neural network-based candidate drug efficacy prediction and selection method, medium and equipment

By constructing a drug efficacy prediction model based on graph neural networks and utilizing causal chains and graph neural networks for feature diffusion, the problem of integrating the relationship between drugs, targets, and diseases was solved, enabling efficient drug discovery and personalized drug efficacy prediction, and improving the efficiency and accuracy of drug development.

CN120998541AActive Publication Date: 2025-11-21BEIJING FRIENDSHIP HOSPITAL CAPITAL MEDICAL UNIV +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511520675.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2025-11-21
Estimated Expiration
2045-10-23

AI Technical Summary

Technical Problem

Existing technologies cannot effectively integrate and explore the multidimensional relationships between drugs, targets, and diseases, resulting in low efficiency and high costs in drug development, and failing to meet the demand for rapid and precise development of new drugs.

Method used

A drug efficacy prediction model based on graph neural networks is constructed. A biomedical directed graph is built through causal chains, and feature diffusion updates are performed using graph neural networks to generate drug, relationship, and target triplet. The drug efficacy prediction results are determined based on the differential protein type.

Benefits of technology

It enables efficient integration of multidimensional relationships between drugs, targets, and diseases, reduces the difficulty of drug discovery, shortens the research and development cycle, and supports personalized efficacy prediction and high-throughput virtual screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998541A_ABST
    Figure CN120998541A_ABST
Patent Text Reader

Abstract

The invention relates to a candidate drug efficacy prediction and selection method based on a graph neural network, a medium and equipment, and the method comprises the steps: firstly obtaining a biological medicine causal chain, and then constructing and training a drug efficacy prediction model according to the causal chain. Then, after the model is constructed, determining a target differential protein, and determining a relationship according to the type of the differential protein; and finally, constructing a triple by taking the differential protein as a tail entity, the relationship as an edge and circularly traversing the candidate drugs as a head entity, inputting the triple into a prediction model, and outputting a drug effect prediction result of each candidate drug. On the whole, the problems that in the prior art, the multi-dimensional incidence relation among drugs, targets and diseases cannot be efficiently integrated and excavated, an interpretable, extensible and updatable knowledge network cannot be constructed, the drug discovery difficulty cannot be lowered, and the research and development cycle cannot be shortened are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of graph neural network technology, and in particular to a method, medium, and device for predicting and selecting the efficacy of candidate drugs based on graph neural networks. Background Technology

[0002] Drug development is a long-term, costly, and highly unsuccessful systematic project. On the one hand, the pathogenesis of known diseases is complex and diverse, with thousands of possible causes, and different causes often act on distinctly different molecular targets within cells. On the other hand, the number of candidate drug entities (including small molecule compounds, peptides, proteins, nucleic acids, etc.) available for screening or optimization is in the tens of thousands. If traditional in vitro and in vivo pharmacological experiments are conducted for every drug-target or even target-disease combination, it would not only require a massive investment of manpower, reagents, animals, and clinical resources, but also face bottlenecks such as long experimental cycles, delayed data updates, and poor scalability, resulting in low development efficiency and high costs. More seriously, with increasingly refined disease subtyping and an exponential expansion of the candidate molecule space, relying solely on trial-and-error experimental verification can no longer meet the urgent need for rapid and precise development of new drugs.

[0003] Chinese patent application CN 114420221 A discloses a "multi-task drug screening method and system based on knowledge graph assistance." However, the knowledge graph in this application only contains entities for drugs, compounds, targets, and proteins, completely lacking the entity for "disease," thus failing to establish a correlation between drug efficacy and disease treatment. Furthermore, it primarily trains cross-entropy classifiers for "drug-target interaction" and "compound-protein interaction," averaging the probabilities of the two paths as the final score. It extracts and trains features from each class individually, without effectively modeling the relationships within the knowledge graph. It lacks modeling and utilization of relationship direction and biological regulatory polarity (such as up / down regulation, activation / inhibition), as well as individualized weighted aggregation among multiple targets. This application is disconnected from the core need in actual drug development—"finding effective drugs for diseases"—and cannot directly support clinical medication decisions.

[0004] Chinese patent application CN 112562791 A discloses a "deep learning prediction system for drug target interaction based on knowledge graph, computer equipment, and storage medium." However, the knowledge graph only contains two types of entities: "drug" and "target," and the relationship is limited to a fixed "Drug X has Target Y" (single interaction relationship). It completely lacks entities such as "disease" and "differential protein," as well as heterogeneous logical relationships such as regulation / causation, making it unable to support directional biological semantics. Furthermore, its use of the symmetric bilinear DistMult method results in a lack of asymmetry and directionality in node and relationship embeddings, unlike biomedical networks. This makes it difficult to drive multi-node, multi-hop pharmacodynamic reasoning using edge attributes as a core factor. This singular structure limits the application to predicting drug-target interactions at the single-molecule level, failing to connect to the complete omics background of diseases and unable to answer the core question of "whether the drug is effective for a certain disease," thus severely disconnecting it from the scenario of "drug development centered on disease pathogenesis."

[0005] More importantly, existing methods extract individual features from drugs and targets, and then fuse these features for training, rather than training and updating the knowledge graph itself. They fail to address the intricate relationships and influences among these elements. Therefore, efficiently integrating and mining the multidimensional relationships between drugs, targets, and diseases to construct an interpretable, scalable, and updatable knowledge network to reduce the difficulty of drug discovery and shorten the development cycle has become a core technical problem urgently needing to be solved in this field. Summary of the Invention

[0006] To address at least one of the aforementioned technical problems, this invention provides a method for predicting the efficacy of candidate drugs based on graph neural networks, comprising: To obtain the causal chain of biomedicine; Based on the causal chain, a drug efficacy prediction model is constructed and trained, including: an input layer, used to construct a biomedical directed graph based on the causal chain and generate drug, relationship, and target triples; the nodes of the biomedical directed graph include at least three medical entities: drug, target, and disease; edges represent the relationships between entities; a graph neural network is used to perform feature diffusion updates on the entities and relationships between entities in the biomedical directed graph to obtain updated entities and relationships between entities; and an output layer is used to score the triples based on the updated entities and relationships between entities to output the drug efficacy prediction results. Identify the target differentially expressed protein and determine the relationship based on the type of the target differentially expressed protein; By constructing triplet pairs using differentially expressed proteins as tail entities, relations as edges, and candidate drugs as head entities through iterative iteration, the prediction results for each candidate drug are output as the prediction results for each candidate drug.

[0007] Furthermore, the input layer includes node embedding units and edge embedding units to construct a biomedical directed graph; Node embedding unit, used to assign a unique identifier to each medical entity according to category, in order to define the nodes of the biomedical directed graph; Edge embedding units are used to extract different relationship categories, relationship properties, and relationship flows between entities based on causal chains, in order to define the types, polarities, and directions of different edges and construct a directed graph for biomedicine.

[0008] Furthermore, the input layer includes a triple generator to generate drug-relationship-target triples; including: Positive sample input unit, used to input positive samples of triplet of drug, relationship, and target; The negative sample generation unit is used to replace the head entity and / or tail entity of the positive sample to generate a negative sample.

[0009] Furthermore, the positive sample input unit is also used for: The baseline strength of each positive sample is determined based on its data source. The additional strength of each positive sample is determined based on the number of data points for each positive sample. The strength of evidence for each positive sample is determined by comprehensively considering its basic strength and additional strength. Based on the strength of evidence for each positive sample, a probability bias mechanism is introduced, so that the probability of a positive sample with high evidence strength is higher than the probability of a positive sample with low evidence strength.

[0010] Furthermore, the negative sample generation unit is also used for: Based on the strength of evidence of positive samples, a quantity bias mechanism is introduced, so that the number of negative samples generated by positive samples with high evidence strength is greater than the number of negative samples generated by positive samples with low evidence strength.

[0011] Furthermore, graph neural networks are used for: Several nodes and edges in a biomedical directed graph are converted into initial embedding vectors using unique identifiers, resulting in initial node embedding vectors and initial edge embedding vectors, thus enabling sparse feature modeling. The information of adjacent nodes and edges is aggregated, and the initial node embedding vector and the initial edge embedding vector are updated layer by layer to obtain the updated node embedding vector and edge embedding vector.

[0012] Furthermore, the output layer is used for: Based on the updated entities and the relationships between them, a boundary loss function is constructed to maximize the score difference between positive and negative samples. Solve the boundary loss function to obtain the drug efficacy prediction results.

[0013] Furthermore, relationships are determined based on the types of target differentially expressed proteins, including: Depending on the type of the target differentially expressed protein, negative relationships are used as edges for highly expressed proteins and positive relationships are used as edges for low-expressed proteins.

[0014] On the other hand, the present invention also provides a candidate drug selection method based on graph neural networks, comprising: Identify the target set of differentially expressed proteins for the current disease; If the target differential protein set contains only one element, then any of the above methods are used to determine the efficacy prediction results of each drug for the target, and the drugs are ranked accordingly, and several candidate drugs are selected in sequence. If the target differential protein set includes several elements, then any of the above methods are used to determine the efficacy prediction results of each drug for each target, and the weighted sum is used to rank the drugs, and several candidate drugs are selected in sequence.

[0015] Furthermore, it also includes: Based on the efficacy prediction results, the drugs corresponding to each target are ranked, and several candidate drugs are selected in sequence. The target differentially expressed proteins are divided into upregulated proteins and downregulated proteins, resulting in a set of drugs for upregulated proteins and a set of drugs for downregulated proteins. The expression fold change of each target differential protein under disease and normal conditions was obtained, and the drug efficacy prediction results were weighted and averaged with the absolute value of the fold change to obtain the comprehensive score of each drug. The drugs were sorted according to their overall scores, and candidate drugs were selected from the sets of drugs that upregulate and downregulate proteins in sequence.

[0016] On the other hand, the present invention also provides a computer storage medium storing executable program code; the executable program code is used to execute any of the above-mentioned candidate drug efficacy prediction methods or any of the above-mentioned candidate drug selection methods.

[0017] On the other hand, the present invention also provides a terminal device, including a memory and a processor; the memory stores program code that can be executed by the processor; the program code is used to execute any of the above-mentioned candidate drug efficacy prediction methods or any of the above-mentioned candidate drug selection methods.

[0018] This invention provides a method, medium, and device for predicting and selecting candidate drug efficacy based on graph neural networks. First, a biomedical causal chain is obtained. Then, based on the causal chain, an efficacy prediction model is constructed and trained: an input layer is used to construct a directed biomedical graph based on the causal chain and generate triples for drug, relation, and target; a graph neural network is used to perform feature diffusion updates on the entities and relationships between entities in the directed biomedical graph, obtaining updated entities and relationships; an output layer is used to score the triples based on the updated entities and relationships, outputting the efficacy prediction results. Next, after model construction, target differentially expressed proteins are identified, and relationships are determined based on the type of differentially expressed protein. Finally, using differentially expressed proteins as tail entities, relationships as edges, and candidate drugs as head entities, triples are constructed, input into the prediction model, and the efficacy prediction results for each candidate drug are output.

[0019] The key aspects of this invention are: 1. Based on causal chains, drugs, targets, and diseases are uniformly modeled as semantically rich and fully expressive directed graphs, which conforms to the directed multi-level association nature of biomedical systems, helps to reveal complex disease mechanisms, and clearly defines the relationships between entities (such as "drug-target", "target-disease", "drug-disease", etc.), providing interpretable structured prior knowledge for subsequent modeling; 2. Based on graph neural networks, through message passing... (Passing) Iteratively propagates the association features of drugs, targets, and diseases in the knowledge graph to achieve multi-hop reasoning. Simultaneously, it explicitly utilizes "relationship type, polarity, and direction" as edge attributes to enhance the model's understanding of biological semantics. It performs feature diffusion on entities and relationships within the knowledge graph. The key difference lies in the fact that feature diffusion is a diffusion of features across the graph structure, a diffusion of overall relationships and causal chains, rather than an update of individual features. This is its biggest difference from existing technologies. The result is a predictive model that takes the relationship chain of head entity, relation, and tail entity as input and the drug efficacy result as output. 3. The unique encoding of entities and relations allows learning the association relationships between different entities without extracting complex individual features of each entity. This invention effectively addresses the feature sparsity problem faced by large-scale data. 4. Based on this, it identifies the differentially expressed protein set for diseases; filters out proteins significantly related to the disease, removes irrelevant signals, reduces graph redundancy, and improves model focus; it uses differentially expressed proteins as tail entities, relationships as edges, and iteratively traverses candidate drugs as head entities to construct model input, inputting it into the prediction model and outputting the efficacy prediction results for each candidate drug. It can batch-traverse all candidate drugs, achieving high-throughput virtual screening and significantly improving efficiency. Simultaneously, each prediction result can be traced back to a specific (drug → relationship → protein) path, supporting mechanism explanation and experimental validation design. Furthermore, it can generate personalized efficacy predictions for differentially expressed protein profiles of different patients / subtypes, moving towards precision medicine. Overall, this invention solves the problem that existing technologies cannot efficiently integrate and mine the multidimensional relationships between drugs, targets, and diseases, constructing an interpretable, scalable, and updatable knowledge network to reduce the difficulty of drug discovery and shorten the research and development cycle. Attached Figure Description

[0020] Figure 1 This is a flowchart of an embodiment of a candidate drug efficacy prediction method based on graph neural networks; Figure 2 A schematic diagram of a structural embodiment of a drug efficacy prediction model; Figure 3 This is a schematic diagram illustrating the types of inhibitors in one embodiment; Figure 4 This is a schematic diagram illustrating the types of activators in one embodiment; Figure 5This is a schematic diagram illustrating the intersection of an upregulated protein inhibitor and a downregulated protein activator in one embodiment. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0022] It should be noted that if the embodiments of the present invention involve directional indications, such as up, down, left, right, front, back, etc., these directional indications are only used to explain the relative positional relationships and movement of the components in a specific posture. If the specific posture changes, the directional indications will also change accordingly. Furthermore, if the embodiments of the present invention involve descriptions such as "first," "second," "S1," "S2," "step one," "step two," etc., these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance, or implicitly indicating the number of technical features indicated or the order of method execution. Those skilled in the art will understand that anything that does not violate the inventive concept and is within the scope of the present invention should be included in the protection scope of the present invention.

[0023] like Figure 1 As shown, this invention provides a method for predicting the efficacy of candidate drugs based on graph neural networks, comprising: S1: Obtaining the causal chain of biomedicine; Specifically, optional but not limited to integrating professional databases such as DrugBank, GNBR, and STRING, we can track and obtain causal chains between various types of entities such as drugs, targets, and diseases, in order to subsequently construct multi-source heterogeneous biomedical directed graphs.

[0024] S2: Based on the causal chain, construct and train a drug efficacy prediction model, such as... Figure 2 As shown, it includes: (a) Input layer The input layer is used to construct a biomedical directed graph based on causal chains and generate drug, relationship, and target triples. The nodes of the biomedical directed graph include at least three medical entities: drug, target, and disease. Edges represent the relationships between the entities. Specifically, based on the obtained causal chain, nodes of the biomedical directed graph are defined using biomedical entities such as drugs (DrugBank ID), targets (UniProt ID), and diseases (MeSH ID); and edges of the biomedical directed graph are defined based on the relationships between the entities to construct the biomedical directed graph and generate drug-relationship-target triples.

[0025] Preferably, the input layer includes node embedding units and edge embedding units to construct a biomedical directed graph; (1) Node embedding unit, used to set a unique identifier for each medical entity according to category, so as to define the nodes of the biomedical directed graph; Specifically, taking the three medical entities of drugs, targets, and diseases as examples, each type may have tens of thousands of entities. The available drugs are named from a1 to aN, where N represents the number of drugs. Each drug entity is encoded with a unique identifier and embedded in the biomedical directed graph as a single node.

[0026] Example: Information and quantities of drugs, targets, and diseases can be obtained from any publicly available professional database as shown in Table 1. At the same time, a unified naming convention (such as UniProt ID for proteins, DrugBank ID for drugs, and MeSH ID for diseases) is adopted to resolve naming conflicts of heterogeneous data.

[0027] Table 1: Number of Drug Targets Based on Each Database

[0028] (2) Edge embedding unit, used to extract different relationship categories, relationship properties and relationship flow directions between entities according to the causal chain, so as to define the type, polarity and direction of different edges and construct a biomedical directed graph.

[0029] Specifically, taking the pairwise relationships between drugs, targets, and diseases as an example, including inhibition, activation, upregulation, downregulation, pathogenesis, and association, as shown in Table 2, different text, colors, and line types can be defined for different relationship categories. Different labels can be assigned to different edges to clearly define which type of relationship each edge represents, thus defining the type of edge. When the relationship has opposing properties such as "positive / negative," such as inhibition / activation, upregulation / downregulation, "polarity" can be used to supplement the edges, such as using "positive / negative values" or "binary labels." The core is to distinguish the "tendency" of the relationship. Positive polarity: represents a relationship that is "positive, supportive, and beneficial"; negative polarity: represents a relationship that is "negative." The relationship is "negative, opposing, and harmful." When there is a flow in the relationship, for example, when a drug acts on a target but the target does not act on the drug, there is an initiator (starting point) and a receiver (ending point) in the relationship. Optionally, the edges can be supplemented with "direction," with arrows pointing from the starting node to the ending node, representing "a certain relationship between the starting point and the ending point, such as the drug's effect on the target." This relationship is unidirectional (unless an additional reverse edge is added). Reversing the direction may completely change the semantics of the relationship. To express a "bidirectional relationship," two edges with opposite directions (instead of one undirected edge) need to be added between the two nodes to construct a directed graph in biomedicine.

[0030] Table 2: Relationships between Entities

[0031] This embodiment presents a preferred embodiment for constructing a biomedical directed graph: 1. Each medical entity is assigned a unique identifier by category to define the nodes of the biomedical directed graph. Instead of introducing the inherent characteristics of drugs, targets, and diseases into the graph to diffuse and update their complex features, a unique identifier is assigned to simply encode the nodes. Subsequent feature diffusion and updates are based solely on the node identifier's position and connectivity within the graph, achieving sparse feature modeling and significantly reducing model complexity and improving accuracy. 2. Based on causal chains, different relationship categories, relationship properties, and relationship flows between entities are extracted to define the types, polarities, and directions of different edges, thus constructing a biomedical directed graph. This is a progressive definition from "basic" to "refined," first defining individual nodes, then defining the relationships between nodes. The "type," "polarity," and "direction" of edges are not independent but progressively layered "semantic supplements," collectively forming a complete edge definition. This fully reflects the relationships between entities, constructing a semantically rich and fully expressive biomedical directed graph.

[0032] Although the types, polarities, and directions of nodes and edges in a directed graph are fundamental concepts, existing medical knowledge graphs primarily fuse the characteristics of drugs, targets, and diseases for training, rather than using a directed graph of nodes and edges for fusion training. This neglects these nuanced issues, which is the key to this invention. It extracts the causal chains of biomedicine to construct a biomedical directed graph. The diversity of types, the expression of polarities, and the completeness of directions make the overall semantics of the directed graph richer and its expression more complete. Building on this, this invention differs from existing technologies by not extracting single features from a node for similarity comparison and feature fusion, but by performing feature diffusion updates on the entities and relationships between them in the directed graph. This allows for the extraction and mining of multidimensional relationships among drugs, targets, and diseases, improving prediction accuracy and precision.

[0033] More preferably, the input layer includes a triple generator to generate drug-relationship-target triples; including: (3) Positive sample input unit, used to input positive samples of drug, relationship, and target triples; Preferably, the positive sample input unit is also used for: Based on the data source of each positive sample, determine the baseline strength of each positive sample. For example, the data source can be any type of data source: clinical trial data samples, such as patient samples from randomized controlled trials (RCTs) showing "drug group + positive target marker + positive efficacy," or samples from real-world studies (RWS) showing the correlation between target-related indicators and efficacy after drug administration. These directly reflect human effects and have the highest baseline strength, thus assigning the highest probability. Animal model data samples, such as gene knockout (KO) mouse validation samples (e.g., drug failure after target knockout, proving the necessity of the target), or positive drug efficacy samples from disease models, can observe tissue / organ-level effects by simulating the in vivo physiological environment. However, due to species differences between animals and humans, the baseline strength is not as high as that directly affecting humans, thus assigning a secondary probability. In vitro experimental data, such as positive samples from drug-target binding experiments (SPR, ITC), cell pathway activation / Suppressing positive samples can quickly verify molecular-level causality, but it lacks the complexity of the in vivo environment and can be assigned secondary probability. In addition, this data source can also be selected from the data provider, such as the base weight of positive samples given by some authoritative institutions and authoritative journal papers, which is higher than the base weight of positive samples given by some newspapers and small institutions. The additional strength of each positive sample is determined based on the number of data points in each positive sample. For example, when extracting positive sample data chains, if a positive sample appears repeatedly and multiple evidence chains point to that sample, its evidentiary strength is greater. The additional strength of that positive sample can be determined based on the number of data points pointing to it, thus supplementing its strength. From the perspective of random samples, it can also be proven that the more numerous the identical samples, the easier it is to select them. The strength of evidence for each positive sample is determined by comprehensively considering its basic strength and additional strength. Based on the strength of evidence for each positive sample, a probability bias mechanism is introduced, so that the probability of a positive sample with high evidence strength is higher than the probability of a positive sample with low evidence strength.

[0034] (4) Negative sample generation unit, used to replace the head entity and / or tail entity of the positive sample to generate a negative sample.

[0035] Preferably, the negative sample generation unit also introduces a quantity bias mechanism based on the evidence strength of the positive samples, so that the number of negative samples generated by positive samples with high evidence strength is greater than the number of negative samples generated by positive samples with low evidence strength.

[0036] This embodiment provides a preferred implementation for generating drug-relationship-target triples: 1. Receive positive samples and use a strategy of randomly sampling head or tail entities to replace the head or tail entities of the positive samples to generate negative samples; 2. Determine the evidence strength of the positive samples from multiple perspectives and introduce a probability bias mechanism, so that the probability of selecting positive samples with high evidence strength is higher than the probability of selecting positive samples with low evidence strength; 3. Further, based on the evidence strength of the positive samples, introduce a quantity bias mechanism, so that the number of negative samples generated by positive samples with high evidence strength is greater than the number of negative samples generated by positive samples with low evidence strength. This not only enriches the model training data, but more importantly, it makes the model training closer to positive samples with high evidence strength, improving the accuracy of subsequent drug efficacy prediction.

[0037] (II) Graph Neural Networks Graph neural networks are used to perform feature diffusion updates on entities and relationships between entities in biomedical directed graphs, resulting in updated entities and relationships between entities. Preferably, for the complex graph structure with multiple edge types mentioned above, a relation-aware graph neural network can be used, including an L-layer message passing and updating mechanism. Different message propagation mechanisms are defined to capture different types of relation information. Specifically, the state update of each node depends not only on the states of its neighboring nodes but also on the type of edges connecting these nodes. Preferably, when aggregating neighbor node information, a weight matrix for each type of relation is introduced, allowing the model to learn the different effects of different types of relations on the representation of the central node.

[0038] Preferably, this step includes: (1) Convert several nodes and edges in the biomedical directed graph into initial embedding vectors through unique identifiers to obtain initial node embedding vectors and initial edge embedding vectors, thereby realizing sparse feature modeling; Specifically, since GNNs (such as R-GCN, CompGCN, and GAT) can learn low-dimensional embeddings of entities and relationships and capture high-order neighbor information, breaking through the linear limitations of traditional KG embeddings (such as TransE), in the process of biomedical directed graph embedding modeling, graph neural networks (GNNs) can be used to perform feature diffusion on entities and relationships between entities in the knowledge graph, i.e., nodes and edges. At the same time, it can be used for drug prediction tasks. In the form of one-hot encoding, an initial embedding vector is generated for each node and edge through a unique identifier, which serves as the input feature of the model to achieve sparse feature modeling.

[0039] More specifically, the original embedding vectors are constructed based on one-hot encoding of nodes and edges in the knowledge graph. For the head entity, relation, and tail entity triples... The original embedding vectors of the head and tail entities can be represented as: The original embedding vector of the relation can be represented as By learning and optimizing these embedding vectors, the model can capture deep semantic information between entities and relationships, thereby improving the accuracy of drug prediction.

[0040] (2) Aggregate the information of adjacent nodes and edges, and update the initial node embedding vector and the initial edge embedding vector layer by layer to obtain the updated node embedding vector and edge embedding vector. Specifically, in graph neural networks, a feature propagation method based on message passing mechanism can be adopted to iteratively propagate the features of drugs, targets, and diseases in the graph, realizing multi-hop reasoning (e.g., drug → target A → pathway → target B → disease). At the same time, the "relationship type" is explicitly used as the edge attribute (e.g., "inhibition", "activation", etc., "upregulation" and "downregulation" relationships) to enhance the model's understanding of biological semantics. The information of adjacent nodes is aggregated, and the node embedding vector is updated layer by layer to obtain the updated node embedding vector and edge embedding vector. By aggregating the adjacent information, the model can learn the local structural features of nodes in the graph, enhancing the representational ability of the embedding vector. As the number of layers increases, the node embedding vector can integrate information from more distant neighbors, thereby improving the model's expressive ability. The mathematical expression for feature propagation can be as shown in Equation (1). At the same time, since the edge embedding vector is only related to the type and is not related to the initial representation content, the edge embedding update method can be as shown in Equation (2). (1) (2) in, Indicates the first Layer node characteristics, For trainable parameters, The sigmoid is a non-linear activation function. Represents a node Neighbors Indicates the first Layer edge features. Through this formula, the features of a node not only contain its own information, but also effectively aggregate the features of its neighboring nodes, thereby capturing the local structure and global semantics in the knowledge graph. Graph neural networks can propagate and aggregate the features of nodes and edges layer by layer, so that the structural and semantic information in the knowledge graph can be fully expressed.

[0041] This embodiment presents a preferred implementation of a graph neural network. The model assigns trainable initial vectors to each type of entity and relation, and then aggregates neighbor information layer by layer on the graph using a message-passing-based graph neural network. Each layer combines the entity's own features with "messages" transmitted through different relation types, and then undergoes learnable transformations and non-linear activations to obtain progressively context-rich representations. To express directionality and semantic differences, inhibition, activation, and association relations participate in propagation as independent relation vectors, enabling the model to explicitly distinguish different edge types and directions when updating node representations. As the number of layers increases, nodes not only absorb local information from first-order neighbors but also integrate a wider range of structural and semantic cues.

[0042] III. Output Layer The output layer is used to score the triples based on the updated entities and the relationships between them, in order to output the drug efficacy prediction results.

[0043] Specifically, this step includes: (1) Based on the updated entities and the relationships between them, such as the node embedding vector and edge embedding vector mentioned above, construct a boundary loss function to maximize the score difference between positive and negative samples. (2) Solve the boundary loss function to obtain the drug efficacy prediction results.

[0044] Specifically, to optimize knowledge graph embedding learning, the classic TransE (Translating Embeddings) framework can be used as the representation target. This maps entities (such as drugs, targets, and diseases) and relations in the knowledge graph to a low-dimensional continuous vector space, guiding the triplets (head entity, relation, tail entity) to satisfy... To ensure that the knowledge graph retains its semantic information after being embedded into a low-dimensional vector space, the embedding of the knowledge graph in the low-dimensional vector space preserves its original semantic information. Then, a boundary loss function is constructed based on the updated node embedding vector and edge embedding vector to maximize the score difference between positive and negative samples. By minimizing the boundary loss (the negative of the score difference), the model can better distinguish between positive and negative samples and improve the accuracy of prediction. Finally, the boundary loss function is solved to obtain the drug efficacy prediction result.

[0045] For example, during model training, a boundary loss function can be defined to maximize the score difference between positive and negative samples. The specific form of the loss function can be as shown in equations (3)-(5): (3) (4) (5) in, and Let them represent the positive sample set and the negative sample set, respectively. For positive sample scoring functions, For negative sample scoring functions, For hyperparameters; , These are the embeddings of the head entity, relation, and tail entity of the positive sample, respectively. , These are the embeddings of the negative sample head entity, relation, and tail entity, respectively.

[0046] In this embodiment, a preferred embodiment of the output layer is given. During the training phase, a distance-based objective is used to constrain entities and relationships. For each real sample, the head or tail entity is randomly replaced to generate a negative sample. The score difference between the real and negative samples is calculated and the corresponding ranking loss is minimized, which enables the model to continuously bring the real association closer and distance the erroneous association.

[0047] S3: Identify the target differentially expressed protein and determine the relationship based on the type of differentially expressed protein; Specifically, since differentially expressed proteins are often pathogenic drivers or therapeutic response biomarkers, providing a biological basis for drug action, differential expression analysis methods can be used, but are not limited to, in combination with transcriptomic and proteomic data to reflect disease-specific molecular characteristics, screen for protein molecules that change significantly in disease states, filter out irrelevant signals, reduce map redundancy, improve model focus, and verify their reliability through statistical models (such as t-tests or analysis of variance) to construct target differentially expressed proteins for the disease. If the disease includes multiple target differentially expressed proteins, a set of target differentially expressed proteins can be constructed to provide key inputs for subsequent target identification and mechanism analysis.

[0048] Taking the CryAB (R120G) mutation as an example, when screening and feature extraction of differentially expressed proteins to determine the differentially expressed protein set for a disease, proteomics analysis of CryAB (R120G) mutant and wild-type mouse heart tissues can be performed to identify differentially expressed proteins. The limma package was used for differential analysis, and the screening conditions were as follows: And FDR < 0.05, the target differential protein set is obtained.

[0049] More preferably, based on the target protein type, negative relationships are used as edges for highly expressed proteins and positive relationships are used as edges for low-expressed proteins; specifically, negative relationships can be, but are not limited to, the aforementioned downregulation, inhibition, etc.; positive relationships can be, but are not limited to, the aforementioned upregulation, activation, action, etc.

[0050] More preferably, when there are multiple relations, the same head and tail entities will be assembled into different triples, such as (head entity, relation 1, tail entity) and (head entity, relation 2, tail entity), and the two triples will be fed into the model as independent samples for calculation.

[0051] S4: Construct triples by treating differentially expressed proteins as tail entities, relationships as edges, and iterating through candidate drugs as head entities. Input these triples into the prediction model and output the efficacy prediction results for each candidate drug.

[0052] Specifically, optional but not limited to using differentially expressed proteins from the differentially expressed protein set obtained in step S3 as tail entities, and aligning the expression direction of differentially expressed proteins with the drug action relationship to achieve predictions consistent with biological logic. For highly expressed proteins, negative relationships are used as edges, meaning the drug should "inhibit" the protein or its upstream; for low-expressed proteins, positive relationships are used as edges, meaning the drug should "activate" the protein or its pathway. Then, candidate drugs are iteratively traversed as head entities to construct model inputs, which are then input into the prediction model to calculate triplet correlations, thereby batch traversing all candidate drugs to achieve high-throughput virtual screening and significantly improve efficiency. Finally, the efficacy prediction results of each candidate drug are output. At this point, each efficacy prediction result can be traced back to a specific (drug → relationship → protein) path, supporting mechanism explanation and experimental validation design. Furthermore, personalized efficacy predictions can be generated for differentially expressed protein profiles of different patients / subtypes, moving towards precision medicine.

[0053] More preferably, depending on the number of elements in the target differential protein set, a combined analysis using single-target and multi-target strategies can be employed during the drug screening process.

[0054] This embodiment presents a method for predicting the efficacy of candidate drugs based on graph neural networks according to the present invention. First, a biomedical causal chain is obtained. Then, based on the causal chain, an efficacy prediction model is constructed and trained: an input layer is used to construct a directed biomedical graph based on the causal chain and generate triples of drug, relation, and target; a graph neural network is used to perform feature diffusion updates on the entities and relationships between entities in the directed biomedical graph, obtaining updated entities and relationships between entities; an output layer is used to score the triples based on the updated entities and relationships between entities to output the efficacy prediction results. Next, after model construction, target differentially expressed proteins are identified, and relationships are determined based on the type of differentially expressed proteins. Finally, using differentially expressed proteins as tail entities, relationships as edges, and candidate drugs as head entities in a loop, triples are constructed, input into the prediction model, and the efficacy prediction results for each candidate drug are output.

[0055] The key aspects of this invention are: 1. Based on causal chains, drugs, targets, and diseases are uniformly modeled as semantically rich and fully expressive directed graphs, which conforms to the directed multi-level association nature of biomedical systems, helps to reveal complex disease mechanisms, and clearly defines the relationships between entities (such as "drug-target", "target-disease", "drug-disease", etc.), providing interpretable structured prior knowledge for subsequent modeling; 2. Based on graph neural networks, through message passing... (Passing) Iteratively propagates the association features of drugs, targets, and diseases in the knowledge graph to achieve multi-hop reasoning. Simultaneously, it explicitly utilizes "relationship type, polarity, and direction" as edge attributes to enhance the model's understanding of biological semantics. It performs feature diffusion on entities and relationships within the knowledge graph. The key difference lies in the fact that feature diffusion is a diffusion of features across the graph structure, a diffusion of overall relationships and causal chains, rather than an update of individual features. This is its biggest difference from existing technologies. The result is a predictive model that takes the relationship chain of head entity, relation, and tail entity as input and the drug efficacy result as output. 3. The unique encoding of entities and relations allows learning the association relationships between different entities without extracting complex individual features of each entity. This invention effectively addresses the feature sparsity problem faced by large-scale data. 4. Based on this, it identifies the differentially expressed protein set for diseases; filters out proteins significantly related to the disease, removes irrelevant signals, reduces graph redundancy, and improves model focus; it uses differentially expressed proteins as tail entities, relationships as edges, and iteratively traverses candidate drugs as head entities to construct model input, inputting it into the prediction model and outputting the efficacy prediction results for each candidate drug. It can batch-traverse all candidate drugs, achieving high-throughput virtual screening and significantly improving efficiency. Simultaneously, each prediction result can be traced back to a specific (drug → relationship → protein) path, supporting mechanism explanation and experimental validation design. Furthermore, it can generate personalized efficacy predictions for differentially expressed protein profiles of different patients / subtypes, moving towards precision medicine. Overall, this invention solves the problem that existing technologies cannot efficiently integrate and mine the multidimensional relationships between drugs, targets, and diseases, constructing an interpretable, scalable, and updatable knowledge network to reduce the difficulty of drug discovery and shorten the research and development cycle.

[0056] Preferably, after obtaining the efficacy prediction results of each candidate drug, the method further includes: S5: Convert the drug efficacy prediction results into a probability representation to obtain the optimized drug efficacy prediction results.

[0057] Specifically, since the efficacy prediction result obtained in step S4 may be any real number (such as -4.3 ~ 6.7), and the scale of different batches and different disease models is inconsistent, it can be converted into a probability representation and compressed into the [0,1] interval to eliminate the difference in dimensions and distribution, so as to obtain the optimized efficacy prediction result, which makes the efficacy prediction results of different types of drugs directly comparable and more intuitive than the output of the original model.

[0058] For example, targeting a specific target and relationships N candidate drugs The efficacy score was Optionally, the correlations of all drugs can be normalized to probabilities using softmax, as shown in Equation 3-5: 3-5 On the other hand, the present invention also provides a candidate drug selection method based on graph neural networks, comprising: P1: Identify the target differentially expressed protein set for the current disease; Specifically, the same method as step S3 can be used to determine the set of target differentially expressed proteins for the current disease. This can be optional, but not limited to, determining how many elements are included in the set, i.e., how many target differentially expressed proteins the corresponding disease includes. This provides a logical basis for subsequent steps, ensuring that single-target diseases use traditional single-target optimization strategies, while multi-target diseases use multi-target synergistic strategies, thus improving R&D efficiency.

[0059] P2: If the target differential protein set contains only one element, then use any of the above methods to determine the efficacy prediction results of each drug for the target, and rank the drugs accordingly, and select several candidate drugs in sequence. Specifically, when the corresponding disease contains only one differentially expressed protein, the efficacy prediction results of various drugs for the differentially expressed protein can be obtained by using a candidate drug efficacy prediction method based on graph neural networks. The drugs are then ranked according to the efficacy prediction results, and several candidate drugs are selected in sequence as the final drug set for further screening or use by staff.

[0060] P3: If the target differential protein set includes several elements, then any of the above methods are used to determine the efficacy prediction results of each drug for each target, and the weighted sum is used to rank the drugs, and several candidate drugs are selected in sequence.

[0061] Specifically, when a corresponding disease involves multiple differentially expressed proteins, weights can be assigned to each target based on its specific location or density in the human body. The efficacy prediction results for each target are then weighted and summed to obtain the final score for each drug. Finally, the drugs are ranked according to the final score, and several candidate drugs are selected in order to form the final drug set for further screening or use by the staff.

[0062] For example, in multi-target screening, drugs can be comprehensively evaluated by combining differentially expressed protein information related to each target. First, for each target, the top 50 candidate drugs with the highest scores are screened, and duplicates are removed to obtain a set of drugs. Then, the predicted scores of the drugs are weighted according to the importance weight of the differentially expressed proteins. For a certain candidate drug d, its comprehensive score can be calculated as shown in equation (6): (6) in, The weights of the targets are represented by a weighting based on the fold change in differentially expressed proteins, reflecting their importance in the disease progression. The prediction score reflects the drug's performance. It reflects the amount of protein that is upregulated or downregulated.

[0063] Preferably, since the expression of different differentially expressed proteins may differ or even be opposite in the same disease state—for example, some proteins may be abnormally overexpressed (upregulated proteins), while others may be underexpressed or have low function (downregulated proteins)—it is necessary to differentiate them when selecting drugs for targeted treatment. Preferably, step P3 also includes: P31: Based on the drug efficacy prediction results, the drugs corresponding to each target are ranked, and several candidate drugs are selected in sequence. P32: Differentially regulated proteins are divided into upregulated proteins and downregulated proteins, resulting in a drug set for upregulated proteins and a drug set for downregulated proteins; Specifically, differentially regulated proteins can be categorized into upregulated proteins and downregulated proteins. Based on the candidate drugs obtained in step P31, an upregulated protein drug set and a downregulated protein drug set can be obtained to facilitate targeted treatment of the upregulated and downregulated proteins.

[0064] For example, inhibitors (such as antagonists or degraders) can be selected to reduce the activity or expression level of upregulated proteins, while agonists or expression enhancers can be selected to restore the function or expression level of downregulated proteins. This ensures that the direction of drug action is consistent with the molecular mechanism of the disease and avoids "reverse regulation" that leads to poor efficacy or side effects. At the same time, complex diseases (such as cancer and neurodegenerative diseases) often involve multiple target abnormalities. By classifying drugs according to the direction of target regulation, synergistic combination strategies can be systematically designed. For example, simultaneously inhibiting oncoproproteins (upregulation) and activating tumor suppressor proteins (downregulation) can achieve "bidirectional regulation" to enhance efficacy.

[0065] P33: Obtain the fold change in expression of each protein under disease and normal conditions, and calculate the weighted average of the drug efficacy prediction results and the absolute value of the fold change to obtain the comprehensive score of each drug; Specifically, since the degree of protein change can reflect the severity of the disease to some extent, the expression change fold of each protein under disease and normal conditions can be obtained as a weight for each drug, and the drug efficacy prediction results are weighted and averaged with the absolute value of the change fold to obtain the comprehensive score Sd for each drug.

[0066] P34: Sort the drugs according to their overall scores, and select a number of drugs from the sets of drugs that upregulate protein and those that downregulate protein in sequence to form the final drug set.

[0067] Specifically, from the sets of drugs that upregulate proteins and the sets of drugs that downregulate proteins, several candidate drugs with the highest scores can be selected to obtain drug sets that act on upregulated proteins and drug sets that act on downregulated proteins, respectively. These sets of drugs provide a reliable theoretical basis for subsequent experimental verification.

[0068] This embodiment presents a candidate drug selection method based on graph neural networks. By determining whether a corresponding disease contains multiple differentially expressed proteins, a branching logic is provided for subsequent steps. Early differentiation of disease complexity (single protein vs. multiple proteins) simplifies the subsequent drug screening process and improves R&D efficiency. If not, drugs are ranked according to efficacy prediction results, and several candidate drugs are selected sequentially as the final drug set. Since only a single target needs to be considered, there are fewer interfering factors in drug screening, so the top-ranked drugs can be directly selected as candidates, shortening the drug development cycle. If yes, the efficacy prediction results for each target are weighted and summed to obtain the final score for each drug. Several candidate drugs are then ranked sequentially as the final drug set. By weighting multi-dimensional efficacy data (such as activity, selectivity, and toxicity prediction), a more robust drug score is obtained. This improves the R&D efficiency of existing technologies and reduces R&D costs.

[0069] For example, taking drug retargeting screening for CryAB-mutant cardiomyopathy as an example, a target-drug interaction network can be constructed based on 27 differentially expressed proteins obtained from the screening. The network construction rules are as follows: for 16 upregulated proteins, establish "downregulation" relationship edges with potential inhibitors; for 11 downregulated proteins, establish "upregulation" relationship edges with potential activators. Using the above-mentioned candidate drug efficacy prediction method based on graph neural networks, the regulatory matching degree of drug molecules to the target network is calculated through a pre-trained vector space model (using marginal loss as the loss function). Specifically, for each upregulated protein, compounds that can significantly inhibit its function are predicted; for each downregulated protein, compounds that can effectively activate its activity are predicted. Finally, the top 20 inhibitors with the highest matching degree are screened, such as... Figures 3 to 4 As shown: Optional deep learning-based multi-target drug relocation screening. Weighted ranking of upregulated protein inhibitors: Based on the TransE deep learning model's transfer learning of known drug-target relationships, the top 20 candidate drugs that can inhibit differentially upregulated proteins are screened, with efficacy scores reflecting targeted regulatory potential; Weighted ranking of downregulated protein activators: For differentially downregulated proteins, the scores of the top 20 drugs that can promote their function are calculated, with bar chart height representing the strength of multi-target synergistic effects.

[0070] like Figure 5 As shown in Table 3, there are 11 overlapping drugs between upregulating protein inhibitors and downregulating protein activators. Further analysis reveals that these overlapping drugs include chemotherapeutic agents and antiparasitic drugs, whose effects may simultaneously correct protein expression imbalances through multi-target effects. Notably, raloxifene (a selective estrogen receptor modulator) and digoxin were both predicted as key co-regulatory drugs. The former may affect lipid metabolism through a synergistic effect of PPARγ mediated by estrogen receptor β, while the latter may enhance myocardial contractility by altering intracellular calcium homeostasis, which compensates for the previously discovered inhibitory phenotype of the contractile pathway.

[0071] Table 3: Intersecting Drugs Table

[0072] In this specific example, the present invention employs the aforementioned drug efficacy prediction method, constructing a biomedical directed graph by integrating multi-source biomedical databases, and combining it with graph neural network technology to successfully develop a drug efficacy prediction model for CryAB mutation-related dilated cardiomyopathy (DRC). This model is based on the PyTorch framework and efficiently trained using an NVIDIA A100 GPU, enabling rapid discovery of potential therapeutic drugs. The training process uses a batch size of 32 and a total of 20 training rounds. After each training round, the validation set is evaluated to monitor model performance and prevent overfitting. Compared to traditional drug screening methods, this invention significantly improves the efficiency and accuracy of drug discovery, accurately identifying candidate drugs that can intervene in abnormal CryAB protein aggregation, alleviate cardiomyocyte stress damage, and energy metabolism disorders. Its beneficial effects lie in providing a novel targeted treatment strategy for CryAB-mutant DRC, for which there is currently no cure, potentially promoting the development of personalized treatment, improving patient prognosis, reducing the risk of heart failure progression, and possessing significant clinical translational value and social significance.

[0073] On the other hand, the present invention also provides a computer storage medium storing executable program code; the executable program code is used to execute any of the above-mentioned candidate drug efficacy prediction methods or selection methods based on graph neural networks.

[0074] On the other hand, the present invention also provides a terminal device, including a memory and a processor; the memory stores program code that can be executed by the processor; the program code is used to execute any of the above-mentioned candidate drug efficacy prediction methods or selection methods based on graph neural networks.

[0075] For example, the program code can be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the program code in the terminal device.

[0076] The terminal device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the terminal device may also include input / output devices, network access devices, buses, etc.

[0077] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0078] The memory can be an internal storage unit of the terminal device, such as a hard drive or RAM. The memory can also be an external storage device of the terminal device, such as a plug-in hard drive, SmartMedia Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, the memory can include both internal and external storage units of the terminal device. The memory is used to store the program code and other programs and data required by the terminal device. The memory can also be used to temporarily store data that has been output or will be output.

[0079] The above-mentioned candidate drug selection method, computer storage medium, and terminal device are created based on the above-mentioned candidate drug efficacy prediction method. Their technical functions and beneficial effects will not be elaborated here. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0080] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.

Claims

1. A method for predicting the efficacy of candidate drugs based on graph neural networks, characterized in that, include: To obtain the causal chain of biomedicine; Based on the causal chain, a drug efficacy prediction model is constructed and trained, including: an input layer, used to construct a biomedical directed graph based on the causal chain and generate drug, relationship, and target triples; the nodes of the biomedical directed graph include at least three medical entities: drug, target, and disease; edges represent the relationships between entities; a graph neural network is used to perform feature diffusion updates on the entities and relationships between entities in the biomedical directed graph to obtain updated entities and relationships between entities; and an output layer is used to score the triples based on the updated entities and relationships between entities to output the drug efficacy prediction results. Identify the target differentially expressed protein and determine the relationship based on the type of the target differentially expressed protein; By constructing triplet pairs using differentially expressed proteins as tail entities, relations as edges, and candidate drugs as head entities through iterative iteration, the prediction results for each candidate drug are output as the prediction results for each candidate drug.

2. The method for predicting the efficacy of candidate drugs according to claim 1, characterized in that, The input layer includes node embedding units and edge embedding units to construct a biomedical directed graph; Node embedding unit, used to assign a unique identifier to each medical entity according to category, in order to define the nodes of the biomedical directed graph; Edge embedding units are used to extract different relationship categories, relationship properties, and relationship flows between entities based on causal chains, in order to define the types, polarities, and directions of different edges and construct a directed graph for biomedicine.

3. The method for predicting the efficacy of candidate drugs according to claim 2, characterized in that, The input layer includes a triple generator to generate drug-relationship-target triples; including: Positive sample input unit, used to input positive samples of triplet of drug, relationship, and target; The negative sample generation unit is used to replace the head entity and / or tail entity of the positive sample to generate a negative sample.

4. The method for predicting the efficacy of candidate drugs according to claim 3, characterized in that, The positive sample input unit is also used for: The baseline strength of each positive sample is determined based on its data source. The additional strength of each positive sample is determined based on the number of data points for each positive sample. The strength of evidence for each positive sample is determined by comprehensively considering its basic strength and additional strength. Based on the strength of evidence for each positive sample, a probability bias mechanism is introduced, so that the probability of a positive sample with high evidence strength is higher than the probability of a positive sample with low evidence strength.

5. The method for predicting the efficacy of candidate drugs according to claim 4, characterized in that, The negative sample generation unit is also used for: Based on the strength of evidence of positive samples, a quantity bias mechanism is introduced, so that the number of negative samples generated by positive samples with high evidence strength is greater than the number of negative samples generated by positive samples with low evidence strength.

6. The method for predicting the efficacy of candidate drugs according to claim 1, characterized in that, Graph neural networks are used for: Several nodes and edges in a biomedical directed graph are converted into initial embedding vectors using unique identifiers, resulting in initial node embedding vectors and initial edge embedding vectors, thus enabling sparse feature modeling. The information of adjacent nodes and edges is aggregated, and the initial node embedding vector and the initial edge embedding vector are updated layer by layer to obtain the updated node embedding vector and edge embedding vector.

7. The method for predicting the efficacy of candidate drugs according to claim 1, characterized in that, Output layer, used for: Based on the updated entities and the relationships between them, a boundary loss function is constructed to maximize the score difference between positive and negative samples. Solve the boundary loss function to obtain the drug efficacy prediction results.

8. The method for predicting the efficacy of candidate drugs according to any one of claims 1 to 7, characterized in that, Relationships are determined based on the type of the target differentially expressed protein, including: Depending on the type of the target differentially expressed protein, negative relationships are used as edges for highly expressed proteins and positive relationships are used as edges for low-expressed proteins.

9. A candidate drug selection method based on graph neural networks, characterized in that, include: Identify the target set of differentially expressed proteins for the current disease; If the target differential protein set includes only one element, the method described in any one of claims 1-8 shall be used to determine the efficacy prediction results of each drug on the target, and the drugs shall be ranked accordingly, and several candidate drugs shall be selected in sequence. If the target differential protein set includes several elements, the method described in any one of claims 1-8 is used to determine the efficacy prediction results of each drug for each target, and the weighted sum is used to rank the drugs, and several candidate drugs are selected in sequence.

10. The candidate drug selection method according to claim 9, characterized in that, Also includes: Based on the efficacy prediction results, the drugs corresponding to each target are ranked, and several candidate drugs are selected in sequence. The target differentially expressed proteins are divided into upregulated proteins and downregulated proteins, resulting in a set of drugs for upregulated proteins and a set of drugs for downregulated proteins. The expression fold change of each target differential protein under disease and normal conditions was obtained, and the drug efficacy prediction results were weighted and averaged with the absolute value of the fold change to obtain the comprehensive score of each drug. The drugs were sorted according to their overall scores, and candidate drugs were selected from the sets of drugs that upregulate and downregulate proteins in sequence.

Citation Information

Patent Citations

  • Drug target effect deep learning prediction system based on knowledge graph, computer equipment and storage medium

    CN112562791A

  • Multi-task drug screening method and system based on knowledge graph assistance

    CN114420221A

  • Hypergraph-based drug-target-disease interaction prediction method

    CN113066526A

  • Metalearning drug-target interaction prediction system and prediction method

    CN113140254A

  • Drug ATCCode prediction method based on graph transformation network

    CN114420310A