An ad complex target prediction method based on a knowledge graph

By constructing a knowledge graph of the disease-symptom-prescription association of Alzheimer's disease and an improved restart random walk algorithm, and integrating multi-source data, the problem of traditional drug target prediction methods failing to fully consider the biological context is solved, and efficient prediction and enhanced interpretability of Alzheimer's disease compound targets are achieved.

CN119864128BActive Publication Date: 2025-11-07GUANGZHOU UNIVERSITY OF CHINESE MEDICINE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411924667.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-11-07
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

Traditional drug target prediction methods fail to fully consider the biological context of drug-target relationships, resulting in low prediction accuracy, especially when traditional Chinese medicine compound treatments for Alzheimer's disease fail to reveal the mechanism of action.

Method used

We employ a knowledge graph-based approach to construct a knowledge graph of the disease-symptom-prescription associations for Alzheimer's disease. By combining a large model and an improved restart random walk algorithm, we integrate multi-source data, capture the complex interactions between drugs and targets, and enhance interpretability through visualization graphs.

Benefits of technology

It achieves efficient prediction of compound targets for Alzheimer's disease, takes into account biological context factors, improves prediction accuracy and enhances interpretability, and is applicable to the biomedical field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119864128B_ABST
    Figure CN119864128B_ABST
Patent Text Reader

Abstract

The application discloses an AD compound target prediction method based on a knowledge graph, and relates to the biological medicine field.The method comprises the following steps: establishing a basic structure of a disease-syndrome-recipe correlation knowledge graph of Alzheimer's disease; extracting information from text data based on a data model to obtain data features; constructing a knowledge graph of Alzheimer's disease based on the data features; extracting features from the knowledge graph based on a preset model to obtain graph features; and predicting in the knowledge graph based on the graph features by using an improved restart random walk algorithm to obtain potential targets. By using the application, biological context factors of drug-target relationships can be comprehensively considered, and the target prediction accuracy is improved. The application can be widely applied in the biological medicine field.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of biological medicine, and in particular to an AD compound target prediction method based on a knowledge graph. BACKGROUND

[0002] Alzheimer's disease (AD) is a common neurodegenerative disease that severely affects the cognitive function and quality of life of patients. Currently, there are limited treatment options for AD, and drug development faces many challenges.

[0003] With the widespread use of traditional Chinese medicine (TCM) worldwide, especially in the treatment of complex diseases such as AD, TCM compounds have attracted much attention due to their unique multi-component, multi-target, and multi-pathway regulation mechanisms. However, this complexity also makes it difficult to predict the targets and evaluate the efficacy of TCM compounds. Traditional experimental methods are often time-consuming and labor-intensive, and are difficult to fully reveal the pharmacological mechanisms of TCM compounds.

[0004] Traditional drug target prediction methods mainly rely on a single data source, making it difficult to fully capture the complex interactions between drugs and targets. Many traditional methods only focus on the direct interactions between drugs and targets, ignoring the biological context in which these interactions occur, such as cellular environment, signaling pathways, and molecular functions, resulting in low prediction accuracy. SUMMARY

[0005] In view of the above, in order to solve the technical problem that the biological context factors of drug-target relationships are not fully considered in existing drug target prediction methods, resulting in low prediction accuracy, the present application proposes an AD compound target prediction method based on a knowledge graph, which comprises the following steps:

[0006] Information extraction is performed on the text data based on a data model to extract information such as target nodes, relationships, and attributes, obtaining data features;

[0007] A basic structure of a disease-syndrome-recipe association knowledge graph for Alzheimer's disease is constructed, and a knowledge graph is constructed in combination with the data features;

[0008] Feature extraction is performed on the knowledge graph based on a preset model to capture complex interactions between "drugs and targets", obtaining graph features;

[0009] Based on the graph features, an improved restart random walk algorithm is used to make predictions on the knowledge graph, and potential targets are obtained based on node scores.

[0010] In some embodiments, in the step of performing feature extraction on the knowledge graph based on a preset model to capture complex interactions between "drugs and targets" and obtaining graph features, it specifically comprises:

[0011] Encode the entity relationship in the knowledge graph, and embed the knowledge graph into a machine-recognizable vector form;

[0012] The LLM model is used to score the "gene" (target) similarity, the "protein interaction correlation" strength score between "genes", and the affinity of "traditional Chinese medicine ingredients-genes" in the embedded knowledge graph, and the obtained scoring results are output in the form of a matrix.

[0013] Meanwhile, the traditional Chinese medicine ingredients are divided into multiple communities.

[0014] In some embodiments, based on the graph features, a modified restart random walk algorithm is used to make predictions in the knowledge graph, and the step of obtaining potential targets according to node scores includes:

[0015] Random walk simulation is performed on the constructed knowledge graph, the starting node of the improved restart random walk algorithm is selected, and the related parameters are initialized;

[0016] A transition matrix is constructed based on the relationship and transition probability between the community nodes divided based on the knowledge graph, and the transition matrix is used to reflect the accuracy of the node relationship and simulate the transition mode between nodes;

[0017] Using the constructed transition matrix, random walk is performed according to the set transition probability and restart probability to simulate the transmission process of the nodes and explore the network structure, record the node access times, and obtain the steady-state distribution result;

[0018] Based on the steady-state distribution result, the node score is calculated, and the potential target is predicted according to the node score.

[0019] Based on the above scheme, the present application provides an AD compound target prediction method based on a knowledge graph, which integrates multiple data sources, uses a large model, a knowledge graph and an improved random walk algorithm, and realizes efficient prediction of AD compound targets. Compared with traditional methods, the data integration angle of the present application is more comprehensive, not only considering the drug-target relationship, but also considering the biological context factors, which can more comprehensively capture the complex interaction between drugs and targets; at the same time, the visual graph is used to enhance the explainability, which can be better applied in the field of biological medicine. The present application provides a new perspective and strategy for AD compound target prediction method. BRIEF DESCRIPTION OF DRAWINGS

[0020] Figure 1 is an implementation flowchart of an AD compound target prediction method based on a knowledge graph of the present application;

[0021] Figure 2 is a step flowchart of an AD compound target prediction method based on a knowledge graph of the present application

[0022] Figure 3 is a process diagram of extracting information to construct a knowledge graph.

[0023] Figure 4 is a relationship diagram of data in the knowledge graph.

[0024] Figure 5 is a flow diagram of feature extraction.

[0025] Figure 6 is a flow diagram of improved random walk prediction. DETAILED DESCRIPTION

[0026] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0027] It should be noted that, for the convenience of description, only the parts related to the present application are shown in the drawings. The embodiments in the present application and the features in the embodiments can be combined with each other without conflict.

[0028] It should be understood that the "system", "device", "unit" and / or "module" used in the present application is a method for distinguishing different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, the words can be replaced by other expressions.

[0029] As shown in the present application and claims, unless the context clearly indicates otherwise, "one", "a", "an" and / or "the" do not refer to the singular, but also include the plural. Generally speaking, the terms "include" and "contain" only indicate the inclusion of the steps and elements clearly identified, and these steps and elements do not constitute an exclusive list, and the method or device can also include other steps or elements. The element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, product or device including the element.

[0030] In the description of the embodiments of the present application, "a plurality of" means two or more than two. The following terms "first", "second" are only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features.

[0031] In addition, flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or subsequent operations are not necessarily performed in sequence. Instead, the steps can be processed in reverse order or simultaneously. Meanwhile, other operations can be added to these processes, or one or more steps can be removed from these processes.

[0032] Figure 1 To implement the flowchart of the present application, four blocks including knowledge modeling, large model text extraction, network feature extraction, and random walk prediction are included.

[0033] Referring to Figure 2 The flowchart of an optional example of the AD compound target prediction method based on the knowledge graph proposed in the present application can be applied to a computer device. The prediction method proposed in the present application can include but is not limited to the following steps:

[0034] Step S0, establish the basic structure of the disease-syndrome-recipe association knowledge graph of Alzheimer's disease. The data sources include ancient Chinese medical books, modern medical literature, clinical data, and public gene databases, etc.

[0035] Step S1, information extraction of text data based on a data model to obtain data features;

[0036] Step S2, constructing a knowledge graph of Alzheimer's disease based on the data features;

[0037] Step S3, feature extraction of the knowledge graph based on a preset model to obtain graph features;

[0038] Step S4, based on the graph features, using an improved restart random walk algorithm to predict in the knowledge graph to obtain potential targets.

[0039] In some feasible embodiments, the step S0 specifically includes:

[0040] Construct a knowledge model. Through pharmacological analysis, establish the basic structure of the disease-syndrome-recipe association knowledge graph of AD. The data sources are explained as shown in Table 1. Define the structure, nodes and relationships of knowledge, thereby establishing the initial framework of the process, defining the basic structure, and laying the foundation for large model text extraction. The entity structure and relationship structure of the knowledge graph are shown in Tables 2 and 3.

[0041] Table 1 Data source explanation

[0042]

[0043] Table 2 Knowledge graph entity structure

[0044]

[0045] Table 3 Knowledge graph relationship structure

[0046]

[0047] In some feasible embodiments, step S1 specifically includes:

[0048] Use large model prompt engineering (Prompt) to extract key information such as labels, nodes and relationships from a large amount of text data, and provide data basis for knowledge graph construction.

[0049] S1.1, prepare the basic data source, extract data from different sources (such as database, file system), and aggregate the data source. The data source contains a large number of triple relationship.

[0050] S1.2, call the corresponding large model API to prepare for information extraction of large model;

[0051] S1.3, store the collected data in the selected storage system for subsequent use.

[0052] S1.4, fine-tune the large model through the corresponding design and specific optimization prompt, train the large model to generate the required output, effectively extract and integrate traditional Chinese medicine compound, target and other related information;

[0053] S1.5, Prompt performs multiple rounds of self-optimization learning, cyclically optimizes information extraction results, adjusts output information, and obtains the final result.

[0054] Reference Figure 3 , the process diagram for data extraction and knowledge graph construction.

[0055] In some feasible embodiments, step S2 specifically includes:

[0056] S2.1, use the large model to structure and structure the extracted effective information again, and lay the foundation for knowledge graph construction.

[0057] Use text mining, cleaning and feature extraction technology.

[0058] S2.2, construct a knowledge graph based on the processed data, and establish entity relationships in the graph. Add attributes to the entity to improve the detailed information of the graph.

[0059] S2.3, periodically verify and update the data through the large model, repeatedly optimize the graph, and improve its accuracy and application efficiency. It can be executed independently, or it can be run regularly to keep the knowledge base accurate and up-to-date.

[0060] Reference Figure 4 for the relationship diagram of data in the knowledge graph.

[0061] In some feasible embodiments, step S3 specifically includes:

[0062] Referring to Figure 5 , network graph feature extraction implemented by large models.

[0063] S3.1, query node and relationship information from Neo4j and embed entities;

[0064] Including:

[0065] S3.1.1, query data;

[0066] S3.1.2, node embedding;

[0067] ① Traditional Chinese medicine ingredients, build Morgan fingerprint embedding through SMILES attribute:

[0068] z Morgan = MorganFingerprint(SMILES, radius=2, nBits=2048)

[0069] Where radius is a parameter used when building Morgan fingerprint, which determines the range of considering atomic neighborhood. Radius of 2 means considering the relationship between atoms and their two "levels" adjacent atoms. nBits is the number of bits of Morgan fingerprint, which determines the length of the fingerprint vector. 2048 bits means that the generated Morgan fingerprint is a 2048-dimensional binary vector.

[0070] ② Gene, ESM-2 is a dedicated protein sequence pre-training model, which can generate high-dimensional embedding (650M model outputs 1280-dimensional embedding), protein embedding based on gene sequence:

[0071] z Gene = ESM-2(Sequence)

[0072] Where Sequence is the input gene sequence, which can be any protein or gene sequence. ESM-2 will convert the gene sequence into a high-dimensional vector representation (for example, 1280 dimensions), which captures the biological function and structural information of the gene.

[0073] ③ Molecular function, biological process, cellular component, Onto2Vec embedding Gene Ontology (GO) node:

[0074] Z GO = Onto2Vec(GO-Annotations)

[0075] where G0-Annotations is the input gene annotation information, describing the function, process or cellular location of the gene. Onto2Vec model converts these annotations into embedding vectors, helping to learn the similarity between GO annotations.

[0076] ④AD endophenotypes, Chinese medicine, AD prescriptions, using BioBERT embedding descriptive text information:

[0077] z Entity = BioBERT(Description)

[0078] where Description is the input descriptive text, containing description information related to drugs, diseases, genes, etc. BioBERT model converts text into a vector that can capture professional information and relationships in the biomedical field. The output embedding vector is a feature representation for subsequent tasks.

[0079] S3.2, TransE embeds all relationships, including embedding hypotheses and optimization objectives.

[0080] S3.3, Constructing multi-modal input to large model Llama.

[0081] S3.4, Input embedding mapping, embedding multi-modal input into a unified mapping to the LLama embedding space.

[0082] S3.5, use LLM model to score preset indicators in embedded knowledge graph, get scoring results, the preset indicators include gene similarity, protein interaction correlation strength between genes and affinity of Chinese medicine ingredients-genes.

[0083] Gene similarity calculation:

[0084] According to the embedding of molecular function, biological process and cellular component, the gene similarity is calculated:

[0085]

[0086] where, represents the embedding vector of gene g i and g j . These embeddings are generated by large models (such as ESM-2, Onto2Vec), representing the function, structure, etc. of the gene. embedding vector of gene g i , capturing the biological characteristics of the gene. embedding vector of gene g j , capturing the biological characteristics of the gene.

[0087] Protein interaction strength between genes:

[0088] w'(g i , g j ) = w(g i , g j ) + β · Sim(g i , g j )

[0089] where w(g i , g j ) is the original protein interaction strength between genes g i and g j . β is a tuning factor that controls the degree of enhancement of the similarity Sim(g i , g j ) to the interaction strength.

[0090] Affinity between traditional Chinese medicine ingredients and genes:

[0091] Construct multi-modal features through the affinity values on existing paths, and complete the affinity of all paths through a large model:

[0092]

[0093] where, is the predicted affinity value, representing the interaction strength between traditional Chinese medicine ingredients and genes (e.g. the binding affinity of a drug to a target gene). Z Drug is the embedding vector of the traditional Chinese medicine ingredient. Z Gene is the embedding vector of the gene. Z Relation represents the embedding of the relationship between traditional Chinese medicine ingredients and genes.

[0094] The fine-tuning target is the mean square error loss:

[0095]

[0096] where, is the affinity value predicted by the model. y i is the true affinity value. That is, in known experimental data, the affinity value between genes and traditional Chinese medicine ingredients. N: the total number of samples, i.e. the number of samples in the data set for which the traditional Chinese medicine ingredient-target affinity is known.

[0097] Also includes: dividing traditional Chinese medicine ingredient communities by K-Means.

[0098] In some feasible embodiments, step S4 specifically includes:

[0099] Referring to Figure 6 , the process schematic diagram of the random walk simulation algorithm.

[0100] S4.1, random walk simulation is performed on the constructed AD disease-syndrome-recipe association knowledge graph, the starting node of the restart random walk algorithm is selected, and the related parameters are initialized.

[0101] The disease-syndrome-recipe association knowledge graph of AD can be represented by a set G = [G1, G1,..., G d ], wherein G is a multi-relation network composed of multiple relation sub-networks, G = [V, E], wherein V represents a set of nodes, and E represents a set of different types of edges. The node G i , G j The edge between nodes G i , G j ) is denoted as (G i , G i ), and the neighbor node set of node G i is denoted as Γ(i), and the degree of G i is denoted as k j .

[0102] S4.2, the relationship and transition probability between the community nodes divided based on the knowledge graph are used to construct the transition matrix;

[0103] On the basis of MHRW (an unbiased sampling algorithm MHRW proposed by Gjoka et al. in combination with RW), an improved algorithm MHRWER (Metropolis-Hasting Random Walk with Extended Restart) is proposed. The influence of the degree of the current node and its neighbor nodes on the transition probability is considered, the path weight is assigned to the neighbor nodes based on the path access history, and the global importance of the global node is considered. By dynamically adjusting the path weight, the ability to explore unknown areas is enhanced while reducing repeated access to known areas.

[0104] The transition probability defined by MHRWER is:

[0105]

[0106] wherein k i represents the degree of node i (i.e. the number of edges adjacent to node i), which is used to measure the connectivity of the node, G j is a node label, indicating that node j is one of the target nodes that node i can transfer to when starting from node i in the random walk process, and F(i) represents the neighbor set of node i. w ij is the path weight, and the calculation formula is as follows:

[0107]

[0108] x ij is the path access frequency of node i to node j, and π ijAffinity or protein interaction correlation strength between Traditional Chinese Medicine ingredient i and gene j (according to different path conditions, corresponding to the results calculated in the previous step).

[0109] S4.3, based on the transition matrix, random walk is performed according to the transition probability and the restart probability, the transfer process of the nodes is simulated and the network structure is explored, and the node access frequency is recorded.

[0110] RWR (Random Walk with Restart) is a basic random walk similarity index. It is assumed that the probability of a particle randomly selecting the next neighbor node for transition is α, and the probability of returning to the initial node is 1-α. According to this assumption, the transition probability matrix P of the network can be defined, where the elements P ij represent the probability of transition from node i to node j. The probability vector of the particles in the next step of transition to other nodes is as follows:

[0111]

[0112] where, represents the access probability vector of node i. c represents the restart probability. represents the weighted transition matrix, representing the probability of transition from one node to other nodes. represents the access probability vector of the current node. is a unit vector, representing the probability distribution of the initial state.

[0113] Iterative update w ij , dynamically adjust the node transition probability.

[0114] S4.4, calculate the steady state distribution, i.e. the stable probability distribution of the nodes.

[0115] Steady state distribution calculation formula

[0116] y = (1-μ)(I-μM) -1 y0

[0117] where μ is the restart parameter, representing the probability of returning to the initial node from the current state. y0 is the initial probability distribution. The transition matrix M is defined as follows:

[0118]

[0119] W is a weight matrix based on path weight and global importance adjustment, and the element is w ij . A is the adjacency matrix, representing the direct connection relationship between nodes. D is a diagonal matrix, defined as:

[0120]

[0121] This ensures that M is a normalized transition probability matrix.

[0122] In the iteration process, the transition probability matrix M is updated step by step, and it is judged whether the convergence condition of the algorithm or the preset maximum iteration number is reached, if yes, the final stationary distribution probability y is calculated. Otherwise, return to the calculation of path weight for iterative update.

[0123] The steady state distribution is applied to the actual knowledge graph analysis for identifying important nodes or analyzing the characteristics of the knowledge graph.

[0124] S4.5, based on the steady state distribution, in-depth analysis of network characteristics, calculate the score of each node, evaluate its possibility as a potential target.

[0125] According to the score ranking, the high-potential target nodes are screened out.

[0126] In some possible embodiments, further comprising:

[0127] With the help of visualization tools, the random walk path is displayed, and the potential target probability is visualized in the form of a heat map, etc. to explain the results, so that it can better play a role in the field of biology and medicine.

[0128] An AD compound target prediction system based on a knowledge graph, comprising:

[0129] A text processing module extracts information from text data based on a data model to obtain data features;

[0130] A graph construction module constructs a knowledge graph of Alzheimer's disease based on the data features;

[0131] A feature extraction module extracts features from the knowledge graph based on a preset model to obtain graph features;

[0132] A random walk module uses an improved restart random walk algorithm to predict in the knowledge graph based on the graph features to obtain potential targets.

[0133] The contents in the above method embodiments are all applicable to the system embodiments, the system embodiments specifically realize the same functions as the above method embodiments, and achieve the same beneficial effects as the above method embodiments.

[0134] An AD compound target prediction device based on a knowledge graph:

[0135] At least one processor;

[0136] At least one memory for storing at least one program;

[0137] When the at least one program is executed by the at least one processor, the at least one processor implements the AD compound target prediction method based on a knowledge graph as described above.

[0138] The contents in the method embodiments are applicable to the device embodiments, the device embodiments specifically implement the same functions as the method embodiments, and achieve the same beneficial effects as the method embodiments.

[0139] A storage medium, wherein the storage medium stores processor-executable instructions, and the processor-executable instructions, when executed by a processor, are used to implement the AD compound target prediction method based on a knowledge graph as described above.

[0140] The contents in the method embodiments are applicable to the storage medium embodiments, the storage medium embodiments specifically implement the same functions as the method embodiments, and achieve the same beneficial effects as the method embodiments.

[0141] The above is a specific description of the preferred embodiments of the application, but the application is not limited to the embodiments described above. Those skilled in the art can make various equivalent modifications or replacements without departing from the spirit of the application. These equivalent modifications or replacements are all included in the scope defined by the claims of the present application.

Claims

1. A method for predicting AD compound targets based on a knowledge graph, characterized in that, The method comprises the following steps: information extraction is performed on the text data based on a data model to obtain data features; a knowledge graph of Alzheimer's disease is constructed based on the data features; feature extraction is performed on the knowledge graph based on a preset model to obtain graph features; based on the graph features, an improved restart random walk algorithm is used to perform prediction on the knowledge graph to obtain potential targets; the step of performing feature extraction on the knowledge graph based on the preset model to obtain graph features specifically comprises: encoding and embedding entity relationships in the knowledge graph into vector form; scoring preset indicators in the embedded knowledge graph using an LLM model to obtain a scoring result; the preset indicators include gene similarity, protein interaction correlation strength between genes, and affinity between traditional Chinese medicine components and genes; gene similarity is calculated according to embedding of molecular functions, biological processes, and cellular components: wherein, represents the embedding vector of the gene g i and g j . calculation of protein interaction correlation strength between genes: w'(g i ,g j ) = w(g i ,g j ) + β · Sim(g i ,g j ) where w(g i ,g j ) is the original protein-protein interaction strength between genes g i and g j ; β represents a regulatory factor that controls the degree of enhancement of the similarity Sim(g i ,g j ) to the interaction strength; calculation of affinity between traditional Chinese medicine components and genes: wherein, is a predicted affinity value representing the interaction strength between the traditional Chinese medicine ingredient and the gene; Z Drug is an embedding vector of the traditional Chinese medicine ingredient; Z Gene is an embedding vector of the gene; Z Relation represents the embedding of the relationship between the traditional Chinese medicine ingredient and the gene.

2. The AD compound target prediction method based on a knowledge graph according to claim 1, characterized in that, further comprising: displaying random walk paths and probabilities of potential targets. 3.The AD compound target prediction method based on a knowledge graph according to claim 1, wherein, The step of performing information extraction on the text data based on a data model to obtain data features specifically comprises: acquiring text data from a data source; calling and setting an API of a data model; performing information extraction on the text data based on the data model to obtain data features; storing the data features in a graph database; the data features include nodes, relationships, and attributes. 4.The AD compound target prediction method based on a knowledge graph according to claim 1, wherein, further comprising: dividing traditional Chinese medicine components into communities.

5. The AD compound target prediction method based on a knowledge graph according to claim 1, characterized in that, The step of performing prediction on the knowledge graph based on the graph features using an improved restart random walk algorithm to obtain potential targets specifically comprises: performing random walk simulation on the knowledge graph, selecting a starting node and initializing related parameters; constructing a transition matrix based on relationships and transition probabilities between community nodes divided based on the knowledge graph; based on the transition matrix, performing random walk according to transition probability and restart probability, simulating the transmission process of nodes and exploring the network structure, recording the number of node visits, and generating a steady-state distribution result; calculating node scores based on the steady-state distribution result, and predicting potential targets according to the node scores.

6. The AD compound target prediction method based on a knowledge graph according to claim 5, characterized in that, The formula of the transition probability is as follows: where w ij is the path weight; k i denotes the degree of node i; G j is a node label indicating that node j is one of the target nodes that node i can be transferred to when starting from node i in the random walk process; and Γ(i) denotes the neighbor set of node i. 7.A knowledge graph-based AD compound target prediction system, characterized in that, A knowledge graph-based AD compound target prediction method as claimed in claim 1 is used to perform the method, comprising: a text processing module that performs information extraction on text data based on a data model to obtain data features; a graph construction module that constructs a knowledge graph of Alzheimer's disease based on the data features; a feature extraction module that performs feature extraction on the knowledge graph based on a preset model to obtain graph features; a random walk module that performs prediction on the knowledge graph based on the graph features using an improved restart random walk algorithm to obtain potential targets. 8.A device for predicting AD compound targets based on a knowledge graph, characterized in that, comprise: at least one processor; at least one memory for storing at least one program; when the at least one program is executed by the at least one processor, the at least one processor implements a knowledge graph-based AD compound target prediction method as claimed in any one of claims 1-6.

Citation Information

Patent Citations

  • Drug target prediction method based on multiple similarity network walk

    CN108520166A

  • Drug target prediction method based on multi-channel graph convolutional neural network

    CN113053457A