Rare disease drug redirection path mining method based on knowledge graph

By constructing a knowledge graph and combining it with R-GCN and a large language model, multi-source data is integrated to mine redirection paths for rare disease drugs. This solves the problems of data dispersion and structural heterogeneity in rare disease drug development and improves the reliability and interpretability of predictions.

CN121983166APending Publication Date: 2026-05-05DALIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DALIAN UNIV OF TECH
Filing Date
2025-12-22
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

In existing technologies, the development of rare disease drugs suffers from problems such as scattered data sources, heterogeneous structures, and difficulty in direct application, resulting in low credibility and interpretability of rare disease drug retargeting prediction results.

Method used

We construct a knowledge graph-based method for mining redirection paths of rare disease drugs. By integrating multi-source biomedical databases and clinical data, we employ entity standardization, synonym merging, and low-confidence filtering mechanisms, combined with R-GCN and a large language model, to predict disease-drug links and score pathways.

Benefits of technology

It enables efficient discovery of candidate drug pathways in rare disease drug retargeting, improves the reliability and interpretability of predictions, and provides a systematic and interpretable intelligent analysis method for rare disease drug discovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121983166A_ABST
    Figure CN121983166A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge graph-based rare disease drug redirection path mining method, which comprises the following steps of: collecting data from a public database and a clinical environment, dividing the collected data into structured data and unstructured data, constructing a preset rule to process the structured data to obtain a triple, obtaining a knowledge graph based on the triple data, and mining a rare disease drug redirection path based on the knowledge graph. Extracting a positive sample and a negative sample from the knowledge graph, training a loss function optimization model, extracting all paths conforming to the template based on the knowledge graph, and inputting the paths conforming to the template into the model to obtain path scores; the purposes of improving the consistency and reliability of the atlas and realizing explainable reasoning of the potential relationship between the disease and the drug are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of biomedical informatics and artificial intelligence, specifically to a method for mining rare disease drug retargeting pathways based on knowledge graphs. Background Technology

[0002] With the rapid development of biomedical data and artificial intelligence technologies, drug retargeting has become an important means to improve the efficiency of new drug development. Rare diseases, due to the limited number of patients and scarcity of clinical samples, suffer from long and costly traditional drug development cycles, necessitating data-driven methods to discover potential therapeutic drugs. Currently, various biomedical databases exist, containing multidimensional information on diseases, genes, proteins, drugs, and clinical trials, providing a foundation for systematic research. However, these data sources are scattered and structurally heterogeneous, making direct application difficult. Knowledge graph technology, by integrating multi-source data and constructing semantic association networks, offers a new approach to mining disease-drug relationships. Meanwhile, graph neural networks (GNNs) excel in modeling graph-structured data and can predict potential disease-drug associations through link prediction. However, existing research largely focuses on common diseases, while rare disease data is sparse and lacks sufficient evidence, resulting in low reliability and interpretability of prediction results. Therefore, a rare disease drug retargeting method based on knowledge graphs and graph neural networks is urgently needed to integrate multi-source data, optimize model performance, and improve the interpretability of results, thus supporting rare disease drug discovery.

[0003] In the present invention, R-GCN comes from SEJR SCHLICHTKRULL M, N. KIPF T, BLOEM P, etc.Modeling Relational Data with Graph Convolutional Networks[C] / / 2018 EuropeanSemantic.Web.Conference.Springer,Cham,2018:593-607.https: / / www.microsoft.com / en-us / research / publication / modeling-relational-data-with-graph-convolutional-networks / .

[0004] The meta-path template is from ZHANG ML, ZHAO BW, SU XR, etc. RLFDDA: a meta-path based graph representation learning model for drug–disease associationprediction[J].BMC.Bioinformatics,2022,23(1):516.DOI:10.1186 / s12859-022-05069-z.

[0005] The Large Language Model (LLM) is from ZHOU S, YU S. High-throughput biomedical relation extraction for semi-structured web articles empowered by large language models[J]. BMC.Medical.InformaticsandDecisionMaking,2025,25(1):351.DOI:10.1186 / s12911-025-03204-3.

[0006] DrugBank comes from WISHART DS, KNOX C, GUO AC, etc. DrugBank: acomprehensive resource for in silico drug discovery and exploration[J]. Nucleic Acids Research, 2006, 34(suppl_1): D668-D672. DOI:10.1093 / nar / gkj067.

[0007] ChEMBL from GAULTON A, BELLIS LJ, BENTO AP, etc. ChEMBL: a large-scale bioactivity database for drug discovery[J]. Nucleic Acids Research, 2011, 40(D1): D1100-D1107. DOI:10.1093 / nar / gkr777.

[0008] UniProt is from THE UNIPROT CONSORTIUM. UniProt: the universal protein knowledgebase[J]. Nucleic Acids Research, 2016, 45(D1): D158-D169. DOI:10.1093 / nar / gkw1099。

[0009] HGNC is from EYRE T A, DUCLUZEAU F, SNEDDON T P, etc. The HUGO Gene Nomenclature Database, 2006 updates[J]. Nucleic Acids Research, 2006, 34(suppl_1): D319-D321. DOI:10.1093 / nar / gkj147。

[0010] DisGeNET is from PIÑERO J, QUERALT-ROSINACH N, BRAVO À, etc. DisGeNET: a discovery platform for the dynamical exploration of human diseases and their genes[J]. Database, 2015, 2015: bav028. DOI:10.1093 / database / bav028。

[0011] OMIM is from MCKUSICK V. Online Mendelian inheritance in man (OMIM) database [J]. Bethesda: National Center for Biotechnology Information for the National Institute of Health, 2004。

[0012] Orphanet is from https: / / www.orpha.net / 。

[0013] ClinicalTrials.gov is from https: / / clinicaltrials.gov / 。

[0014] PubMed is from https: / / pubmed.ncbi.nlm.nih.gov / 。

[0015] Reactome comes from CROFT D, O'KELLY G, WU G, etc. Reactome: a database of reactions, pathways and biological processes[J]. Nucleic Acids Research, 2010, 39(suppl_1): D691-D697. DOI:10.1093 / nar / gkq1018.

[0016] KEGG is from https: / / www.genome.jp / kegg / .

[0017] Pathway Commons comes from CERAMI EG, GROSS BE, DEMIR E, etc. PathwayCommons, a web resource for biological pathway data[J]. Nucleic AcidsResearch, 2010, 39(suppl_1): D685-D690. DOI:10.1093 / nar / gkq1039.

[0018] SIDER from KUHN M, LETUNIC I, JENSEN LJ, etc. The SIDER database ofdrugs and side effects[J]. Nucleic Acids Research, 2015, 44(D1): D1075-D1079.DOI:10.1093 / nar / gkv1075. Summary of the Invention

[0019] The purpose of this invention is to solve the problems of scattered data sources, heterogeneous structures, and difficulty in direct application in the prior art.

[0020] To address the above problems, this invention provides a rare disease drug retargeting path mining method based on knowledge graphs, including: Step 1: Knowledge Graph Construction; Step 1-1: Data Collection; Public databases, including DrugBank, CheEMBL, UniProt, HGNC, DisGeNET, OMIM, Orphanet, ClinicalTrials.gov, PubMed, Reactome, KEGG, PathwayCommons, and SIDER, covering multi-dimensional data on drug information, target proteins, genes, diseases, pathways, and adverse reactions; Clinical data, including: real-world data related to Marshall syndrome, including concise structured information from electronic medical records, including diagnosis, Medication records and examination results; Steps 1-2: Construct triples and metadata; Divide the collected data into structured and unstructured data. Structured data includes: JSON files, tabular data, and database tables; unstructured data includes: literature text, clinical report abstracts, and database annotation text. Construct preset rules to process the structured data to obtain triples; Steps 1-3: Import all triples and metadata into the Neo4j graph database to construct a preliminary knowledge base to be optimized. Perform structure optimization on the knowledge base to be optimized, including: entity standardization and synonym entity merging, and low-confidence edge filtering. After the knowledge base to be optimized completes the structure optimization, a knowledge graph is obtained; Step 2: Construct a drug redirection path mining model based on knowledge graph; Step 2-1: Extract positive examples from the knowledge graph. and negative samples Positive samples are known disease-drug associations; negative samples are randomly selected irrelevant drug-drug pairs. Step 2-2: For each disease-drug pair, extract all paths that match the template from the knowledge graph to form a candidate path set. This includes: disease-gene-protein-drug, disease-pathway-protein-drug, disease-gene-pathway-drug, and disease-protein-signaling pathway-drug; [The text abruptly ends here, likely due to an incomplete sentence or a formatting error.] Candidate Path Set Taking relevant metadata as input, the output is the first... Path score and Explanation Information The formula is: in, It includes node types, relationship types, and statistical information within the path. Includes source databases, literature evidence, and confidence levels; path scoring. Paths with scores above a set threshold of 0.7 to 0.8 are considered candidate results; the final output is a path score. The advanced pathway provides support for subsequent clinical validation.

[0021] In the preferred approach, pre-defined rules are constructed to process structured data. The specific rules are as follows: Extract field content from a JSON file by iterating through key-value pairs and generating triples; For tabular data, generate triples by extracting the required column fields row by row; For a database table, select a specified field by executing an SQL statement and convert it into a triple according to rules; Large Language Model (LLM) is used to extract entities and relations to obtain triples. Each triple contains corresponding metadata, including: entity source database, relation source, data timestamp, original text fragment, and document ID.

[0022] In the preferred approach, entity standardization and synonym merging include: Key words were extracted based on expert opinions. Regular expression matching was used to replace key words. Based on synonymous entities in the knowledge graph, regular expression matching rules were manually formulated to replace the original words for standardized naming. The remaining unstandardized entities were used to calculate semantic vector similarity using Sentence-BERT. Entities with similarity greater than 0.85 were considered synonyms and merged. Low-confidence side filtering includes: A triple, or edge, in a knowledge graph contains the two entities it connects and their corresponding relationships. The confidence of an edge is calculated based on the entity's source database and the relationship's source database recorded in the metadata, using the following formula: in, Let e ​​represent the confidence level of edge e (0-1); This indicates the credibility of the entity's source database: 1.0 for authoritative databases, 0.9 for secondary databases, and 0.8 for clinical data sources. The reliability of the relation source is indicated by: 1.0 for relations manually compiled / annotated by experts, 0.9 for relations from authoritative database records, and 0.8 for relations from literature. and These are weights for the credibility of the data source and the support of the evidence, respectively. Based on experience, such as Typically, 0.8 is taken as the low confidence threshold. Then remove edges; the entity source database includes: authoritative databases DrugBank, ChemBL, UniProt, HGNC, DisGeNET, OMIM, and Orphanet, secondary databases Reactome, KEGG, Pathway Commons, SIDER, ClinicalTrials.gov, and clinical data.

[0023] In the preferred embodiment, step 2 further includes: Let the constructed knowledge graph be: in, For a set of nodes, For a set of relation types, Let be the set of edges, then the update formula for R-GCN node embedding is: in, Indicates the first Layer nodes Embedded, Represents nodes There is a set of neighbor nodes of relation type r. Represents a node Neighbor nodes of relation type r Embedded, This is the weight matrix for the corresponding relation type. For the first Layer self-loop weight matrix, The normalization constant is Let be a nonlinear activation function; let the embedding of the node in the last layer (Lth layer) of R-GCN be . The probability is predicted by maximizing positive sample edges and minimizing negative sample edges, as shown in the formula: in, Embedded for disease nodes, For drug node embedding, This is the matrix transpose operator; the training loss uses cross-entropy, and the formula is: .

[0024] In the preferred approach, positive examples are known disease-drug associations including: atenolol can improve diastolic function of the cardiovascular system in patients with Mascheranos syndrome, losartan can reduce the risk of aortic dilation, and propranolol can control heart rate abnormalities.

[0025] In the preferred approach, negative samples are randomly selected irrelevant drug-disease pairs, including: Mascheranos syndrome—Amoxicillin, Mascheranos syndrome—Ibuprofen, and Mascheranos syndrome—Omeprazole.

[0026] The beneficial effects of this invention are as follows: Addressing the problems of data sparsity, knowledge heterogeneity, and insufficient interpretability in rare disease drug retargeting, this invention proposes a rare disease drug retargeting path mining method based on knowledge graphs. By integrating multi-source databases and clinical data, a unified semantic knowledge graph is constructed, realizing the systematic association of multi-layered information such as diseases, genes, proteins, drugs, and pathways. This invention improves upon traditional techniques in two key aspects: (1) In the knowledge integration stage, standardized naming, synonym entity merging, and confidence filtering mechanisms are adopted to improve the consistency and reliability of the graph; (2) In the association mining stage, R-GCN is introduced for link prediction, combined with meta-path template search and large model scoring to achieve interpretable reasoning about the potential relationship between diseases and drugs. This method can efficiently discover candidate drug paths in a knowledge-sparse environment and provide quantitative scoring and interpretation results, improving the credibility and traceability of predictions, and providing a systematic and interpretable intelligent analysis method for rare disease drug retargeting. Attached Figure Description

[0027] Figure 1 This is a schematic diagram of the knowledge graph construction process; Figure 2 This is a schematic diagram of the R-GCN training process; Figure 3 This is a schematic diagram of the overall reasoning process of the model. Detailed Implementation

[0028] Example 1: This invention relates to a knowledge graph-based method for mining drug redirection paths in rare diseases, using Marshall syndrome as a specific research subject. This method integrates multi-source biomedical databases and clinical data to construct a semantically unified knowledge graph, and utilizes a graph neural network model to predict potential associations between diseases and drugs. Simultaneously, it combines meta-path template search and path scoring mechanisms to generate interpretable drug redirection paths.

[0029] In practice, structured and unstructured information is first extracted from biomedical databases and clinical data related to Marshall syndrome, and triples are generated using rules or large models and imported into a graph database. Subsequently, the graph structure is optimized through naming standardization, synonym entity merging, and low-confidence edge filtering. The optimized knowledge graph is used to train a relational graph convolutional network for disease-drug link prediction and provides a structural foundation for meta-path template search. Each round of link prediction and path generation results is combined with graph metadata for path scoring and interpretation, thereby outputting reliable drug redirection candidate results. Existing knowledge graph methods often focus on static structures or single data sources, while this method achieves systematic mining of rare disease drug redirection paths through multi-source data integration and structural optimization. Specific steps include: Step 1: As Figure 1 As shown, the purpose of knowledge graph construction is to integrate multi-source biomedical data into a unified graph structure to facilitate intelligent reasoning about disease-drug associations.

[0030] Step 1-1: The collected data mainly includes two categories: public database data and clinical data; Public databases include: DrugBank, CheEMBL, UniProt, HGNC, DisGeNET, OMIM, Orphanet, ClinicalTrials.gov, PubMed, Reactome / KEGG / Pathway Commons, and SIDER, covering multi-dimensional data on drug information, target proteins, genes, diseases, pathways, and adverse reactions; For example, the drug "Atenolol" and its target protein "ADRB1" were extracted from DrugBank, the "FBN1" gene related to Mascherano's syndrome was obtained from DisGeNET, and the "elastic fiber formation pathway" involved by this gene was found in the Reactome database. Clinical data includes: real-world data related to Marshall syndrome, including: concise structured information in electronic medical records, such as diagnosis, medication records and test results; For example, records of patients using beta-blockers to control their heart rhythm can be identified in electronic medical record data. This electronic medical record data can be used to uncover potential pathways not reflected in public databases, such as (Madison syndrome—FBN1—elastic fiber formation pathway—ADRB1—atenolol).

[0031] Steps 1-2: After data collection is complete, proceed to the triple and metadata construction phase; The collected data can be divided into two categories: structured data and unstructured data. Structured data includes JSON files, tabular data, or database tables, all with standardized formats and clearly defined information fields. Different preset rules can be designed to process and extract triples for different data structures. For JSON files, key-value pairs are iterated to extract field content and triples are generated according to a template; for tabular data, the required columns are extracted row by row to generate triples; for database tables, SQL statements are executed to select specified fields and convert them into triples according to rules. For example, for disease-related data, triples can be generated (disease: Marshall syndrome, gene: FBN1, association type: related); for pathway participation information, triples can be generated (gene / protein: FBN1, pathway: elastic fiber formation pathway, participation relationship: involved); for clinical intervention data, triples can be generated (disease: Marshall syndrome, drug: losartan, treatment effect: improved aortic dilation).

[0032] Unstructured data includes: literature text, clinical report abstracts, and database annotation text. Large Language Model (LLM) is used to extract entities (including: diseases, genes, proteins, drugs, pathways) and relationships (including: association types, treatment relationships, and participation effects) to obtain triples. By providing prompts to the large language model, the model is guided to extract triples. For example, the description "FBN1 mutation is significantly associated with abnormal phenotypes of Marshall syndrome" can be extracted from PubMed literature abstracts and converted into triples (disease: Marshall syndrome, gene: FBN1, relation: gene mutation associated). Each triple also includes corresponding metadata: entity source database, relation source, data timestamp, original text fragment, and literature ID, which are used for subsequent graph optimization and path scoring.

[0033] Steps 1-3: All triples and metadata are imported into the Neo4j graph database to build a preliminary knowledge base to be optimized.

[0034] The knowledge base to be optimized then enters the structural optimization stage, which mainly includes two parts: entity standardization and synonym entity merging, and low confidence edge filtering. The purpose of entity standardization and synonym entity merging is to standardize the entities in triples.

[0035] Key keywords were extracted based on expert opinions. Regular expression matching was used to replace these keywords. Based on synonymous entities in the knowledge graph, manual regular expression matching rules were developed to replace the original terms (this was determined based on the specific circumstances; for example, aspirin might be acetylsalicylic acid or ASA in different data sources. These possible aliases need to be listed and uniformly replaced with Aspirin) to standardize the naming. The remaining unstandardized entities had their semantic vector similarity calculated using Sentence-BERT. Entities with a similarity greater than 0.85 were considered synonyms and merged.

[0036] The low-confidence edge filtering section aims to remove triples (i.e., edges in the knowledge graph) with low confidence. It assesses the confidence of each edge based on its source and evidence, removing edges below a set threshold to ensure the reliability and data quality of the graph. A triple, or edge, in the knowledge graph contains the two entities it connects and their corresponding relationship. The edge's confidence is calculated based on the entity source database and relationship source recorded in the metadata, using the following formula:

[0037] in, Let e ​​represent the confidence level of edge e (0-1); This indicates the credibility of the database from which the entity originates (e.g., authoritative database 1.0, secondary database 0.9, clinical data source 0.8). This indicates the reliability of the relation's source (e.g., relations that have been manually compiled or annotated by experts are 1.0, relations from authoritative database records are 0.9, and relations from literature are 0.8). and These are weights for the credibility of the data source and the support of the evidence, respectively. Based on experience, such as Typically, 0.8 is taken as the low confidence threshold. Then remove edges; the entity source database includes: authoritative databases (DrugBank, ChemBL, UniProt, HGNC, DisGeNET, OMIM, Orphanet), secondary databases (Reactome, KEGG, Pathway Commons, SIDER, ClinicalTrials.gov); clinical data (real-world data).

[0038] Once optimized, the knowledge graph can be used for subsequent graph neural network training and meta-path search, providing a structured intelligent analysis foundation for rare disease drug retargeting.

[0039] Step 2: As Figure 2 , Figure 3As shown, a drug redirection path mining model based on knowledge graphs; This invention proposes using R-GCN combined with meta-path templates and a large model scoring mechanism to address the challenge of mining potential associations for drugs in rare diseases. The knowledge graph contains various entities such as diseases, drugs, genes, proteins, and pathways, along with their diverse relationships. The goal of the link prediction task is to determine whether a potential association exists between a disease and a candidate drug, and to provide candidate endpoints for path search.

[0040] Step 2-1: During the training phase, first extract positive examples from the knowledge graph. and negative samples ; Positive examples are known disease-drug associations, such as the three reported drug effects: "Atenolol can improve diastolic function of the cardiovascular system in patients with Mars syndrome," "Losartan can reduce the risk of aortic dilation," and "Propranolol can control heart rate abnormalities." Negative examples are randomly selected irrelevant drug-disease pairs, such as "Mars syndrome—Amoxicillin," "Mars syndrome—Ibuprofen," and "Mars syndrome—Omeprazole." The constructed knowledge graph is denoted as: in, For a set of nodes, For a set of relation types, Let be the set of edges, then the update formula for R-GCN node embedding is: in, Indicates the first Layer nodes Embedded, Represents nodes There is a set of neighbor nodes of relation type r. Represents a node Neighbor nodes of relation type r Embedded, This is the weight matrix for the corresponding relation type. For the first Layer self-loop weight matrix, The normalization constant is It is a non-linear activation function; Let the embedding of the last layer (Lth layer) node in the R-GCN be denoted as . The probability is predicted by maximizing positive sample edges and minimizing negative sample edges, as shown in the formula: in, Embedded for disease nodes, For drug node embedding, This is the matrix transpose operator; The training loss uses cross-entropy, and the formula is: .

[0041] Step 2-2: In the inference phase, the input consists of disease nodes related to Maslen syndrome and a set of candidate drugs. R-GCN outputs the potential association probability for each disease-drug pair. For example, the model predicts an association probability of 0.91 for "Maslen syndrome-losartan," 0.87 for "Maslen syndrome-ateninolol," and only 0.05 for "Maslen syndrome-amoxicillin," indicating that the first two drugs have a strong potential therapeutic correlation. If the predicted probability exceeds a set threshold, the drug is considered potentially related to the disease, and it is used as the endpoint for path search.

[0042] Pathway search is performed by defining meta-pathway templates. Common meta-pathways include "disease—gene—protein—drug", "disease—pathway—protein—drug", "disease—gene—pathway—drug", and "disease—protein—signaling pathway—drug". For example, for the predicted candidate drug losartan, the system can search for the path "Marseille syndrome—Fibrillin-1 gene (FBN1)—transforming growth factor β receptor 2 (TGFBR2)—mitogen-activated protein kinase signaling pathway (MAPK)—losartan"; for atenolol, the path "Marseille syndrome—FBN1—elastic fibrous formation pathway—β1-adrenergic receptor gene (ADRB1)—atenolol". Each path represents a potential association between a disease and a candidate drug through different biological entities.

[0043] For each disease-drug pair, the knowledge graph extracts all paths that match the template, forming a candidate path set. This includes: disease-gene-protein-drug, disease-pathway-protein-drug, disease-gene-pathway-drug, and disease-protein-signaling pathway-drug. The first Candidate Path Set Taking relevant metadata as input, the output is the first... Path score and Explanation Information The formula is: in, It includes node types, relationship types, and statistical information within the path. The study included source databases, literature evidence, and confidence levels. For example, after large-scale model analysis, the comprehensive score of the "Madison syndrome—FBN1—TGFBR2—MAPK pathway—Losartan" pathway was 0.92, indicating that its mechanism is related to the inhibition of the TGF-β signaling pathway and has high potential for drug retargeting. The score of the "Madison syndrome—FBN1—elastic fiber formation pathway—ADRB1—Atenolol" pathway was 0.83, and it was evaluated as an adjunctive treatment pathway.

[0044] Finally, after sorting and filtering, highly reliable and interpretable drug retargeting pathways are output to support subsequent clinical validation. Typically, pathway scores can be generated. Paths exceeding a set threshold (e.g., 0.7–0.8) are considered candidate results. This threshold can be adjusted based on actual data distribution or clinical needs to balance the number of candidate drugs with prediction reliability.

[0045] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the present invention. Various changes and modifications can be made to the present invention without departing from its spirit and scope. All such changes and modifications fall within the scope of the present invention as claimed, which is defined by the appended claims and their equivalents.

Claims

1. A method for mining rare disease drug retargeting paths based on knowledge graphs, characterized in that, include: Step 1: Knowledge Graph Construction; Step 1-1: Data Collection; Public databases, including DrugBank, CheEMBL, UniProt, HGNC, DisGeNET, OMIM, Orphanet, ClinicalTrials.gov, PubMed, Reactome, KEGG, Pathway Commons, and SIDER, covering multi-dimensional data on drug information, target proteins, genes, diseases, pathways, and adverse reactions; Clinical data, including: real-world data related to Marshall syndrome, including brief structured information from electronic medical records, including diagnosis, medication records, and examination results; Step 1-2: Constructing Triples and Metadata; The collected data is divided into structured and unstructured data. Structured data includes: JSON files, tabular data, and database tables; unstructured data includes: literature text, clinical report abstracts, and database annotation text. Pre-defined rules are used to process the structured data to obtain triples; Steps 1-3: Import all triples and metadata into the Neo4j graph database to build a preliminary knowledge base to be optimized. Then, perform structural optimization on the knowledge base to be optimized, including entity standardization and merging of synonymous entities, and low-confidence edge filtering. After the structural optimization of the knowledge base to be optimized is completed, a knowledge graph is obtained. Step 2: Construct a drug redirection path mining model based on knowledge graph; Step 2-1: Extract positive examples from the knowledge graph and negative samples Positive samples are known disease-drug associations; negative samples are randomly selected irrelevant drug-drug pairs. Step 2-2: For each disease-drug pair, extract all paths that match the template from the knowledge graph to form a candidate path set. This includes: disease-gene-protein-drug, disease-pathway-protein-drug, disease-gene-pathway-drug, and disease-protein-signaling pathway-drug; [The text abruptly ends here, likely due to an incomplete sentence or a formatting error.] Candidate Path Set Taking relevant metadata as input, the output is the first... Path score and Explanation Information The formula is: in, It includes node types, relationship types, and statistical information within the path. Includes source databases, literature evidence, and confidence levels; path scoring. Paths with scores above a set threshold of 0.7 to 0.8 are considered candidate results; the final output is a path score. The advanced pathway provides support for subsequent clinical validation.

2. The method for mining rare disease drug retargeting paths based on knowledge graphs according to claim 1, characterized in that, Establish preset rules to process structured data. The specific rules are as follows: Extract field content from a JSON file by iterating through key-value pairs and generating triples; For tabular data, generate triples by extracting the required column fields row by row; For a database table, select a specified field by executing an SQL statement and convert it into a triple according to rules; Large Language Model (LLM) is used to extract entities and relations to obtain triples. Each triple contains corresponding metadata, including: entity source database, relation source, data timestamp, original text fragment, and document ID.

3. The method for rare disease drug retargeting path mining based on knowledge graphs according to claim 1, characterized in that, Entity standardization and synonym merging include: Key words were extracted based on expert opinions. Regular expression matching was used to replace key words. Based on synonymous entities in the knowledge graph, regular expression matching rules were manually formulated to replace the original words for standardized naming. The remaining unstandardized entities were used to calculate semantic vector similarity using Sentence-BERT. Entities with similarity greater than 0.85 were considered synonyms and merged. Low-confidence side filtering includes: A triple, or edge, in a knowledge graph contains the two entities it connects and their corresponding relationships. The confidence of an edge is calculated based on the entity's source database and the relationship's source database recorded in the metadata, using the following formula: in, Let e ​​represent the confidence level of edge e (0-1); This indicates the credibility of the entity's source database: 1.0 for authoritative databases, 0.9 for secondary databases, and 0.8 for clinical data sources. The reliability of the relation source is indicated by: 1.0 for relations manually compiled / annotated by experts, 0.9 for relations from authoritative database records, and 0.8 for relations from literature. and These are weights for the credibility of the data source and the support of the evidence, respectively. Based on experience, such as Typically, 0.8 is taken as the low confidence threshold. Then remove edges; the entity source database includes: authoritative databases DrugBank, ChemBL, UniProt, HGNC, DisGeNET, OMIM, and Orphanet, secondary databases Reactome, KEGG, Pathway Commons, SIDER, ClinicalTrials.gov, and clinical data.

4. The method for mining rare disease drug retargeting paths based on knowledge graphs according to claim 1, characterized in that, Step 2 also includes: Let the constructed knowledge graph be: in, For a set of nodes, For a set of relation types, Let be the set of edges, then the update formula for R-GCN node embedding is: in, Indicates the first Layer nodes Embedded, Represents nodes There is a set of neighbor nodes of relation type r. Represents a node Neighbor nodes of relation type r Embedded, This is the weight matrix for the corresponding relation type. For the first Layer self-loop weight matrix, The normalization constant is It is a non-linear activation function; Let the embedding of the last layer (Lth layer) node in the R-GCN be denoted as . The probability is predicted by maximizing positive sample edges and minimizing negative sample edges, as shown in the formula: in, Embedded for disease nodes, For drug node embedding, This is the matrix transpose operator; The training loss uses cross-entropy, and the formula is: 。 5. The method for rare disease drug retargeting path mining based on knowledge graphs according to claim 1, characterized in that, Positive examples are known disease-drug associations including: atenolol, which improves diastolic function of the cardiovascular system in patients with Mascheranos syndrome; losartan, which reduces the risk of aortic dilation; and propranolol, which controls heart rate abnormalities.

6. The method for rare disease drug retargeting path mining based on knowledge graphs according to claim 1, characterized in that, Negative samples are randomly selected irrelevant drug-disease pairs, including: Amoxicillin for Mascheranos syndrome, Ibuprofen for Mascheranos syndrome, and Omeprazole for Mascheranos syndrome.