Diabetes special disease potential drug discovery method and system based on interpretable semantic reasoning

By constructing a semantic knowledge graph specific to diabetes and introducing disease stage labels and sequence templates, the problem of inconsistent interpretation of disease stages in diabetes drug discovery was solved, achieving higher clinical consistency and credible explanatory drug discovery.

CN121983282APending Publication Date: 2026-05-05INST OF INFORMATION ON TRADITIONAL CHINESE MEDICINE CACMS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST OF INFORMATION ON TRADITIONAL CHINESE MEDICINE CACMS
Filing Date
2026-01-23
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing knowledge graph-based drug discovery protocols in diabetes suffer from difficulties in providing explanations that meet the semantic and sequential constraints of disease progression, resulting in unclear or inconsistent explanation chains and a lack of clinical consistency and credible explanatory power.

Method used

We construct a semantic knowledge graph for diabetes, introduce disease stage labels and sequence templates, and use an explicit disease stage semantic reasoning mechanism to filter and calculate path contribution and output an interpretable semantic reasoning chain.

Benefits of technology

It improves the clinical consistency and semantic credibility of potential drug discovery, enhances the medical auditability and interpretability of the drug discovery process, and outputs a perceptible semantic interpretation chain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121983282A_ABST
    Figure CN121983282A_ABST
Patent Text Reader

Abstract

The invention provides an interpretable semantic reasoning-based potential drug discovery method and system for a special diabetes disease, and relates to the technical field of semantic reasoning of drug discovery. The method comprises the following steps: performing entity relationship extraction on clinical guidelines, electronic medical records and medical literatures, and constructing a special diabetes disease semantic knowledge graph containing entities such as diseases, complications, drugs, targets and pathways; establishing a disease course stage label set and a disease course sequence template library, and endowing related entities with disease course stage labels to form a knowledge graph with labels; a multi-hop candidate reasoning path is generated on the map by taking a diabetes entity as an end point, the path contribution degree is calculated according to the disease course sequence consistency and an evidence field, multiple consistent paths of the same medicine are aggregated to obtain candidate scores, candidate medicine sorting and explainable reasoning chain output along with disease course stage labels are achieved, and the accuracy of the reasoning chain is improved. Therefore, clinical consistency and interpretability of potential drug discovery results are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of semantic reasoning technology for drug discovery, and in particular to a method and system for discovering potential drugs for diabetes based on interpretable semantic reasoning. Background Technology

[0002] In recent years, medical knowledge graphs have been used to structurally represent entities and relationships in clinical guidelines, expert knowledge, electronic medical records, and medical literature, and to provide semantic retrieval and reasoning capabilities. For example, in "Construction and Application Research of Knowledge Graph for Blood Glucose Management in Diabetic Patients," Shao Minhui et al. extracted semantic entities and relationships based on clinical guidelines, expert experience, and hospital electronic medical records, constructed a diabetes knowledge graph using a graph database, and conducted application verification. Meanwhile, Ye Yajuan et al., in "Construction of Entity and Relationship Annotated Corpus for Diabetic Electronic Medical Records," established an entity and relationship classification system around diabetes electronic medical records and formed an annotated corpus, providing a data foundation for disease-specific entity and relationship extraction and subsequent graph construction. Han Pu et al., in "Research on Multimodal Knowledge Graph Construction Method for Chinese Electronic Medical Records," further proposed a construction method for multimodal data organization of Chinese electronic medical records, reflecting the continuous advancement of disease-specific knowledge organization and graph-based representation.

[0003] In potential drug discovery, one trend is the integration of multi-source biomedical data into knowledge graphs, using relational reasoning on the graph to achieve drug relocation and evidence output. For example, Zhang Han, An Xinyu, and Liu Chunhe, in their paper "Drug Knowledge Discovery Based on Multi-Source Semantic Knowledge Graphs: Empirical Evidence from Drug Relocation," conduct drug knowledge discovery research based on multi-source semantic knowledge graphs, demonstrating that knowledge graphs are becoming an important data organization form for drug relocation. Meanwhile, Hou Mengwei et al., in their paper "A Review of Knowledge Graph Research and Its Applications in the Medical Field," point out that the combination of knowledge graphs with big data and deep learning technologies is driving the development of applications such as intelligent semantic retrieval in medicine, question answering, and clinical decision support, indicating that interpretable semantic reasoning and knowledge utilization will become important directions for evolution.

[0004] Existing drug discovery schemes based on knowledge graphs often use link prediction or multi-hop paths as reasoning carriers, and the explanations usually remain at the level of reachable related paths on the graph. However, diabetes is a chronic metabolic disease with staged characteristics such as pathogenesis, metabolic abnormalities and the evolution of complications (Mao Fumin et al., "Advances in the Application of Knowledge Graphs in the Health Management of Diabetic Patients", a review of diabetes health management and knowledge graph applications in 2025). If the reasoning path only satisfies topological reachability but lacks semantic and stage sequence constraints for the disease course, it is easy to have explanation chains with unclear stages or inconsistent stage sequences, making it difficult for the explanations to be reviewed and reused by the disease-specific diagnosis and treatment logic. Furthermore, Hou Mengwei et al. also pointed out that medical knowledge graphs still have common problems in terms of efficiency, constraints and scalability, which further amplifies the difficulty of implementing usable and reliable explanations for disease-specific reasoning. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide a method and system for discovering potential drugs for diabetes based on interpretable semantic reasoning, so as to explicitly introduce a semantic reasoning mechanism for disease stages, so that the results of potential drug discovery have clinical consistency and interpretability in the semantic space of the specific disease.

[0006] To achieve the above objectives, the present invention provides the following solution: A method for discovering potential drugs specific to diabetes based on interpretable semantic reasoning, comprising: Entity relationships are extracted from clinical guidelines, electronic medical records, and medical literature to form a set of triplets containing drug entities, target entities, pathway entities, diabetes disease entities, and complication entities. A semantic knowledge graph for diabetes is then constructed based on the set of triplets. Construct a set of disease course stage labels and a disease course sequence template library; the set of disease course stage labels includes pathogenesis stage labels and complication stage labels; the disease course sequence template library represents the allowed sequential constraints between disease course stage labels; Based on the disease progression template library, disease progression stage labels are assigned to entities related to diabetes, complications, and target and pathway entities that are associated with diabetes or complications in the triplet set, resulting in a labeled knowledge graph. In the labeled knowledge graph, a set of candidate drug entities is determined, and with the diabetes disease entity as the reasoning endpoint, a set of multi-hop candidate reasoning paths is obtained by searching for each candidate drug entity in the candidate drug entity set. For each multi-hop candidate inference path in the multi-hop candidate inference path set, the disease stage sequence is parsed and matched with the disease sequence template library. Paths that do not meet the preset consistency conditions are removed to obtain a consistent path set. The path contribution of each consistent path in the consistent path set is calculated based on the relationship reliability parameter and the consistency with the disease sequence. The relationship reliability parameter is used to characterize the credibility of the relationship in the path. The candidate scores are obtained by aggregating the path contributions corresponding to the same candidate drug entity. The candidate drug ranking results are output based on the candidate scores, and at least one consistent path with the highest contribution and the disease stage labels corresponding to each entity in the consistent path are output as interpretable semantic reasoning chains.

[0007] Preferably, entity relationships are extracted from clinical guidelines, electronic medical records, and medical literature to form a set of triplets containing drug entities, target entities, pathway entities, diabetes disease entities, and complication entities. Based on this set of triplets, a semantic knowledge graph for diabetes is constructed, including: Construct a diabetes-specific entity dictionary and relation type set; the diabetes-specific entity dictionary covers the standard names and synonyms of drug entities, target entities, pathway entities, diabetes disease entities, and complication entities; Based on a diabetes-specific entity dictionary, entity recognition and standardization are performed on clinical guidelines, electronic medical records, and medical literature to obtain an entity set. Based on the set of relation types, relation extraction is performed on the entity set to generate entity pairs and relations, and the entity pairs and relations are combined to form a set of triples.

[0008] Preferably, the formation of the triplet set includes: Each triplet in the triplet set is assigned a relation type identifier; the relation type identifier includes at least one or more of the following: therapeutic effect relation, action target relation, pathway involvement relation, and complication evolution relation. Each triplet in the triplet set is assigned an evidence field; the evidence field is used to indicate that the triplet originates from at least one of clinical guidelines, electronic medical records, and medical literature. The duplicate triples are standardized based on the evidence field to obtain a set of deduplicated triples.

[0009] Preferably, a set of disease stage labels and a disease sequence template library are constructed, including: Construct a set of labels for pathogenesis stage labels and a set of labels for complication stage labels, and form a set of disease course stage labels based on the set of labels for pathogenesis stage labels and the set of labels for complication stage labels; Construct a set of stage sequence templates; a stage sequence template is a template sequence composed of disease stage labels in sequence. Configure the allowed sequence constraint rules for each stage sequence template in the stage sequence template set; the allowed sequence constraint rules are used to limit the sequential relationship between adjacent stages and across stages in the template sequence; The process sequence template library is composed of the set of stage sequence templates and the rules for allowing sequential constraints.

[0010] Preferably, based on the disease progression template library, disease progression stage labels are assigned to entities related to diabetes, complications, and target and pathway entities that are associated with diabetes or complications in the triplet set, resulting in a labeled knowledge graph, including: Based on the disease stage label set, assign corresponding disease stage labels to the diabetes disease entity and complication entity; Retrieve target entities and pathway entities that are associated with diabetes disease entities or complication entities from the triple set, and assign disease stage labels corresponding to diabetes disease entities or complication entities that are associated with target entities or pathway entities in the triple set, or assign adjacent stage labels that comply with the allowed sequence constraint rules. The entities with disease stage labels are combined with sets of triples to construct a labeled knowledge graph.

[0011] Preferably, a set of candidate drug entities is determined in the labeled knowledge graph, and a set of multi-hop candidate reasoning paths is obtained for each candidate drug entity in the candidate drug entity set, using the diabetes disease entity as the reasoning endpoint, including: Using the diabetes disease entity as the target node, an reachability search is performed in the labeled knowledge graph to obtain drug entities that have a connection relationship with the diabetes disease entity, and the obtained drug entities are used to form a candidate drug entity set. For each candidate drug entity in the candidate drug entity set, a path search is performed starting from the candidate drug entity and ending with the diabetes disease entity, resulting in a multi-hop candidate reasoning path set formed by sequentially connecting entities and relations. Perform path deduplication and loop path elimination on the multi-hop candidate inference path set to obtain the multi-hop candidate inference path set for subsequent parsing.

[0012] Preferably, for each multi-hop candidate inference path in the multi-hop candidate inference path set, a disease stage sequence is parsed and matched with a disease sequence template library. Paths that do not meet the preset consistency conditions are removed to obtain a consistent path set, including: Read the disease stage labels of entities in each multi-hop candidate inference path; The disease stage tags read are arranged into a disease stage sequence according to the path order; The disease course stage sequence is matched with the stage sequence template set to obtain the matching results; Set the preset consistency condition to meet the following conditions simultaneously: the disease stage sequence meets the allowed sequence constraint rule, and the same disease stage label is allowed to appear consecutively in the disease stage sequence; Based on the matching results, determine whether the preset consistency conditions are met, and form a set of consistent paths by the paths that meet the preset consistency conditions.

[0013] Preferably, the path contribution of each consistent path in the consistent path set is calculated based on the consistency between the relationship reliability parameter and the disease course sequence. The relationship reliability parameter is used to characterize the credibility of the relationship in the path, including: The reliability parameters of each relation in the consistency path are determined based on the evidence field of the triples contained in the consistency path within the triple set. The consistency of the disease course sequence of the consistency path is determined based on the matching results between the disease course stage sequence corresponding to the consistency path and the allowable sequence constraint rules. Based on the preset path contribution calculation rules, the path contribution of the consistent path is calculated based on the consistency of the relationship reliability parameter and the disease course sequence.

[0014] Preferably, candidate scores are obtained by aggregating the path contributions corresponding to the same candidate drug entity. Based on these scores, a ranking of candidate drugs is output, along with at least one consistent path with the highest contribution and the disease stage labels corresponding to each entity within that consistent path, serving as an interpretable semantic reasoning chain. Aggregate the path contributions in the set of consistent paths corresponding to the same candidate drug entity to obtain the candidate score of the candidate drug entity; The candidate drug entity set is sorted according to the candidate scores to obtain the candidate drug ranking results; For each candidate drug entity, select at least one consistency path with the highest path contribution from the corresponding consistency path set, and output the disease stage label corresponding to each entity in the consistency path to form an interpretable semantic reasoning chain.

[0015] A potential drug discovery system for diabetes based on interpretable semantic reasoning, comprising: The entity relation extraction and knowledge graph construction unit is used to extract entity relations from clinical guidelines, electronic medical records and medical literature, forming a set of triplets containing drug entities, target entities, pathway entities, diabetes disease entities and complication entities, and constructing a diabetes-specific semantic knowledge graph based on the set of triplets. The disease course semantic modeling unit is used to construct a set of disease course stage labels and a disease course sequence template library. The set of disease course stage labels includes pathogenesis stage labels and complication stage labels. The disease course sequence template library represents the allowed sequential constraints between disease course stage labels. The disease course labeling unit is used to assign disease course stage labels to entities related to diabetes, complication entities, and target entities and pathway entities that are associated with diabetes or complication entities in the triple set, based on the disease course sequence template library, to obtain a labeled knowledge graph. The candidate drug path generation unit is used to determine the set of candidate drug entities in the labeled knowledge graph, and to search for a set of multi-hop candidate reasoning paths for each candidate drug entity in the candidate drug entity set, with the diabetes disease entity as the reasoning endpoint. The disease course consistency screening and path contribution calculation unit is used to parse each multi-hop candidate inference path in the multi-hop candidate inference path set to obtain the disease course stage sequence and match it with the disease course sequence template library. Paths that do not meet the preset consistency conditions are eliminated to obtain a consistent path set. The path contribution of each consistent path in the consistent path set is calculated based on the relationship reliability parameter and the consistency of the disease course sequence. The relationship reliability parameter is used to characterize the credibility of the relationship in the path. The candidate score and interpretable chain output unit is used to aggregate the path contribution of the same candidate drug entity to obtain the candidate score, output the candidate drug ranking result based on the candidate score, and output at least one consistency path with the highest contribution and the disease stage label corresponding to each entity in the consistency path as an interpretable semantic reasoning chain.

[0016] The present invention discloses the following technical effects: (1) Based on the semantic knowledge expression of diabetes disease, this invention introduces a semantic knowledge graph of diabetes disease to unify the organization of entity relationships such as drugs, targets, pathways, diabetes disease and complications. This enables the potential drug discovery process to no longer rely on isolated datasets or weak semantic associations, but to complete reasoning in a unified semantic space. This helps to improve the accessibility and semantic interpretability of the association recognition between disease-related entities and overcomes the problem of lack of structured knowledge and semantic constraints in the background technology.

[0017] (2) This invention models the disease-specific disease course semantics of diabetes based on the disease course stage label set and disease course sequence template library, and explicitly encodes the sequential relationship between the pathogenesis stage and the complication stage, thereby extending "path reachability" to "semantic reachability", solving the defect that the disease course stages and clinical sequence of diabetes cannot be distinguished based solely on graph structure or representation learning methods, and improving the clinical consistency and semantic credibility of potential drug reasoning results.

[0018] (3) The present invention determines the candidate drug entity set and generates multi-hop candidate reasoning paths based on the labeled knowledge graph, so that the potential drug discovery process can carry out path reasoning operations around the diabetes disease entity, rather than just based on disease classification or co-occurrence statistics, making potential drug discovery more focused on disease-specific mechanisms and avoiding interference from irrelevant paths, thereby improving the effectiveness and convergence of candidate drug identification.

[0019] (4) Based on the disease stage sequence and disease sequence template library, the present invention performs disease consistency screening on candidate paths and calculates path contribution. The path contribution evaluation basis is jointly constructed by the disease sequence consistency and relationship reliability parameters, so that the output results can distinguish between semantically reasonable paths and only structurally reasonable paths. This overcomes the problem that the path interpretation in the background technology does not have stage and pathological causality, and improves the medical reviewability and verifiability of potential drug interpretation basis.

[0020] (5) The present invention aggregates candidate scores based on path contribution and outputs an interpretable semantic reasoning chain containing disease stage labels, so that potential drug discovery can not only output ranking results, but also output clinically perceptible semantic explanation chains, thereby improving the practical value of interpretive drug discovery and the ability to utilize medical knowledge, and making up for the shortcomings of the explanation output in the background technology, which lacks stage, causal semantics and disease-specific verifiability. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 A flowchart of the method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the semantic knowledge graph relationship for diabetes as provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the stages and prevention levels of diabetes provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the system structure provided in an embodiment of the present invention. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] The purpose of this invention is to provide a method and system for discovering potential drugs for diabetes based on interpretable semantic reasoning, which improves the medical credibility and pathological verifiability of potential drug discovery by constraining the semantics of disease stages and the sequence of disease progression.

[0025] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0026] Figure 1 The method flowchart provided in the embodiments of the present invention is as follows: Figure 1 As shown, this invention provides a method for discovering potential drugs for diabetes based on interpretable semantic reasoning, comprising: Step 100: Extract entity relationships from clinical guidelines, electronic medical records and medical literature to form a set of triplets containing drug entities, target entities, pathway entities, diabetes disease entities and complication entities, and construct a semantic knowledge graph for diabetes based on the set of triplets. Step 200: Construct a set of disease course stage labels and a disease course sequence template library; the set of disease course stage labels includes pathogenesis stage labels and complication stage labels; the disease course sequence template library represents the allowed sequential constraints between disease course stage labels; Step 300: Based on the disease progression template library, assign disease progression stage labels to the entities of diabetes disease, complication, and target and pathway entities that are associated with the entities of diabetes disease or complication in the triplet set, to obtain a labeled knowledge graph. Step 400: Determine the set of candidate drug entities in the labeled knowledge graph, and use the diabetes disease entity as the reasoning endpoint to search for a set of multi-hop candidate reasoning paths for each candidate drug entity in the candidate drug entity set. Step 500: Parse each multi-hop candidate inference path in the multi-hop candidate inference path set to obtain the disease stage sequence and match it with the disease sequence template library. Remove paths that do not meet the preset consistency conditions to obtain a consistent path set. Calculate the path contribution of each consistent path in the consistent path set based on the relationship reliability parameter and the consistency of the disease sequence. The relationship reliability parameter is used to characterize the credibility of the relationship in the path. Step 600: Aggregate the path contribution values ​​corresponding to the same candidate drug entity to obtain the candidate score, output the candidate drug ranking result based on the candidate score, and output at least one consistent path with the highest contribution and the disease stage labels corresponding to each entity in the consistent path as an interpretable semantic reasoning chain.

[0027] Specifically, in this embodiment, when extracting entity relationships from clinical guidelines, electronic medical records, and medical literature, a diabetes-specific entity dictionary and a set of relationship types are first constructed. The diabetes-specific entity dictionary is a terminology table oriented towards the diabetes field, used to store standard names and synonyms for drug entities, target entities, pathway entities, diabetes disease entities, and complication entities. Each entity entry contains at least one standard name and one synonym tag. In a preferred implementation, the number of synonyms is no less than two and no more than ten, to ensure coverage of common clinical expressions while avoiding excessive noise. The entity dictionary is sourced from national or regional diabetes treatment guidelines, electronic medical record samples with a cumulative number of no less than 1000 cases, and medical literature published within the last 10 years. Candidate terms appearing at least three times are selected and added to the entity dictionary through a combination of manual review and automatic statistics. The set of relation types is used to enumerate the semantic relation types that can be extracted between entities, preferably including therapeutic relation, action-target relation, pathway involvement relation, and complication evolution relation. Each relation type corresponds to a unique relation identifier and relation description text in the set, which is used to limit the scope of identifiable relations in the subsequent relation extraction process.

[0028] In this embodiment, after obtaining the diabetes-specific entity dictionary, entity recognition and entity standardization are performed on clinical guidelines, electronic medical records, and medical literature. The entity recognition process employs a word segmentation tool adapted to Chinese medical text and a named entity recognition model. Preferably, each text is divided into analysis units of no more than 500 characters per sentence segment to reduce recognition ambiguity caused by contextual differences. The named entity recognition model scores each candidate segment based on context, selecting segments with scores no lower than a first threshold as entity segments. The first threshold can be set to 80 points or higher on a percentage scale to balance recall and precision. Entity standardization refers to mapping the identified entity segments to standard names in the diabetes-specific entity dictionary. When the same segment can match multiple dictionary entries, the entry with the highest frequency in the entity dictionary is preferably selected as the standard name. When the frequencies are the same and the number of candidate entries does not exceed three, disambiguation can be achieved by prioritizing entries from clinical guidelines, thereby obtaining an entity set. Each entity in the entity set contains three basic pieces of information: standard name, original segment location, and source text identifier.

[0029] In this embodiment, after obtaining the entity set, relation extraction is performed on the entity set based on the relation type set to form a set of triples. During the relation extraction process, this embodiment constructs entity pairs in the order of entity appearance within the same analysis unit, calculates the text distance between entity pairs, and preferably constrains the character distance between entities to no more than 200 characters to improve the semantic relevance of relation determination. For entity pairs that meet the distance constraint, the existence of therapeutic relationship, target relationship, pathway participation relationship, or complication evolution relationship is checked one by one according to the relation type set. When the determination score corresponding to a certain relation type is not lower than a second threshold, the entity pair and the corresponding relation type are combined into a triple. Each triple in the triple set includes a first entity, a second entity, and a relation type identifier. At the same time, an evidence field is set to record the source information of the triple. The evidence field includes at least the source category (clinical guidelines, electronic medical records, or medical literature), source quantity statistics, and evidence level. The evidence level can adopt a three-level structure, setting consistent records from multiple guidelines as the highest level and records appearing only in a single document as the lowest level. For duplicate triples with identical content or the same entity order, this embodiment performs consistency processing by comparing the level of evidence and the number of sources. Triples with a higher level of evidence and at least two sources are preferentially retained as the deduplication result. The final deduplicated triple set is then logically constructed using this triple set, such as... Figure 2 As shown, this provides a unified semantic data foundation for subsequent steps 200 to 600.

[0030] As an example, the data sources in this embodiment are shown in Table 1. These various data sources have a broad basis for use in medical research and clinical practice, and can provide data support for entity relation extraction, evidence field construction, and disease course labeling models.

[0031] Table 1 Data Source Table

[0032] Optionally, in step 200 of this embodiment, when constructing the disease stage label set, the key stage descriptions related to disease progression are first reviewed in diabetes-related clinical guidelines and review literature, and descriptions related to pathogenesis and complications are extracted respectively. Preferably, the pathogenesis stage label is subdivided into no fewer than 3 and no more than 6 label items, such as high-risk stage, prodromal stage, stage with clear glucose metabolism abnormalities, and organ compensation stage; the complication stage label is subdivided into no fewer than 3 and no more than 6 label items, such as early stage of microvascular complications, stage of macrovascular complications, and stage of multi-system complications progression. Each stage label item records three elements in the disease stage label set: stage name, stage type, and stage sequence number. The stage type indicates whether it belongs to the pathogenesis stage label or the complication stage label, and the stage sequence number is an integer from 1 to 9, used to indicate the chronological order in the disease progression.

[0033] In this embodiment, as Figure 3 As shown, to ensure the clinical interpretability and uniformity of the disease progression label set, a quantitative coverage standard was adopted for the selection and merging of label items. Specifically, this embodiment statistically analyzed stage descriptions related to the course of diabetes in at least three domestic guidelines and at least two international guidelines. Stage descriptions appearing in at least two domestic guidelines or in at least one domestic guideline and at least two systematic reviews were considered candidate stages, and then merged into one label item according to the principle of semantic similarity. For example, if "impaired glucose tolerance" and "prediabetes stage" appear multiple times in different literature and point to similar clinical states, they are merged into the same pre-diabetes stage label. For each merged label item, the number of supporting documents was recorded. When the number of supporting documents is not less than three, it is included in the final disease progression label set, thus forming a disease progression label set containing no less than six stage label items.

[0034] In this embodiment, after obtaining the set of disease stage labels, a set of stage sequence templates is constructed. A stage sequence template is defined as a sequence of templates arranged chronologically from multiple disease stage labels, used to describe a clinically common disease path. This embodiment preferably extracts at least three typical disease paths from guidelines, such as a simple path from a high-risk stage to a prodromal stage and then to a stage with confirmed glucose metabolism abnormalities; a complication path from a stage with confirmed glucose metabolism abnormalities to an early stage of microvascular complications and then to a stage of multi-system complications; and a comprehensive path covering the entire disease path. Each disease path is transcribed into a sequence of stage labels of length 3 to 7, with each position in the sequence corresponding to a stage sequence number of a disease stage label, thus forming a set of stage sequence templates. For each stage sequence template, this embodiment records the template number, template type, and template length for subsequent constraint matching.

[0035] In this embodiment, when configuring the allowed sequence constraint rules for the stage sequence template set, firstly, adjacent stage constraints are defined within each stage sequence template. Combinations of two adjacent stage labels in the template are registered as allowed adjacent relationships and recorded in an adjacent relationship table. Preferably, each template contains at least two adjacent relationships. Subsequently, this embodiment counts cross-stage transfers between multiple stage sequence templates. For example, it counts the number of times a description directly connects a high-risk stage to an early stage of microvascular complications occurs. If such a cross-stage transfer occurs less than once in guidelines and less than twice in case reports, this combination is marked as an unacceptable cross-stage relationship. Combining the adjacent relationship table and the unacceptable cross-stage relationships, this embodiment forms allowed sequence constraint rules, stipulating that in the actual disease progression sequence, the stage sequence number can only remain unchanged or increase by one; more than two stages cannot be skipped, and it is not allowed to return from the complication stage to the pathogenesis stage. Finally, this embodiment uniformly registers the stage sequence template set and the corresponding allowed sequence constraint rules into a disease progression template library, providing a quantitative constraint basis for consistency matching and path selection of disease progression sequences in subsequent steps.

[0036] Further, in step 300 of this embodiment, when assigning disease stage labels to the diabetic disease entity and complication entity, the disease stage label set constructed in step 200 is first used to map the diabetic disease entity and complication entity to unique disease stage labels respectively. This embodiment statistically analyzes the co-occurrence information of disease names and complication names with disease stage descriptions in at least three domestic guidelines and at least two international guidelines. When the co-occurrence frequency of a certain stage is not less than 60% of the total co-occurrence frequency, that stage is determined as the corresponding disease stage label. For cases where the co-occurrence distribution is relatively dispersed, it is reviewed by at least three endocrinologists with the title of associate chief physician. The labeling result is determined when at least two physicians agree. This forms a disease stage label set containing multiple disease stage label items, as shown in Table 2. Each stage label item has three core attributes: label name, label type, and stage sequence number. The stage sequence number indicates the disease progression order and ranges from 1 to 9.

[0037] Table 2 Examples of disease stage labels

[0038] In this embodiment, after assigning disease stage labels to disease entities and complication entities, target entities and pathway entities with associations to the aforementioned entities are retrieved from the triplet set. These associations include therapeutic effects, target mechanisms, pathway involvement, and complication progression. For each target entity or pathway entity, this embodiment counts the number of connections it makes with entities at different stages. When the number of connections to a particular stage accounts for more than 50% of its total connections, that stage is determined as the primary stage label for that target entity or pathway entity. For example, if a target entity has three associated triplets with a pathogenesis stage label entity and one associated triplet with a complication stage label entity, then the target entity is assigned a pathogenesis stage label. If no stage meets this ratio condition, the label is determined based on an adjacent stage label selection strategy.

[0039] In this embodiment, the disease progression sequence template library constructed in step 200 is used as the constraint basis when adopting the adjacent stage label selection strategy. Adjacent stage labels refer to stage labels that differ by 1 in stage sequence number. For example, the adjacent stage labels of stage label number 3 are stage labels numbered 2 and 4. This embodiment stipulates that when a target entity or pathway entity has an association with multiple entities at different stage labels and there is no obvious dominant stage, adjacent stage labels can be assigned. However, the difference in stage sequence numbers must not exceed 1; if the difference exceeds 1, the assignment of adjacent stage labels is not performed. The permissible relationships for adjacent stage labels are shown in Table 3. These permissible relationships constrain the propagation direction of disease progression stage labels within the graph structure, prohibiting paths from crossing multiple disease progression stages or regressing to the pathogenesis stage, thereby ensuring clinical rationality.

[0040] Table 3 Examples of Permissible Relationships for Adjacent Stage Labels

[0041] In this embodiment, after assigning disease stage labels to disease entities, complication entities, target entities, and pathway entities, the entities with stage labels are combined with triple sets to construct a labeled knowledge graph. The labeled knowledge graph adds the semantic attribute of disease stage labels to entities on top of the original diabetes-specific semantic knowledge graph, enabling the subsequent path selection and interpretive reasoning processes to retrieve the stage order information of entities along the path. Preferably, each entity records four attributes: entity type, standard name, disease stage label, and stage sequence number, where the stage sequence number corresponds one-to-one with the numbers in Table 2. When the disease stage label cannot be determined during entity assignment, it can be marked as an unknown stage, and paths containing unknown stage labels are given a lower priority during the path selection stage, thus ensuring the robustness of the reasoning results in scenarios with incomplete data.

[0042] Furthermore, in step 400 of this embodiment, when determining the candidate drug entity set in the tagged knowledge graph, the diabetes disease entity is used as the target node, and reachability retrieval is performed on drug entities that have connections with the diabetes disease entity. Reachability retrieval refers to determining whether a drug entity can be reached from the diabetes disease entity through several intermediate entities under a given maximum hop count constraint. In this embodiment, the maximum hop count is preferably set between 3 and 5, and 4 can be used as the default value in specific implementations. To reduce noise from highly connected nodes, this embodiment can stipulate that the number of associated triples for each intermediate entity does not exceed 100; when this number is exceeded, further expansion along that entity is stopped. For drug entities that meet the reachability conditions, the number of paths between the diabetes disease entity and the drug entity is counted. When the number of paths is not less than 1 and preferably not less than 2, the drug entity is included in the candidate drug entity set. Typical parameters for candidate drug entity screening can be found in the example configuration in Table 4.

[0043] Table 4 Examples of candidate drug entity screening parameters

[0044] In this embodiment, after obtaining the candidate drug entity set, a multi-hop candidate reasoning path search is performed on each candidate drug entity in the set. Specifically, starting with a candidate drug entity and ending with a diabetes disease entity, the search expands layer by layer along the connection direction of entities and relationships in the labeled knowledge graph, recording the path sequence from the candidate drug entity, through several intermediate entities, and finally to the diabetes disease entity. The path length is represented by the number of entities or relationships. In this embodiment, paths with 2 to 6 relationships are preferably considered valid multi-hop paths, i.e., containing at least one intermediate entity and at most 5 intermediate entities. During the path expansion process, only predefined combinations of entity types are allowed, such as drug entities, target entities, pathway entities, diabetes disease entities, and complication entities. External entities unrelated to the course of diabetes are preferably not introduced. To control the number of paths, this embodiment can limit each candidate drug entity to a maximum of 100 multi-hop candidate reasoning paths. When the number of path candidates exceeds this, paths passing through key entity type combinations can be retained first. Key entity type combinations can include patterns such as drug entity-target entity-pathway entity-diabetes disease entity, as shown in Table 5.

[0045] Table 5 Examples of Path Search and Pruning Parameters

[0046] After obtaining a preliminary set of multi-hop candidate inference paths, this embodiment performs path deduplication and cyclic path removal on the set. Cyclic paths refer to paths where entities appear repeatedly, meaning at least one entity in the path sequence is visited twice or more. This embodiment considers such paths semantically unsuitable for explaining the unidirectional evolution of the diabetes disease course and therefore excludes them from the multi-hop candidate inference path set for subsequent parsing. Deduplication involves retaining only one representative path from multiple paths where the entity and relation sequences are identical. A statistical field recording the number of occurrences on this representative path can be added, for example, recording that the path appears 3 times in the original set. For paths that differ only in some non-critical intermediate entities, this embodiment can choose to retain or merge them as needed. Preferably, the complete consistency of the entity and relation sequences is used as the deduplication criterion, ensuring that the deduplication rules are simple and clear.

[0047] After deduplication and loop path removal, this embodiment obtains a consistent structural basis for the subsequent step 500 of disease stage sequence parsing, namely, a set of multi-hop candidate reasoning paths for subsequent parsing. Each path in this set is an acyclic path starting from a candidate drug entity and ending at a diabetes disease entity, and satisfies the maximum hop count constraint and entity type constraint. To facilitate the selection of appropriate parameter combinations during implementation, this embodiment provides example configurations for candidate drug entity screening parameters and path search pruning parameters, as shown in Tables 4 and 5. Implementers can adjust the various parameters within the range of values ​​given in Tables 4 and 5 to achieve a balance between interpretability and computational complexity in the path generation results.

[0048] Furthermore, in step 500 of this embodiment, when parsing the pathological stage sequence for each multi-hop candidate inference path in the multi-hop candidate inference path set, the pathological stage tags assigned to each entity are read along the multi-hop candidate inference path in the order of entity appearance, and the read pathological stage tags are arranged sequentially to form a pathological stage sequence. The length of the pathological stage sequence is equal to the number of entities in the multi-hop candidate inference path. In this embodiment, it is preferable to form a pathological stage sequence of length 3 to 7 for paths with 2 to 6 relations. When there are unknown stage tags in the pathological stage sequence, this embodiment registers the unknown stage tags as special tags and marks the pathological stage sequence containing unknown stage tags as a low-confidence sequence, which is used to reduce the priority during subsequent screening, but is not directly eliminated, so as to maintain the coverage of incomplete data.

[0049] In this embodiment, when matching the disease progression sequence with the stage sequence template set, the disease progression sequence is first converted into a stage sequence number sequence, and a consistency check is performed on the number sequence according to the allowed order constraint rules in the disease sequence template library. The consistency check includes at least two types of constraints: one type is the order monotonic constraint, which requires that the number sequence does not show reverse backtracking; the other type is the prohibition of crossing constraints, which requires that the difference between adjacent numbers does not exceed 1, and allows the difference between adjacent numbers to be 0 to support the consecutive appearance of the same disease progression stage label. Only when the disease progression sequence meets the above constraints will it enter template matching; the purpose of template matching is to determine the template that is closest to the disease progression sequence from the stage sequence template set and output the matching result. In this embodiment, matching is preferably performed within a range of no less than 3 and no more than 20 stage sequence templates to ensure template coverage and matching interpretability.

[0050] in, The length of the disease progression sequence; The first stage in the disease progression sequence Each disease stage label corresponds to a stage sequence number, which is derived from the disease stage label set constructed in step 200 and ranges from 1 to 9. This is an indicator function; it takes the value 1 if the condition is true, and 0 otherwise. The number of times the allowed sequence constraint rule has been violated; To ensure consistency in the disease progression sequence, the value ranges from 0 to 1. A larger value indicates that the sequence of disease progression stages conforms more closely to the allowed sequential constraint rules.

[0051] In this embodiment, after obtaining the set of consistent paths, a relationship reliability parameter is determined for each consistent path in the set. The relationship reliability parameter characterizes the credibility of the relationship within the path, and its source is the evidence field of the corresponding triple in the triple set. The evidence field includes at least an evidence level and a number of sources. The evidence level characterizes the strength of the evidence, and the number of sources characterizes the number of times the same triple is supported by independent sources. Preferably, this embodiment sets the evidence level to three levels, corresponding to at least one of guideline evidence, medical record evidence, and literature evidence; and truncates the number of sources with an upper limit to suppress bias caused by abnormally high frequencies. A typical value for the upper limit truncation is 5. For each relationship in the path, the relationship reliability parameter is calculated based on the evidence level and the number of sources, and recorded as the credibility value of that relationship.

[0052] in, This is a relational reliability parameter, with a value range of 0 to... The level of evidence is determined by a value of [value]. Or 3, and determined by the evidence field. For example, 3 can be taken when the triple is supported by both guidelines and literature, 2 can be taken when it is supported only by medical records, and 1 can be taken when it is supported only by a single article. The number of sources is calculated from the evidence field. This indicates that the number of sources is truncated to an upper limit of 5, which is used to prevent the reliability parameter of the relationship from being excessively amplified due to an excessive number of sources.

[0053] In this embodiment, to avoid reliability deviations caused by differences in path length, the reliability parameters of each relation in the consistency path are aggregated at the path level to obtain the path relation reliability. The path relation reliability is used to characterize the overall credibility of the consistency path at the evidentiary level. This embodiment preferably uses a geometric mean for aggregation, so that if the reliability of any relation is too low, it can significantly reduce the overall reliability of the path, thus conforming to the interpretation habit of the weakest link effect in the medical evidence chain. Assuming the consistency path contains 2 to 6 relations, the path relation reliability remains within the range of 0 to 1 and maintains comparability for paths of different lengths.

[0054] in, The path relationship reliability is represented by a value ranging from 0 to... The number of relations in a consistent path is preferably between 2 and 6. For the first in this consistency path The reliability parameters corresponding to each relation are obtained according to the aforementioned method for calculating relation reliability parameters.

[0055] In this embodiment, after obtaining the consistency of disease progression order and the reliability of path relationships, the path contribution of consistent paths is calculated according to preset path contribution calculation rules. Paths meeting preset consistency conditions are then grouped into a consistent path set and output for subsequent ranking. The preset path contribution calculation rules are used to jointly represent the consistency of disease progression order and the reliability of path relationships as a single scalar. This embodiment preferably uses a product rule, ensuring that the path contribution decreases synchronously when any dimension is insufficient, thereby guaranteeing that the output path simultaneously possesses both the rationality of the disease progression order and the credibility of the evidence. The path contribution value range is maintained between 0 and 1, facilitating the subsequent aggregation of multiple consistent paths for the same candidate drug entity.

[0056] in, The path contribution score, with a value ranging from 0 to 1; The reliability of the path relationship is obtained by aggregating the relationship reliability parameters of each relationship in the consistent path. To ensure consistency in the disease progression sequence, the degree to which the sequence of disease stages conforms to the rules governing permissible order is determined; when the consistent path meets the preset consistency conditions, =0 and The value is 1, thus the path contribution is mainly determined by the strength of the evidence field.

[0057] Furthermore, in this embodiment, when aggregating the path contributions corresponding to the same candidate drug entity, the set of consistent paths corresponding to the candidate drug entity is first determined, and the path contribution of each consistent path in the set is read. In this embodiment, paths with a path contribution below a minimum effective threshold are considered low-confidence noise paths and are removed, preferably with a minimum effective threshold of 0.05; after removal, at least one consistent path is retained. This embodiment further sorts the remaining paths from highest to lowest path contribution, and preferably selects the top three paths as a "representative path subset." When the number of consistent paths is less than three, the number of representative path subsets is the number of consistent paths, ensuring that the aggregated score balances stability and interpretability.

[0058] in, The candidate score ranges from 0 to... The value represents the number of path subsets, ranging from 1 to... For the representative path subset, the first Large path contribution The path contribution is calculated using the method given in step 500.

[0059] In this embodiment, when sorting the candidate drug entity set based on candidate scores, the candidate drug ranking results are formed from high to low candidate scores, and the candidate scores are used as the primary sorting key. To avoid unstable ranking due to extremely close candidate scores, this embodiment sets a similarity threshold, preferably 0.01. When the difference between the candidate scores of two candidate drug entities does not exceed 0.01, this embodiment uses the maximum path contribution as the secondary sorting key, that is, comparing the maximum path contribution in the consistency path set corresponding to the candidate drug entity, and the one with the largest contribution is ranked first. If they are still the same, it is preferable to compare the number of paths in the representative path subset, and the one with more paths is ranked first, thereby improving the verifiability of the ranking.

[0060] in, Indicates the order of positions; and For different candidate drug entities; and For the corresponding candidate scores; and These represent the maximum path contribution within the corresponding set of consistent paths; path contribution. Consistent with step 500.

[0061] In this embodiment, when selecting at least one consensus path with the highest contribution for each candidate drug entity, the path with the largest path contribution is selected from the consensus path set corresponding to that candidate drug entity as the main explanation path, and at least this main explanation path is output. Optionally, no more than two paths are output to control the explanation length. When the contribution of the second path is not less than 0.90 of the contribution of the main explanation path, the second path is used as a supplementary explanation path; otherwise, the second path is not output to avoid introducing weak explanation chains that could cause ambiguity.

[0062] in, The main explanation path; This is the set of consistency paths corresponding to the candidate drug entity; For path Path contribution; This indicates the path that maximizes the path contribution.

[0063] In this embodiment, when forming an interpretable semantic reasoning chain, the entity sequence and relation sequence are output for the main interpretation path in the order of entity and relation, and the disease stage labels corresponding to each entity in the path are output simultaneously, forming a multi-segment chain representation of entity-relation-entity. To ensure that the interpretation information can be manually verified, this embodiment preferably controls the number of relations in a single interpretable semantic reasoning chain to be between 2 and 6, and the number of corresponding entities to be between 3 and 7, and outputs the disease stage labels in the order of the path as a disease stage sequence; when there are consecutive identical labels in the disease stage sequence, consecutive appearance is retained to reflect the disease stage stagnation phenomenon. For each relation in the path, this embodiment can also output the evidence summary information corresponding to the relation. The evidence summary information includes at least a discrete description of the evidence level and the number of sources, which is used to support the interpretation basis of credibility, wherein the evidence level is preferably divided into 3 levels, and the number of sources is preferably truncated to 5.

[0064] In this embodiment, when outputting the candidate drug ranking results and interpretable semantic reasoning chains, it is preferable to output the top 20 candidate drug entities as the candidate drug ranking results, and output at least one interpretable semantic reasoning chain for each candidate drug entity; when the total number of candidate drug entities is less than 20, all candidate drug entities are output. This embodiment can also set a lower threshold for candidate scores to suppress weak candidates; the lower threshold for candidate scores is preferably 0.20, and candidates with scores below 0.20 are not included in the candidate drug ranking results. Through the above aggregation, ranking, and interpretation chain output, this embodiment, without introducing too many manually tuned weights, enables the candidate drug ranking to be both comparable and verifiable, and ensures that the ranking basis of each candidate drug entity can be interpreted through at least one consistent path and the corresponding disease stage label.

[0065] Corresponding to the above methods, such as Figure 4 As shown, this embodiment also provides a potential drug discovery system for diabetes based on interpretable semantic reasoning, including: The entity relation extraction and knowledge graph construction unit is used to extract entity relations from clinical guidelines, electronic medical records and medical literature, forming a set of triplets containing drug entities, target entities, pathway entities, diabetes disease entities and complication entities, and constructing a diabetes-specific semantic knowledge graph based on the set of triplets. The disease course semantic modeling unit is used to construct a set of disease course stage labels and a disease course sequence template library. The set of disease course stage labels includes pathogenesis stage labels and complication stage labels. The disease course sequence template library represents the allowed sequential constraints between disease course stage labels. The disease course labeling unit is used to assign disease course stage labels to entities related to diabetes, complication entities, and target entities and pathway entities that are associated with diabetes or complication entities in the triple set, based on the disease course sequence template library, to obtain a labeled knowledge graph. The candidate drug path generation unit is used to determine the set of candidate drug entities in the labeled knowledge graph, and to search for a set of multi-hop candidate reasoning paths for each candidate drug entity in the candidate drug entity set, with the diabetes disease entity as the reasoning endpoint. The disease course consistency screening and path contribution calculation unit is used to parse each multi-hop candidate inference path in the multi-hop candidate inference path set to obtain the disease course stage sequence and match it with the disease course sequence template library. Paths that do not meet the preset consistency conditions are eliminated to obtain a consistent path set. The path contribution of each consistent path in the consistent path set is calculated based on the relationship reliability parameter and the consistency of the disease course sequence. The relationship reliability parameter is used to characterize the credibility of the relationship in the path. The candidate score and interpretable chain output unit is used to aggregate the path contribution of the same candidate drug entity to obtain the candidate score, output the candidate drug ranking result based on the candidate score, and output at least one consistency path with the highest contribution and the disease stage label corresponding to each entity in the consistency path as an interpretable semantic reasoning chain.

[0066] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.

[0067] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for discovering potential drugs specific to diabetes based on interpretable semantic reasoning, characterized in that, include: Entity relationships are extracted from clinical guidelines, electronic medical records, and medical literature to form a set of triplets containing drug entities, target entities, pathway entities, diabetes disease entities, and complication entities. A semantic knowledge graph for diabetes is then constructed based on the set of triplets. Construct a set of disease course stage labels and a disease course sequence template library; the set of disease course stage labels includes pathogenesis stage labels and complication stage labels; the disease course sequence template library represents the allowed sequential constraints between the disease course stage labels. Based on the disease progression template library, disease progression stage labels are assigned to the diabetes disease entity, the complication entity, the target entity and the pathway entity that are associated with the diabetes disease entity or the complication entity in the triple set, thereby obtaining a labeled knowledge graph; In the labeled knowledge graph, a set of candidate drug entities is determined, and with the diabetes disease entity as the reasoning endpoint, a set of multi-hop candidate reasoning paths is obtained by searching for each candidate drug entity in the set of candidate drug entities. For each multi-hop candidate inference path in the multi-hop candidate inference path set, the disease stage sequence is parsed and matched with the disease sequence template library. Paths that do not meet the preset consistency conditions are eliminated to obtain a consistent path set. The path contribution of each consistent path in the consistent path set is calculated based on the relationship reliability parameter and the consistency of the disease sequence. The relationship reliability parameter is used to characterize the credibility of the relationship in the path. The path contribution values ​​corresponding to the same candidate drug entity are aggregated to obtain a candidate score. Based on the candidate score, the candidate drug ranking result is output, and at least one consistent path with the highest contribution and the disease stage label corresponding to each entity in the consistent path are output as an interpretable semantic reasoning chain.

2. The method for discovering potential drugs for diabetes based on interpretable semantic reasoning according to claim 1, characterized in that, Entity relationships are extracted from clinical guidelines, electronic medical records, and medical literature to form a set of triplets containing drug entities, target entities, pathway entities, diabetes disease entities, and complication entities. Based on this set of triplets, a semantic knowledge graph for diabetes is constructed, including: Construct a diabetes-specific entity dictionary and a set of relation types; the diabetes-specific entity dictionary covers the standard names and synonyms of drug entities, target entities, pathway entities, diabetes disease entities, and complication entities; Based on the diabetes-specific entity dictionary, entity recognition and entity standardization are performed on the clinical guidelines, the electronic medical records, and the medical literature to obtain an entity set. Based on the set of relation types, relation extraction is performed on the entity set to generate entity pairs and relations, and the entity pairs and relations are combined to form the set of triples.

3. The method for discovering potential drugs for diabetes based on interpretable semantic reasoning according to claim 1, characterized in that, The methods for forming the set of triples include: Each triplet in the set of triplets is assigned a relation type identifier; the relation type identifier includes at least one or more of the following: therapeutic effect relation, action target relation, pathway involvement relation, and complication evolution relation. Each triple in the set of triples is assigned an evidence field; the evidence field is used to characterize that the triple originates from at least one of the clinical guidelines, the electronic medical records, and the medical literature. The duplicate triples are processed for consistency based on the evidence field to obtain the set of duplicate triples.

4. The method for discovering potential drugs for diabetes based on interpretable semantic reasoning according to claim 1, characterized in that, Construct a set of disease stage tags and a disease sequence template library, including: Construct a set of label items for the pathogenesis stage label and a set of label items for the complication stage label, and form the disease course stage label set based on the set of label items for the pathogenesis stage label and the set of label items for the complication stage label; Construct a set of stage sequence templates; the stage sequence template is a template sequence composed of the disease stage labels in sequence. Configure sequential constraint rules for each stage sequence template in the stage sequence template set; the sequential constraint rules are used to limit the sequential relationship between adjacent stages and across stages in the template sequence; The set of stage sequence templates and the rules for allowing sequential order constraints are combined to form the disease course sequence template library.

5. The method for discovering potential drugs for diabetes based on interpretable semantic reasoning according to claim 1, characterized in that, Based on the disease progression template library, disease progression stage labels are assigned to the diabetes disease entity, the complication entity, and the target entity and pathway entity that are associated with the diabetes disease entity or the complication entity in the triplet set, resulting in a labeled knowledge graph, including: Based on the set of disease stage labels, assign corresponding disease stage labels to the diabetes disease entity and the complication entity; Retrieve the target entity and the pathway entity that are associated with the diabetes disease entity or the complication entity from the triple set, and assign the target entity and the pathway entity the disease stage label corresponding to the diabetes disease entity or the complication entity that are associated with the target entity or the pathway entity in the triple set, or assign adjacent stage labels that conform to the allowed sequence constraint rules. The labeled knowledge graph is constructed by combining the entities that have been labeled with disease stage tags with the set of triples.

6. The method for discovering potential drugs for diabetes based on interpretable semantic reasoning according to claim 1, characterized in that, A set of candidate drug entities is determined in the labeled knowledge graph, and using the diabetes disease entity as the reasoning endpoint, a set of multi-hop candidate reasoning paths is obtained for each candidate drug entity in the candidate drug entity set, including: Using the diabetes disease entity as the target node, an reachability search is performed in the labeled knowledge graph to obtain drug entities that have a connection relationship with the diabetes disease entity, and the obtained drug entities constitute the candidate drug entity set; For each candidate drug entity in the candidate drug entity set, a path search is performed with the candidate drug entity as the starting point and the diabetes disease entity as the reasoning endpoint to obtain the multi-hop candidate reasoning path set formed by sequentially connecting entities and relations. The set of multi-hop candidate inference paths is deduplicated and looped paths are removed to obtain the set of multi-hop candidate inference paths for subsequent parsing.

7. The method for discovering potential drugs for diabetes based on interpretable semantic reasoning according to claim 4, characterized in that, For each multi-hop candidate inference path in the multi-hop candidate inference path set, the disease stage sequence is parsed and matched with the disease sequence template library. Paths that do not meet the preset consistency conditions are removed to obtain a consistent path set, including: Read the disease stage label of the entity in each of the multi-hop candidate inference paths; The disease stage tags read in the path order are used to form the disease stage sequence; The disease course stage sequence is matched with the stage sequence template set to obtain the matching result; The preset consistency condition is set to simultaneously satisfy the following conditions: the disease stage sequence satisfies the allowed sequence constraint rule, and the same disease stage label is allowed to appear consecutively in the disease stage sequence; Based on the matching result, it is determined whether the preset consistency condition is met, and the paths that meet the preset consistency condition constitute the consistency path set.

8. The method for discovering potential drugs for diabetes based on interpretable semantic reasoning according to claim 1, characterized in that, The path contribution of each consistent path in the consistent path set is calculated based on the consistency between the relationship reliability parameter and the disease course sequence. The relationship reliability parameter is used to characterize the credibility of the relationship in the path, including: The reliability parameter of each relation in the consistency path is determined based on the evidence field of the triples contained in the consistency path in the triple set; The consistency of the disease course sequence in the consistency path is determined based on the matching result between the disease course stage sequence corresponding to the consistency path and the allowed sequence constraint rule. Based on the preset path contribution calculation rules, the path contribution of the consistent path is calculated based on the consistency between the relationship reliability parameter and the disease course sequence.

9. The method for discovering potential drugs for diabetes based on interpretable semantic reasoning according to claim 1, characterized in that, The contribution of the paths corresponding to the same candidate drug entity is aggregated to obtain a candidate score. Based on the candidate score, the candidate drug ranking result is output, and at least one consistent path with the highest contribution and the disease stage labels corresponding to each entity in the consistent path are output as an interpretable semantic reasoning chain, including: Aggregate the path contribution values ​​in the consistency path set corresponding to the same candidate drug entity to obtain the candidate score of the candidate drug entity; The candidate drug entity set is sorted according to the candidate scores to obtain the candidate drug ranking result; For each candidate drug entity, at least one consistency path with the highest path contribution is selected from the corresponding consistency path set, and the disease stage label corresponding to each entity in the consistency path is output to form the interpretable semantic reasoning chain.

10. A potential drug discovery system for diabetes based on interpretable semantic reasoning, characterized in that, include: The entity relation extraction and knowledge graph construction unit is used to extract entity relations from clinical guidelines, electronic medical records and medical literature, forming a set of triplets containing drug entities, target entities, pathway entities, diabetes disease entities and complication entities, and constructing a diabetes-specific semantic knowledge graph based on the set of triplets. A disease course semantic modeling unit is used to construct a disease course stage label set and a disease course sequence template library; the disease course stage label set includes pathogenesis stage labels and complication stage labels; the disease course sequence template library represents the allowed sequential constraints between the disease course stage labels. The disease course labeling unit is used to assign disease course stage labels to the diabetes disease entity, the complication entity, and the target entity and the pathway entity that are associated with the diabetes disease entity or the complication entity in the triple set according to the disease course sequence template library, so as to obtain a labeled knowledge graph. The candidate drug path generation unit is used to determine a set of candidate drug entities in the labeled knowledge graph, and to search for a set of multi-hop candidate reasoning paths for each candidate drug entity in the set of candidate drug entities, with the diabetes disease entity as the reasoning endpoint. The disease course consistency screening and path contribution calculation unit is used to parse each of the multi-hop candidate inference paths in the multi-hop candidate inference path set to obtain a disease course stage sequence and match it with the disease course sequence template library, eliminate paths that do not meet the preset consistency conditions to obtain a consistent path set, and calculate the path contribution of each consistent path in the consistent path set based on the relationship reliability parameter and the consistency of the disease course sequence; the relationship reliability parameter is used to characterize the credibility of the relationship in the path; The candidate scoring and interpretable chain output unit is used to aggregate the path contribution of the same candidate drug entity to obtain a candidate score, output the candidate drug ranking result based on the candidate score, and output at least one consistency path with the highest contribution and the disease stage label corresponding to each entity in the consistency path as an interpretable semantic reasoning chain.