Knowledge graph completion method, device and product

By traversing the subgraphs of the knowledge graph, generating candidate triple information using a large generative model and performing enhanced retrieval, challenging the large model to construct counterexamples, and arbitrating the confidence level using an arbitration model, the problem of balancing accuracy and consistency in knowledge graph completion is solved, achieving highly reliable enterprise-level knowledge management.

CN120806103AActive Publication Date: 2025-10-17SHANGHAI COOPERS TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511302201.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-10-17
Estimated Expiration
2045-09-12

AI Technical Summary

Technical Problem

Existing knowledge graph completion methods struggle to balance accuracy and consistency, failing to meet the high reliability requirements of enterprise-level knowledge management.

Method used

By traversing the knowledge graph subgraphs, generating candidate triple information using a large generative model, and performing enhanced retrieval, counterexamples and counterfactual questions are constructed by questioning the large model. Finally, an arbitration model is used to synthesize opinions from all parties and determine the confidence level of the candidate triples to ensure the quality of the completion.

Benefits of technology

It improves the accuracy and reliability of knowledge graphs, making them suitable for various enterprise-level knowledge management systems while balancing semantic richness and structural consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806103A_ABST
    Figure CN120806103A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of information, and discloses a knowledge graph completion method, device and product, and the method comprises the steps: traversing a to-be-completed knowledge graph sub-graph, and obtaining the structure and context information of the to-be-completed sub-graph; constructing a prompt text according to the to-be-complemented sub-graph structure and the context information; generating candidate triple information by generating a large model according to the prompt text; performing retrieval enhancement on the candidate triple information to obtain retrieval information; generating question information for the candidate triple information and the retrieval information through a question large model; calculating a confidence coefficient through an arbitration large model according to the candidate triple information, the retrieval information and the question information; and storing the candidate triple information in response to the condition that the confidence is greater than or equal to the preset threshold, thereby ensuring that each generated triple data is reliable, considering semantic richness and structural consistency, improving the accuracy of the complemented knowledge graph, and being suitable for various enterprise-level knowledge management systems.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of information technology, and in particular to a knowledge graph completion method, device and product. BACKGROUND

[0002] As a structured representation form of entities and their relationships, knowledge graphs have been widely applied in search recommendation, intelligent question answering, risk identification and other fields. However, in the actual construction process, due to the cost of data acquisition and the difficulty of relationship annotation, knowledge graphs often have the problem of missing relationships or incomplete triples, which significantly affects the accuracy and robustness of downstream tasks.

[0003] To alleviate the sparsity of the graph, knowledge graph completion (KGC) has gradually become a research hotspot. Existing methods mainly include: graph embedding-based completion methods (such as TransE, RotatE), which model the potential association between entities and relationships through vector space, but are difficult to capture complex reasoning paths and lack generalization ability; rule and path-based reasoning methods (such as logical rule extraction, path ranking), which have certain interpretability, but rely on explicit paths and are difficult to adapt to large-scale graphs; and pre-trained language model-based methods (such as GPT (Generative Pre-trained Transformer) and Qwen (Qwen), etc.) that have emerged in recent years, which can generate missing relationships using natural language and have shown strong potential, but lack structured constraints and are prone to generate fictitious triples, which lack reliability.

[0004] In summary, there is currently a lack of a general completion method that can balance accuracy, consistency and verifiability, making it difficult to meet the needs of enterprise-level knowledge management for high reliability and controllability. SUMMARY

[0005] An object of the present application is to provide a knowledge graph completion method, device and product, at least to solve the problem that the accuracy and consistency of the knowledge graph completion result are difficult to balance, and it is difficult to meet the high reliability requirements of enterprise-level knowledge management.

[0006] To achieve the above object, some embodiments of the present application provide the following aspects:

[0007] In a first aspect, the present application provides a knowledge graph completion method, comprising:

[0008] traversing a subgraph of a knowledge graph to be completed to obtain a structure of the subgraph to be completed and context information;

[0009] constructing a prompt text according to the structure of the subgraph to be completed and the context information;

[0010] generating candidate triple information according to the prompt text through a large model;

[0011] performing retrieval enhancement on the candidate triple information to obtain retrieval information;

[0012] generating challenge information through a challenge large model by challenging the candidate triple information and the retrieval information;

[0013] calculating a confidence degree through an arbitration large model according to the candidate triple information, the retrieval information and the challenge information;

[0014] in response to the confidence degree being greater than or equal to a preset threshold, saving the candidate triple information.

[0015] In a second aspect, the present application further provides an electronic device, which comprises:

[0016] one or more processors; and

[0017] a memory storing computer program instructions which, when executed, cause the processor to perform the steps of the method of any one of the above aspects.

[0018] In a third aspect, the present application further provides a computer program product comprising computer programs / instructions which, when executed by a processor, implement the steps of the method of any one of the above aspects.

[0019] Compared with the related art, in the scheme provided by the embodiments of the present application, after traversing the subgraph of the knowledge graph, candidate triple information is generated based on the language generation capability of a large model and a prompt text, and after enhancement retrieval, a challenge large model is used to actively construct counterexamples and counterfactual questions, and then an arbitration model is used to integrate opinions from all parties to determine the confidence degree of the candidate triple, thereby ensuring the quality of completion, ensuring that each generated triple data is reliable, taking into account semantic richness and structural consistency, improving the accuracy of the completed knowledge graph, and being applicable to various enterprise-level knowledge management systems. BRIEF DESCRIPTION OF DRAWINGS

[0020] One or more embodiments are illustrated by way of example in the drawings, which are not intended to limit the embodiments and in which the same or similar elements are denoted by the same reference numerals unless otherwise specified, and the drawings are not intended to be to scale.

[0021] Figure 1 a flowchart of a knowledge graph completion method provided for an exemplary embodiment of the present disclosure;

[0022] Figure 2An exemplary structural diagram of the electronic device provided in some embodiments of the present application. DETAILED DESCRIPTION

[0023] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0024] Figure 1 A knowledge graph completion method is provided as an exemplary embodiment of the present disclosure. The knowledge graph completion method includes:

[0025] S101. Traverse the subgraph of the knowledge graph to be completed and obtain the structure and context information of the subgraph to be completed.

[0026] Specifically, the local subgraph of the knowledge graph to be completed is directly traversed to obtain the subgraph topology (nodes, edges, paths), entity type labels (such as "disease", "drug"), relationship semantic constraints (including the domain and range of the relationship in the knowledge graph ontology, that is, the set of allowed head and tail entity types), examples (the tail entity type of the relationship "treatment" should be "disease" or "symptom"), historical completion records, and knowledge version information. The relationship semantic constraints are derived from: the knowledge graph's own ontology (Schema) definition, type rules of external standard knowledge bases (such as UMLS, SNOMED CT, Wikidata), and high-frequency patterns based on statistical induction of graph data.

[0027] Taking medical knowledge graph completion as an example, the scenario is that when a hospital builds a diabetes-related knowledge graph, some treatment and complication information is missing. This step is the starting point of the entire medical knowledge graph completion process. Its core goal is to detect and locate key information such as treatment plans and complication relationships that are missing from the hospital's existing knowledge graph under the "diabetes" theme. By using the "diabetes" entity as an anchor, a graph database (such as Neo4j) or a graph embedding search engine can be used to extract all nodes and edges within a radius of k hops (k ≥ 2). By traversing the relevant subgraph, the subgraph structure and existing triples can be obtained, and the triples that need to be completed can be determined.

[0028] S102: Constructing a prompt text according to the subgraph structure to be completed and context information.

[0029] Specifically, the prompt text is constructed based on the sub-graph information, for example:

[0030] Given: (diabetes, complications, retinopathy), (diabetes, treatment, metformin). Please complete the 'complications' of another 'diabetes'. Output format: (head entity, relationship, tail entity), and attach the reason for the generation.

[0031] The output strictly follows the following format example:

[0032] json

[0033] {

[0034] "triple":["diabetes","complications","xx"],

[0035] "reasoning":"xxxxx",

[0036] "evidence":"xxxxx",

[0037] }

[0038] ”

[0039] S103: Generate candidate triple information by generating a large model according to the prompt text.

[0040] Specifically, the above prompt text is input into a large model to generate a structured output.

[0041] Based on the constructed prompt text, it is input into a large generation model to generate formatted triples. The large generation model can be a large language model (LLM) such as GPT or Qwen. Formatted triples are triples with a specific structure, including a head entity, a relation, and a tail entity, expressed in the format (head entity, relation, tail entity). Based on the prompt text, candidate triples are generated, including the triple T_c, reasoning, and evidence. The reasoning and evidence constitute the generated reason R_gen.

[0042] For example:

[0043] {

[0044] "triple":["diabetes","complications","fatigue"],

[0045] "reasoning":"Fatigue is a common symptom in diabetic patients. Long-term high blood sugar affects the function of the nervous system and leads to abnormal energy metabolism.",

[0046] "evidence":"Some documents indicate that approximately 60% of diabetic patients report chronic fatigue.",

[0047] }

[0048] S104, retrieve enhancement is performed on the candidate triple information to obtain retrieval information.

[0049] Specifically, based on the candidate triple T_c and R_gen, a retrieval query statement is constructed, for example, the above-mentioned candidate triple information can be constructed as "Diabetes Complications Fatigue Latest Research". Then, a knowledge base interface is called to perform retrieval. The latest research results related to T_c are obtained as external verification evidence, and the retrieval results are shown. For example:

[0050] "2023 research shows: 'Chronic Fatigue Syndrome is not a direct complication of type 2 diabetes: a cohort study', indicating that there is a correlation between chronic fatigue syndrome and diabetes, but no direct causal evidence."

[0051] The retrieval information is recorded as R_rag, which is used for subsequent questioning and arbitration.

[0052] S105, the questioning big model generates questioning information for the candidate triple information and the retrieval information.

[0053] Specifically, the questioning big model is a language model for generating counterexamples and questioning information. The candidate triple information is questioned and asked, and the questioning information is output.

[0054] For example, R_gen and R_rag are questioned. For example, "Fatigue is a common symptom in patients with diabetes, long-term high blood sugar affects the function of the nervous system, leading to abnormal energy metabolism. In some documents, it is pointed out that about 60% of patients with diabetes report chronic fatigue. 2023 research shows: 'Chronic Fatigue Syndrome is not a direct complication of type 2 diabetes: a cohort study', indicating that there is a correlation between chronic fatigue syndrome and diabetes, but no direct causal evidence, does it affect the conclusion?"

[0055] Multiple instantiated questioning sentences, R_gen and R_rag are input into the questioning big model to generate the final questioning prompt with coherent semantics and natural tone:

[0056] Fatigue is a common symptom in patients with diabetes, and long-term hyperglycemia affects the function of the nervous system, leading to abnormal energy metabolism. In some documents, it is pointed out that about 60% of patients with diabetes report chronic fatigue. A 2023 study shows that 'Chronic Fatigue Syndrome is not a direct complication of type 2 diabetes: a cohort study', indicating that there is a correlation between chronic fatigue syndrome and diabetes, but no direct causal evidence. Does this affect the conclusion?

[0057] The skeptical model output includes a skepticism strength score and a skeptical reason. The skepticism strength score S_skeptic ∈ [0, 1], for example: S_skeptic = 0.7, reason: research shows that it is not listed as an independent complication; and a new study questions causality.

[0058] S106, according to the candidate triple information, the search information and the skeptical information, calculate the confidence through the arbitration large model.

[0059] Receive candidate triple T_c, generate reason R_gen, RAG search result R_rag, and skeptical information; construct structured arbitration prompt text P_arbitrate and input arbitration large model, for example:

[0060] You are a medical knowledge graph review expert, please evaluate whether the following candidate triple is credible:

[0061]

Candidate Triple

[0062] (Diabetes, Complication, Fatigue)

[0063]

Generated Reason

[0064] Reasoning: Fatigue is a common symptom in patients with diabetes, and long-term hyperglycemia affects the function of the nervous system, leading to abnormal energy metabolism.

[0065] Evidence: In some documents, it is pointed out that about 60% of patients with diabetes report chronic fatigue.

[0066]

RAG search result R_rag

[0067] A 2023 study shows that 'Chronic Fatigue Syndrome is not a direct complication of type 2 diabetes: a cohort study', indicating that there is a correlation between chronic fatigue syndrome and diabetes, but no direct causal evidence.

[0068]

Skeptical Content

[0069] Fatigue is a symptom, not a disease; new research questions causality.

[0070] S_skeptic=0.7

[0071] Please output:

[0072] 1. Comprehensive confidence (between 0 and 1, retaining two decimal places)

[0073] 2. Does it support writing? (Yes / No)

[0074] 3. Reasons for the ruling (explain the basis for the scoring)

[0075] The output strictly follows the following format example:

[0076] {

[0077] "confidence":"xx",

[0078] "decision":"yes or no",

[0079] "reason":"xxxx"

[0080] }

[0081] S107 : In response to the confidence being greater than or equal to a preset threshold, save the candidate triplet information.

[0082] If a threshold C is preset (eg, 0.8), then if the confidence of the candidate triple information is greater than 0.8, the candidate triple information is saved.

[0083] The candidate triples are written into the knowledge graph and synchronized to the evolutionary memory. The stored information includes: the confidence assigned to the candidate triples; the generation time, R_gen, R_rag, and verification path (for the input and output of the entire generation, challenge, and arbitration models).

[0084] The above-mentioned generation big model, questioning big model, and arbitration big model can be the same big model or multiple big models. The types of the above-mentioned big models can be selected from different types of big models. They can also be big models trained for a specific scenario, or general big models such as GPT, Qwen and other large language models (LLM).

[0085] In this embodiment, after traversing the knowledge graph subgraph, the language generation capability of the large model is utilized to generate candidate triple information based on the prompt text, and after enhanced retrieval, the large model is actively constructed to construct counterexample and counterfactual questions, and the arbitration model is used to integrate the opinions of all parties to determine the confidence of the candidate triple, thereby ensuring the quality of the completion, ensuring the reliability of each generated triple data, taking into account the semantic richness and structural consistency, improving the accuracy of the completed knowledge graph, and being suitable for various enterprise-level knowledge management systems.

[0086] In one embodiment, the knowledge graph completion method further comprises:

[0087] performing type consistency verification and / or semantic conflict detection and / or common sense violation detection on the candidate triple information by a reasoning large model, and outputting logical consistency information;

[0088] calculating the confidence by an arbitration large model according to the candidate triple information, the retrieval information, the logical consistency information, and the questioning information.

[0089] Specifically, the reasoning large model checks whether the candidate triple violates the graph logic, and the reasoning large model can be a general large model or a large model trained separately for reasoning. It includes type consistency, semantic conflict detection, and common sense violation.

[0090] For example: type consistency: the tail entity "fatigue" belongs to the "symptom" category, and the value domain constraint of the "complication" relationship is the "disease" category, so the types are inconsistent.

[0091] Semantic conflict detection: Is there a relationship (such as "relieve", "improve", "reduce") that is opposite in semantics to "complication" pointing to the same tail entity "fatigue"? Table 1 shows the representation of semantic opposite relationships.

[0092] For example: if there is (metformin, relieve, fatigue) and (metformin, treat, diabetes) in the graph, it indicates that "fatigue" is an intervenable symptom and may not be suitable as a stable "complication" of "diabetes";

[0093] Common sense violation: Is it against medical common sense?

[0094] The output logical consistency information includes a logical consistency score and a reason. The logical consistency score S_logic ∈ [0, 1], for example: S_logic = 0.2, reason: (type inconsistency: tail entity "fatigue" is of the "symptom" category, but the "complication" relationship requires the "disease" category).

[0095] Through the above steps, the reasoning large model can provide a logical verification mechanism based on logical rules, semantic consistency and common sense constraints, so that the arbitration large model can consider the factors of semantics, logic and facts when calculating the confidence.

[0096] Table 1

[0097]

[0098] In this embodiment, by increasing the logical reasoning verification of the reasoning large model on the candidate triple information, the type, semantics and common sense of the candidate triple can be further constrained, ensuring that each generated triple has "semantic support", taking into account the semantic richness and structural consistency, and improving the accuracy of the completed knowledge graph.

[0099] In one embodiment, after the step of saving the candidate triple information in response to the confidence being greater than or equal to the preset threshold, the knowledge graph completion method further comprises:

[0100] The confidence of the candidate triple information decreases over time;

[0101] In response to the confidence being less than the preset threshold, recalculating the confidence of the candidate triple information;

[0102] In one embodiment, the knowledge graph completion method further comprises recalculating the confidence of the candidate triple information at regular intervals.

[0103] Specifically, the knowledge credit score of all written triples decays over time:

[0104]

[0105] where S(t) is the confidence at time t, is the initial knowledge credit score, t is the current time, is the triple writing time, and λ is the decay coefficient.

[0106] S(t) represents the degree of confidence of the triple at the current time, and the value range is [0, 1]. The lower the score, the more likely the knowledge is outdated and needs to be revalidated. For example, S(2025)=0.5, indicating that the confidence of the triple in 2025 is 0.5. is the confidence assigned when the triple is first validated and written into the graph, . This value is determined by the confidence score of the arbitration large model. For example, if the arbitration large model outputs confidence (confidence) = 0.8, then t is the real-time time when the confidence is calculated or revalidation decision is made, for example, the current system time: August 19, 2025. is the timestamp when the triple is first successfully written into the knowledge graph, accurate to day or hour, e.g. . is the time duration of the triple, the time elapsed since the triple is written, usually in "years". If t and in years, then represents the number of years elapsed. For example, t = 2025, , then years. λ is a parameter that controls the decay speed of the credit score, λ > 0. The larger λ is, the faster the triple becomes outdated. For example, in the medical field: λ ≈ 0.1 / year (about 7 years to decay to 50% of the initial value).

[0107] When the confidence decreases over time, automatically re-verify the triples with credit scores below the threshold, or periodically (e.g., every quarter) re-verify the triples with credit scores below the threshold. The RAG retrieval can be re-executed to query the latest literature to support, update the credit score or mark as "to be manually reviewed".

[0108] In this embodiment, the confidence is associated with time, and the confidence of the triple will naturally decrease over time, avoiding long-term dependence on outdated or invalid knowledge. When the credit score is below the threshold, automatic re-verification is triggered, for example, by RAG retrieval to find the latest papers, databases or authoritative sources, which can automatically discover, update or even replace outdated knowledge without human intervention. This decay + re-verification mechanism upgrades the knowledge graph from "static storage" to "dynamic evolution", ensuring its timeliness, reliability and self-updating ability.

[0109] In an embodiment, the generating, by the large model, of the challenge information based on the candidate triple information and the retrieval information comprises:

[0110] performing type challenge and / or timeliness challenge on the candidate triple information;

[0111] generating, by the large model, a challenge degree score and a challenge reason to output the challenge information.

[0112] Specifically, the large model generates challenge content based on a template library, matches a template from a preset challenge prompt template library according to the relationship type r and the entity category of T_c, and the template is as follows:

[0113] Table Two

[0114]

[0115] Substitute the variables in the matched template with actual values (variable filling):

[0116] T01 template:

[0117] "Does the {tail_entity} belong to the {expected_type} type, but the relationship usually requires the tail entity to be a {expected_type}, is there an error?"

[0118] After filling in:

[0119] "Does the fatigue belong to the symptom type, but the relationship usually requires the tail entity to be a disease, is there an error?"

[0120] Where {tail_entity} represents the tail entity of the triple, i.e. the entity that the relationship points to. For example, in the example, it is "fatigue".

[0121] {type} represents the type of the tail entity currently labeled or identified. For example, in the example, "symptom".

[0122] {expected_type} represents the type requirement of the tail entity for this relationship, i.e. according to the definition of the relationship in the knowledge graph, the tail entity should belong to this type. For example, in the example, "disease", indicating that the "complication" relationship usually requires the tail entity to be a "disease" category.

[0123] T05 template:

[0124] "Is this knowledge based on earlier research? {R_gen}, {R_rag}, does it affect the conclusion?"

[0125] Enhanced with R_gen and R_rag:

[0126] "Is this knowledge based on earlier research? Fatigue is a common symptom in patients with diabetes, and long-term high blood sugar affects the function of the nervous system, leading to abnormal energy metabolism. In some documents, it is pointed out that about 60% of patients with diabetes report chronic fatigue, and a 2023 study shows that 'Chronic Fatigue Syndrome is not a direct complication of type 2 diabetes: a cohort study', indicating that there is a correlation between chronic fatigue syndrome and diabetes, but no direct causal evidence, does it affect the conclusion?"

[0127] Input the instantiated questioning statements, R_gen and R_rag into the questioning large model together, and generate the final questioning prompt with semantic coherence and natural tone:

[0128] Is there an error in the relationship that usually requires the tail entity to be a disease type? In addition, is this knowledge based on earlier research? Fatigue is a common symptom in patients with diabetes, and long-term hyperglycemia affects the function of the nervous system, leading to abnormal energy metabolism. It is pointed out in some documents that about 60% of patients with diabetes report chronic fatigue, and a 2023 study shows that 'Chronic Fatigue Syndrome is not a direct complication of type 2 diabetes: a cohort study', indicating that there is a correlation between chronic fatigue syndrome and diabetes, but no direct causal evidence, does it affect the conclusion?

[0129] The skeptical large model output skepticism score S_skeptic ∈ [0, 1], for example S_skeptic = 0.7, and the reason is that "fatigue" is a symptom rather than a disease, and it is not listed as an independent complication; and a new study questions the causality.

[0130] In this embodiment, by questioning the large model for type questioning (such as whether the entity category matches), structural errors (such as "fatigue" being incorrectly labeled as a disease) can be quickly found. Through the timeliness of the questioning, outdated conclusions or conflicts between new and old research can be found, avoiding reliance on outdated information. This makes subsequent arbitration more accurate, reduces the flow of "illusory knowledge" or incorrect triples into the final knowledge base, and improves overall credibility.

[0131] In one embodiment, the knowledge graph completion method further comprises adding counterexamples to the knowledge base in response to the skepticism score in the skepticism information being greater than a skepticism score threshold.

[0132] Specifically, the counterexample information is used to explicitly mark that the candidate triple may have errors or disputes in the current semantic context or research conclusion. For example, if the skeptical large model identifies that the candidate triple "fatigue → complication → diabetes" has entity type mismatch and causality problems, and outputs a skepticism score S_skeptic = 0.8, which is greater than the preset threshold (such as 0.6), a corresponding counterexample triple "fatigue → non-complication → diabetes" is generated and stored in the knowledge base in the form of labeling. The advantage of this is that by introducing counterexamples, not only can incorrect knowledge be directly adopted, but also can provide comparative evidence for subsequent training, reasoning and arbitration.

[0133] In one embodiment, the knowledge graph completion method further comprises, in response to the timeliness judgment of the skeptical large model, saving the candidate triple information to the knowledge base if the timeliness of the candidate triple information is higher than the timeliness of the search information.

[0134] Specifically, the timeliness is determined by comprehensively comparing the publication time, the number of citations, and the source authority of the candidate triple and the search information. For example, when the candidate triple comes from a latest clinical study in 2024, and the search information mainly comes from a review article in 2018, it is judged that the search information R_rag cites outdated literature, and at this time, the R_gen in the candidate triple information is replaced with the R_rag information in the knowledge base. Thus, new knowledge is not discarded due to interference of old literature or outdated data, and the knowledge base can timely absorb cutting-edge research results, thereby better serving subsequent questioning and arbitration.

[0135] In the embodiment, by automatically generating and storing counterexamples when the questioning score exceeds the threshold value, not only can the incorrect or unreliable candidate triples be effectively prevented from directly entering the knowledge base, but also the content of the knowledge graph can be enriched in a comparative form, so that the subsequent arbitration model and reasoning engine can refer to both “positive examples” and “counterexamples” when judging, thereby improving the accuracy and robustness of reasoning. At the same time, in combination with the timeliness judgment mechanism of the questioning large model, the knowledge items with higher timeliness can be automatically distinguished among multiple candidate evidences and preferentially retained, thereby ensuring that the information stored in the knowledge base is more consistent with the current research progress and fact state. Through the combination of the two aspects, the credibility, timeliness, and explainability of the knowledge base are significantly improved.

[0136] In one embodiment, the knowledge graph completion method further comprises:

[0137] The logical consistency information includes a logical consistency score, and in response to the logical consistency score in the logical consistency information being less than a logical consistency threshold value, a type constraint is added to the logical large model.

[0138] Specifically, the logical consistency score is used to represent the matching degree between the candidate triple and the existing relationship mode and type rule in the knowledge base, and the value range is [0, 1]. When the score is lower than a preset threshold value (for example, 0.5), it indicates that there is obvious inconsistency or conflict in the candidate triple at the logical level. For example, if the candidate triple is “fatigue→belongs to→disease”, and the existing relationship constraint in the knowledge base requires that the tail entity type of the “belongs to (∈)” relationship should be “disease classification”, the logical consistency detection result may determine that the tail entity type of the triple does not match, and output the logical consistency score S_logic=0.3, which is lower than the threshold value 0.5.

[0139] In this case, a type constraint is automatically generated to limit the logical large model from generating similar errors in future knowledge reasoning or triple generation processes. For example, a new rule can be added to the constraint library of the logical large model: "If the relationship r = belongs, then the tail entity type must be {disease classification}". When the logical large model generates a candidate triple involving the "belongs" relationship again, the constraint mechanism will prioritize checking the tail entity type. If it does not meet the preset rule, the confidence will be automatically reduced, or the output will be directly rejected.

[0140] In this embodiment, by introducing a logic consistency score driven type constraint mechanism, the logical large model can gradually learn and solidify logical rules in the process of continuous iteration, thereby reducing the occurrence of repeated errors and ensuring the consistency, rigor, and scalability of the knowledge graph in the semantic reasoning level.

[0141] In one embodiment, the knowledge graph completion method further comprises:

[0142] In response to the confidence being less than a preset threshold, discarding the candidate triple information;

[0143] According to the prompt text, the candidate triple information is regenerated by calling the large model again;

[0144] until the confidence of the candidate triple information is greater than or equal to the preset threshold.

[0145] Specifically, the lower the confidence of the candidate triple information, the lower the credibility. Therefore, the candidate triple with low confidence will be discarded. After discarding the candidate triple, the candidate triple information will be regenerated according to the prompt text again. The regenerated candidate triple information will again perform verification steps such as retrieval, questioning, and arbitration to regenerate the confidence through the arbitration large model. Until the confidence of the candidate triple information is greater than or equal to the preset threshold, the above triple is saved and output.

[0146] In this embodiment, by automatically discarding low-confidence candidate triples and the cyclic regeneration mechanism, unreliable knowledge can be effectively prevented from directly entering the knowledge graph. The candidate triples are then generated again, and are sequentially subjected to the retrieval, questioning and arbitration links to ensure that the generated results are accepted after multiple verifications. This not only significantly improves the accuracy and reliability of the candidate triples, but also establishes a process similar to "adaptive training", enabling the generated large model to gradually converge to the correct output through multiple iterations, thereby improving the stability and robustness of the entire knowledge completion process. In addition, this mechanism not only ensures the quality of the candidate triple data, but also makes the knowledge graph construction more automated and intelligent, and has the ability to self-correct and evolve. Through continuous generation and verification matching, the large model can continuously output triple data, integrating generation, verification, reasoning, questioning, and arbitration into an adaptive pipeline, reducing human involvement, and automatically generating triple data that meets the conditions.

[0147] In one embodiment, the knowledge graph completion method further comprises:

[0148] In response to the number of times of generating the candidate triple information being greater than the maximum number of generations, the generation of the candidate triple information is stopped.

[0149] Specifically, when the number of times of generating the candidate triple information is greater than the maximum number of generations (default 3) and the confidence of the generated candidate triple information is still lower than the preset threshold, the candidate triple information is marked, the generation process is terminated, and infinite loop is prevented. The completion point is marked as a completion failure, and is waiting for manual confirmation for subsequent review or manual annotation.

[0150] In one embodiment, the knowledge graph completion method further comprises:

[0151] In response to the generation time of the candidate triple information being greater than the maximum time threshold, the generation of the candidate triple information is stopped.

[0152] Specifically, by setting the maximum time threshold and the maximum time consumption per round (default 120s, including LLM generation and verification time consumption), the generation is stopped and an error is reported.

[0153] In one embodiment, the knowledge graph completion method further comprises:

[0154] In response to the generation of text units of the candidate triple information being greater than the maximum text unit threshold, the generation of the candidate triple information is stopped.

[0155] Specifically, by setting the maximum token (text unit) threshold and the maximum token consumption per round (default 2M), the generation is stopped and an error is reported.

[0156] In the above embodiments, by limiting the generation times, the generation time and the token consumption, the infinite loop is prevented, the waiting time is saved, and the token consumption is saved.

[0157] Furthermore, some embodiments of the application provide an electronic device. The electronic device can be a variety of forms of digital computers, such as a laptop computer, a desktop computer, a workstation, a personal digital assistant, a server, a blade server, a mainframe computer, and so on. The electronic device can also be a variety of forms of mobile devices, such as a personal digital processing, a cellular phone, a smart phone, a wearable device, and other similar computing devices.

[0158] The electronic device includes one or more processors, and a memory storing computer program instructions which, when executed, cause the processor to perform the steps of the method provided by any one or more embodiments described above. Figure 2 An exemplary structural diagram of the electronic device is disclosed. The electronic device includes one or more processors 1101, a memory 1102, and an interface for connecting components, including a high-speed interface and a low-speed interface. The components are connected to each other by different buses, and can be installed on a common motherboard or otherwise installed as needed. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display graphical information of a GUI on an external input / output device, such as a display device coupled to the interface. In some other embodiments, multiple processors and / or buses can be used with multiple memories and multiple memories, if needed. Also, multiple electronic devices can be connected, each device providing part of the necessary operations. Among them, the components shown herein, their connections and relationships, and their functions are only examples, and are not intended to limit the implementation of the application described and / or claimed herein.

[0159] The electronic device can also include an input device 1103 and an output device 1104. The processor 1101, the memory 1102, the input device 1103, and the output device 1104 can be connected by a bus or other means, and are connected by a bus in the figure.

[0160] The input device 1103 can receive input of a number or character information, and generate a key signal corresponding to a user's setting and function control of the electronic device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 1104 can include a display device, an auxiliary lighting device (e.g., an LED), a haptic feedback device (e.g., a vibration motor), etc. The display device can include, but is not limited to, a liquid crystal display, a light emitting diode display, and a plasma display. In some embodiments, the display device can be a touch screen.

[0161] To provide the interaction with the user, the electronic device can be a computer. The computer has a display device (e.g., a cathode ray tube or an LCD monitor) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse) through which the user can provide input to the computer. Other kinds of devices can also be used for providing the interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback), and the input from the user can be received in any form (e.g., voice input or tactile input).

[0162] In the embodiments of the present application, the computer program / instruction is stored on the computer readable medium, and the computer program / instruction is executed by the processor to implement the steps of the method provided by any one or more of the embodiments. The computer readable medium can be included in the electronic device described in the above embodiments; or can exist separately without being assembled into the device. The computer readable medium carries one or more computer readable instructions.

[0163] The memory 1102 can be used as a non-transitory computer readable storage medium for storing non-transitory software programs, non-transitory computer executable programs and modules. The processor 1101 executes various functions and data processing of the server by running the non-transitory software programs, instructions and modules stored in the memory 1102, so as to implement the program instructions / modules corresponding to the method provided by any one or more of the embodiments in the present application.

[0164] The memory 1102 can include a program region and a data region, where the program region can store an operating system, an application required by at least one function, and the like, and the data region can store data created according to use of the electronic device, and the like. In addition, the memory 1102 can include a high-speed random access memory, and can further include a non-transitory memory such as at least one of a magnetic disk storage device, a flash memory device, or other non-transitory solid state storage device. In some embodiments, the memory 1102 can optionally include a memory disposed remotely with respect to the processor 1101, and these remote memories can be connected to the electronic device through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0165] It should be noted that the computer-readable medium described in the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus.

[0166] The computer-readable medium includes permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory, static random access memory, dynamic random access memory, other types of random access memory, read-only memory, electrically erasable programmable read-only memory, flash memory or other memory technology, read-only optical disc, digital versatile disc or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage device, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0167] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network ("LAN") or a wide area network ("WAN"), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0168] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. For example, the embodiments can be implemented by an application specific integrated circuit, a general purpose computer or any other similar hardware device. In some embodiments, the software programs of the embodiments can be executed by a processor to implement the above steps or functions. Similarly, the software programs of the embodiments (including related data structures) can be stored in a computer readable recording medium, such as a RAM memory, a magnetic or optical drive or a floppy disk and the like. In addition, some steps or functions of the embodiments can be implemented by hardware, such as a circuit cooperating with a processor to perform the steps or functions.

[0169] The computer program product provided by the embodiments of the present application includes one or more computer programs / instructions, which, when executed by a processor, generate all or part of the processes or functions described in the embodiments of the present application. The computer can be a general purpose computer, a special purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center and the like integrated with one or more available media. The available media can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk) and the like.

[0170] The computer program product of the present application can be a computer program embodied on a non-transitory computer readable medium. The computer program product can also be a propagated signal on a carrier wave.

[0171] The scope of the application is defined by the appended claims rather than by the description set forth herein, and therefore the terminology used is for the purpose of describing particular embodiments only and is not intended to be limiting and is not intended to be limiting and is intended to cover all changes and modifications of the application which fall within the scope of the claims, together with all equivalents thereof. Further, it is intended to cover all logical, equivalent, and functional alterations except those which would be impermissible under the patent statutes and rules. In this patent, the use of the singular includes the plural, the use of "or" means "and / or", and the use of "one" means "one or more". It will be further understood that the terms "comprises" and "comprising", when used in this specification, specify the presence of stated features, integers, steps, or components but do not preclude the presence or addition of one or more other features, integers, steps, components, or groups thereof. It will be understood that the drawings are not necessarily drawn to scale and that, where appropriate, the dimensions of components have been exaggerated or minimised to illustrate features clearly.

[0172] The above description is only specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be defined by the protection scope of the claims, and the above-described embodiments should be regarded as exemplary and non-limiting.

Claims

1. A knowledge graph completion method, characterized in that: The knowledge graph completion method includes: Traverse the subgraph of the knowledge graph to be completed and obtain the structure and context information of the subgraph to be completed; Constructing a prompt text according to the subgraph structure to be completed and context information; Generate candidate triple information by generating a large model according to the prompt text; Performing retrieval enhancement on the candidate triple information to obtain retrieval information; Generate query information for the candidate triple information and the search information by querying the large model; Calculating confidence using an arbitration model based on the candidate triple information, the search information, and the question information; In response to the confidence being greater than or equal to a preset threshold, the candidate triplet information is saved.

2. The knowledge graph completion method according to claim 1, characterized in that: The knowledge graph completion method further includes: Performing type consistency verification and / or semantic conflict detection and / or common sense violation detection on the candidate triple information through the inference big model, and outputting logical consistency information; The confidence level is calculated by using an arbitration model according to the candidate triple information, the search information, the logical consistency information and the challenge information.

3. The knowledge graph completion method according to claim 1 or 2, characterized in that: After the step of saving the candidate triple information in response to the confidence being greater than or equal to a preset threshold, the knowledge graph completion method further includes: The confidence of the candidate triple information decreases over time; In response to the confidence being less than a preset threshold, recalculating the confidence for the candidate triplet information; and / or; The confidence level of the candidate triplet information is recalculated periodically.

4. The knowledge graph completion method according to claim 1 or 2, characterized in that: The generating of the query information from the candidate triple information and the search information by querying the large model specifically includes: Conducting type and / or timeliness queries on the candidate triple information; The questioning model generates a questioning score and questioning reasons, and outputs questioning information.

5. The knowledge graph completion method according to claim 4, characterized in that: The knowledge graph completion method further includes: In response to a question score in the question information being greater than a question score threshold, adding a counterexample to the knowledge base; and / or; In response to the timeliness judgment of the challenged large model, the timeliness of the candidate triple information is higher than the timeliness of the search information, and the candidate triple information is saved in the knowledge base.

6. The knowledge graph completion method according to claim 2, characterized in that: The knowledge graph completion method further includes: The logical consistency information includes a logical consistency score. In response to the logical consistency score in the logical consistency information being less than a logical consistency threshold, a type constraint is added to the logical big model.

7. The knowledge graph completion method according to claim 1 or 2, characterized in that: The knowledge graph completion method further includes: In response to the confidence being less than a preset threshold, discarding the candidate triplet information; According to the prompt text, the large model is called again to regenerate the candidate triple information; Until the confidence level of the candidate triplet information is greater than or equal to a preset threshold.

8. The knowledge graph completion method according to claim 1 or 2, characterized in that: The knowledge graph completion method further includes: In response to the number of times the candidate triple information is generated being greater than a maximum number of times, stopping the generation of the candidate triple information; and / or in response to the candidate triplet information generation time being greater than a maximum time threshold, stopping the generation of the candidate triplet information; And / or in response to the text unit consumption of generating the candidate triple information being greater than a maximum text unit threshold, stopping generating the candidate triple information.

9. An electronic device, characterized in that: The electronic device comprises: one or more processors; and A memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the method according to any one of claims 1 to 8.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Collaborative system and method for validating equipment failure models in an analytics crowdsourcing environment

    CA3178658A1

  • Method and device for complementing knowledge graph and computer storage medium

    CN117094395A

  • Large language model knowledge graph completion method based on subgraph structure information enhancement

    CN118674026A

  • Policy knowledge graph construction method and system based on digital human interaction data analysis

    CN120471160A

  • Geological knowledge graph completion method based on graph convolutional network

    CN120578772A