Knowledge graph embedding method, system, equipment and medium

By integrating rule mining and forward chaining inference to generate and instantiate new rules within knowledge graph embeddings, the method addresses the challenges of noise and uncertainty in UKGs, improving accuracy and robustness of knowledge representation.

CN120317352APending Publication Date: 2025-07-15NINGXIA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510441245.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing uncertain knowledge graph embedding model is difficult to effectively process false negative samples during the training stage, resulting in the impact of the accuracy of the knowledge graph embedding model.

Method used

Through rule mining tools, soft rules with different confidence levels are mined from the original uncertain knowledge graph, new rules are generated using forward link inference method, and entity variables are instantiated, new triple instantiated data is generated, and new triple instantiated data is fused with the original triple and inputted into the knowledge graph embedding model for training, optimize triple and rule loss functions to improve model accuracy.

Benefits of technology

Effectively identify and reduce false negative samples in uncertain knowledge graphs, improve the accuracy and reliability of knowledge graph embedding, and enhance the coverage and reasoning ability of knowledge graphs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120317352A_ABST
    Figure CN120317352A_ABST
Patent Text Reader

Abstract

The invention provides a knowledge graph embedding method, system and device and a medium, and belongs to the field of knowledge engineering research in artificial intelligence, and the method comprises the following steps: obtaining an original uncertain knowledge graph; mining soft rules with different confidence degrees from the original uncertain knowledge graph by using a rule mining tool; performing rule reasoning on the soft rule by using a forward link reasoning method to generate a new rule; performing instantiation on the entity variable in the new rule to obtain new triple instantiation data; fusing the new triple instantiated data with an original triple in the original uncertain knowledge graph, inputting fused enhanced data into a knowledge graph embedding model, and training the knowledge graph embedding model; the to-be-processed uncertain knowledge graph is obtained, the to-be-processed uncertain knowledge graph is input into the trained knowledge graph embedding model, the knowledge graph with the embedded representation is obtained, and the knowledge graph embedding accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the research field of knowledge engineering in artificial intelligence, and particularly relates to a knowledge graph embedding method, system, device and medium. Background Art

[0002] Knowledge Graphs (KGs) are the epitome of knowledge engineering in the big data era, the product of the combination of symbolism and connectionism, and an important cornerstone for realizing cognitive intelligence. However, the knowledge graph represented in symbolic form faces challenges of low computational efficiency and strong data sparsity. To solve this problem, in recent years, researchers have proposed methods for knowledge graph representation learning, using deep learning techniques to solve the modeling problems of entities and relationships in the knowledge base from multiple perspectives. By mapping entities and relationships to a low-dimensional feature vector space, efficient representations of entities, relationships and their complex semantics are achieved. The research on knowledge graph representation learning is of great significance for its applications in fields such as vertex classification, recommendation systems, link prediction, semantic search, and machine question answering.

[0003] However, there are also noises and errors that are difficult to eliminate in the knowledge graph. The construction of Uncertain Knowledge Graphs (UKGs) relaxes the assumption that all knowledge must be accurate and correct. Uncertainty is an inherent feature of knowledge graph construction. The expression of uncertain relational facts not only enriches the structure and content of the knowledge graph, but also retains more potentially valuable information, improves the coverage of the knowledge graph, and promotes the evolution and update of the knowledge graph. In addition, some knowledge itself is difficult to express in a deterministic way. In many fields, such as medicine, law and finance, much knowledge often highly relies on experience. If these attributes are ignored and traditional graphs are directly used to represent this knowledge, it is very inaccurate or even wrong. Therefore, some open knowledge graphs with uncertain information describe their uncertainty by adding confidence scores to each triple.

[0004] Compared with the research on embedding deterministic KGs, representing UKGs requires encoding additional confidence to maintain uncertainty. Most existing knowledge graph embedding models for uncertain knowledge graph representation learning draw on traditional knowledge graph embedding models and are unable to effectively capture and represent the inherent uncertainty of uncertain knowledge graphs. Therefore, how to design and train a knowledge graph embedding model for uncertain knowledge graphs is an important challenge. On the one hand, due to the existence of a large number of low-confidence triples, the noise content in uncertain knowledge graphs is relatively high. For example, in the open knowledge graph ConceptNet, 7% of the triples have a confidence level lower than 0.2, and these low-confidence triples are very likely to be noise. Traditional knowledge graph embedding models assume that all triples are correct. Therefore, in a high-noise environment, inaccurate graph representations may be learned, leading to incorrect reasoning results. On the other hand, under the Open World Assumption (OWA), facts that do not exist in the knowledge graph may still be true. Due to the existence of a large number of low-confidence facts, the density of uncertain knowledge graphs is greater, so negative sampling often introduces a more serious false negative sample problem, that is, during the training process, the correct facts missing in the knowledge graph are wrongly regarded as negative examples with a confidence level of 0.

[0005] In summary, existing uncertain knowledge graph embedding models are difficult to effectively process false negative samples during the training stage, which affects the accuracy of the knowledge graph embedding model. Summary of the Invention

[0006] To overcome the deficiencies of the above-mentioned existing technologies, the present invention provides a knowledge graph embedding method, including the following steps: Obtain the original uncertain knowledge graph; Use a rule mining tool to mine soft rules with different confidence levels from the original uncertain knowledge graph; use the forward chaining inference method to perform rule inference on the soft rules to generate new rules; instantiate the entity variables in the new rules to obtain new triple instantiation data corresponding to the entity variables; Fuse the new triple instantiation data with the original triples in the original uncertain knowledge graph, and then input the fused enhanced data into the knowledge graph embedding model to train the knowledge graph embedding model to obtain a trained knowledge graph embedding model; Obtain the uncertain knowledge graph to be processed, and input the uncertain knowledge graph to be processed into the trained knowledge graph embedding model to obtain the knowledge graph after embedding representation.

[0007] Using the forward chaining inference method to perform rule inference on the soft rules to generate new rules includes the following steps: Use the conflict resolution strategy to select the rules that need to be triggered among the soft rules; Activate the triggered rules to generate new rules; Add the generated new rules to the rule candidate set; Repeat the above three steps until no rules are triggered.

[0008] The calculation formula for the confidence is: g ( l ) = φ ( dr ( h , t )); Among them, φ (·) is a conversion function, d r ( h , t ) is the distance function of any given triple, ([[]] h , t ) is the entity set, h is the head entity, t is the tail entity.

[0009] Preferably, when training the knowledge graph embedding model, the following loss function is specifically adopted: L = L tr + L rule ; Among them, L tr is the loss function of the triple, L rule is the loss function of all rules.

[0010] Preferably, the loss function of the triple is as follows: ; Among them, L pos is the positive example loss function, is the final negative sampling loss function, is the number of generated negative samples.

[0011] Preferably, the loss function of all rules is as follows: ; Among them, is the loss function of the new triple f i , L rcis the confidence loss function of the rule, f r is a soft rule, c r is the rule confidence, F is a set of soft rules.

[0012] The present invention also provides a knowledge graph embedding system, including: A data acquisition module, configured to acquire an original uncertain knowledge graph; A data processing module, configured to use a rule mining tool to mine soft rules with different confidences from the original uncertain knowledge graph; use a forward chaining inference method to perform rule inference on the soft rules to generate new rules; instantiate the entity variables in the new rules to obtain new triple instantiation data corresponding to the entity variables; A knowledge graph embedding model training module, configured to fuse the new triple instantiation data with the original triples in the original uncertain knowledge graph, and then input the fused enhanced data into a knowledge graph embedding model to train the knowledge graph embedding model to obtain a trained knowledge graph embedding model; A knowledge graph embedding module, which acquires an uncertain knowledge graph to be processed, and inputs the uncertain knowledge graph to be processed into the trained knowledge graph embedding model to obtain a knowledge graph after embedding representation.

[0013] The present invention also provides a computer device, including a memory and a processor; the memory stores a computer program, and the processor is configured to run the computer program in the memory to execute the knowledge graph embedding method.

[0014] The present invention also provides a computer-readable storage medium, which stores a computer program, and the computer program is suitable for being loaded by a processor to execute the knowledge graph embedding method.

[0015] The knowledge graph embedding method provided by the present invention has the following beneficial effects: Through the rule mining tool, the present invention can mine reasonable soft rules with a certain confidence, and then use the forward chaining inference method to perform rule inference on the soft rules to generate new rules; by instantiating the entity variables in the new rules, new triple instantiation data corresponding to the entity variables is obtained. This process can infer missing new relationships or facts, that is, new triples, from known information, so as to be able to identify previously unrecognized true positive samples and reduce false negative samples in the uncertain knowledge graph.

[0016] In the present invention, the enhanced data obtained by fusing the newly instantiated triple data with the original triples in the original uncertain knowledge graph is input into the knowledge graph embedding model, and the knowledge graph embedding model is trained to obtain a trained knowledge graph embedding model; the trained knowledge graph embedding model can effectively process false negative samples and improve the accuracy of knowledge graph embedding. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] To more clearly illustrate the embodiments of the present invention and their design schemes, the accompanying drawings required for the present embodiments will be briefly introduced below. The accompanying drawings in the following description are only partial embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 is a flowchart of the knowledge graph embedding method according to the embodiment of the present invention; Figure 2 is a flowchart of forward chaining inference; Figure 3 is the confidence prediction result of CN15k; Figure 4 is the confidence prediction result of NL27k; Figure 5 is the confidence prediction result of PPI5k. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To enable those skilled in the art to better understand the technical solutions of the present invention and be able to implement them, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.

[0020] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "axial", "radial", "circumferential", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the technical solutions of the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present invention.

[0021] In addition, the terms "first", "second", etc. are for descriptive purposes only and should not be construed as indicating or implying relative importance. In the description of the present invention, it should be noted that unless otherwise clearly specified or limited, the terms "connected" and "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. In the description of the present invention, unless otherwise stated, the meaning of "a plurality of" is two or more, which will not be elaborated herein.

[0022] Embodiment The present invention provides a knowledge graph embedding method, specifically as Figure 1As shown in the figure, a part of the uncertain knowledge graph is presented, which contains multiple entities and the relationships between them. For example, there is a "provide" relationship between "diet" and "nutrition" with a confidence of 0.8; there is a "consume" relationship between "exercise" and "calories" with a confidence of 0.6, etc. These entities and relationships together constitute the uncertain knowledge graph, which is used to represent the associations and interactions between different concepts. The confidence reflects the reliability degree of the relationship. By mining the data in the uncertain knowledge graph, some soft rules are obtained, such as the rules shown in the figure: (x, provide, y) ∧ (y, promote, z) → (x, promote, z) with a confidence of 0.9; (x, consume, y) ∧ (y, affect, z) → (x, affect, z) with a confidence of 0.8, etc. These soft rules reflect the indirect relationships and potential logical laws between entities, and can help better understand and reason about the information in the knowledge graph. Based on the mined soft rules, forward chaining reasoning is carried out to generate new relation triples. For example, according to the rule (diet, provide, nutrition) ∧ (nutrition, promote, health) → (diet, promote, health), it can be deduced that there is a "promote" relationship between "diet" and "health" with a confidence of 0.8; according to the rule (exercise, consume, calories) ∧ (calories, affect, weight) → (exercise, affect, weight), it can be deduced that there is an "affect" relationship between "exercise" and "weight" with a confidence of 0.9, etc. In this way, the content in the knowledge graph can be continuously expanded to discover new knowledge and relationships. The mined soft rules are instantiated to generate specific rule embeddings. For example, for the rule (x, provide, y) ∧ (y, promote, z) → (x, promote, z), through specific entity instantiation, instantiated rules such as (diet, provide, nutrition) ∧ (nutrition, promote, health) → (diet, promote, health) can be obtained. Instantiated rule embeddings can combine abstract rules with specific entities, making the rules more operable and practical, helping to improve the effect and accuracy of embedding learning, and at the same time being able to better explain and understand the embedding results, providing more valuable support for the application of the knowledge graph. The entities and relationships in the uncertain knowledge graph are embedded into a low-dimensional vector space, and the semantic information and structural features of the entities and relationships are captured through vector representation. In this process, the uncertainty and confidence information in the knowledge graph need to be considered so that the embedded vectors can better reflect the reliability and accuracy of the knowledge. The uncertain knowledge graph embedding and the instantiated rule embedding are jointly trained by simultaneously optimizing the triple loss and the rule loss. The triple loss is used to measure the difference between the embedded vectors and the actual knowledge graph data, and the rule loss is used to measure the difference between the instantiated rule embedding and the actual rules. By minimizing the joint loss, the model can take into account both the structure of the knowledge graph and the constraints of the rules during the learning process, improving the accuracy and reliability of the embedding results.

[0023] The specific process includes the following steps: Step 1: Obtain the original uncertain knowledge graph.

[0024] Definition 1 (uncertain knowledge graph): The uncertain knowledge graph is denoted as G , and contains a set of fact quadruples {( h , r , t , s ) ∈ T | h , t ∈ E ; r ∈ R 1; s ∈ [0, 1]}, where E represents the entity set, R 1 represents the relation set, T represents the set of quadruples (usually, the term "triple" is used to represent the combination of ( h , r , t ), and in the present invention, it is extended to include the confidence score, called "quadruple"), h , r , t and s respectively represent the head entity, relation, tail entity and confidence score of the quadruple.

[0025] The confidence score s represents the degree of certainty of the relationship h between the head entity t and the tail entity r . Therefore, each quadruple ( h , r , t , s ) contains both the relation fact and the associated uncertainty level.

[0026] Definition 2 (uncertain knowledge graph embedding): The uncertain knowledge graph embedding is denoted as f , and is a mapping that projects the entity set E and the relation set R 1 into the vector space R k , and also includes the confidence of the triple. f The framework of f is as follows: E , R 1 → R k ; f s : R k →s ∈ [0, 1]; Among them, E is the entity set; R 1 is the relation set; R k is the vector space; f are the embedding functions of entities and relations respectively, used to convert entities and relations into k vector representations in the -dimensional space; f s is a function that assigns a confidence score to each triple component within the range [0, 1] s , used to represent the confidence in the validity of the triple.

[0027] Definition 3 (Horn Rules): Horn rules are a special case of first-order logic (FOL) rules, usually formalized as {B1 ∧ B2 ∧ … ∧ Bn} → C, where B1 ∧ B2 ∧ … ∧ Bn is called the premise or body of the rule, C is called the conclusion or head of the rule, the propositions (B1 ∧ B2 ∧ … ∧ Bn) in the body of the rule are connected by logical AND (∧), and the body of the rule and the head of the rule are connected by logical implication (→).

[0028] Definition 4 (Probabilistic Soft Logic): Probabilistic Soft Logic (PSL) uses logical rules to build a graphical knowledge graph embedding model for variables with soft truth values, converting it into a convex optimization problem to make reasoning easier. In PSL, the soft truth value is a real number between 0 and 1, indicating the likelihood that a predicate is true.

[0029] To determine the degree of implementation of the instantiated rules, PSL uses the Łukasiewicz t-norm to represent the truth value of the instantiated rules as a combination of the truth values of its constituent triples, achieved by applying specific logical connectives, including conjunction (˄), disjunction (˅), negation ( ) and implication (→), and the function descriptions corresponding to the logical operators are as follows: ; Among them, a and b are logical expressions, π (·) is the truth value of the logical expression.

[0030] Step 2: Use a rule mining tool to mine soft rules with different confidence levels from the original uncertain knowledge graph; use the forward chaining inference method to perform rule reasoning on the soft rules to generate new rules; instantiate the entity variables in the new rules to obtain new triple instantiation data corresponding to the entity variables.

[0031] In order to better solve the problem of false negative samples, build more complete soft rules, remove incorrect conclusions, and improve the quality of embeddings by increasing the confidence of the rules, the present invention processes the data used for training the knowledge graph embedding model. The processing process includes three stages: the soft rule mining stage, the rule reasoning stage, and the rule instantiation stage. In the soft rule mining stage, first, the Association Rule Mining under Incomplete Evidence (AMIE3) is used to automatically extract soft rules with different confidence levels from the uncertain knowledge graph. Then, these soft rules and the fact set are sent together into the rule engine system Apache Jena to enter the rule reasoning stage, where forward chaining reasoning is performed to generate new rules. Next, the entity variables in the new rules are instantiated to obtain new triple instantiation data corresponding to the entity variables. The specific process is as follows:

[0032] (9) Mining soft rules.

[0033] Soft rules are logical rules with probabilities or confidence levels extracted from the uncertain knowledge graph. They evaluate the quality of the rules through statistical methods (such as support and PCA confidence), rather than relying on absolute logic. Soft rules can handle the uncertainty of data and improve the flexibility and fault tolerance of reasoning by indirectly enhancing the coverage of the knowledge graph. Using AMIE3 to automatically extract soft rules from the uncertain knowledge graph and adopting pruning strategies and optimization algorithms can mine the rules of the uncertain knowledge graph and indirectly enhance the data.

[0034] AMIE3 uses support supp ( R ) to measure the quality of the rules R , which refers to the number of true predictions made in the knowledge base x . The support p ( supp ) is calculated as follows: R ; ; In the formula, x is the knowledge base; p is the number of true predictions, that is, the number of r ( h , t ). r ( h , t ) represents the number of fact pairs( x , r ) that satisfy the relationship h , t in the knowledge baseR as a rule.

[0035] The Probabilistic Confidence Assessment ( PCA ) is adopted to evaluate the reliability of Horn rules, and its calculation formula is shown as follows: ; where, PCA conf (B→r ( h , t ) represents the confidence of the given Horn rule B→r ( h , t ), PCA support degree of the Horn rule supp ( B→r ( h , t )) represents the Horn rule B→r ( h , t ), B represents the premise of the Horn rule, ([[]] h , t ) represents the fact pair of the Horn rule, ∣{([[]] h , t ): ∃ t ′: B ∧ r ( h , t ′)}∣ represents the number of all possible ( B , h , t ) pairs that satisfy the premise t ′ is t 's one possible value, indicating that given the premise B holds, r ( h , t ) can have different values.

[0036] The following is an example mined by AMIE from the CN15k dataset: ([[]] X , BornInCountry , Y ) → ([[]] X , Nationality , Y ), PCA the confidence score is 0.954, and this rule indicates that if X is born in Y country, it can be inferred with high confidence that X 's nationality is Y .

[0037] The rule selection process includes manually selecting rules with high confidence and discarding rules with a confidence lower than 0.8. The soft rule set is represented as F ={( f r , c r )}, f r is a soft rule mined by AMIE3. c r is the confidence of the rule. To illustrate this, consider an example where the confidence interval of a specific rule is defined as [0.9, 1], indicating a confidence factor of 0.9. Suppose there is a mined rule (a ˄ b → c, 0.85) with an original confidence score of 0.85. To adjust the final confidence of this rule, multiply its original confidence by the confidence factor (i.e., 0.85 × 0.9), resulting in a corrected confidence score of 0.765. Apply the same calculation to all other rules to ensure corresponding adjustments are made to them. By adopting this method, it is ensured that only the most reliable rules are retained and their confidences are adjusted in a consistent manner, thereby improving the overall quality and reliability of the rule set.

[0038] (2)Rule inference.

[0039] Forward Chaining Inference is a rule-driven inference method that starts from known facts and gradually derives new conclusions by applying rules. It uses existing triples and rules to derive new triples. This inference process may discover new entities and relationship types that are not explicitly represented in the original dataset. Incorporating these discoveries into the knowledge graph can enrich the semantic structure of the knowledge graph. In addition, the inferred rules can also provide explanations for the relationships between entities in the knowledge graph, increasing the interpretability of the data.

[0040] The flowchart of the forward chaining inference process is as Figure 2As shown. At the start of each iteration, all the rules in the rule base are systematically examined to identify those rules that satisfy the current state of the working memory, determine the rules that may be triggered, and these rules are subsequently included in the rule candidate set (conflict set). In the case of a rule to be triggered being found, a conflict resolution strategy (such as polling, priority, or most specific match) is used to determine which rule to activate. Once the selected rule is found, the rule is executed and integrated into the working memory (in the rule base or working memory). This integration may involve adding new facts, modifying existing facts, or performing other specified operations. If no rule satisfies the conditions, the loop terminates, indicating the completion of the forward chaining inference process. The whole process constitutes a continuous loop in which the system continuously evaluates, selects, and triggers rules until no more rules can be activated. This process mimics the human cognitive process in which new facts and information are continuously added to the working memory until a certain conclusion or decision is reached.

[0041] The Rete algorithm, as the cornerstone of the rule engine Apache Jena, accelerates rule matching through caching of matching results and pattern sharing of rules. The Rete algorithm starts by constructing a Rete network, a directed acyclic graph network anchored by a root node. In this network, all nodes except the root node store intermediate matching results. Various parts of the network can share different types of data. The conditional sequence in the rule plays a key role in determining the number of shared nodes in the Rete network. Therefore, the existence of shared nodes can speed up rule matching and reduce the memory for building the network. However, it is not without limitations. Under the same hardware conditions and the same rule matching algorithm, the efficiency of the rule engine still stagnates. To further optimize the Rete algorithm and improve the efficiency of rule inference, the conditions within the rules must be pre-sorted. The rule engine first scans the entire rule base and decomposes the conditions embedded in each rule into different terms according to logical characters. During this decomposition process, it simultaneously lists the frequency of each condition and sorts them in descending order. After completing the initial scan and statistical analysis, a secondary scan is performed to re-order the conditions in each rule according to the occurrence frequency. After completing these steps, the preprocessing stage of the rule engine is completed.

[0042] (3) Rule instantiation.

[0043] The extracted rules contain entity variables, and these variables are instantiated with specific entity instances to establish a specific rule basis. More precisely, the variables in the rule body are filled by replacing entities that are consistent with the relational triples and the rule body criteria. Subsequently, instantiating the rule head generates a corresponding new triple, and this instantiated rule is called a ground, which serves as a logical path connecting the existing triples and the newly generated triples. For example, X can be instantiated with "nutrition" and Y with "immunity", and the resulting new triple ("nutrition", "enhance", "immunity") adds to the rationale behind the knowledge base. During the rule establishment phase, a constraint is enforced that only allows using the triples in the current knowledge graph to instantiate the rule body, thus ensuring the robustness of the reasoning process. In addition, the inferred triples generated by instantiating the rule head must not exist in the current knowledge graph to ensure that the inferred candidate triples are new to the current knowledge graph.

[0044] Step 3: Fuse the new triple instantiation data with the original triples in the original uncertain knowledge graph, and then input the fused enhanced data into the knowledge graph embedding model to train the knowledge graph embedding model, obtaining the trained knowledge graph embedding model.

[0045] Jointly perform embedding learning on the instantiated rule data and the original uncertain knowledge graph, integrate the newly obtained triples into the original triples for data enhancement, and then use the enhanced data to train the knowledge graph embedding model. Through this training process, the knowledge encoded in the logical rules is effectively transferred and embodied in the learned embeddings, ultimately generating a more robust and richer information representation.

[0046] The training process of the knowledge graph embedding model is as follows: Use I ( l ) to represent the soft truth value (between 0 and 1) of an atom. For the observed triples, the soft truth value of the atom is equal to the confidence of the triple. For the unobserved triples, following the open-world assumption principle, if a fact does not exist in the knowledge base, the system does not immediately deny the existence of this fact but maintains an open attitude, considering that this fact may be true or not yet recognized. This assumption is more in line with the real-world situation because decisions and inferences often need to be made under incomplete information. At this time, the soft truth value of the atom is equal to the confidence score predicted based on the embedding, as follows: I ( l ) = s , l ∈L + ; I ( l ) =g ( l ), l ∈L − ; where L + is the observed triple, as a positive example triple; L − is the unobserved triple. In fact, all unobserved triples can be regarded as unlabeled triples, but in the present invention, only those encoded in the instantiated soft rule conclusion are considered. Randomly select other entities to replace the head entity or the tail entity of the triple to construct the negative example triple L − , s is the true confidence score of the observed triple, g ( l ) represents the confidence score predicted based on the embedding, l is the triple.

[0047] Currently, the benchmark model adopted by the Uncertain Knowledge Graph Embedding (UKGE) is the Distributed Representation for Multi-relational Data (DistMult). Although DistMult uses a diagonal matrix to represent the relationship, reducing the number of parameters of the knowledge graph embedding model, it cannot model asymmetric relationships. UKGE-SRDA (the knowledge graph embedding model of the present invention) selects the Rotation-based Embedding Model (RotatE) as the basic model when modeling triples. It processes symmetric and asymmetric relationships in the knowledge graph by mapping entities and relationships to the complex space. For any given triple ( h , r , t ) ∈ E × R 1 × E , its distance function is: ; where d r ( h , t ) is the distance function of any given triple, h is the embedding vector of the head entity, r is the embedding vector of the relationship, t is the embedding vector of the tail entity, is the multiplication of the corresponding elements in h and r, h, r, t ∈ Ck , C is the complex number field, that is, the set of all complex numbers, k is the dimension of the complex space, and the modulus length .

[0048] The confidence score based on embedding prediction is obtained by a transformation function φ (·) that maps d r ( h , t ) to the range [0, 1], as shown in the following formula: g ( l ) = φ ( d r ( h , t )); Among them, φ (·) is the transformation function, d r ( h , t ) is the distance function of any given triple, ( h , t ) is the entity set, h is the head entity, t is the tail entity.

[0049] Among them, the transformation function φ (·) has the following two options: (1) sigmoid function, ; (2) bounded rectifier function, φ ( x ) = min(max( x , 0), 1). The transformation function φ (·) of the present invention adopts sigmoid function.

[0050] However, simply replacing the scoring function of RotatE with the new scoring function may lead to performance degradation. The transformation function does not work well for scoring functions based on translational distances, where the score is the negative of the distance and ranges from [−∞, 0]. Therefore, to concentrate the range of scores, the following score mapping function with an additional hyperparameter γ = 2 is applied to the confidence score based on embedding prediction.

[0051] ; Among them, represents the weight, b represents the bias value, d r ( h , t ) represents the distance function of any given triple, γ represents the additional hyperparameter.

[0052] The training objective is to minimize the root mean square error (RMSE) between each triple l ∈L+ s and g ( l ), as follows: ; where L pos is the positive example loss function, g ( l ) is the confidence score based on the embedding prediction, L + is the observed triple, l is the triple.

[0053] The negative triples are sampled using self-adversarial negative sampling, sampling negative triples from the following distribution: ; where α is the sampling coefficient, is the scoring function of the candidate negative triple , is the head entity in the candidate negative triple, is the relation, is the tail entity of the candidate negative triple, is the scoring function of the triple in the given knowledge graph, is the set of triples in the given knowledge graph under the condition that the candidate negative triple has a probability.

[0054] The loss function L1 of the randomly sampled negative sample L − is defined as follows: ; where g ( l ) is the confidence score based on the embedding prediction, l is the triple.

[0055] The final negative sampling loss function is as follows: where σ is the sigmoid function, is the i th negative triple, g ( l ) is the confidence score based on the embedding prediction.

[0056] The loss function of the original triple is defined as: ; Among them, L pos is the positive example loss function, is the final negative sampling loss function, is the number of generated negative samples.

[0057] Use the scoring function of the RotatE knowledge graph embedding model F i as the scoring function for the new triple f r deduced by the rule f i : ; Among them, is the head entity of the new triple, is the tail entity of the new triple, is the distance function of the new triple.

[0058] Since f i also has a confidence level, define the loss f i of the new triple as follows: ; Among them, f i is the new triple, F i is the scoring function of the new triple f i , s is the confidence score, is the head entity of the new triple, is the relation of the new triple, is the tail entity of the new triple.

[0059] Confidence loss function of the rule L rc is based on the scoring function F i , the confidence score of the new triple s and the confidence level c r of the rule to calculate. Define the confidence loss function L rc as follows: ; Among them, f r is the soft rule, is a weight parameter used to balance the rule confidencec r and the triple confidence score s The influence, when is close to 1, the loss function depends more on the confidence of the rule; when is close to 0, the loss function depends more on the confidence of the triple, F i is the scoring function of the new triple f i and F is the soft rule set.

[0060] The loss function of all rules L rule can be defined as: ; wherein, is the loss function of the new triple f i and L rc is the confidence loss function of the rule, f r is the soft rule, c r is the rule confidence, F is the soft rule set.

[0061] The jointly trained triple set L (including positive and negative samples) and the soft rule set F , the loss function used is: L = L tr + L rule ; wherein, L tr is the triple loss function, L rule is the loss function of all rules.

[0062] Although the triples in the dataset already come with some form of confidence, the knowledge graph embedding model needs to learn representations that can reflect these confidences. The confidence calculated by the knowledge graph embedding model is a prediction of the relationships between entities existing in the knowledge graph and needs to match the actual confidence. That is, the knowledge graph embedding model needs to accurately calculate the confidence of each triple and compare the prediction results with the true scores.

[0063] Step 4: Obtain the uncertain knowledge graph to be processed, and input the uncertain knowledge graph to be processed into the trained knowledge graph embedding model to obtain the knowledge graph after embedding representation.

[0064] To verify the accuracy of the method of the present invention, the following comparative experiments were conducted in the present invention.

[0065] Three datasets were used in the experiment: CN15k, NL27k, and PPI5k, which were extracted from ConceptNet, NELL, and STRING respectively. The statistical data are shown in Table 1, where Std(s) and Avg(s) are the standard deviation and average value of the confidence scores respectively, the training set represents the number of relational facts in the training set, the validation set represents the number of relational facts in the validation set, and the test set represents the number of relational facts in the test set.

[0066] Table 1 Statistical information of the datasets The experimental environment configuration of the present invention is shown in Table 2.

[0067] Table 2 Environment configuration During the process of training the knowledge graph embedding model UKGE-SRDA, the Adam optimizer was used for gradient descent optimization and soft rules were instantiated. The learning rate l r ∈ {0.005, 0.002, 0.001, 0.005}, the dimension of the embedding vector d ∈ {32, 64, 128, 256, 512}, the batch size was set to b ∈ {256, 512, 1024}, L The L2 regularization coefficient λ was 0.0005. The minimum mean square error early stopping strategy was used to stop training on the validation set, and the minimum mean square error was calculated every 10 epochs. The optimal parameter settings of the present invention are as follows: the learning rate l r = 0.005, the batch size b = 1024, the dimension of the embedding vector d = 512, in CN15k b = 1024, in PPI5k d = 256.

[0068] Analysis of experimental results: The confidence prediction task takes the graph triple as the input and outputs the confidence of the triple, which can be used to verify the embedding effect of the knowledge graph embedding model on semantic structure information and confidence information.

[0069] Evaluation metrics: For each uncertain relation fact in the test set, use the knowledge graph embedding model UKGE-SRDA to predict its confidence score, and calculate the mean squared error (MSE) and mean absolute error (MAE). The smaller the values of these metrics, the better the performance of the knowledge graph embedding model. The calculation formulas are as follows:

[0070] ; ; where MSE is the mean squared error, which is used to measure the average squared error between the model prediction value and the true value, is the total number of uncertain relation facts in the test set, is the prediction confidence score of the model for the -th triple, is the true confidence score of the -th triple, is the error between the model prediction value and the true value. MAE is the mean absolute error, which is used to measure the average absolute error between the model prediction value and the true value, is all triples in the test set, is the prediction function of the model.

[0071] Table 3 shows the experimental results of UKGE-SRDA and the comparison models in the confidence prediction task. Among them, the comparison models include the Rectangle Embedding Uncertain KnowledgeGraphs (UKGErect), the Logistic Embedding UncertainKnowledge Graphs (UKGElogi), the ComplEx EmbeddingUncertain Knowledge Graphs (UComplEx), and the Fast Confidence Prediction of Uncertainty based on Knowledge GraphEmbedding (UKG sE), Probabilistic Box Embeddings for Uncertain Knowledge Graph Reasoning (BEUrRE), UnceRtain Graph Embedding (URGE), Uncertain knowledge graph embedding - an effective method combining multi-relation and multi-path (MUKGE), and UKGE-SRDA (the method of the present invention). As can be seen from Table 3, the best results are highlighted in bold, and the dash (——) indicates that the experiment was not performed on that specific dataset.

[0072] Table 3 Comparison of Model Confidence Prediction Evaluation Results Among them, UKGE and BEUrRE used the original data. It can be analyzed from Table 3 that UKGE-SRDA achieved the best results in the other two datasets except for the PPI5k dataset. Comparing with BEUrRE, the best model in the CN15k dataset, the MSE of UKGE-SRDA of the present invention decreased by 0.27, and the MAE decreased by 0.1. Similarly, in the NL27k dataset, compared with UComplEx, the best model among the comparison models, the MSE of UKGE-SRDA decreased by 0.02, and the MAE decreased by 0.03. In the PPI5k dataset, compared with UKGErect, the MSE of UKGE-SRDA decreased by 0.64, and the MAE decreased by 1.81. This shows that the accuracy of UKGE-SRDA in confidence prediction evaluation has been improved compared with BEUrRE, UKGE, MUKGE, UKGsE, and UComplEx, and it can better fit the confidence values in the uncertain knowledge graph.

[0073] To deeply elaborate the difference between the predicted confidence and the actual confidence of UKGE-SRDA of the present invention, scatter plots were drawn for the above three datasets, as Figure 3 、 Figure 4 、 Figure 5 shown. The x axis in the figure represents the exponent of the quadruple in the test data, yThe axis represents confidence, the black dots represent the actual confidence, and the red dots represent the predicted confidence. By observing the scatter plot, it can be found that in CN15k, the actual confidence is mainly distributed in the interval of 0.1 to 1.0, while the predicted confidence is more concentrated in the interval of 0.2 to 0.9. Generally speaking, UKGE-SRDA shows relatively excellent prediction ability. On NL27k, it is found that the distribution of the actual confidence is relatively tight, and the proximity of the predicted confidence to the actual value is relatively high, which further confirms the efficiency of UKGE-SRDA in prediction. On PPI5k, the distribution of the actual confidence is mainly in the interval of 0.1 to 0.6 and the region close to 0.9, and the distribution of the predicted confidence is almost the same as that of the actual confidence, which indicates that the effect of UKGE-SRDA in prediction is very significant. These observation results are mutually verified with the data in Table 3.

[0074] The purpose of the relational fact ranking task is to evaluate the accuracy of predicting the ranking order of an uncertain knowledge graph. The specific approach of this task is to extract the head entity, relation, tail entity, and confidence score from some triples in the uncertain knowledge graph, use the confidence score to predict and sort the tail entities, and compare the predicted order of the real tail entities through the evaluation method of normalized discounted cumulative gain ( NDCG ), and measure the performance of its ranking.

[0075] Using the NDCG of linear gain NDCG and the NDCG of exponential gain

[0076] where is the head entity, is the relation, is the tail entity, is the confidence score, is the triple set, is the triple in the current ranking position, is the triple in the ideal ranking position, is the discount factor of the current ranking, which is used to punish the triples with lower rankings. The lower the ranking, the larger the discount factor, and the smaller the contribution to NDCG , is the discount factor of the ideal ranking (sorted in descending order of relevance).

[0077] Based on exponential gain NDCG The calculation is as follows: Wherein, is the gain, is the head entity, is the relation, is the tail entity, is the confidence score, is the triple set, is the triple in the position in the current ranking, is the triple in the position in the ideal ranking, is the discount factor of the current ranking, used to penalize the triples ranked at the end. The lower the ranking, the larger the discount factor, and the smaller the contribution to NDCG is, is the discount factor of the ideal ranking.

[0078] Table 4 Comparison of UKGE-SRDA relation fact ranking evaluation results Table 4 reports the experimental results of UKGE-SRDA and the comparison models in the relation fact ranking task. UKGE-SRDA has better representation effects on larger-scale datasets. The three deterministic knowledge graph representation learning models, Distmult, ComplEx, and TransE, cannot model uncertain knowledge graphs. UKGE performs poorly among all uncertain knowledge graph representation learning models. Although probability soft logic is adopted, it still cannot reason invisible fact relations well. The reason is that UKGE adopts Distmult which can only model symmetric relations. UComplEx adopts Beta embedding and two embedding spaces to represent entities and relations. Therefore, UComplEx has obtained good results in terms of performance. However, on the NL27K dataset with more relation transmissions, none of them can significantly improve the performance. In contrast, MUKGE adopts a multi-path reasoning method and can handle various logical relations in uncertain knowledge graphs. However, this model cannot perform local reasoning on sparse long-tail data in uncertain knowledge graphs. Therefore, this model only shows better performance improvement on the NL27K dataset. Due to the introduction of enhanced rules and relation embeddings, UKGE-SRDA further improves the model's effect. Compared with UKGE, the linear gain of UKGE-SRDA has increased by 12.2%, 1.5%, and 1.9% respectively, and the exponential gain has increased by 12.1%, 1.3%, and 1.8% respectively.

[0079] The task of tail entity prediction is to predict the correct tail entity given the head entity and relation, and rank all possible tail entities.

[0080] All entities are regarded as potential tail entities, combined with the given head entity and relation to form potential triples, and then they are ranked according to the scores of these triples. The evaluation of the ranking effect uses NDCG. Since there is uncertainty in the context of the uncertain knowledge graph, it is particularly important to select an evaluation metric that can comprehensively consider both ranking and uncertainty. NDCG can take into account the ranking order and give a smaller weight to the correct predictions that are ranked later through a discount function, which is consistent with the uncertainty of predictions in the uncertain knowledge graph.

[0081] The experimental results are shown in Table 5. UComplEx performs relatively consistently on all datasets. And it performs best on the CN15k dataset, with an NDCG reaching 0.296. This indicates that UComplEx

[0082] Table 5 Comparison of Evaluation Results of Tail Entity Prediction in UKGE-SRDA May perform better when dealing with datasets with complex relations and high-dimensional features. UKGsE performs slightly better than UKGE on the CN15k dataset rect and UKGE logi , but performs worse than the former two on the NL27k and PPI5k datasets, which may indicate that UKGsE may be more effective when dealing with certain specific types of datasets. BEUrRE performs equivalently to UKGsE on the CN15k dataset, but performs slightly worse on the NL27k and PPI5k datasets. This may indicate that BEUrRE needs further optimization when dealing with datasets of different scales and types. MUKGE performs worse than other models on the CN15k dataset, but performs best on the NL27k dataset, which may indicate that MUKGE has an advantage when dealing with specific types of datasets (such as NL27k).

[0083] UKGE-SRDA outperforms other models in tail entity prediction and performs best on the NL27k dataset, with an NDCG reaching 0.753. This result reveals its significant advantage in dealing with complex datasets, being able to well capture the uncertainty of entities and relations and effectively handle complex one-to-many relations. Secondly, by combining logical rules, the model's ability in tail entity prediction is enhanced, which indicates that introducing logical rules helps reduce the impact of false negative samples. The introduction of false negative samples may cause the correct candidate tail entities to be wrongly ranked later, which has a significant negative impact on the model's performance.

[0084] The present invention uses AMIE3 to mine some triple soft rules from CN15k, as shown in Table 6 specifically, and predicts the confidence scores of unobserved facts.

[0085] Table 6 Some triple soft rules mined from CN15k Through inference, the predicted confidence is very close to the calculated confidence, as shown in Table 7. For example, the unobserved relational fact (Diet, Promote, Health) is deduced from the rule (A, Provide, B) ∧ (B, Promote, C) → (A, Promote, C). The confidence of (A, Provide, B) is 0.813, the confidence of (B, Promote, C) is 0.816, and the theoretical value of the confidence of (Diet, Promote, Health) should be 0.815. The predicted confidence of UKGE-SRDA is 0.810, which is very close to the theoretical value, with only a difference of 0.005. In addition, the predicted confidence is consistent with human common sense, further demonstrating the excellent performance of the proposed UKGE-SRDA in reasoning using unobserved knowledge.

[0086] Table 7 Some unobserved relational facts and their confidences The UKGE-SRDA proposed by the present invention solves the false negative problem in uncertain knowledge graph embedding by integrating logical rules, which enriches the understanding of data, improves the prediction accuracy and the interpretability of the model. In addition, the map mapping sorting technique is used to enhance the Rete algorithm, which preprocesses the rule conditions, reduces shared nodes and redundant matches, thereby optimizing the memory.

[0087] The present invention also provides a knowledge graph embedding system, including: A data acquisition module for acquiring the original uncertain knowledge graph; A data processing module for mining soft rules with different confidences from the original uncertain knowledge graph using a rule mining tool; performing rule inference on the soft rules using a forward chaining inference method to generate new rules; instantiating the entity variables in the new rules to obtain new triple instantiation data corresponding to the entity variables consistently; A knowledge graph embedding model training module for fusing the new triple instantiation data with the original triples in the original uncertain knowledge graph, and then inputting the fused enhanced data into the knowledge graph embedding model to train the knowledge graph embedding model to obtain a trained knowledge graph embedding model; The knowledge graph embedding module obtains an uncertain knowledge graph to be processed, and inputs the uncertain knowledge graph to be processed into a trained knowledge graph embedding model to obtain the knowledge graph after embedding representation.

[0088] The present invention also provides a computer device, including a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute the knowledge graph embedding method.

[0089] The present invention also provides a computer-readable storage medium, which stores a computer program, and the computer program is suitable for being loaded by a processor to execute the knowledge graph embedding method.

[0090] The above-described embodiments are only preferred specific embodiments of the present invention, and the protection scope of the present invention is not limited thereto. Any simple changes or equivalent replacements of technical solutions that can be obviously obtained by those skilled in the art within the technical scope disclosed by the present invention all belong to the protection scope of the present invention.

Claims

1. A knowledge graph embedding method, characterized in that, It includes the following steps: Obtain the original uncertain knowledge graph; Use a rule mining tool to mine soft rules with different confidence levels from the original uncertain knowledge graph; use the forward chaining inference method to perform rule inference on the soft rules to generate new rules; instantiate the entity variables in the new rules to obtain new triple instantiation data corresponding to the entity variables; Fuse the new triple instantiation data with the original triples in the original uncertain knowledge graph, and then input the fused enhanced data into the knowledge graph embedding model to train the knowledge graph embedding model to obtain a trained knowledge graph embedding model; Obtain the uncertain knowledge graph to be processed, and input the uncertain knowledge graph to be processed into the trained knowledge graph embedding model to obtain a knowledge graph after embedding representation.

2. The knowledge graph embedding method according to claim 1, wherein The step of using the forward chaining inference method to perform rule inference on the soft rules to generate new rules includes the following steps: Use a conflict resolution strategy to select the rules that need to be triggered in the soft rules; Activate the triggered rules to generate new rules; Add the generated new rules to the rule candidate set; Repeat the above three steps until no rules are triggered.

3. The knowledge graph embedding method according to claim 1, wherein The calculation formula for the confidence level is: g ( l )= φ ( dr ( h , t )); Among them, φ (·) is a conversion function, d r ( h , t ) is the distance function of any given triple, ([[]] h , t ) is the entity set, h is the head entity, t is the tail entity.

4. The knowledge graph embedding method according to claim 1, characterized in that When training the knowledge graph embedding model, the following loss function is specifically used: L = L tr + L rule ; Among them, L tr is the loss function of the triple, L rule is the loss function of all rules.

5. The knowledge graph embedding method according to claim 4, wherein The loss function for the triples is as follows: ; Among them, L pos is the positive example loss function, is the final negative sampling loss function, is the number of generated negative samples.

6. The knowledge graph embedding method according to claim 4, wherein The loss function for all rules is as follows: ; Among them, is the new triple f i 's loss function, L rc is the confidence loss function of the rule, f r is the soft rule, c r is the rule confidence, F is the soft rule set.

7. A knowledge graph embedding system, characterized in that It includes: A data acquisition module for obtaining the original uncertain knowledge graph; A data processing module for using a rule mining tool to mine soft rules with different confidence levels from the original uncertain knowledge graph; Use the forward chaining inference method to perform rule inference on the soft rules to generate new rules; instantiate the entity variables in the new rules to obtain new triple instantiation data corresponding to the entity variables; A knowledge graph embedding model training module for fusing the new triple instantiation data with the original triples in the original uncertain knowledge graph, and then inputting the fused enhanced data into the knowledge graph embedding model to train the knowledge graph embedding model to obtain a trained knowledge graph embedding model; A knowledge graph embedding module for obtaining the uncertain knowledge graph to be processed and inputting the uncertain knowledge graph to be processed into the trained knowledge graph embedding model to obtain a knowledge graph after embedding representation.

8. A computer device, characterized in that, It includes a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute the knowledge graph embedding method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by the processor to execute the knowledge graph embedding method according to any one of claims 1-6.