A Multi-Space Semantic Data Stream Inference Method Based on Rule Embedding Representation
Through a multi-space semantic data flow inference method based on rule embedding representation, combined with the joint embedding representation of rules and facts, the problem of inability to adapt to high-speed updates and weak rule constraints in the prior art is solved, and efficient and accurate knowledge reasoning and extensive knowledge acquisition are achieved.
Patent Information
- Application Number
- CN202210497722.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-09
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2042-05-09
AI Technical Summary
The existing semantic data flow inference methods cannot be applied to high-speed update scenarios, cannot achieve a balance of accuracy and timeliness, and cannot conduct extensive knowledge inference under the constraints of weak rules.
A multi-space semantic data flow inference method based on rule embedding representation is adopted. By combining the joint embedding representation of rules and facts, different dynamic inference processes are designed, including dynamic inference based on parallel semantic space, context perception and rule generalization learning, to adapt to the semantic data flow requirements of different scenarios.
It realizes efficient and accurate knowledge inference in the high-speed updated semantic data flow scenario, which can improve the inference efficiency when sacrificing certain accuracy, and obtain more comprehensive inference results under the constraints of weak rules.
Smart Images

Figure CN115438789B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of knowledge graph and data stream processing. Specifically, it uses knowledge representation learning methods for dynamic knowledge reasoning tasks to achieve semantic data stream reasoning technology in high-speed knowledge update scenarios. Background Art
[0002] With the development of technologies such as mobile Internet, big data, Internet of Things, and social networks, there is an increasing need to analyze and process high-speed real-time data streams. Traditional data stream processing means mainly conduct research on data stream aggregation and classification according to the data structure or value in the data stream [1] , and cannot analyze and reason about data streams at the semantic and knowledge levels. With the support of knowledge graph technology, the Resource Description Framework (RDF) can be used to endow data streams with semantics and knowledge, obtaining semantic data streams with better interoperability and inferability. Semantic data streams can be combined with static background knowledge for querying and reasoning, and have broad application value.
[0003] In recent years, there have been many progresses in the field of knowledge reasoning, but the existing parallel rule reasoning mechanisms still have deficiencies in dealing with large-scale graphs [2] . RDFox [3] proposed a high-performance knowledge graph query and reasoning mechanism, and used a combination of forward and backward methods for deletion consistency checking. The literature [4] proposed a parallel forward reasoning algorithm optimized based on single-way derivability, and its performance is better than that of RDFox. However, the above knowledge reasoning research mainly focuses on static data and is not applicable to semantic data streams. The domestic research on semantic data streams mainly focuses on window and query optimization [5] , event processing [6] and clustering analysis [7] , and there is less research on the real-time reasoning of semantic data streams. The existing semantic data stream reasoning schemes are mainly based on deterministic rule derivation, and uncertain knowledge reasoning cannot be applied to semantic data streams at present. Earlier uncertain statistical reasoning mainly used means such as Markov logic network [8] and probabilistic soft logic [9] to analyze ontology frequent patterns, constraints and paths to obtain reasoning results, but these methods still need to rely on instantiation, so the computability is poor
[10] .
[0004] In 2013, Mikolov et al. proposed word2vec
[11] After the word representation learning model, the way of representation learning has received extensive attention in the field of knowledge graphs. Subsequently, Bordes et al. proposed the TransE knowledge representation learning model
[12] . This model regards the relationship in the triple as a translation vector, representing the translation between the head entity and the tail entity. Because the parameters of the TransE model are simple and the computational complexity is low, a large number of knowledge representation learning models have emerged that improve and supplement the TransE model, such as TransH
[13] , TransR
[14] and TranSpare
[15] etc. Although the above models have made corresponding improvements to the TransE model, they are all based on the vectorization of factual triples without using rules. Therefore, KALE
[16] established a unified representation framework for triples and rules on the basis of the TransE model. However, KALE cannot perform incremental updates. In addition, KALE also cannot learn and generalize the rules themselves to meet the reasoning needs under weak rule constraints.
[0005] Dynamic Graph Embedding methods study the reasoning of dynamically changing graphs on the basis of knowledge representation learning. Dynamic embedding methods capture the changing characteristics of the graph by introducing temporal regularization in the cost function
[17] or adding a time-layer-based attention mechanism
[18] and other methods, and are often used in news events, social networks or wireless network analysis fields. However, this method still requires full-scale training. Therefore, it mainly studies improving the prediction accuracy for slowly changing (update time ranging from several days to several years) graphs, rather than timeliness. By dynamically adjusting the learning rate, the training time can be shortened
[19] , but this method depends on data characteristics and thus has poor transferability.
[0006] To sum up, the existing knowledge representation reasoning cannot be applied to dynamically updated semantic data streams, cannot achieve high-efficiency complex reasoning by sacrificing a certain degree of accuracy in scenarios with poor accuracy requirements, and cannot obtain more comprehensive and extensive reasoning results through rule generalization and rule learning under weak rule constraints.
[0007] [1] Zhang Xiaolong, Zeng Wei. New Advances in the Research of Real-Time Data Stream Clustering [J]. Computer Engineering and Design, 2009, 30(9): 2177 - 2181, 2186.
[0008] [2]Antoniou G, Batsakis S, Mutharaju R, et al. A survey of large-scale reasoning on the Web of data[J]. The Knowledge Engineering Review, 2018, 33.
[0009] [3]Yavor Nenov, Robert Piro, Boris Motik, et al. RDFox: A Highly-Scalable RDF Store[C] / / International Semantic Web Conference. Springer International Publishing, 2015.
[0010] [4]Zhangquan Zhou, Guilin Qi, Birte Glimm: Parallel tractability of ontology materialization: Technique and practice. J.Web Semant. 52-53:45-65(2018)
[0011] [5]Gu Yu, Li Xiaojing, Xu Jia, et al. Modeling and query optimization of sliding window joins for data streams supporting complex semantics[J]. Journal of Northeastern University (Natural Science Edition) (11): 34-37.
[0012] [6]Wang Weifang. Research on complex event processing method based on semantic data stream in semantic Internet of Things[D]. Liaoning: Dalian Maritime University, 2017.
[0013] [7]Chen Fengjiao, Fu Haidong, Wu Gang, et al. Research on partitioning algorithm for massive semantic data streams based on heuristic strategy[J]. Systems Engineering - Theory & Practice, 2014, 34(s1): 248-254.
[0014] [8]Shangpu Jiang, Daniel Lowd, Dejing Dou. Learning to Refine an Automatically Extracted Knowledge Base Using Markov Logic[C] / / IEEE International Conference on Data Mining. IEEE, 2012.
[0015] [9]Pujara J,Miao H,Getoor L,et al.Knowledge Graph Identification[C] / / International Semantic Web Conference.Springer Berlin Heidelberg,2013.
[0016]
[10] Guan S P,Jin X L,Jia Y T,et al.Research Progress of Knowledge Reasoning for Knowledge Graph[J].Journal of Software,2018, 29(10):74-102.
[0017]
[11] Mikolov T,Sutskever I,Chen K,et al.Distributed representations ofwords and phrases and their compositionality[J].Advances in NeuralInformation Processing Systems,2013.
[0018]
[12] Bordes A,Usunier N,Garcia-Duran A,et al.Translating embeddingsfor modeling multi-relational data[C] / / In Proceedings of the 27th AnnualConference on Neural Information Processing Systems,2013:2787–2795.
[0019]
[13] Wang Z,Zhang J,Feng J,et al.Knowledge Graph Embedding byTranslating on Hyperplanes[C] / / In Proceedings of the 28th AAAI Conference onArtificial Intelligence.2014.
[0020]
[14] Lin Y, Liu Z, Sun M, et al. Learning entity and relation embeddings for knowledge graph completion[C] / / In Proceedings of the 29th AAAI Conference on Artificial Intelligence, 2015:2181–2187.
[0021]
[15] Ji G, Liu K, He S, et al. Knowledge Graph Completion with Adaptive Sparse Transfer Matrix[C] / / Thirtieth Aaai Conference on Artificial Intelligence. AAAI Press, 2016.
[0022]
[16] Guo S, Wang Q, Wang L, et al. Jointly embedding knowledge graphs and logical rules[C] / / Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, 2016:192–202.
[0023]
[17] Jérémie Rappaz, Dylan Bourgeois, and Karl Aberer. 2019. A Dynamic Embedding Model of the Media Landscape. In The World Wide Web Conference (WWW’19). Association for Computing Machinery, New York, NY, USA, 1544–1554.
[0024]
[18] Yang L., Xiao Z., Jiang W., et al. (2020) Dynamic Heterogeneous Graph Embedding Using Hierarchical Attentions. In: Jose J. et al. (eds) Advances in Information Retrieval. ECIR 2020. Lecture Notes in Computer Science, vol 12036. Springer, Cham
[0025]
[19] Minervini P, Fanizzi N, D'Amato C, et al. Scalable Learning of Entity and Predicate Embeddings for Knowledge Graph Completion[C] / / 2015 IEEE 14th International Conference on Machine Learning and Applications(ICMLA). IEEE, 2015. Summary of the Invention
[0026] Aiming at the deficiencies of existing semantic data stream reasoning methods, the present invention proposes a multi-space semantic data stream reasoning method based on rule embedding representation. The purpose of the present invention is to provide an uncertainty knowledge reasoning method with high accuracy and good timeliness on the basis of existing semantic data stream query and reasoning, so as to realize dynamic real-time knowledge reasoning in the scenario of a triple data stream with high-speed updates.
[0027] In order to solve the above-mentioned deficiencies existing in the prior art, the present invention proposes a dynamic knowledge reasoning method based on rule embedding and context awareness. First, the joint embedding representation of rules and facts is realized, and then different dynamic reasoning processes are set according to different semantic data stream reasoning requirements and scenarios, including rule embedding dynamic reasoning based on parallel semantic space reasoning, rule embedding dynamic reasoning based on context awareness, and dynamic reasoning based on rule generalization learning;
[0028] The rule-embedding dynamic reasoning based on parallel semantic space reasoning is for the reasoning scenario with a head-to-tail entity ratio of 1:1. It is processed by combining a semantic data flow processing platform and knowledge representation learning. The implementation process includes: First, expand the joint embedding model KALE to implement a unified representation reasoning framework for complex rules and fact tuples; then, construct and select the embedding space; finally, implement dynamic reasoning training based on the joint embedding model, and perform real-time knowledge reasoning according to the trained output model.
[0029] The rule-embedding dynamic reasoning based on context awareness is for the reasoning scenario with a head-to-tail entity ratio of 1-to-many or many-to-many. It performs knowledge reasoning with context-aware multi-semantic space fusion representation. The implementation process includes: First, expand on the basis of the multi-space joint embedding model to implement a context-aware dynamic reasoning framework; second, model and represent the entity relationship, its type, and the rule-associated context; third, integrate the context vector space and the rule embedding space; finally, support incremental update and cumulative offset awareness.
[0030] The dynamic reasoning based on rule generalization learning is for the reasoning scenario with weak rule constraints or incomplete rules. It conducts rule learning based on rule examples to generalize the rules and expand the reasoning results. The implementation process includes conducting rule generalization research based on the real-time reasoning of knowledge representation learning and combining it with a semantic data flow processing platform to implement uncertain real-time knowledge reasoning under weak rule constraints.
[0031] Moreover, the rule-embedding dynamic reasoning based on parallel semantic space reasoning includes the following steps:
[0032] Step 1, construct the overall training framework of the dynamic joint embedding model;
[0033] Step 2, generate a sub-embedding space based on the method combining semantics and structure association, covering the semantic and structure association space within 2-hop range;
[0034] Step 3, select the sub-embedding space based on triple indexing, and select the updated self-embedding space based on the new triple and its associated rules;
[0035] Step 4, formulate the objective function and solution based on the margin model;
[0036] Step 5, implement the extension and real-time reasoning of the dynamic joint embedding model based on the C-SPARQL engine.
[0037] Moreover, the rule-embedding dynamic reasoning based on context awareness includes the following steps:
[0038] Step 1, construct a subgraph of entities and their type contexts covering a 2-hop distance;
[0039] Step 2, construction of a relationship context subgraph covering the same-direction path, parent relationship type, and rule association within a 2-hop distance;
[0040] Step 3, multi-vector integration and rule joint embedding;
[0041] Step 4, incremental update based on the context subgraph, and dynamically maintain the cumulative vector offset. When the cumulative offset exceeds the limit value, perform a global embedding space update.
[0042] Moreover, the dynamic reasoning of the rule generalization learning includes the following steps:
[0043] Step 1, a generalization learning model for OWL rules;
[0044] Step 2, NRE reasoning based on multi-space incremental training.
[0045] Compared with the prior art, the present invention has the following advantages: 1) Using a deep learning-based method to achieve uncertainty knowledge reasoning applicable to semantic data streams, sacrificing a certain degree of accuracy to obtain high efficiency in reasoning through vector calculations, thereby achieving high-throughput dynamic knowledge reasoning; 2) Using a method based on rule generalization learning to achieve uncertainty knowledge reasoning applicable to semantic data streams, avoiding the complexity of manual rule formulation through learning the rules themselves, and thus obtaining more comprehensive real-time reasoning results.
[0046] The implementation of the present invention is simple and convenient, with strong practicability. It solves the problems of low practicability and inconvenience in actual application existing in the related technology, can improve the user experience, and has important market value. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 Schematic diagram of the real-time reasoning architecture of the extended joint embedding model according to an embodiment of the present invention;
[0048] Figure 2 Schematic diagram of the specific implementation of the NRE atomic operation and tree structure according to an embodiment of the present invention.
[0049] Figure 3 Schematic diagram of the semantic-related embedding space triple generation strategy according to an embodiment of the present invention;
[0050] Figure 4 Schematic diagram of the structure-related embedding space triple generation strategy according to an embodiment of the present invention;
[0051] Figure 5 Schematic diagram of the semantic structure fusion-related embedding space triple generation strategy according to an embodiment of the present invention;
[0052] Figure 6 Schematic diagram of the mapping of entities and relationships to the embedding space according to an embodiment of the present invention;
[0053] Figure 7 Flowchart of embedding space selection based on index mapping according to an embodiment of the present invention;
[0054] Figure 8 Basic structure diagram of the KALE-CTX model according to an embodiment of the present invention;
[0055] Figure 9 Schematic diagram of context subgraph selection according to an embodiment of the present invention;
[0056] Figure 10 Schematic diagram of relationship context selection according to an embodiment of the present invention;
[0057] Figure 11 Structural diagram of context-based entity and relationship joint embedding according to an embodiment of the present invention;
[0058] Figure 12 Schematic diagram of incremental update according to an embodiment of the present invention. Detailed implementation manners
[0059] The technical solution of the present invention will be specifically described below in conjunction with the accompanying drawings and embodiments.
[0060] An embodiment of the present invention discloses a multi-space semantic data stream reasoning method based on rule embedding representation, the goal of which is to study the reasoning of massive high-speed semantic data streams and reduce its reasoning latency. The present invention improves the joint embedding model integrated with rule learning to adapt to the reasoning queries of high-speed changing semantic data streams. Specifically, in order to achieve incremental update of knowledge in the data stream and reduce the reasoning time delay, the present invention introduces parallel multi-embedding spaces, and proposes three embedding space generation algorithms for the generation method of the embedding space; for the selection of the embedding space, the present invention proposes an embedding space selection algorithm based on index mapping to speed up its selection speed. In addition, in order to further improve the relevance between the complex and diverse data and queries in the embedding space, the present invention introduces a knowledge representation and embedding method based on context information, and proposes an embedding space generation method based on entity and relationship contexts; finally, the present invention proposes a semantic data stream reasoning method based on rule generalization learning, so as to obtain more comprehensive reasoning results under weak rule constraints. The present invention integrates the improved rule embedding representation learning method into the data stream reasoning engine CSPARQL-engine to make it have the real-time reasoning ability based on multi-mode knowledge representation learning. This solution can improve the complex reasoning speed on the premise of ensuring high reasoning accuracy compared with the traditional stream reasoning engine. Its knowledge embedding representation method is applicable to 1-1 to N-N entity relationship embedding scenarios, and its rule learning method is applicable to dynamic reasoning scenarios under weak rule constraints.
[0061] The embodiments of the present invention are improved based on the classical TransE translation model. First, the joint embedding representation of rules and facts is realized, and then different dynamic reasoning schemes are designed for different semantic data stream reasoning requirements and scenarios, mainly including three parts: dynamic reasoning of rule embedding based on parallel semantic space reasoning, dynamic reasoning of rule embedding based on context awareness, and dynamic reasoning based on rule generalization learning.
[0062] (1) For the reasoning scenario where the ratio of head and tail entities is 1:1 in many cases, the present invention designs a solution combining a semantic data stream processing platform and knowledge representation learning. First, the extension of the joint embedding model (KALE) is studied to realize a unified representation and reasoning framework for complex rules and fact tuples; then, the construction and selection methods of the embedding space are designed; finally, a dynamic reasoning training scheme based on the joint embedding model is realized, and real-time knowledge reasoning is achieved according to the trained output model.
[0063] (2) For the reasoning scenario where the ratio of head and tail entities is 1 to many or many to many, the present invention designs a knowledge reasoning solution based on context-aware multi-semantic space fusion representation. First, it is extended on the basis of the multi-space joint embedding model to realize a context-aware dynamic reasoning framework; secondly, the entity relationship, its type, and the rule-associated context are modeled and represented; thirdly, the integration of the context vector space and the rule embedding space is carried out; finally, the support for incremental update and cumulative offset perception is realized.
[0064] (3) For the reasoning scenario with weak rule constraints or incomplete rules, it is necessary to learn rules based on rule examples to generalize the rules and expand the reasoning results. The present invention conducts research on rule generalization based on the real-time reasoning of knowledge representation learning, and combines it with the semantic data stream processing platform to realize uncertain real-time knowledge reasoning under weak rule constraints.
[0065] 1. The semantic data stream reasoning scheme based on the joint embedding representation of rules and facts provided by the embodiments of the present invention is realized as follows:
[0066] The present invention is extended on the basis of the joint embedding model. First, the support range of rule embedding is extended to OWL rule types and the corresponding truth value calculation method is designed, and then it is made applicable to the dynamically updated semantic data stream, and a prediction template generation algorithm based on triple indexing is constructed to make it applicable to the knowledge reasoning task. Its overall reasoning architecture is as Figure 1As shown in Figure 1. The architecture mainly consists of three parts: the KR reasoning mechanism, the query engine, and the query parsing. The KR reasoning mechanism first uses the RDF data stream and rule set, as well as the range selection information of the stream in the query, to build an RDF index structure, and generates triples to be predicted based on the index structure. The predicted triples are then fed into the PU-KALE model for triple classification based on the knowledge representation model. The classification results, i.e., the inferred additional triples, are fed into the query engine together with the original RDF stream for graph matching, and a result stream is generated. The query parsing model mainly parses the continuous RDF stream query, for example Figure 1 The query in the lower left corner is intended to find information about students and their advisors (?x advisor?y) in the data stream, where the advisor's title is professor (?y type Professor). The query range is the past ten minutes, with a sliding window of 2 minutes (win=10, step=2). The stream selection portion of the query (win=10, step=2) is used to generate predicted triples in conjunction with the index, and the SPARQL query pattern is submitted to the query engine for graph matching. Implementing this architecture requires the following steps.
[0067] Step 1: Construct an overall training framework: Construct an overall training framework for a dynamic joint embedding model; the present invention proposes a PU-KALE (Parallel Universe KALE, parallel space rule joint embedding model) solution for constructing a knowledge reasoning framework under a dynamic update background. PU-KALE borrows the idea of parallel semantic space and explicitly generates multiple embedding spaces based on the KALE model to learn newly added fact triples. During the change process of the initial knowledge graph model, PU-KALE divides different subspaces and can match the sliding of the query window, that is, each window update content is used as a new vector space for training. During each training process, the data in the sliding window will establish a connection with the newly added semantic space, which is convenient for distinguishing different fact triples in multiple semantic spaces.
[0068] In the training preparation phase, the present invention generates multiple embedding spaces based on pre-set parameters such as the total number of training set triplets and the sub-embedding size, and the joint embedding model is trained separately in multiple embedding spaces. When the triples are incrementally updated, if the new entities and relationships are in the existing embedding space, the corresponding sub-embedding space is updated and retrained; conversely, if there are still remaining embedding positions in the current multiple subspaces (i.e., the subspace capacity has not reached the limit value), the new triples are added to the unfilled subspace and retrained; if all subspaces are full, a new sub-embedding space is created. Specifically, it includes the following sub-steps:
[0069] Step 1.1: Prepare to partition the embedding space, including obtaining a set of triples T = {(h, r, t)}, a set of entities E, a set of relations R, the vector dimension d, and the number of training iterations δ; where h, r, and t represent the head entity, relation, and tail entity respectively.
[0070] Step 1.2: Read the rule set ζ into the model;
[0071] Step 1.3: For each sub-embedding space Execute steps 1.4 - 1.11 respectively; where i represents the i-th sub-space;
[0072] Step 1.4: Generate random hyperparameters for the sub-embedding space;
[0073] Step 1.5: Initialize the entities Ei and the relation matrix Ri of the current embedding space, as well as the set of triples Ti;
[0074] Step 1.6: Infer new triples according to the rules and add them to Ti as positive examples;
[0075] Step 1.7: When the number of training iterations is less than the preset value δ (it is recommended to take 20, which can be adjusted according to the actual effect), execute steps 1.8 - 1.10. When the number of training iterations is greater than or equal to δ, enter step 1.12;
[0076] Step 1.8: Select random triples to generate negative triples and add them to Ti as negative examples;
[0077] Step 1.9: Select rules to generate negative rules and add them to Ti as rule negative examples;
[0078] Step 1.10: Update the matrix parameters;
[0079] Step 1.11: Increment the iteration count by 1, and then return to step 1.7;
[0080] Step 1.12: Save the entities and the relation matrix of the embedding space.
[0081] Step 2: Generation of sub-embedding spaces: Based on a method for generating sub-embedding spaces that combines semantic and structural associations, covering the semantic and structural association spaces within 2-hop range; for the generation of the embedding space in the PUKALE model, the present invention implements multiple strategies for generating embedding spaces. The first strategy is a semantically related generation method, the second strategy is a structurally related method, and the third strategy is an organic combination of the two.
[0082] In the first semantically related strategy, when partitioning triples into an embedding space, triples with the same relation tend to be partitioned into the same embedding space. While triples with different relations tend to be partitioned into other embedding spaces.
[0083] The second structure-related strategy. The idea is that when creating an embedding space, first randomly select a triple and add it to the current embedding space. Then, starting from the head entity and the tail entity of the current triple, select triples in the out-degree and in-degree directions and add them to the current embedding space. After that, continue to start from the head entity and the tail entity of the newly added triples and continue to add triples in the out-degree and in-degree directions, and so on, until the number of added triples reaches the limit value of the embedding space.
[0084] The generation strategy of the third embedding space is a combination of the first method and the second method. The main idea of generating the embedding space is to first randomly select a triple, add it to the current embedding space, and record the predicate relationship of the current triple. Then, starting from the head and tail entities of the current triple, select triples within two-hop ranges in the out-degree and in-degree directions and add them to the embedding space.
[0085] The subspace generation method of the present invention mainly adopts the third strategy, which includes the following sub-steps.
[0086] Step 2.1: Prepare the training set triples T = {(h, r, t)}, the limit const on the number of triples in the embedding space, and save the triples to be embedded in the space
[0087] Step 2.2: When the number of triples in is greater than 0, execute Steps 2.3 - 2.9; when the number of triples in is less than or equal to 0, terminate the training;
[0088] Step 2.3: Construct the i-th embedding space
[0089] Step 2.4: Construct the triple set tripleSet, entity set entitySet, and relationship set relationSet of the current embedding space;
[0090] Step 2.5: When the number in tripleSet is less than the preset threshold const, execute Steps 2.6 - 2.9, otherwise stop constructing the subspace;
[0091] Step 2.6: Randomly select a triple (h, r, t) and add it to tripleSet;
[0092] Step 2.7: Add the head and tail entities h, t to the set tmpList;
[0093] Step 2.8: For each element e in tmpList, execute Step 2.9. After all the Step 2.9 for e is completed, enter 2.10;
[0094] Step 2.9: Obtain all triples V within 2-hop range of the e in-degree direction and add V to tripleSet;
[0095] Step 2.10: Obtain and output all sub-embedding spaces
[0096] Step 3: Sub-embedding space selection: Select sub-embedding spaces based on triple indices and select updated self-embedding spaces based on newly added triples and their associated rules; When performing incremental updates and retraining of sub-embedding spaces, the selection of the sub-embedding spaces to be updated or retrained is very important. A relatively simple approach is to linearly traverse each embedding space and calculate the energy score of the current triple in the current embedding space (calculated through the existing semantic matching energy model SME). Finally, select the one with the highest energy score among the eligible embedding spaces. This method requires traversal calculations and consumes a lot of time. The present invention designs an optimized selection method based on indices. First, establish an index mapping from entities and relationships to each embedding space, as shown in Figure 6 shown. For example, entity e1 and relationship r2 are mapped to 3 sub-spaces E1, E3, and E6 through index mapping, and e3 is mapped to 3 sub-spaces E2, E3, and E6; Then, when a new triple arrives, based on its entity and relationship, as well as the triple entity relationships that may trigger logical rules, search for the corresponding index list and visit the eligible embedding spaces. As in Figure 7 , first find the corresponding sub-spaces E2, E3, and E6 through the head and tail entities (Head / tail), then find E1, E3, and E6 according to the relationship index (relation), and then take the intersection to get 2 sub-spaces E3 and E6. The embedding space algorithm steps of the present invention include the following sub-steps.
[0097] Step 3.1: Initialize the trained embedding space, prepare the test tuple T = {(e, r)}, and obtain the index I e , I r ; where e and r are the entities and relationships to be tested, and I e , I r represent their index values respectively;
[0098] Step 3.2: For each triple t = (e, r) in T, find the set of eligible embedding spaces by searching for entities and relationships in I e , I r ;
[0099] Step 3.3: For each triple t = (e, r), retrieve all rule sets, find the triples generated by the rules that it may trigger, and use the same method as in Step 3.2 to find the set of embedding spaces that meet the conditions by looking up the triple index, and supplement them into the set of embedding spaces.
[0100] Step 3.4: For each space in calculate the energy score s of t t , and store s t in the score set S t ;
[0101] Step 3.5: Take the embedding space corresponding to max(S t ) as the selected embedding space. Among them, max(S ) represents to be filled t
[0102] Step 4: Define the objective function and scheme for model training; an objective function and scheme based on the margin model can be defined. In the embodiments of the present invention, based on the loss functions in the margin model and the joint embedding model, a training scheme is defined, which specifically includes the following sub-steps.
[0103] Step 4.1: Design the loss function as follows:
[0104]
[0105] In formula (1): L represents the margin distance between positive and negative triples, S+ and S- are sets of positive and negative triples, I represents the true value calculated by the model, and γ is the distance hyperparameter; f + is a positive example, f - is a negative example, I(f + ) is the true value of the positive example, I(f - ) is the true value of the negative example, and max() is the maximum value function.
[0106] Step 4.2: After the model starts training, first randomly initialize the vector representations (h, r, t) of the entities and relationships of the triples;
[0107] Step 4.3: Generate negative triples for the training samples. The generation method is to randomly replace the head entity or the tail entity of the sample to construct negative examples (h′, r, t) or (h, r, t′); where h′ is the negative example after the change of the head entity, and t′ is the negative example after the change of the tail entity;
[0108] Step 4.4: Calculate the score for each triple according to formula (1). For a correct triple, its expected score is lower; while for a negative triple, its expected score is higher. In this way, positive and negative samples are distinguished by the scores of the training samples.
[0109] Step 5: Implement the dynamic federated embedding model extension and real-time reasoning based on C-SPARQL; The present invention embeds the PUKALE model based on multiple embedding spaces into the semantic data stream platform CSPARQL-engine, enabling complex rule reasoning to be mapped to a low-dimensional vector space for simple vector calculations, reducing the reasoning latency after the continuous increase of the graph scale, and making it adapt to the real-time reasoning requirements in the scenario of large-scale streaming data. Specifically, it includes the following steps:
[0110] Step 5.1: Through the real-time semantic annotation module, call the Jena library method to format the input data into an RDF graph;
[0111] Step 5.2: Add a timestamp to each triple in the triple set generated by Jena to form a quadruple input acceptable by C-SPARQL;
[0112] Step 5.3: Read and parse the CSPARQL query to obtain the time step of the query and the RDF stream window size;
[0113] Step 5.4: Retrieve the underlying physical window data from the ESPER engine and provide it to both the CSPARQL engine and the PUKALE model simultaneously;
[0114] Step 5.5: Execute the PUKALE inference module, use the prediction results to generate a set of inference result triples, and input them into the query engine;
[0115] Step 5.6: The CSPARQL engine merges the original triples and the inference triples into the input stream, executes the query statement, and gives the result.
[0116] Among them, Jena is an open-source knowledge processing framework under APACHE, RDF is the Resource Description Framework, and ESPER is an open-source data stream processing engine.
[0117] Second, the knowledge reasoning solution based on context-aware multi-semantic space fusion representation provided by the embodiment of the present invention is implemented as follows:
[0118] PUKALE implements incremental updates on the semantic data stream platform CSPARQL - engine, but there are still the following problems: 1) The scoring function in TransE is still used in PUKALE to calculate the score of each triple, which cannot well model the 1 - N, N - 1, and N - N relationships in the knowledge graph; 2) PUKALE lacks the representation of type information; 3) PUKALE uses multiple semantic spaces. Although it avoids retraining the incremental update model on the entire dataset and instead selects the corresponding embedding space for retraining, reducing the number of triples that need to be retrained, it still lacks finer - grained control over the triples that need to be retrained.
[0119] To address the above problems, based on the KALE model, the present invention uses a context - aware dynamic graph embedding method, improves the generation method of entity and relationship context sub - graphs in the graph, and proposes the (Context - Joint Embedding Model Reasoner) KALE - CTX model. Then, the model is embedded into the data stream engine CSPARQL - engine. The basic structure diagram of the model is as Figure 8 shown. In KALE - CTX, each entity and relationship is represented by 2 vectors: the self - entity vector representation formed by the knowledge embedding method and the context sub - graph vector of the entity. After fusing the two parts of vectors of all entities, an initial vector space is formed. After incremental updates, the affected (added, deleted, modified) triples are located, their context embeddings are updated, and incremental retraining is implemented. The implementation of this model includes the following steps.
[0120] Step 1: Selection of entity and its type context: Construction of entity and its type context sub - graphs covering 2 - hop distances; for the context sub - graphs of entities in the graph network, the present invention makes selections from two aspects. First, all entities within 2 - hop ranges in all in - degree and out - degree directions adjacent to the entity are selected; then, all pattern - layer triples within 2 - hop ranges of the entity type are selected. The above - mentioned entity and its type - related triples form the entity context sub - graph. The reason for not introducing other entities with a farther distance is, on the one hand, to limit the scale of nodes in the sub - graph and reduce complexity; on the other hand, in experiments, after including entities with multi - hop distances, the accuracy of the model does not increase significantly, but instead, the training time increases. For example, in Figure 9 part (a), if (fleet, owns, ship 2) is added to the graph, the context sub - graph of the entity fleet consists of entities ship 1, ship 2, ship 3, and the fleet entity itself, as shown in Figure 9 part (b). The method for selecting the entity context sub - graph specifically includes the following sub - steps:
[0121] Step 1.1: Obtain all triples within one - hop in the in - degree and out - degree directions of the target entity;
[0122] Step 1.2: Save the selected triples and save the tail entities of these one-hop triples;
[0123] Step 1.3: Traverse in a loop to find triples with a target of two-hop;
[0124] Step 1.4: Select the triples that start from the tail entity of the two-hop triples and point to the target entity;
[0125] Step 1.5: Select the schema-layer triples within the 2-hop range of the entity type and save them;
[0126] Step 1.6: Set all the saved triples as the context space to perform embedding.
[0127] Step 2: Selection of relationship and its type context: Construction of relationship context subgraphs covering co-directional paths, parent relationship types, and rule-associated relationship contexts within a 2-hop distance; For the context of a certain relationship in the graph network, it is difficult to select appropriate adjacent entities or relationships to form a context subgraph because relationships in the graph may appear many times. For relationship context, the present invention makes selections from three aspects:
[0128] 1) Context information based on relationship paths: Select the relationships that are in the same direction as the relationship and are on the path connecting to the same entity pair as the components of the relationship context. In order to obtain the structural information between these relationships, the present invention maps both the relationship and the corresponding relationship path into vertices of a graph, adds undirected edges between them. If two relationship paths starting from the same entity point to the same end point, an undirected edge is also added between the vertices corresponding to the two paths. In this way, a context subgraph of the relationship can be constructed. Considering that one can reach the end point through multi-hop paths starting from an entity vertex, in order to control the number of relationship paths and reduce the complexity, the relationship paths selected by the present invention do not exceed two hops;
[0129] 2) Context information based on relationship types: Find the parent type of the relationship through subproperty association and construct a context subgraph based on relationship types using the same method as in 1);
[0130] 3) Context information based on rules: Logical rules also provide important information in the selection of relationships. Therefore, the present invention also combines rules for relationship context selection. For example, in Figure 10In part (a), the relationship "same combat group" is added between the physical ship 3 and the physical ship 1, and the relationship path related to the "same combat group" is p1 = (subordinate to, has destroyer), and this two-hop path connects the physical ship 3 and the physical ship 1. In addition, there is a logical rule for the "same combat group", that is, (x, same combat group, y) -> (x, same fleet, y). Therefore, the relationship "same fleet" can also be introduced into the context subgraph of the "same combat group", that is, the vertex p2 = (same fleet). The context subgraph of the "same combat group" can be obtained as shown in part (b) of Figure 10. The relationship context selection specifically includes the following sub-steps:
[0131] Step 2.1: Obtain the head and tail entities of the target triple;
[0132] Step 2.2: Select the triples whose paths connecting to the same entity pair are within the 2-hop range;
[0133] Step 2.3: Select the triples within the 2-hop range of the parent type relationships of all relationships on the 2-hop path;
[0134] Step 2.4: Retrieve the logical rule set and add all the rule context associated triples of all relationships in 2.2 - 2.3;
[0135] Step 2.5: Save the relationship context subgraph formed by 2.2 - 2.4.
[0136] Step 3: Multi-vector integration and joint embedding of rules; After obtaining the context subgraph of the entity or relationship, the subgraph can be encoded by the neural network in DKGE to obtain a context vector, and the entity or relationship itself also has a rule joint embedding vector. Figure 11 A multi-vector joint embedding structure diagram is given. Among them, subgraph(h) and subgraph(t) are the head and tail entity subgraphs, and subgraph(r) refers to the relationship subgraph. The head and tail entity subgraphs pass through the entity graph neural network (entity GCN) to obtain the entity's own vectors h, t and their context vectors c(h), c(t); the relationship subgraph passes through the relationship graph neural network (relation GCN) to obtain its own vector r and the relationship context vector c(r); after fusing their own and context vectors, the fused vectors h*, r* and t* are formed and integrated to form a complete model. The specific implementation of this step includes the following sub-steps:
[0137] Step 3.1: For the joint embedding of the entity or relationship context vector and the knowledge vector, an average operation is taken to obtain the result vector, as shown in Equation (2):
[0138]
[0139] Among them, o represents an entity or relationship in the graph, and o k represents the knowledge embedding corresponding to the entity or relationship, context(o) represents the context embedding corresponding to the entity or relationship, and o * represents the combined embedding vector of the two.
[0140] Step 3.2: For rule embedding, the method of rule embedding in KALE is used. Through specific t-norm-based logical connectives, the rule is converted into the calculation of truth values. For the calculation of the truth value of a triple, KALE-CTX uses the truth value calculation function shown in Equation (3).
[0141]
[0142] Among them, h * , r * , t * are calculated by formula (2), ||.|| l1 is the l1 norm, d is the vector dimension, and the initialization of the entity and relationship vectors during the training process follows a uniform distribution The value range of the triple truth value I is [0, 1], and the greater the truth value corresponding to the correct triple.
[0143] Step 3.3: Define the distance-based loss function as follows:
[0144] L = ∑ (h,r,t)∈5 ∑ (M,r,t,)∈s , max(0, I(h, r, t) + γ - I(h′, r, t′)) (4)
[0145] Among them, S is the set of correct triples in the graph, S′ is the set of incorrect triples, and the way to generate incorrect triples is the new triple (h′, r, t′) formed by replacing the head entity and tail entity in the correct triple. I is the truth value corresponding to the triple, and the max() function is used to represent taking the non-negative truth value gap (take 0 if less than 0).
[0146] Step 4: Incremental update and cumulative offset awareness: Based on the incremental update of the context subgraph, dynamically maintain the cumulative vector offset. When the cumulative offset exceeds the limit value, perform a global embedding space update; The embedding space of the knowledge graph in the present invention actually includes two types. The first type is the global embedding space, which is generated by the knowledge-axiom learning embedding (KALE) method, and the second type is the context embedding space, which is generated through the entity and relationship context generation steps in Steps 1-2. When small-scale dynamic updates occur in the graph, the knowledge representations of entities or relationships that are not directly affected remain unchanged, and the context vector representations also remain unchanged. They can still ensure that the triple satisfies h*+r*≈t*. For entities or relationships directly affected, only the knowledge embeddings and context embeddings of these elements and the newly added elements need to be retrained, which greatly reduces the number of triples that need to be retrained.
[0147] Take Figure 12 as an example. As shown in the figure, a new triple (e5, r5, e1) is added to the graph, where e5 and r5 are the newly added entity and relationship, and e1 is an existing entity. For the triples (e2, r2, e3) and (e4, r4, e3), the global embeddings and context embeddings of their corresponding entities and relationships have not changed. Therefore, the current update has no impact on them, and the triple constraint h * +r * ≈t * still holds and does not need to be adjusted. The affected triples in the original graph are (e1, r1, e2) and (e4, r3, e1). Because the context of entity e1 has changed, only the knowledge embeddings and context embeddings of e1 affected in the original graph, as well as the knowledge embeddings and context embeddings of the newly added entity e5 and relationship r5, need to be re-learned. That is to say, the model only needs to retrain the three triples (e1, r1, e2), (e4, r3, e1), and (e5, r5, e1), rather than retraining the entire graph, which can well meet the rapid update of the knowledge graph in the semantic data stream scenario.
[0148] Obviously, when there are many cumulative triple updates, only updating the context subgraph embedding is not sufficient to provide accurate inference results. At this time, there is a large offset between the initial global embedding space and the actual global embedding space, and it is necessary to perform retraining. To achieve the performance stability of the inference results in the long term or in a large number of update scenarios, it is necessary to provide an offset awareness function. In response to this problem, the present invention counts the offset δ of the relevant entity relationship vectors during each context subgraph update and dynamically maintains the total offset Δ = ∑δ. When the total offset Δ is greater than or equal to the threshold Δ limWhen this happens, global offline retraining is triggered to ensure inference accuracy. The threshold can be determined simply by using empirical methods or by performing a step-by-step search to plot the accuracy-vector offset curve to find the optimal value.
[0149] III. An embodiment of the present invention provides a semantic data stream inference solution based on a neural rule engine, which is implemented as follows:
[0150] The neural rule engine (NRE) is a newly proposed new model that deeply integrates rules and neural networks and applies them to natural language processing tasks. The characteristic of NRE is that after rewriting the rules in the form of regular expressions into a rule tree with nodes being one of the 6 basic text matching operations, end-to-end sequence annotation and training are performed on the structure or layout of the tree to obtain the generalization of the rules. Furthermore, for the leaf operation nodes that require neural networks, learning and training based on convolutional neural networks are carried out, and the trained model is called based on the generalized rule layout during the text classification task stage for text annotation and classification. Currently, NRE cannot be used for knowledge inference, and obviously cannot be used for real-time data stream processing. The present invention is extended on this basis to meet the requirements of real-time knowledge inference in weakly rule-constrained scenarios. The main technical solutions include two steps: improving the rule layout training model and combining it with PU-KALE.
[0151] Step 1: Rule generalization learning for OWL rules; in the original NRE, the rule tree formed by regular expressions is expressed in reverse Polish notation and parsed into a unique sequence form. The present invention can perform similar processing on the query tree in knowledge inference. However, the operations of SPARQL queries and regular matching are different, so it is necessary to redefine the atomic operations and their parameters. The atomic operations and usage scopes of NRE are as Figure 2 shown. Among them, Find_Pos (finding positive rule matching) and Find_Neg (finding negative rule matching) can be mapped to mapping (i.e., knowledge query pattern matching) and unbound() (i.e., filtering unmatched variables) operations in SPARQL; since RDF is unordered, and_order (ordered conjunction) and and_unorder (unordered conjunction) can be combined into a join (variable connection) operation, and or (disjunction) can be mapped to union (pattern merging) or optional (optional pattern) operations. In addition, the parameters of the operations also need to be redefined. The operation parameters of the original NRE are mainly the position information of the text, and the present invention intends to use the graph structure relationship to replace these parameters. After the atomic operations and parameters are defined, the parsing and generalization learning of the layout can be carried out using a sequence annotation model similar to the original NRE.
[0152] Step 2: NRE Inference Based on Multi-Space Incremental Training; To address the applicability issue of the improved NRE to dynamically updated semantic data streams, the present invention uses a scheme similar to PU-KALE, i.e., adopts a multi-space approach for incremental training. The model thereof will not be elaborated here. Considering that NRE needs to train both the query tree layout and atomic operations simultaneously, there may be a problem of too long training time. The present invention intends to use three schemes to attempt to achieve real-time knowledge inference on NRE: In addition to inferring in the PU-NRE manner similar to PU-KALE, it also adopts the method of offline rule generalization + online operation training or offline learning + online query for inference.
[0153] In specific implementation, the method proposed by the technical solution of the present invention can be automatically run by those skilled in the art using computer software technology. The system device for implementing the method, such as a computer-readable storage medium storing the corresponding computer program of the technical solution of the present invention and a computer device including running the corresponding computer program, should also be within the protection scope of the present invention.
[0154] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the technical field to which the present invention pertains can make various modifications or supplements to the described specific embodiments or use similar ways to substitute, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. A dynamic knowledge reasoning method based on rule embedding and context awareness, characterized in that: First, realize the joint embedding representation of rules and facts, and then set different dynamic reasoning processes for different semantic data flow reasoning requirements and scenarios, including the dynamic reasoning of rule embedding based on parallel semantic space reasoning, the dynamic reasoning of rule embedding based on context awareness, and the dynamic reasoning based on rule generalization learning; The dynamic reasoning of rule embedding based on parallel semantic space reasoning is for the reasoning scenario where the head and tail entity ratio is 1:1, and is processed in combination with the semantic data flow processing platform and knowledge representation learning. The implementation process includes the following steps: Step 1, construct the overall training framework of the dynamic joint embedding model; Step 2, generate a sub-embedding space method based on the combination of semantics and structure association, covering the semantic and structure association space within 2-hop range; Step 3, select the sub-embedding space based on the triple index, and select the updated self-embedding space based on the newly added triples and their associated rules; Step 4, formulate the objective function and scheme based on the margin model; Step 5, realize the extension and real-time reasoning of the dynamic joint embedding model based on the C-SPARQL engine; The dynamic reasoning of rule embedding based on context awareness is for the reasoning scenario where the head and tail entity ratio is one-to-many or many-to-many, and conducts knowledge reasoning on the multi-semantic space fusion representation based on context awareness. The implementation process includes the following steps: Step 1, construct the entity and its type context subgraph covering a 2-hop distance; Step 2, construct the relationship context subgraph covering the same-direction paths, parent relationship types, and rule associations within a 2-hop distance; Step 3, multi-vector integration and rule joint embedding; Step 4, perform incremental update based on the context subgraph, and dynamically maintain the cumulative vector offset. When the cumulative offset exceeds the limit value, perform global embedding space update; The dynamic reasoning based on rule generalization learning is for the reasoning scenario with weak rule constraints or incomplete rules. Rule learning is carried out based on rule examples to generalize rules and expand the reasoning results. The implementation process includes conducting rule generalization research on the basis of the real-time reasoning of knowledge representation learning, and combining with the semantic data flow processing platform to realize the uncertain real-time knowledge reasoning under weak rule constraints.
2. The dynamic knowledge reasoning method based on rule embedding and context awareness according to claim 1, characterized in that: The dynamic reasoning of rule generalization learning includes the following steps: Step 1, a generalization learning model for OWL rules; Step 2, NRE reasoning based on multi-space incremental training.
Citation Information
Patent Citations
Knowledge graph link prediction model, method and device fusing context semantics
CN113535972A
Multi-layered knowledge base system and processing method thereof
US20210192372A1