Industrial knowledge graph reasoning method fusing ontology and neighborhood semantic information
Through the industrial knowledge graph inference method that integrates ontology and neighborhood semantic information, the problem of unsatisfactory automatic completion of industrial knowledge graphs is solved, and more accurate entity and relationship embedding representations are achieved, which improves the inference and completion capabilities of industrial knowledge graphs.
Patent Information
- Application Number
- CN202510071407.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
AI Technical Summary
The industrial knowledge graph has not been effective in automatic completion, and there are problems of sparse data and under-explored entity relationships, which is difficult to meet the needs of rapid evolution in industrial scenarios.
Using an industrial knowledge graph inference method that integrates ontology and neighborhood semantic information, an industrial knowledge graph inference model is constructed through cross-view aggregation model and neighborhood information enhancement model, and a joint training is performed using pre-trained entity and relationship embedding representations to improve the accuracy of entity and relationship embedding.
By integrating ontology and neighborhood semantic information, an embedded representation of entity relationships that encapsulate richer information can be obtained, which improves the accuracy of industrial knowledge graph inference and completion, and adapts to the complexity and rapid changes of industrial scenarios.
Smart Images

Figure CN119990316A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of knowledge graph reasoning technology, and specifically to an industrial knowledge graph reasoning method that integrates ontology and neighborhood semantic information. Background Art
[0002] Knowledge graph is a structured knowledge representation method that aims to organize the rich and diverse knowledge of human beings into a structured form so that computers can better understand and process this knowledge. It presents various concepts in the real world and the connections between them by organizing entities, attributes and relationships into a large-scale knowledge graph network. With the continuous release of large-scale general knowledge graphs and specific domain knowledge graphs, as well as the continuous development of intelligent applications of information services, knowledge graphs have been widely used in the fields of intelligent question answering, semantic search, personalized recommendation, reasoning and decision-making. According to the different knowledge fields and scopes, knowledge graphs can be divided into general knowledge graphs and domain knowledge graphs. General knowledge graphs cover a wide range of knowledge and usually contain a large amount of common sense knowledge in the real world; while domain knowledge graphs are oriented to a specific field and have higher requirements for the depth and accuracy of knowledge.
[0003] As industrial production gradually moves towards informatization and intelligence, the automation level of key industrial production equipment, as the driving force of production, is increasing. However, with the explosive growth of newly generated operation and maintenance records and maintenance data during the operation of these production equipment, the challenges faced by knowledge maintenance are becoming more and more severe. Because these operation and maintenance work orders are often scattered in different databases, electronic files and paper documents, and the degree of structuring is low, resulting in low efficiency of information retrieval and utilization. In addition, traditional expert systems rely on manual acquisition and maintenance of knowledge, which is not only costly, but also difficult to cope with frequently updated fault diagnosis information, and cannot meet the needs of large-scale and rapidly evolving industrial scenarios. Faced with these challenges, industrial knowledge graphs, as an emerging technical means, have shown great potential. Compared with traditional expert systems, industrial knowledge graphs can standardize the storage of expert knowledge and other unstructured data, and have more flexible forms of expression and better scalability. By building a framework based on knowledge graphs, scattered operation and maintenance data can be effectively integrated to achieve structured storage and intelligent analysis of data.
[0004] Taking the situation in the steel industry as an example, the steel production line has the characteristics of complex production process and harsh environmental conditions. Some equipment is under high load and harsh conditions for a long time, resulting in the steel production process being frequently affected by equipment failure. Equipment failure may cause the production line to stop working temporarily, affecting production efficiency and product quality. In severe cases, it may cause safety accidents and even threaten the lives of workers, resulting in significant economic losses. In order to reduce the losses caused by equipment failure and improve the stability and reliability of the steel production line, it is necessary to improve the accuracy and efficiency of equipment fault diagnosis to ensure that the production line can quickly recover from the fault state. Traditional steel production line fault diagnosis methods mainly rely on the knowledge and experience of experts in the field, which places high demands on the professional quality and experience accumulation of technicians. In addition, the process of manual detailed analysis and diagnosis is often time-consuming, resulting in long-term fault shutdown of the production line, which is difficult to adapt to the urgent need for rapid response of modern steel production lines. In the long-term operation and maintenance process, the steel production line generates a large number of documents such as fault reports and maintenance logs. These documents contain rich equipment operation and maintenance knowledge, which are of great value for improving the effect of fault diagnosis and prediction. Constructing a steel fault diagnosis knowledge graph based on this knowledge is of great significance for solving problems such as knowledge management and reuse difficulties faced by industrial scenarios and improving the intelligence level of production lines.
[0005] However, currently, knowledge graphs are usually constructed manually or semi-automatically, so there are inevitably problems of insufficient completeness and data sparsity. In vertical fields such as industry, on the one hand, the application scenarios in these fields are often complex and changeable, and the available corpus resources are relatively scarce. On the other hand, the specialized characteristics of domain knowledge graphs require a deep understanding of the knowledge in the field when constructing the graph, and the ability to accurately extract and describe the complex relationships between entities. This makes the incompleteness and sparsity of domain knowledge graphs more serious. That is, there are situations where the implicit relationships between entities are not fully mined. These implicit relationships contain rich and in-depth semantic information, which is of great significance for improving the effectiveness of knowledge-driven tasks. However, the problem of incomplete knowledge graphs has always restricted its application effect.
[0006] At present, the completion of knowledge graphs relies on knowledge graph reasoning methods. However, due to the above-mentioned characteristics of industrial knowledge graphs, the existing knowledge graph reasoning models and methods are not ideal for automatic completion of industrial knowledge graphs and need to be improved.
[0007] The knowledge graph completion task can be abstracted as a triple prediction task, that is, predicting the missing parts in the triple, which is specifically divided into head entity prediction, tail entity prediction and relationship prediction. The present invention only considers the head and tail entity prediction problems of the data layer instance triples, that is, for the partially unknown instance triples (h, r, ?) or (?, r, t), predict the instance corresponding to ?. In the knowledge graph reasoning model, a scoring function f(h, r, t) based on embedding representation is usually used to evaluate the rationality of the triple (h, r, t). Through model training, the score of the real triple is higher than that of the non-real triple, thereby completing the triple prediction task.
[0008] The core of knowledge graph completion technology is to obtain more accurate entity and relationship embedding representations, which not only helps to fill the data gaps in the knowledge graph, but also provides a more solid foundation for downstream tasks such as reasoning and decision-making. Compared with general knowledge graphs, industrial knowledge graphs have a more complete and rigorous ontology design. As the top-level design of domain knowledge graphs, ontology mines domain knowledge patterns and models them, which contains entity corresponding concepts and hierarchical relationships between concepts, and plays a guiding role in the entities in the instance graph. Therefore, for industrial knowledge graphs with sparsity problems, the cross-view connection between concepts and instance entities in the ontology can determine the approximate position of sparse nodes in the embedding space, which is actually a kind of information supplement. In addition, the complex and changeable scenes of industrial knowledge graphs lead to complex entity associations. Some entities have rich neighborhood structures, and there are attribute triples containing rich semantic information in these entity neighborhoods, which can help obtain more accurate entity embedding representations. In view of this, the present invention proposes a knowledge graph reasoning method that integrates ontology and neighborhood semantic information to obtain entity relationship embedding representations that encapsulate richer information, thereby improving the accuracy of industrial knowledge graph reasoning and completion. Summary of the invention
[0009] The present invention is made to solve the above problems, and aims to provide a method that can more accurately reason about industrial knowledge graphs, so as to achieve better completion effects. The present invention adopts the following technical solutions:
[0010] The present invention provides an industrial knowledge graph reasoning method that integrates ontology and neighborhood semantic information. The method has the following technical features and includes the following steps: step S1, obtaining an industrial knowledge graph derived from actual production, and performing data preprocessing operations on the industrial knowledge graph to obtain a data set; step S2, dividing the data set into a training set and a test set; step S3, inputting the data of the training set into a knowledge graph embedding model for pre-training to obtain pre-trained entity embedding representations and relationship embedding representations; step S4, constructing an industrial knowledge graph reasoning model based on a cross-view aggregation model and a neighborhood information enhancement model, and using the data of the training set to jointly train the cross-view aggregation model and the neighborhood information enhancement model, wherein the industrial knowledge graph reasoning model uses the entity embedding representation and the relationship embedding representation obtained in step S3 as initial vectors; step S5, using the trained industrial graph reasoning model to output the reasoning result.
[0011] The industrial knowledge graph reasoning method that integrates ontology and neighborhood semantic information provided by the present invention may also have such a technical feature, wherein the industrial knowledge graph dataset includes an ontology view knowledge graph and an instance view knowledge graph, the ontology view knowledge graph is composed of ontology triples, the ontology triples contain ontology concepts and their relationships, the instance view knowledge graph is composed of instance triples, the instance triples contain instance entities and their relationships, and the instance entities have a corresponding relationship with the ontology concepts. In step S1, for the industrial knowledge graph dataset, the instance entities that are missing the ontology concepts are completed with predetermined placeholder concepts, and then the set of ontology concepts, the set of instance entities and the set of corresponding relationships are extracted from the ontology view knowledge graph and the instance view knowledge graph respectively to form the dataset.
[0012] The industrial knowledge graph reasoning method that integrates ontology and neighborhood semantic information provided by the present invention may also have such a technical feature, wherein in step S3, the knowledge graph embedding model is a TransE model, and the entity embedding representation and the relationship embedding representation after pre-training of the instance view knowledge graph and the ontology view knowledge graph are obtained based on the score function f(h,r,t)=||h+rt||2 of the TransE model, wherein h, r, t are the head entity, relationship, and tail entity of the ontology triplet or the instance triplet, respectively.
[0013] The industrial knowledge graph reasoning method for integrating ontology and neighborhood semantic information provided by the present invention may also have such a technical feature, wherein step S4 includes the following sub-steps: step S4-1, initializing the parameters of the industrial knowledge graph reasoning model; step S4-2, for the initial entity embedding representation and the initial relationship embedding representation obtained in step S3, mapping the information of the entity embedding space to the ontology embedding space, so that the embedding dimensions are aligned; step S4-3, inputting the aligned initial entity embedding representation and the initial relationship embedding representation into the cross-view aggregation model, which outputs an entity embedding vector representation and a concept embedding vector representation; step S4-4, mapping the entity embedding vector representation and the concept embedding vector representation. The neighborhood information enhancement model is input for training to obtain an entity embedding vector representation of the aggregated entity neighborhood information and a corresponding relationship embedding vector representation; step S4-5, based on the entity embedding vector representation of the aggregated entity neighborhood information and the corresponding relationship embedding vector, the loss is calculated, and the cross-view aggregation model and the neighborhood information enhancement model are trained by minimizing the loss function to obtain an entity embedding vector representation and a relationship embedding vector representation of the fused ontology information and entity neighborhood information. In step S5, based on the entity embedding vector representation and the relationship embedding vector representation of the fused ontology information and entity neighborhood information, the scores of all corresponding candidate triples are calculated for the data missing position, and the triple with the highest score is obtained as the completion result.
[0014] The industrial knowledge graph reasoning method for integrating ontology and neighborhood semantic information provided by the present invention may also have such a technical feature, wherein in step S4-2, the information of the entity embedding space is mapped to the ontology embedding space using a nonlinear variation function:
[0015] e c =f T (e)
[0016] f T (e) = σ(W0·e+b)
[0017] In the formula, e c Represents the embedding vector of entity e after nonlinear transformation, and its embedding dimension σ(·) represents the nonlinear function LeakyReLU, W0 is the weight matrix for linear change, b is the bias, and in step S4-3, the total loss function of the cross-view aggregation model is:
[0018]
[0019] Where S represents the set of positive associations between entities and concepts, and (e,c′) represents the negative samples after the concepts are randomly replaced by the positive associations between entities and concepts.
[0020] The industrial knowledge graph reasoning method for integrating ontology and neighborhood semantic information provided by the present invention may also have such a technical feature, wherein step S4-4 includes the following sub-steps: step S4-4-1, using the entity embedding vector representation and the relationship embedding vector representation after training the cross-view aggregation model as the initial vector of the neighborhood information enhancement model; step S4-4-2, for the entity in the instance view knowledge graph, concatenating and linearly transforming the entities and relationships in the neighborhood triplet with the entity as the head entity to obtain the embedding vector representation of the neighborhood triplet; step S4-4-3, calculating the embedding vector of the neighborhood triplet. Attention value represented by the representation, and based on the attention value, an entity embedding vector representation of preliminary aggregated neighborhood information is obtained; step S4-4-4, linearly changing the initial entity embedding representation to make its dimension consistent with the entity embedding vector representation after preliminary aggregated neighborhood information; step S4-4-5, adding the initial entity embedding representation after linear change to the entity embedding vector representation after preliminary aggregated neighborhood information to obtain the entity embedding vector representation after aggregated neighborhood information; step S4-4-6, linearly changing the initial relationship embedding representation to make its dimension consistent with the dimension of the entity embedding vector representation to obtain the relationship embedding vector representation.
[0021] The industrial knowledge graph reasoning method that integrates ontology and neighborhood semantic information provided by the present invention may also have such a technical feature, wherein, in step S4-4-3, the embedding vector representation of the neighborhood triplet is linearly changed and the absolute attention value of the neighborhood triplet is calculated through a nonlinear function, and then the absolute attention value of the neighborhood triplet is activated to obtain the relative attention value of each of the neighborhood triples, and the entity embedding vector representation of the preliminary aggregated neighborhood information is obtained based on the relative attention value.
[0022] The industrial knowledge graph reasoning method for integrating ontology and neighborhood semantic information provided by the present invention may also have such a technical feature, wherein in step S4-4-2, the formula for splicing and linear change is:
[0023]
[0024] In the formula, c ijk Represents e i Neighborhood triplet of the head entity The embedding vector representation of , W1 is the linear change matrix, Represent the entities h after cross-view aggregation operation i ,t j and the relationship k The embedding representation of || represents the concatenation operation of the vector. In step S4-4-3, the calculation formula of the absolute attention value is:
[0025] b ijk =σ(W2c ijk )
[0026] Where W2 is the weight matrix and σ(·) is the LeakyReLU nonlinear function. The calculation formula of the relative attention value is:
[0027]
[0028] In the formula, Represents entity e i The set of neighborhood entities, Represents the connection entity e i and e n The entity embedding vector of the preliminary aggregated neighborhood information is expressed as:
[0029]
[0030] In step S4-4-5, the entity embedding vector after aggregating neighborhood information is expressed as:
[0031]
[0032] In the formula, It is entity e i The embedding representation after the cross-view aggregation operation. In step S4-4-6, the relationship embedding vector is represented as:
[0033]
[0034] Where W4 is the linear transformation matrix, It is a relationship i Relation embedding vector representation after cross-view aggregation operation.
[0035] The industrial knowledge graph reasoning method for integrating ontology and neighborhood semantic information provided by the present invention may also have such a technical feature, wherein, in step S4-4-3, M independent attention mechanisms are used to calculate the attention scores of the neighborhood triples, and the entity embedding vector of the preliminary aggregated neighborhood information is expressed as:
[0036]
[0037] Where M is the number of independent attention mechanisms, and They are the relative attention coefficient and neighborhood triplet embedding representation under the m-th attention mechanism, respectively.
[0038] The industrial knowledge graph reasoning method for integrating ontology and neighborhood semantic information provided by the present invention may also have such a technical feature, wherein in step S4-5, the loss function is a hinge loss function:
[0039]
[0040] In the formula, γ NIE >0 is the marginal distance hyperparameter, (h I ,r I ,t I ) represents a positive triple, (h′ I ,r′ I ,t′ I ) represents the negative triplet generated by randomly replacing the head entity or the tail entity with the positive triplet, f NIE (h I ,r I ,t I ) is the translation score function of TransE:
[0041]
[0042] Functions and Effects of the Invention
[0043] According to the industrial knowledge graph reasoning method for integrating ontology and neighborhood semantic information provided by the present invention, the steps of preprocessing the industrial knowledge graph data set, obtaining the initial vector through the knowledge graph embedding model, constructing the industrial knowledge graph reasoning model and training it, and using the trained model to input the completion result. As mentioned above, the core of the knowledge graph completion technology is to obtain more accurate entity and relationship embedding representations. Since the industrial knowledge graph has a more complete and strict ontology design compared to the general knowledge graph, the ontology, as the top-level design of the domain knowledge graph, mines the domain knowledge model and models it, which contains the hierarchical relationship between the entity corresponding concept and the concept, and plays a guiding role in the entity in the instance graph. Therefore, for the industrial knowledge graph with sparsity problems, the cross-view connection between the concept and the instance entity in the ontology can determine the approximate position of the sparse node in the embedding space, which is actually a kind of information supplement. In addition, the complex and changeable scenes of the industrial knowledge graph lead to complex entity associations. Some entities have rich neighborhood structures, and there are attribute triples containing rich semantic information in these entity neighborhoods, which can help obtain more accurate entity embedding representations. Therefore, the reasoning method of the present invention can obtain an entity relationship embedding representation that encapsulates richer information, thereby improving the accuracy of industrial knowledge graph reasoning and completion. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1is a flow chart of an industrial knowledge graph reasoning method integrating ontology and neighborhood semantic information in an embodiment of the present invention;
[0045] Figure 2 It is a schematic diagram of the structure of the instance view knowledge graph and the ontology view knowledge graph in an embodiment of the present invention;
[0046] Figure 3 is a schematic diagram of the principle of a cross-view aggregation model in an embodiment of the present invention;
[0047] Figure 4 is a flow chart of step S4 in an embodiment of the present invention;
[0048] Figure 5 Schematic diagram of the principle of the neighborhood information enhancement model in an embodiment of the present invention;
[0049] Figure 6 is a flow chart of step S4-4 in an embodiment of the present invention. DETAILED DESCRIPTION
[0050] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the industrial knowledge graph reasoning method that integrates ontology and neighborhood semantic information of the present invention is specifically described below in combination with embodiments and drawings.
[0051] <Example>
[0052] Figure 1 It is a flowchart of the industrial knowledge graph reasoning method that integrates ontology and neighborhood semantic information in this embodiment.
[0053] like Figure 1 As shown, the method comprises the following steps:
[0054] Step S1, obtain the industrial knowledge graph from actual production, and perform data preprocessing operations on the industrial knowledge graph data to obtain a data set.
[0055] Step S2, dividing the preprocessed data set into a training set and a test set.
[0056] Step S3: input the training set data into the knowledge graph embedding model for pre-training to obtain the pre-trained entity embedding representation and relationship embedding representation as the initial vector of the subsequent model.
[0057] Step S4: construct an industrial knowledge graph reasoning model based on the cross-view aggregation model and the neighborhood information enhancement model, and use the training set to jointly train the cross-view aggregation model and the neighborhood information enhancement model.
[0058] Step S5: Use the trained industrial graph inference model to output the completed result.
[0059] For ease of explanation, the following will first briefly describe the composition of the industrial knowledge graph, and then specifically describe the method of this embodiment in combination with its composition.
[0060] The industrial knowledge graph includes two parts: the instance-view knowledge graph and the ontology-view knowledge graph. In this embodiment, the method is used to complete the industrial knowledge graph constructed from the actual equipment fault diagnosis data of the steel industry.
[0061] Figure 2 It is a structural diagram of the instance view knowledge graph and the ontology view knowledge graph in this embodiment. Figure 2 The part above the horizontal dotted line is the ontology view knowledge graph, and each rectangular box represents a concept in the ontology. The part below the horizontal dotted line is the instance view knowledge graph, and each oval box represents an instance entity.
[0062] like Figure 2 As shown in the figure, the ontology view knowledge graph contains the abstract concept information of instance data, represented as G O ={C,R O}, where C represents the concept set in the ontology view knowledge graph, R o Represents a set of relations. The instance view knowledge graph consists of instance triples, denoted by G I ={E,R I}, where E represents the entity set, R I Represents a set of relations.
[0063] The knowledge in both the instance view knowledge graph and the ontology view knowledge graph exists in the form of triples. o ,r o ,t o )∈G O and (h I ,r I ,t I )∈G I Represent concept triples and instance triples respectively, where h o ,t o ∈C,h I ,t I ∈E,r o ∈R O ,r I ∈R I . Bold h I ,r I ,t I Represents the head entity h in the instance view knowledge graphI , relationship I and tail entity t I The embedding vector of . Similarly, the bold h o ,r o ,t o Indicates the corresponding concepts and relations in the ontology view knowledge graph. In addition to the above symbols, there is a corresponding relationship between the concepts in the ontology view knowledge graph and the entities in the instance view knowledge graph, that is, for any entity e∈E in the instance view knowledge graph, there is a corresponding concept c∈C in the ontology view knowledge graph, indicating that e is an instance of c. (e,c)∈S is used to represent the cross-view connection between entities and concepts, where S represents the associated set of entities and concepts. Similarly, the bold e and c represent the embedding vectors corresponding to the entity and concept, respectively. Figure 2 The dashed arrows across the instance view and the ontology view represent cross-view connections between entities and concepts. In this embodiment, it is assumed that the connection type is unique and therefore is no longer represented separately.
[0064] Concepts in an ontology provide a high-level summary of their instances, which is very helpful when the corresponding instances are sparsely distributed in the knowledge graph.
[0065] Step S1, obtain the industrial knowledge graph from actual production, and perform data preprocessing operations on the industrial knowledge graph data to obtain a data set.
[0066] Among them, for the obtained industrial knowledge graph, firstly, the instance entities that are missing the ontology concepts are completed with the predetermined placeholder concepts, such as the concept of "unknown". Then, the instance entity set E and the relationship set R are extracted from the ontology view knowledge graph and the instance view knowledge graph respectively. I , R O ,The ontology concept set C, constitutes the data set for subsequent training.
[0067] Step S2, dividing the preprocessed data set into a training set and a test set.
[0068] In this embodiment, 70% of the data set after the S1 preprocessing operation is used as a training set and 30% is used as a test set. In an alternative solution, the data set can also be divided according to other proportions.
[0069] Step S3, input the training set data into the knowledge graph embedding model for pre-training, and obtain the pre-trained entity embedding representation and relationship embedding representation as the initial vector of the subsequent model. For the convenience of description, the pre-trained entity embedding representation and relationship embedding representation are respectively recorded as the initial entity embedding representation and the initial relationship embedding representation.
[0070] To realize the utilization of the ontology view knowledge graph, it is first necessary to embed the concepts and relationships therein. In this embodiment, the TransE model based on the principle of translation invariance is used to pre-train the ontology view knowledge graph and the instance view knowledge graph respectively, with the aim of learning the original structural information of each view in the two embedding spaces corresponding to the ontology view knowledge graph and the instance view knowledge graph. Among them, since the semantic levels of entities and relationships in the instance view knowledge graph and the ontology view knowledge graph are different, each view should be processed separately instead of combining the two views into the same embedding space, so as to retain the semantic information of each view to the greatest extent.
[0071] In the TransE model, for a given triple (h, r, t), assuming that the condition h+r≈t holds, the scoring function is as follows: the lower the score, the higher the degree of fit of the triple.
[0072] f(h,r,t)=∥h+rt∥2
[0073] Where ∥·∥2 means calculating the L2 norm distance.
[0074] Based on the above scoring function, the pre-trained entity relationship embedding representations of the instance view knowledge graph and the ontology view knowledge graph are obtained respectively.
[0075] Step S4: construct an industrial knowledge graph reasoning model based on the cross-view aggregation model and the neighborhood information enhancement model, and use the training set to jointly train the cross-view aggregation model and the neighborhood information enhancement model.
[0076] Among them, the inventor named the industrial knowledge graph reasoning model ION-GAT.
[0077] Figure 4 It is a flow chart of step S4 in this embodiment.
[0078] like Figure 4 As shown, in step S4, the training of the model includes the following sub-steps:
[0079] Step S4-1, initialize the model parameters of the industrial knowledge graph reasoning model, that is, initialize the number of neurons and network connection weights of the graph neural network in ION-GAT.
[0080] Step S4-2, for the entity embedding representation and relationship embedding representation obtained in step S3, map the information of the entity embedding space to the ontology embedding space so that their embedding dimensions are aligned.
[0081] Among them, after obtaining the entity relationship vector representation of the pre-trained instance view and ontology view in their respective spaces based on the TransE model, before performing the cross-view aggregation operation, it is necessary to map the information of the entity embedding space to the ontology embedding space to align their embedding dimensions. Therefore, we introduce the following nonlinear transformation function f for the entity e∈E in the instance view T (·).
[0082] e c =f T (e)
[0083] f T (e) = σ(W0·e+b)
[0084] In the formula, e c Represents the embedding vector of entity e after nonlinear transformation, and its embedding dimension σ(·) represents the nonlinear function LeakyReLU, W0 is the weight matrix used for linear transformation, and b is the bias. This transformation embeds the two views into the same space, thereby unifying the embedding dimensions of the entity e∈E and the concept c∈C.
[0085] In step S4-3, the aligned entity embedding representation and relationship embedding representation are input into the cross-view aggregation model for training to obtain the entity embedding vector representation and concept embedding vector representation that integrate the ontology information.
[0086] The goal of the cross-view aggregation model is to establish the association between the entity embedding space and the ontology embedding space based on the cross-view connection of entities and concepts in the industrial knowledge graph, and use the ontology to guide the aggregation of entities. To achieve this goal, any instance e∈E is forced to be close to its corresponding concept c∈C on the basis that the entity and the concept have been embedded in the same space.
[0087] Figure 3 Schematic diagram of the principle of the cross-view aggregation model in this embodiment.
[0088] like Figure 3 As shown in Figure 1, multiple entities corresponding to the same concept are clustered around the concept. To achieve this goal, for a given set of entity and concept associations (e, c) ∈ S, the classification loss is the distance between the embedding vectors of e and c and the boundary distance γ CG′ The difference between , and the classification loss is defined as:
[0089]
[0090] Where, [x] + is the positive part of the input x, i.e. [x] + = max{x,0}. This penalizes the embedding of entity e to be centered around the embedding of the corresponding concept c. CG'For situations outside the radius neighborhood, the cross-view aggregation model has a strong clustering effect, making the entity embedding eventually close to the concept embedding corresponding to the entity embedding.
[0091] Based on the cross-view alignment function and the classification loss function, the total loss of the cross-view aggregation model is as follows. The total loss function aims to maximize the difference in the aggregation scores between positive samples and negative samples to enhance the clustering effect.
[0092]
[0093] In the formula, S represents the set of positive associations between entities and concepts, and (e,c′) represents the negative samples after the concepts are randomly replaced by the positive associations between entities and concepts. The cross-view aggregation model optimizes its graph neural network according to the changes in the calculated loss to train the model.
[0094] For a given entity and concept association (e,c)∈S, its embedding representation is input into the trained cross-view aggregation model, and the model output is the embedding vector representation e of entity e CG The embedding vector representation of concept c is c CG Specifically, for the triples (h I ,r I ,t I )∈G I , the entity h output by the cross-view aggregation model I and t I The embedding representation of and Relationship I The embedding representation of
[0095] Step S4-4, input the embedded vector representation of the entity and the embedded vector representation of the concept into the neighborhood information enhancement model for training, and obtain the entity embedded vector representation and the corresponding relationship embedded vector representation that further aggregate the entity neighborhood information.
[0096] Figure 5 Schematic diagram of the principle of the neighborhood information enhancement model in this embodiment.
[0097] Figure 6 It is a flowchart of step S4-4 in this embodiment.
[0098] like Figure 5 and Figure 6 As shown, step S4-4 specifically includes the following sub-steps:
[0099] Step S4-4-1, using the entity embedding vector representation and the relationship embedding vector representation trained by the cross-view aggregation model as the initial vector of the neighborhood information enhancement model.
[0100] Step S4-4-2, for the entities in the instance view knowledge graph, concatenate and linearly transform the entities and relations in the neighborhood triples with the entity as the head entity to obtain the embedded vector representation of the neighborhood triples.
[0101] Specifically, for an entity e in the instance view knowledge graph i ∈E, in order to aggregate e i The neighborhood information of e i Neighborhood triplet of the head entity The entities and relationships in the splicing and linear transformation (e i and h i The key here is to consider the triple as a whole, rather than processing the relations or entities independently, so as to capture more complete neighborhood information. The specific formula is as follows:
[0102]
[0103] In the formula, c ijk Represents e i Neighborhood triplet of the head entity The embedding vector representation of , W1 is the linear change matrix, Represent the entities h after cross-view aggregation operation i ,t j and the relationship k The embedding representation of . || represents the concatenation operation of vectors.
[0104] Step S4-4-3, calculate the attention value of the embedding vector representation of the neighborhood triplet, and obtain the embedding vector representation of the entity that preliminarily aggregates the neighborhood information based on the attention value.
[0105] Among them, the weight matrix W2 is used to ijk Perform linear changes and calculate the neighborhood triplet through the LeakyReLU nonlinear function σ(·) The absolute attention value b ijk .
[0106] b ijk =σ(W2c ijk )
[0107] In the formula, the absolute attention value b ijk Represents a neighborhood triplet For the central entity e i importance.
[0108] Then, for e iThe absolute attention values of all neighborhood triplets are softmaxed to obtain the relative attention value a of each neighborhood triplet ijk :
[0109]
[0110] In the formula, Represents entity e i The set of neighborhood entities, Represents the connection entity e i and e n Based on this, we can use the relative attention value to preliminarily obtain the entity embedding vector representation of the aggregated neighborhood information The updated entity embedding vector representation is obtained by weighted summing of neighborhood triplets and their corresponding relative attention values, which can be expressed as:
[0111]
[0112] In addition, in order to more comprehensively identify the different effects of different neighborhood node features on the central node, a multi-head attention mechanism is introduced. M independent attention mechanisms are used to calculate the attention scores of neighborhood triplets, and multiple sets of independent entity embedding representations of fused neighborhood information are obtained. The embedding representations obtained by multiple sets of attention mechanisms are averaged to obtain the entity embedding representation, which can be expressed as:
[0113]
[0114] Where M is the number of independent attention mechanisms, and They are the relative attention coefficient and neighborhood triplet embedding representation under the m-th attention mechanism, respectively.
[0115] Step S4-4-4, linearly transform the initial entity embedding representation so that its dimension is consistent with the entity embedding vector representation after preliminary aggregation of neighborhood information.
[0116] Step S4-4-5, add the initial entity embedding representation after linear change and the entity embedding vector representation after preliminary aggregation of neighborhood information to obtain the final embedding vector representation of the entity after aggregation of neighborhood information.
[0117] Among them, since the entity will lose its initial embedding information in the process of aggregating neighborhood information, we use the weight matrix W3 to linearly change the initial embedding vector of the entity to make its dimension consistent with the entity embedding vector after preliminary aggregation of neighborhood information, and add the two to get the entity e i The final embedding representation It can be expressed as:
[0118]
[0119] In the formula, It is entity e i The embedded representation after the cross-view aggregation operation is used as the initial embedding representation of the neighborhood information enhancement model.
[0120] Step S4-4-6, linearly change the initial embedding vector representation of the relationship to make its dimension consistent with the dimension of the entity embedding vector representation, and obtain the final relationship embedding vector representation.
[0121] Among them, for the relationship between entities r I ∈R I , as a bridge connecting entities, it mainly plays the role of information transmission in the neighborhood information enhancement model. In addition, when fusing neighborhood information, the relationship has been taken into consideration, and we no longer perform repeated information enhancement operations on the relationship. Therefore, we only linearly change the initial embedding vector of the relationship to make it consistent with the dimension of the entity embedding vector, and obtain the final relationship embedding vector representation
[0122]
[0123] Where W4 is the linear transformation matrix, It is a relationship i The relation embedding vector representation after the cross-view aggregation operation is used as the initial embedding vector representation of the relation for the neighborhood information enhancement model.
[0124] Step S4-5, calculate the loss based on the entity embedding vector representation of the aggregated entity neighborhood information and the corresponding relationship embedding vector, and train the cross-view aggregation model and the neighborhood information enhancement model by minimizing the loss function to obtain the entity embedding vector representation and the relationship embedding vector representation that fuse the ontology information and the entity neighborhood information.
[0125] Among them, the translation score function of TransE is used. For the triplet of instance view (h I ,r I ,t I )∈G I , the embedding vector represents and The corresponding score function of the embedded representation after the neighborhood information enhancement model operation is:
[0126]
[0127] In this embodiment, the cross-view aggregation model and the neighborhood information enhancement model are trained by minimizing the following hinge loss function:
[0128]
[0129] In the formula, γ NIE >0 is the marginal distance hyperparameter, (h I ,r I ,t I ) represents a positive triple, (h′ I ,r′ I ,t′ I ) represents the negative triplet generated by randomly replacing the head entity or the tail entity with the positive triplet. The model optimizes its graph neural network according to the change in loss. After training, the embedded vector representation e that integrates the ontology information and the neighborhood information is obtained. NIE and r NIE .
[0130] Step S5: Use the trained industrial graph inference model to output the completed result.
[0131] Among them, the embedding vector representation e obtained based on the final training NIE and r NIE For any given triple (h, r, t) (i.e., the location where data is missing), according to the formula Calculate the scores of all corresponding candidate triples, and take the triple with the highest score as the completion result.
[0132] Functions and Effects of the Embodiments
[0133] According to the industrial knowledge graph reasoning method that integrates ontology and neighborhood semantic information provided by this embodiment, the steps include preprocessing the industrial knowledge graph data set, obtaining the initial vector through the knowledge graph embedding model, constructing the industrial knowledge graph reasoning model and training it, and using the trained model input to complete the result. The core of the knowledge graph completion technology is to obtain more accurate entity and relationship embedding representations. Since the industrial knowledge graph has a more complete and strict ontology design compared to the general knowledge graph, the ontology, as the top-level design of the domain knowledge graph, mines the domain knowledge model and models it, which contains the hierarchical relationship between the entity corresponding concepts and the concepts, and plays a guiding role in the entities in the instance graph. Therefore, for the industrial knowledge graph with sparsity problems, the cross-view connection between the concept and the instance entity in the ontology can determine the approximate position of the sparse node in the embedding space, which is actually a kind of information supplement. In addition, the complex and changeable scenes of the industrial knowledge graph lead to complex entity associations. Some entities have rich neighborhood structures, and there are attribute triples containing rich semantic information in these entity neighborhoods, which can help obtain more accurate entity embedding representations. In view of this, the reasoning method of this embodiment can obtain an entity relationship embedding representation that encapsulates richer information, thereby improving the accuracy of industrial knowledge graph completion.
[0134] In an embodiment, for the cross-view aggregation model, a total loss function is constructed based on the cross-view alignment function and the classification loss function, which aims to maximize the difference in aggregation scores between positive samples and negative samples, thereby enhancing the clustering effect and making the entity embedding representation closer to its corresponding concept embedding representation, thereby improving the accuracy of completion.
[0135] Furthermore, for the neighborhood information enhancement model, multiple independent attention mechanisms are used to calculate the attention scores of neighborhood triplets. This can more comprehensively identify the different impacts of different neighborhood node characteristics on the central node, thereby achieving a better neighborhood information aggregation effect.
[0136] <Comparative Example>
[0137] This comparative example provides a variety of models and methods in the prior art for comparison with the models and methods provided in the above embodiments.
[0138] Several existing models are: TransE, TransH, RotatE, ConvE, ConvKB, CompGCN, RelEns-DSC. The corresponding method is to use these existing models to complete the industrial knowledge graph.
[0139] In this comparative example, the industrial knowledge graph to be completed is an industrial knowledge graph IFD constructed based on the equipment fault diagnosis data of a steel company. The industrial knowledge graph IFD is artificially constructed based on the knowledge captured on the Internet and the equipment fault report information provided by the enterprise, and has complete ontology and instance data, including common concepts and expert experience. In addition to the basic production lines, equipment, components and their related information, the industrial knowledge graph IFD also contains equipment failures and their operation and maintenance knowledge based on the historical facts of the enterprise's operation and expert experience, including equipment abnormalities, causes of failures, types of failures, scope of impact, inspection methods, solutions, etc. The IFD industrial knowledge graph dataset contains a total of 19 relationships, 781 entities and 2448 triples in the instance graph, and 6 relationships, 85 concepts and 253 triples in the ontology view, as well as 588 cross-view connections.
[0140] For evaluation indicators, MR (Mean Rank), MRR (Mean Reciprocal Rank) and Hits@n, which are commonly used in knowledge graph completion research, are selected as evaluation indicators in this comparative example. MR stands for mean ranking, which replaces the head entity or the tail entity in the test triple with the candidate entity in turn, and calculates the score and the average ranking of the correct entity. The smaller the MR indicator, the better the performance of the model. MRR stands for mean reciprocal ranking. The larger the MRR indicator, the better the performance of the model. Hits@n refers to the proportion of correct entities ranked in the top n among all candidate entities. The larger the Hits@n indicator, the better the performance of the model. In this comparative example, n is selected as 1, 3 and 10.
[0141] The same industrial knowledge graph IFD is completed by the above seven existing models and methods and the ION-GAT model and method of the above embodiment, and the completion effects of the seven existing models and the model of the embodiment are evaluated by the above evaluation indicators. The results are shown in Table 1 below:
[0142] Table 1 Comparison of the completion effects of various models on industrial knowledge graphs
[0143]
[0144]
[0145] As shown in the data in Table 1, compared with seven similar models and methods in the prior art, the ION-GAT model and method of the embodiment achieved optimal values in five evaluation indicators, and each evaluation indicator was significantly better than the corresponding suboptimal value, demonstrating its effectiveness and superiority in completing industrial knowledge graphs.
[0146] The above embodiments are only used to illustrate the specific implementation of the present invention, and the present invention is not limited to the description scope of the above embodiments. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and descriptions are only to illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.
[0147] For example, in the above embodiments and comparative examples, the industrial knowledge graph constructed based on equipment fault diagnosis data of the steel industry is taken as an example for specific description. It can be understood that the method of the present invention is not limited to this, and can also be used for automatic completion of industrial knowledge graphs constructed based on data from other industries.
Claims
1. An industrial knowledge graph reasoning method integrating ontology and neighborhood semantic information, characterized in that: The following steps are involved: Step S1, obtaining an industrial knowledge graph derived from actual production, and performing data preprocessing operations on the industrial knowledge graph to obtain a data set; Step S2, dividing the data set into a training set and a test set; Step S3, inputting the data of the training set into the knowledge graph embedding model for pre-training to obtain pre-trained entity embedding representation and relationship embedding representation; Step S4, constructing an industrial knowledge graph reasoning model based on the cross-view aggregation model and the neighborhood information enhancement model, and jointly training the cross-view aggregation model and the neighborhood information enhancement model using the data of the training set, wherein the industrial knowledge graph reasoning model uses the entity embedding representation and the relationship embedding representation obtained in step S3 as initial vectors; Step S5: Use the trained industrial graph reasoning model to output the reasoning result.
2. The industrial knowledge graph reasoning method integrating ontology and neighborhood semantic information according to claim 1 is characterized by: in, The industrial knowledge graph dataset includes an ontology view knowledge graph and an instance view knowledge graph. The ontology view knowledge graph is composed of ontology triples, which contain ontology concepts and their relationships. The instance view knowledge graph is composed of instance triples, each of which contains instance entities and their relationships, and each of which has a corresponding relationship with the ontology concept. In step S1, for the industrial knowledge graph dataset, a predetermined placeholder concept is used to complete the instance entity that lacks the ontology concept, and then the set of ontology concepts, the set of instance entities and the set of corresponding relationships are extracted from the ontology view knowledge graph and the instance view knowledge graph respectively to form the dataset.
3. The industrial knowledge graph reasoning method integrating ontology and neighborhood semantic information according to claim 2 is characterized by: in, In step S3, the knowledge graph embedding model is a TransE model, and the entity embedding representation and the relationship embedding representation after pre-training of the instance view knowledge graph and the ontology view knowledge graph are obtained based on the score function f(h,r,t)=||h+rt||2 of the TransE model, where h, r,t are the head entity, relationship, and tail entity of the ontology triplet or the instance triplet, respectively.
4. The industrial knowledge graph reasoning method integrating ontology and neighborhood semantic information according to claim 3, Features: Wherein, step S4 includes the following sub-steps: Step S4-1, initializing the parameters of the industrial knowledge graph reasoning model; Step S4-2, for the initial entity embedding representation and the initial relationship embedding representation obtained in step S3, map the information of the entity embedding space to the ontology embedding space so that their embedding dimensions are aligned; Step S4-3, inputting the aligned initial entity embedding representation and the initial relationship embedding representation into the cross-view aggregation model, and the model outputs an entity embedding vector representation and a concept embedding vector representation; Step S4-4, inputting the entity embedding vector representation and the concept embedding vector representation into the neighborhood information enhancement model for training, to obtain the entity embedding vector representation of the aggregated entity neighborhood information and the corresponding relationship embedding vector representation; Step S4-5, calculating the loss based on the entity embedding vector representation of the aggregated entity neighborhood information and the corresponding relationship embedding vector, and training the cross-view aggregation model and the neighborhood information enhancement model by minimizing the loss function to obtain the entity embedding vector representation and the relationship embedding vector representation that fuse the ontology information and the entity neighborhood information, In step S5, the scores of all candidate triples corresponding to the data missing position are calculated based on the entity embedding vector representation and the relationship embedding vector representation that integrate the ontology information and the entity neighborhood information, and the triple with the highest score is obtained as the completion result.
5. The industrial knowledge graph reasoning method integrating ontology and neighborhood semantic information according to claim 4 is characterized in that: in, In step S4-2, the information of the entity embedding space is mapped to the ontology embedding space using a nonlinear transformation function: And c =f T (And) f T (e)=σ(W0·e+b) In the formula, e c Represents the embedding vector of entity e after nonlinear transformation, and its embedding dimension σ(·) represents the nonlinear function LeakyReLU, W0 is the weight matrix used for linear changes, b is the bias, In step S4-3, the total loss function of the cross-view aggregation model is: In the formula, S represents the positive example association set of entities and concepts, (e,c ′ ) represents the negative examples after the concept is randomly replaced by the positive association between the entity and the concept.
6. The industrial knowledge graph reasoning method integrating ontology and neighborhood semantic information according to claim 4, Features: Wherein, step S4-4 includes the following sub-steps: Step S4-4-1, using the entity embedding vector representation and the relationship embedding vector representation trained by the cross-view aggregation model as initial vectors of the neighborhood information enhancement model; Step S4-4-2, for an entity in the instance view knowledge graph, concatenate and linearly transform the entities and relations in the neighborhood triplet with the entity as the head entity to obtain an embedded vector representation of the neighborhood triplet; Step S4-4-3, calculating the attention value of the embedding vector representation of the neighborhood triplet, and obtaining the entity embedding vector representation of the preliminary aggregated neighborhood information based on the attention value; Step S4-4-4, linearly changing the initial entity embedding representation so that its dimension is consistent with the entity embedding vector representation after preliminary aggregation of neighborhood information; Step S4-4-5, adding the initial entity embedding representation after linear change to the entity embedding vector representation after preliminary aggregation of neighborhood information to obtain the entity embedding vector representation after aggregation of neighborhood information; Step S4-4-6, linearly change the initial relationship embedding representation so that its dimension is consistent with the dimension of the entity embedding vector representation, and obtain the relationship embedding vector representation.
7. The industrial knowledge graph reasoning method integrating ontology and neighborhood semantic information according to claim 6 is characterized by: in, In step S4-4-3, a linear change is performed on the embedding vector representation of the neighborhood triplet and the absolute attention value of the neighborhood triplet is calculated by a nonlinear function. Then, an activation operation is performed on the absolute attention value of the neighborhood triplet to obtain the relative attention value of each neighborhood triplet. An entity embedding vector representation of the preliminary aggregated neighborhood information is obtained based on the relative attention value.
8. The industrial knowledge graph reasoning method integrating ontology and neighborhood semantic information according to claim 7 is characterized by: in, In step S4-4-2, the formula for splicing and linear change is: In the formula, c ijk Represents e i Neighborhood triplet of the head entity The embedding vector representation of , W1 is the linear change matrix, Represent the entities h after cross-view aggregation operation i ,t j and the relationship k The embedding representation of || represents the concatenation operation of the vector. In step S4-4-3, the calculation formula of the absolute attention value is: b ijk =σ(W2c ijk ) Where W2 is the weight matrix, σ(·) is the LeakyReLU nonlinear function, The calculation formula of the relative attention value is: In the formula, Represents entity e i The set of neighborhood entities, Represents the connection entity e i and e n The relationship set, The entity embedding vector of the preliminary aggregated neighborhood information is expressed as: In step S4-4-5, the entity embedding vector after aggregating neighborhood information is expressed as: In the formula, It is entity e i The embedded representation after cross-view aggregation operation, In step S4-4-6, the relationship embedding vector is expressed as: Where W4 is the linear transformation matrix, It is a relationship i Relation embedding vector representation after cross-view aggregation operation.
9. The industrial knowledge graph reasoning method integrating ontology and neighborhood semantic information according to claim 8 is characterized by: in, In step S4-4-3, M independent attention mechanisms are used to calculate the attention scores of the neighborhood triples, and the entity embedding vector of the preliminary aggregated neighborhood information is expressed as: Where M is the number of independent attention mechanisms, and They are the relative attention coefficient and neighborhood triplet embedding representation under the m-th attention mechanism, respectively.
10. The industrial knowledge graph reasoning method integrating ontology and neighborhood semantic information according to claim 1 is characterized by: in, In step S4-5, the loss function is a hinge loss function: In the formula, γ NIE >0 is the marginal distance hyperparameter, (h I ,r I ,t I ) represents a positive triple, (h I ′ ,r I ′ ,t I ′ ) represents the negative triplet generated by randomly replacing the head entity or the tail entity with the positive triplet, f NIE (h I ,r I ,t I ) is the translation score function of TransE:
Citation Information
Cited By
Hierarchical temporal ontology for industrial facilities
US12736935B1