Double-view-angle super-relation knowledge graph completion method, system and equipment and storage medium
Through the two-perspective hyperrelational knowledge graph completion method, geometric space representation learning and ontology hierarchical information are used to solve the modeling problem of multiple relationships in the hyperrelational knowledge graph, and more accurate knowledge representation and reasoning are achieved.
Patent Information
- Application Number
- CN202510393672.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art is difficult to effectively model and represent multiple relationships and additional key-value pairs in the hyperrelational knowledge graph, resulting in semantic ambiguity and information loss, and the existing methods lack the expression of semantic information and hierarchical information of the hyperrelational knowledge graph.
The two-view hyperrelational knowledge graph completion method is used to calculate the intersection of box embedding through the representation learning method of geometric space, and combine the pre-trained model to enrich the semantic information of the ontology, establish the semantic connection between the hyperrelational facts and the ontology embedding, and model it using contrast learning minimized loss function.
Effectively modeling the complex structure of the hyperrelational knowledge graph improves the performance of the hyperrelational knowledge graph completion task, solves the problems of semantic ambiguity and information loss, and enhances the richness and accuracy of knowledge expression.
Smart Images

Figure CN120278247A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge graphs, and in particular to a dual-view hyper-relation knowledge graph completion method, system, device, and storage medium. Background Art
[0002] In recent years, knowledge graphs, as a structured form of human knowledge, have attracted extensive research attention in both academia and industry. Knowledge graphs are centered around entities, relations, and semantic descriptions, and structurally represent facts in the real world. Currently, most research on knowledge graphs focuses on triples in the form of binary relations connecting head and tail entities. However, the expression of binary relations is not comprehensive enough to directly represent complex relations involving multiple entities. In actual semantic texts, in addition to binary relations, multi-relation facts involving more than two entities are also prevalent in reality. Using the triple representation method in traditional knowledge graphs oversimplifies the complexity of the data stored in the knowledge graph. A hyper-relation knowledge graph is a knowledge graph used to represent multi-relation facts, that is, each fact not only contains a basic triple but also contains related key-value pairs. Different from traditional knowledge graphs that represent facts as {entity-relation-entity} triples, hyper-relation knowledge graphs allow triples to be associated with additional {relation-entity} pairs, that is, key-value pairs, to convey more complex information. Hyper-relation knowledge graphs are knowledge graphs that can efficiently aggregate complex multi-relation facts, and such hyper-relation data is ubiquitous in knowledge graph datasets. Taking the creative commons website Freebase as an example, more than 30% of the entities involve such hyper-relation facts. The information of these additional key-value pairs can be used to eliminate ambiguities between triples or restrict the validity of facts.
[0003] Knowledge graphs have addressed challenges such as the difficulty of fusing multi-source heterogeneous data and the dynamic evolution of data patterns in the real world, and have been widely applied in semantic understanding, information retrieval, natural language processing, and large model enhancement fields. Hyper-relation knowledge graphs have evolved on the basis of knowledge graphs and simultaneously address the following three obvious limitations in knowledge graphs:
[0004] (1) Limited semantic expression ability: Triples can only express binary relations between two entities and cannot directly represent complex relations involving multiple entities. For example, a multi-relation such as "A completed project D in cooperation with C under the guidance of B" cannot be fully expressed by a single triple.
[0005] (2) Information loss: In actual scenarios, many facts require additional context information for supplementation or qualification. For example, in the fact "A served as the CEO of Company B in 2023", the time information "2023" is a key qualification condition, but traditional triples cannot directly contain the corresponding information.
[0006] (3) Ambiguity problem: The triple representation method may lead to semantic ambiguity. For example, the relationship "A is the founder of B" may have different meanings depending on time, place, or context, but traditional knowledge graphs cannot effectively distinguish these nuances.
[0007] Therefore, hyper-relation knowledge graphs can represent and reason about complex semantic relationships in the real world more comprehensively and accurately. They can not only effectively represent multi-relational and context information but also eliminate semantic ambiguity, enhancing the richness and accuracy of knowledge representation. This ability gives hyper-relation knowledge graphs significant advantages in practical applications such as question answering systems, recommendation systems, and semantic search, enabling them to support more complex reasoning tasks such as temporal reasoning, conditional reasoning, and multi-hop reasoning. Currently, most research focuses on binary relationships composed of triples, while in actual scenarios, relationships are complex and diverse, including one-to-many, many-to-one, and many-to-many relationships. Hyper-relation knowledge graphs expand the expressive power of traditional knowledge graphs and can represent more complex knowledge structures, such as facts with additional attributes (e.g., time, place, condition). However, this expansion also brings many new difficulties and challenges. Based on this, combined with the current research status, the following two challenges are summarized for the representation and reasoning of hyper-relation knowledge graphs:
[0008] 1) Hyper-relation knowledge graphs have a more complex structure compared to traditional knowledge graphs, making it difficult to effectively model hyper-relation facts. As shown in Figure 1(a) is the structure of a knowledge graph, and Figure 1(b) is the structure of a hyper-relation knowledge graph. The additional key-value pairs in the hyper-relation knowledge graph increase the complexity of modeling. The additional key-value pairs may be structured (such as time, numerical value) or unstructured data (text description). How to effectively fuse heterogeneous data, capture the semantic relationships between them, and at the same time achieve effective interaction between the additional key-value pairs and the main triples is one of the challenges at the present stage. Therefore, it is necessary to propose a representation method that can represent multi-relational and achieve effective interaction between the additional key-value pairs and the main triples.
[0009] 2) Semantic diversity of hyper-relation knowledge graphs. By counting the proportion of the number of different semantics corresponding to the same entity in three hyper-relation datasets, it is found that the proportion is more than 83% in the datasets, and the average value of the number of different semantics corresponding to the same entity is greater than three. This phenomenon indicates that the same entity expresses different semantic information in different situations, that is, the semantic diversity in hyper-relation knowledge graphs. For example, in the dataset "(Apple, price increase, 10%)", it is impossible to distinguish whether Apple refers to Apple Inc. or a kind of fruit through the information of the triple, which will cause semantic ambiguity. Therefore, the model should be able to judge the semantics of the entities in hyper-relation facts based on the additional information provided by the additional key-value pairs in the hyper-relation facts and the semantic information of the main triples. Summary of the Invention
[0010] The object of the present invention is to provide a dual - perspective hyper - relational knowledge graph completion method, system, device and storage medium for the above - mentioned problems in the existing technology, effectively modeling the complex structure of the hyper - relational knowledge graph and improving the performance of the hyper - relational knowledge graph completion task.
[0011] To achieve the above object, the present invention has the following technical solutions:
[0012] In a first aspect, a dual - perspective hyper - relational knowledge graph completion method is provided, including:
[0013] Based on the representation learning method in geometric space, for each hyper - relational instance, the head - entity relation and the additional key - value pairs are respectively box - embedded and intersected. The representation of the entity is updated by constraining the expected answer entity to fall inside the box - intersection operation. The expression of the main triple is enhanced through the interaction between the box of the additional key - value pair and the box of the main triple, and the hyper - relational fact representation loss is obtained.
[0014] For the ontology corresponding to the hyper - relational instance, the semantic information of the ontology is enriched through the initialization representation by a pre - trained model, and then the ontology is boxed in the same way to naturally express the hierarchical information between ontologies. By constraining the intersection of the expected prediction answer set and the ontology box, the ontology - constrained instance is realized, and the ontology representation loss is obtained.
[0015] A semantic connection is established between the hyper - relational fact embedding and the corresponding ontology embedding. By constraining the distance between the entity and the corresponding ontology vector, a metric function is defined to measure the distance between the ontology box and the hyper - relational instance box. Contrastive learning is used to minimize the loss function, and the dual - perspective joint - modeling loss is obtained.
[0016] The jointly obtained hyper - relational fact representation loss, ontology representation loss and dual - perspective joint - modeling loss are balanced to complete the hyper - relational knowledge graph completion.
[0017] As a preferred solution, the hyper - relational knowledge graph is \(G=(E,R,F)\), where \(E\) is the set of entities, \(R\) is the set of relations, and \(F\) is the set of hyper - relational facts \((f_1,f_2,\cdots,f\) n ), where \(f\) j \(\in E\times R\times E\times P(R\times E)\), and \(P\) in the formula represents the power set; the hyper - relational fact \(f\) i \(\in F\) is represented as a tuple where \(h,t,v\) i \(\in E\); \(r,k\) i \(\in R\). In the above tuple, \((h,r,t)\) is the main triple, is the additional key - value pair corresponding to the main triple;
[0018] Hyper-relation knowledge graph completion refers to predicting missing elements from hyper-relation facts. Given an incomplete hyper-relation knowledge graph G, the task objectives include: First, predicting missing entities when some entities and relations are known. In hyper-relation facts where h, t, v i ∈ E; r, k i ∈ R, the missing entities that can be predicted are h, r, and v i ; Second, predicting the relation r in the missing main triple or the relation v in the attached key-value pair when entities are known i .
[0019] As a preferred solution, for the representation learning method based on geometric space, the step of taking the intersection of the box embeddings of the head entity relation and the attached key-value pair for each hyper-relation instance is implemented using box embedding (Box Embedding), including:
[0020] Perform node initial boxing in the following way:
[0021] According to the characteristics of the Box, there is: B d (cen, off) = {X ∈ R d | cen(X) - off(X) ≤ x i ≤ cen(X) + off(X), i = 1,... d}, where cen(X) ∈ R d represents the center vector of the Box, and off(X) ∈ R d represents the offset vector of the Box; then initialize each source node x as a set containing a single entity, and model it using a box with the center vector cen(x) and the offset vector off(x). The initial box embedding is (cen(x), 0), where cen(x) ∈ R d is the center vector, and 0 is a d-dimensional all-zero vector; for relation nodes, r = (cen(r), off(r)) ∈ R 2d , where off(r) ≥ 0;
[0022] Perform hyper-relation fact box mapping in the following way:
[0023] Since each relation r ∈ R is embedded as r = (cen(r), off(r)) ∈ R 2d , where off(r) ≥ 0, then given an entity embedding box x, r acts as a spatial transformation of x:
[0024] f r : cen(f r (x)) = cen(x) + τ r
[0025]
[0026] In the formula, is the Hadamard product, and (τ r, , Δ r ) are relationship-specific parameters;
[0027] Using the above box mapping method, the head entity and relationship of the hyper-relation fact are mapped into a specific box space, and each set of additional key-value pairs where k i ∈R, v i ∈E, or is mapped to k i for the box space transformation of v i ;
[0028] The box intersection of the hyper-relation fact is obtained as follows:
[0029] Design a hierarchical attention mechanism and a differential processing of offset contraction, and perform intersection modeling on a set of boxes {p1,... p n} after mapping a hyper-relation fact, where p inter =(cen(p inter ), off(p inter )) The specific processing is as follows:
[0030] In the hierarchical attention mechanism, attention weights are separately assigned to the main triple boxes to ensure their dominance, and secondary attention weights are shared among the additional key-value pair boxes as supplementary constraints. The expression is as follows:
[0031] cen(p inter ) = α main ·cen(p main ) + α aux ·cen(p aux )
[0032]
[0033] Use two independent MLPs (MLP main , MLP aux ) to process the boxes of the main triple and the additional key-value pairs respectively, and restrict the attention weight α main ≥α aux .
[0034] As a preferred solution, the steps of updating the representation of the entity by constraining the expected answer entity to fall inside the box intersection operation, and enhancing the expression of the main triple through the interaction between the boxes of the additional key-value pairs and the main triple boxes to obtain the loss of the hyper-relation fact representation include:
[0035] The offset of the additional key-value pair is constrained and shrunk by the Sigmoid function:
[0036]
[0037] s = DeepSets([off(p main );off(p aux ,1);...;off(p aux ,k)])
[0038] σ(s) = Sigmoid(s)+γ(γ ∈ [0,1])
[0039] In the formula, is the Hadamard product, and DeepSets(·) is a permutation-invariant deep architecture that updates the position and size of the box of the super-relation fact obtained through the above two transformation operations in space:
[0040] p inter = (cen(p inter ), off(p inter ))
[0041] Calculate the distance from the entity to the box as follows: For the box after the joint contraction of the main triple and the additional key-value pair as the query space, check whether the distance from the entity e to the corresponding box conforms to the constraints of the super-relation fact. Given the contracted box q = (cen(q), off(q)) ∈ R 2d and the embedding vectors of the candidate entity v ∈ R d , the distance expression is calculated as follows:
[0042] dist box (v; q) = dist outside (v; q)+α · dist inside (v; q)
[0043] Where:
[0044] dist outside = ||max(v - q max , 0)+max(q min - v, 0)||1
[0045] dist inside = ||cen(q)-clip(v, q min , q max )||1
[0046] q max = cen(q)+off(q)
[0047] qmin = cen(q) - off(q)
[0048] The distance between the external corresponding entity and the nearest corner or edge of the box; the internal distance corresponds to the distance between the center of the box and the corresponding side or corner, or the entity itself if the entity is inside the box.
[0049] The distance within the box is shrunk by using 0 < α < 1.
[0050] Set the training objective to learn entity embeddings and geometric projection and intersection operators. Given a training set of missing and incomplete hyper-relation facts and answers, use negative sampling loss to optimize the distance-based model:
[0051]
[0052] In the formula, γ represents the margin hyperparameter, v + represents the correct answer entity in the input missing hyper-relation fact q, v - is an answer that is not in the missing answers and is an incorrect negative instance, and k is the number of negative instances.
[0053] As a preferred solution, for the ontology corresponding to the hyper-relation instance, enrich the semantic information of the ontology through initial representation by a pre-trained model, then boxify the ontology in the same way to naturally express the hierarchical information between ontologies, and realize ontology-constrained instances by constraining the intersection of the expected predicted answer set and the ontology box. The steps for obtaining the ontology representation loss include:
[0054] Perform ontology box embedding representation in the following way. The hierarchy of the ontology is represented by the position and size of the box in space, and integrate the semantic information from the ontology layer and the relation r through the following transformation function:
[0055] h [CLS] = BERT(O text + ′SEP′ + r text )
[0056] cen(f r (BOX(O))), off(f r (BOX(O))) = f proj (h [CLS] )
[0057] Connect the information in the ontology with the text content of the relation, input a pre-trained model BERT to capture the semantics of the text; obtain the embedding by getting the hidden vector at the [CLS] token, and use a projection function f proj to map the hidden vector h [CLS] ∈ R l to R 2da space, where l is the hidden dimension of the BERT encoder and d is the dimension of the box space; f proj is implemented as a two-layer multi-perceptron layer with a non-linear activation function, and the hidden vector is split into two equal-length vectors as the center and offset of the transformed box; for calculating the rationality of the ontology, the approximate conditional probability is calculated using the following formula:
[0058]
[0059] where represents the box transformation transformed by the relation r O ; for the correct ontology triple it is close to 1, and for the wrong ontology triple it is close to 0;
[0060] The goal of ontology representation training is to correspond the semantics of the ontology with the box. Each ontology concept is represented as a multi-dimensional box, and the corresponding position and size are represented by the center vector and the offset vector. The geometric range of the box directly reflects the semantic coverage range of the concept. Therefore, the loss of ontology representation consists of two parts: the ontology triple rationality loss and the ontology hierarchical inclusion loss;
[0061] The ontology triple rationality loss is as follows:
[0062]
[0063] In the formula, negative instances constructed by using replacement of correct head and tail ontology triples;
[0064] The ontology hierarchical inclusion loss is obtained by constraining the size of the ontology hierarchical box. The range of the parent ontology box contains the range of the child ontology box. Therefore, the loss function is as follows:
[0065]
[0066] In the formula, O sub is the box of the child ontology, which consists of two vectors, the minimum inflection point and the maximum inflection point (O sub,min , O sub,max ); O super is the box of the parent ontology, which consists of two vectors, the minimum inflection point and the maximum inflection point (O super,min , O super,max );
[0067] The ontology representation loss is obtained as:
[0068] L ontolgy = L BCE + L contain .
[0069] As a preferred solution, the steps of establishing the semantic connection between the hyper-relation fact embedding and the corresponding ontology embedding, defining a metric function to measure the distance between the ontology box and the hyper-relation instance box by constraining the distance between the entity and the corresponding ontology vector, and using contrastive learning to minimize the loss function to obtain the dual-view joint modeling loss include:
[0070] Establish the link distance between the entity and the ontology layer:
[0071] f d (e, o) = dist out (e, o) + α · dist in (e, o)
[0072] dist out (e, o) = ||Max(e - μ M , 0) + Max(μ m - e, 0)||1
[0073] dist in (e, o) = ||Cen(Box(c)) - Min(μ m , Max(μ m , e))||1
[0074] In the formula, dist out is the distance between the entity and the nearest corner of the box, dist in corresponds to the distance between the center of the box and the side of the box or the entity itself. 0 < α < 1 is a fixed scalar to balance the distance between the inside and the outside. If the target entity falls outside the ontology box, more penalties are imposed. If the entity is inside the box, the penalty is reduced. The distance from the entity to the box is non-negative;
[0075] The dual-view joint modeling loss is calculated by the following expression:
[0076]
[0077] where γ cross represents a fixed scalar margin.
[0078] As a preferred solution, the calculation expression for the balance of the jointly obtained hyper-relation fact representation loss, ontology representation loss, and dual-view joint modeling loss is as follows:
[0079] L = L fact + λ1L ontolgy + λ2L cross
[0080] In the formula, L fact is the hyper-relation fact representation loss, L ontolgy is the ontology representation loss; Lcross is the loss for dual - perspective joint modeling; λ1, λ2 > 0 are two positive hyperparameters to balance the three losses.
[0081] In a second aspect, a dual - perspective hyper - relational knowledge graph completion system is provided, including:
[0082] A hyper - relational fact representation module, which, based on the representation learning method in geometric space, performs intersection of box embeddings on the head - entity relationship and additional key - value pairs for each hyper - relational instance respectively, updates the entity representation by constraining the expected answer entity to fall within the intersection operation of the boxes, and enhances the expression of the main triple through the interaction between the box of the additional key - value pair and the box of the main triple, obtaining the hyper - relational fact representation loss;
[0083] An ontology representation module, which, for the ontology corresponding to the hyper - relational instance, initializes the representation through a pre - trained model to enrich the semantic information of the ontology, then boxes the ontology in the same way, naturally expressing the hierarchical information between ontologies, and realizes ontology - constrained instances by constraining the intersection of the expected predicted answer set and the ontology box, obtaining the ontology representation loss;
[0084] A dual - perspective joint modeling module, which is used to establish the semantic connection between the hyper - relational fact embedding and the corresponding ontology embedding, defines a metric function to measure the distance between the ontology box and the hyper - relational instance box by constraining the distance between the entity and the corresponding ontology vector, and uses contrastive learning to minimize the loss function, obtaining the dual - perspective joint modeling loss;
[0085] A loss balancing module, which is used to balance the jointly obtained hyper - relational fact representation loss, ontology representation loss and dual - perspective joint modeling loss, and complete the hyper - relational knowledge graph completion.
[0086] In a third aspect, an electronic device is provided, including a processor and a memory. The processor is used to execute a computer program stored in the memory to implement the dual - perspective hyper - relational knowledge graph completion method.
[0087] In a fourth aspect, a computer - readable storage medium is provided. The computer - readable storage medium stores at least one instruction, and when the at least one instruction is executed by a processor, the dual - perspective hyper - relational knowledge graph completion method is implemented.
[0088] Compared with the prior art, the present invention has at least the following beneficial effects:
[0089] Through a geometric space-based representation learning method, for each hyper-relation instance, the head entity relation and additional key-value pairs are respectively subjected to box embedding and intersection. The representation of the entity is updated by constraining the expected answer entity to fall within the box intersection operation. The interaction between the box of the additional key-value pair and the box of the main triple enhances the expression of the main triple, effectively modeling the additional key-value pair to achieve effective interaction with the main triple. Using the geometric space-based representation learning method can explicitly express hierarchical relationships, and the corresponding geometric characteristics can directly reflect semantic relationships. For the ontology corresponding to the hyper-relation instance, the semantic information of the ontology is enriched by initializing the representation through a pre-trained model, and then the ontology is boxed in the same way. The boxed boxes can naturally express the hierarchical information between ontologies. By constraining the intersection of the expected prediction answer set and the ontology box, ontology-constrained instances are realized, and the semantic ambiguity of entities is solved through the semantic information and hierarchical information expressed by the ontology. The goal of the dual-perspective joint modeling is to establish a semantic connection between the hyper-relation fact embedding and the corresponding ontology embedding through cross-view links. By constraining the distance between the entity and the corresponding ontology vector, a metric function is defined to measure the distance between the ontology box and the hyper-relation instance box, and contrastive learning is used to minimize the loss function. Finally, the present invention achieves a balance through the jointly obtained hyper-relation fact representation loss, ontology representation loss, and dual-perspective joint modeling loss. In the face of the limitation that it is difficult to represent multiple relationships in the hyper-relation knowledge graph, the present invention introduces a box embedding method to model the hyper-relation knowledge graph, which can utilize the natural spatial characteristics of the box to express the hyper-relation knowledge graph completion task; in the face of the diversity of semantic information in the hyper-relation knowledge graph, the current method lacks the expression of semantic information and hierarchical information of the hyper-relation knowledge graph. Introducing the ontology layer can distinguish the meanings expressed by entities and at the same time fully discover the parent-child relationships between entities, enhancing the semantic information in the hyper-relation facts. The dual-perspective hyper-relation knowledge graph completion method of the present invention can effectively model the complex structure of the hyper-relation knowledge graph to improve the performance of the hyper-relation knowledge graph completion task. BRIEF DESCRIPTION OF THE DRAWINGS
[0090] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0091] FIG. 1(a) is a schematic diagram of the knowledge graph structure;
[0092] FIG. 1(b) is a schematic diagram of the hyper-relation knowledge graph structure;
[0093] FIG. 2(a) is a schematic diagram of the Box Embedding box mapping operation;
[0094] Figure 2(b) is a schematic diagram of the Box Embedding box intersection operation;
[0095] Figure 3 is a schematic diagram of the corresponding relationship between the fact layer and the ontology layer of the hyper-relation knowledge graph;
[0096] Figure 4 is a schematic diagram of the framework of the dual-view hyper-relation knowledge graph model based on Box Embedding according to an embodiment of the present invention;
[0097] Figure 5 is a flowchart of the dual-view hyper-relation knowledge graph completion method according to an embodiment of the present invention. Detailed implementation manners
[0098] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0099] In a hyper-relation knowledge graph G=(E, R, F), E is a set of entities, R is a set of relations, and F is a set of hyper-relation facts (f1, f2,..., f n ), where f j ∈E×R×E×P(R×E), and P in the formula represents the power set; the hyper-relation fact f i ∈F is represented as a tuple where h, t, v i ∈E; r, k i ∈R. In the above tuple, (h, r, t) is the main triple, is the additional key-value pair corresponding to the main triple;
[0100] Hyper-relation knowledge graph completion means: predicting missing elements from hyper-relation facts; given an incomplete hyper-relation knowledge graph G, the task objectives include: first, predicting missing entities when some entities and relations are known. In the hyper-relation facts where h, t, v i ∈E; r, k i ∈R, the missing entity that can be predicted can be h, r, and v i ; second, predicting the relation r in the missing main triple or the relation v in the additional key-value pair when the entity is known i .
[0101] Please refer to Figure 4 With Figure 5, an embodiment of the present invention proposes a dual - perspective hyper - relation knowledge graph completion method based on Box Embedding, which mainly includes the following steps:
[0102] S1. Based on the representation learning method in geometric space, for each hyper - relation instance, perform box embedding on the head entity relation and the additional key - value pairs respectively and then take the intersection. Update the representation of the entity by constraining the expected answer entity to fall inside the box intersection operation. Enhance the expression of the main triple through the interaction between the box of the additional key - value pair and the box of the main triple, and obtain the hyper - relation fact representation loss;
[0103] S2. For the ontology corresponding to the hyper - relation instance, initialize the representation through a pre - trained model to enrich the semantic information of the ontology, and then box the ontology in the same way to naturally express the hierarchical information between ontologies. Realize ontology - constrained instances by constraining the intersection of the expected prediction answer set and the ontology box, and obtain the ontology representation loss;
[0104] S3. Establish the semantic connection between the hyper - relation fact embedding and the corresponding ontology embedding. Define a metric function to measure the distance between the ontology box and the hyper - relation instance box by constraining the distance between the entity and the corresponding ontology vector, and use contrastive learning to minimize the loss function to obtain the dual - perspective joint modeling loss;
[0105] S4. Balance the jointly obtained hyper - relation fact representation loss, ontology representation loss and dual - perspective joint modeling loss to complete the hyper - relation knowledge graph completion.
[0106] In a possible implementation manner, in the representation learning method in geometric space described in step S1, the step of performing box embedding on the head entity relation and the additional key - value pairs respectively and then taking the intersection for each hyper - relation instance is implemented using box embedding BoxEmbedding, and includes the following steps:
[0107] Perform node initial boxing in the following way:
[0108] According to the characteristics of the Box, we have: B d (cen, off) = {X ∈ R d |cen(X) - off(X) ≤ x i ≤ cen(X)+off(X), i = 1,...d}, where cen(X) ∈ R d represents the center vector of the Box, and off(X) ∈ R d represents the offset vector of the Box; then initialize each source node x as a set containing a single entity, and model it using a box with a center vector cen(x) and an offset vector off(x). The initial box embedding is (cen(x), 0), where cen(x) ∈ R dis the central vector, and 0 is the all-zero vector of dimension d; for a relational node, r = (cen(r), off(r)) ∈ R 2d , where off(r) ≥ 0;
[0109] Perform the hyper-relational fact box mapping in the following way:
[0110] Since each relation r ∈ R is embedded as r = (cen(r), off(r)) ∈ R 2d , where off(r) ≥ 0, then given an entity embedding box x, r serves as a spatial transformation of x:
[0111] f r :cen(f r (x)) = cen(x) + τ r
[0112]
[0113] In the formula, is the Hadamard product, and (τ r, , Δ r ) are relation-specific parameters;
[0114] Since the Box has the characteristic of spatial region by itself, it can naturally enclose the entities in a set. By analogy, in the real-world scenario where there are multiple answers to a question, these answers can fall within the closed region of the Box space. Therefore, using the above box mapping method, the head entity and relation of the hyper-relational fact are mapped into a specific box space, and each set of additional key-value pairs where k i ∈ R, v i ∈ E, or is mapped to the box space transformation of k i for v i , as shown in Fig. 2(a).
[0115] Perform the box intersection of the hyper-relational fact in the following way:
[0116] For the operation of box intersection of the hyper-relational fact, design the following two objectives:
[0117] Principal triple guidance: Since in the hyper-relational knowledge graph, the principal triple occupies a more important position and has a more important impact on prediction, therefore, in the design, the contribution of the principal triple to the intersection should be higher than that of the additional key-value pairs to highlight the importance of the principal triple and avoid treating the relationship between the principal triple and the additional key-value pairs equally.
[0118] Semantic hierarchy: The geometric constraint of the core relation should be stronger than that of the additional key-value pairs.
[0119] Therefore, by designing a hierarchical attention mechanism and differential processing of offset contraction, a set of box embeddings {p1,...p n} obtained after mapping a hyper-relation fact are used to model the intersection, and p inter =(cen(p inter ), off(p inter )) is obtained. The specific processing is as follows:
[0120] In the hierarchical attention mechanism, attention weights are separately assigned to the main triple boxes to ensure their dominance, and secondary attention weights are shared among the additional key-value pair boxes as supplementary constraints. The expression is as follows:
[0121] cen(p inter ) = α main ·cen(p main ) + α aux ·cen(p aux )
[0122]
[0123] Two independent MLPs (MLP main , MLP aux ) are used to process the boxes of the main triple and the additional key-value pairs respectively, and the attention weight α main of the main triple is restricted such that α aux ≥α
[0124] In the differential processing of offset contraction, the offset of the main triple retains more original information and excessive contraction should be avoided. The offset of the additional key-value pairs is constrained and contracted through the Sigmoid function, that is:
[0125]
[0126] s = DeepSets([off(p main ) ; off(p aux , 1) ;... ; off(p aux , k)])
[0127] σ(s) = Sigmoid(s) + γ (γ ∈ [0, 1])
[0128] In the formula, ° is the Hadamard product, and DeepSets(·) is a permutation-invariant deep architecture that The position and size of the box updated for a hyper-relation fact in space are obtained through the above two transformation operations:
[0129] p inter =(cen(p inter ), off(p inter))
[0130] Meanwhile, the above formula can also reflect that the additional information in the hyper-relation is a kind of restriction on the triple information, and adding additional key-value pairs is a restriction on the main triple box space, which can explicitly express the process of adding additional key-value pairs to narrow down the answer set, conforming to the restrictive hypothesis, as shown in Figure 2(b).
[0131] Calculate the distance from the entity to the box in the following way:
[0132] For the box after the joint contraction of the main triple and the additional key-value pair as the query space, whether the distance from the entity e to the corresponding box conforms to the constraints of the hyper-relation fact. Given the contracted box q = (cen(q), off(q)) ∈ R 2d and the embedding vector of the candidate entity v ∈ R d The distance expression is calculated as follows:
[0133] dist box (v; q) = dist outside (v; q) + α · dist inside (v; q)
[0134] Where:
[0135] dist outside = ||max(v - q max , 0) + max(q min - v, 0)||1
[0136] dist inside = ||cen(q) - clip(v, q min , q max )||1
[0137] q max = cen(q) + off(q)
[0138] q min = cen(q) - off(q)
[0139] The distance between the external corresponding entity and the nearest corner or edge of the box. Similarly, the internal distance corresponds to the distance between the center of the box and its side or corner (if the entity is inside the box, it is the entity itself). The key here is to shrink the distance within the box by using 0 < α < 1. This means that as long as the entity vectors are within the box, they are considered "close enough" to the query center, that is, the external distance is 0, and the internal distance is scaled by α. When α = 1, the distance from the entity to the box is the ordinary L1 distance, that is, ||cen(q) - v||1, and at this time there is the traditional TransE model.
[0140] Set the training objective to learn entity embeddings and geometric projection and intersection operators. Given the training set of missing and incomplete hyper-relation facts and answers, use negative sampling loss to optimize the distance-based model:
[0141]
[0142] where γ represents the margin hyperparameter, v + represents the correct answer entity in the input missing hyper-relation fact q, and v - is an answer not in the missing answers and is an incorrect negative instance, and k is the number of negative instances.
[0143] In a possible implementation, the correspondence between the ontology layer and the hyper-relation fact layer is as Figure 3 shown. When statistically analyzing the semantics corresponding to the same entity in the hyper-relation knowledge graph, first match the entity to its ontology and then look up its parent ontology through the current ontology. Eventually, the corresponding relationship of the entity-ontology list is formed. The ontology can reflect the semantic information of an entity. It is found that the same entity may correspond to different semantic information. For example, in the following example, "The price of apples has increased by 10%". At this time, it is impossible to clearly distinguish whether "apples" refers to Apple Inc. or a kind of fruit, resulting in ambiguity. Therefore, the current module aims at the semantic ambiguity phenomenon existing in the hyper-relation knowledge graph and solves the above problems through the ontology representation module.
[0144] Step S2 performs ontology box embedding representation in the following way. Since the embedding method of Box Embedding has the characteristics of spatial regions, the hierarchy of the ontology can be represented by the position and size of the box in space, and the hierarchical structure and transitivity of the ontology can be naturally expressed. If the ontology "chemist" is a subclass of "scientist", then the box of "chemist" should be completely contained inside the box of "scientist". At the same time it can implicitly support the reasoning of the ontology hierarchy.
[0145] To further integrate the semantic information from the ontology layer and the relation r, the embodiment of the present invention proposes a new transformation function:
[0146] h [CLS] = BERT(O text + 'SEP' + r text )
[0147] cen(f r (BOX(O))), off(f r (BOX(O))) = f proj (h [CLS] )
[0148] Connect the text content of the information and relationships in the ontology, input a pre-trained model BERT to capture the semantics of the text; obtain the embedding by taking the hidden vector at the [CLS] token, and use a projection function f proj Map the hidden vector h [CLS] ∈R l to R 2d space, where l is the hidden dimension of the BERT encoder and d is the dimension of the box space; implement f proj as a two-layer multi-layer perceptron (MLP) with a non-linear activation function, and further split the hidden vector into two equal-length vectors as the center and offset of the transformed box; for calculating the rationality of the ontology, use the following formula to calculate the approximate conditional probability:
[0149]
[0150] where represents the box transformation through the relationship r O ; for the correct ontology triples is close to 1, and for the wrong ontology triples is close to 0;
[0151] The goal of ontology representation training is to correspond the semantics of the ontology with the box. Each ontology concept is represented as a multi-dimensional box, and the corresponding position and size are represented by the center vector and the offset vector. The geometric range of the box directly reflects the semantic coverage range of the concept; therefore, the loss of ontology representation consists of two parts: ontology triple rationality loss and ontology hierarchical inclusion loss;
[0152] The ontology triple rationality loss is as follows:
[0153]
[0154] In the formula, negative instances constructed by using replacement of correct head and tail ontology triples;
[0155] The ontology hierarchical inclusion loss is obtained by constraining the size of the ontology hierarchical box. The range of the parent ontology box contains the range of the child ontology box. Therefore, the loss function is as follows:
[0156]
[0157] In the formula, O sub is the box of the child ontology, which consists of two vectors, the minimum inflection point and the maximum inflection point (O sub,min , O sub,max ); O super is the box of the parent ontology, which consists of two vectors, the minimum inflection point and the maximum inflection point (O super,min , O super,max );
[0158] The loss of the ontology representation is obtained as follows:
[0159] L ontolgy = L BCE + L contain .
[0160] In a possible implementation manner, the steps of establishing a semantic connection between the hyper-relation fact embedding and the corresponding ontology embedding in step S3, defining a metric function to measure the distance between the ontology box and the hyper-relation instance box by constraining the distance between the entity and the corresponding ontology vector, and obtaining the dual-perspective joint modeling loss by using contrastive learning to minimize the loss function include:
[0161] The goal of the dual-perspective joint modeling is to establish a semantic link between the entity embedding and the concept embedding, hoping that the ontology at the hyper-relation fact layer is close to its corresponding ontology information. This setting is because an entity can be regarded as a category with only one instance in the ontology. Therefore, the entity should also satisfy the hierarchical relationship between ontologies and belong to the smallest level in the hierarchy.
[0162] Therefore, the link distance between the entity and the ontology layer is established according to the following formula:
[0163] f d (e, o) = dist out (e, o) + α · dist in (e, o)
[0164] dist out (e, o) = ||Max(e - μ M , 0) + Max(μ m - e, 0)||1
[0165] dist in (e, o) = ||Cen(Box(c)) - Min(μ m , Max(μ m , e))||1
[0166] In the formula, dist out is the distance between the entity and the closest corner of the box, dist in corresponds to the distance between the center of the corresponding box and the side of the box or the entity itself. 0 < α < 1 is a fixed scalar to balance the distance between the inside and the outside. If the target entity falls outside the ontology box, more penalties are imposed. If the entity is inside the box, the penalty is reduced. The distance from the entity to the box is non-negative;
[0167] The dual-perspective joint modeling loss is calculated through the following expression:
[0168]
[0169] where γ cross represents a fixed scalar interval.
[0170] In a possible implementation, the computational expression for the balance of the hyper-relation fact representation loss, ontology representation loss, and dual-perspective joint modeling loss jointly obtained in step S4 is as follows:
[0171] L = L fact + λ1L ontolgy + λ2L cross
[0172] In the formula, L fact is the hyper-relation fact representation loss, L ontolgy is the ontology representation loss; L cross is the dual-perspective joint modeling loss; λ1, λ2 > 0 are two positive hyperparameters to balance the three losses.
[0173] Existing research mainly focuses on binary relations in knowledge graphs, and there is less research on the multi-relation structure of hyper-relation knowledge graph modeling. Existing binary relation embedding models cannot well represent the complex structure of hyper-relation knowledge graphs. To solve the limitation of the existing multi-relations being difficult to characterize, the present invention introduces Box Embedding for hyper-relation knowledge graph modeling. At the same time, the current method lacks the semantic information and structural information for expressing hyper-relation knowledge graphs. The present invention introduces an ontology perspective to enhance the semantic information of entities and relations in hyper-relation facts through the ontology perspective layer, further improving the ability of hyper-relation knowledge graph completion.
[0174] Another embodiment of the present invention also proposes a dual-perspective hyper-relation knowledge graph completion system, including:
[0175] A hyper-relation fact representation module, which, based on the representation learning method in geometric space, performs box embedding on the head entity relation and additional key-value pairs for each hyper-relation instance respectively to find the intersection, updates the entity representation by constraining the expected answer entity to fall inside the box intersection operation, and enhances the expression of the main triple through the interaction between the box of the additional key-value pair and the box of the main triple, obtaining the hyper-relation fact representation loss;
[0176] An ontology representation module, which, for the ontology corresponding to the hyper-relation instance, initializes the representation through a pre-trained model to enrich the semantic information of the ontology, then boxes the ontology in the same way to naturally express the hierarchical information between ontologies, and realizes ontology-constrained instances by constraining the intersection of the expected prediction answer set and the ontology box, obtaining the ontology representation loss;
[0177] The dual - perspective joint modeling module is used to establish the semantic connection between the hyper - relation fact embedding and the corresponding ontology embedding. By constraining the distance between the entity and the corresponding ontology vector, a metric function is defined to measure the distance between the ontology box and the hyper - relation instance box, and contrastive learning is used to minimize the loss function to obtain the dual - perspective joint modeling loss;
[0178] The loss balancing module is used to balance the hyper - relation fact representation loss, ontology representation loss, and dual - perspective joint modeling loss jointly obtained, and complete the hyper - relation knowledge graph completion.
[0179] Compared with the prior art, on the one hand, in the face of the limitation that it is difficult to represent the multiple relations in the hyper - relation knowledge graph, this invention introduces the embedding method of Box Embedding to model the hyper - relation knowledge graph, and can use the natural spatial characteristics of the box to express the hyper - relation knowledge graph completion task; on the other hand, in the face of the diversity of semantic information in the hyper - relation knowledge graph, the current methods lack the expression of semantic information and hierarchical information of the hyper - relation knowledge graph. Introducing the ontology layer can not only distinguish the meanings expressed by entities, but also fully discover the parent - child class relationships between entities, enhancing the semantic information in the hyper - relation facts.
[0180] Another embodiment of this invention also proposes an electronic device, including a processor and a memory. The processor is used to execute the computer program stored in the memory to implement the dual - perspective hyper - relation knowledge graph completion method.
[0181] Another embodiment of this invention also proposes a computer - readable storage medium. The computer - readable storage medium stores at least one instruction, and when the at least one instruction is executed by the processor, the dual - perspective hyper - relation knowledge graph completion method is implemented.
[0182] The computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer - readable storage medium can include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read - only memory, random - access memory, electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer - readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer - readable medium does not include electrical carrier signals and telecommunication signals. For the sake of convenience of description, only the parts related to the embodiments of this invention are shown above. For the specific technical details not disclosed, please refer to the method part of the embodiments of this invention. This computer - readable storage medium is non - transitory and can be stored in the storage devices formed by various electronic devices, and can implement the execution process recorded in the method of the embodiments of this invention.
[0183] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0184] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.
[0185] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that realizes the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.
[0186] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.
[0187] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.
Claims
1. A dual-viewpoint hyper-relationship knowledge graph completion method, characterized in that Including: Based on the geometric space representation learning method, for each hyper-relation instance, the head entity relation and the additional key-value pairs are respectively subjected to box embedding to find the intersection. By constraining the expected answer entity to fall inside the box intersection operation, the representation of the entity is updated. The interaction between the box of the additional key-value pairs and the box of the main triple enhances the expression of the main triple, and the hyper-relation fact representation loss is obtained; For the ontology corresponding to the hyper-relation instance, the semantic information of the ontology is enriched by initializing the representation through a pre-trained model, and then the ontology is boxed in the same way, which naturally expresses the hierarchical information between ontologies. By constraining the intersection of the expected prediction answer set and the ontology box, the ontology constraint instance is realized, and the ontology representation loss is obtained; Establish the semantic connection between the hyper-relation fact embedding and the corresponding ontology embedding. By constraining the distance between the entity and the corresponding ontology vector, a metric function is defined to measure the distance between the ontology box and the hyper-relation instance box, and contrastive learning is used to minimize the loss function to obtain the dual-view joint modeling loss; The hyper-relation fact representation loss, the ontology representation loss, and the dual-view joint modeling loss obtained jointly reach a balance, and the hyper-relation knowledge graph completion is completed.
2. The method for complementing a dual-viewpoint hyper-relation knowledge graph according to claim 1, wherein The hyper-relation knowledge graph is \(G=(E, R, F)\), where \(E\) is the set of entities, \(R\) is the set of relations, and \(F\) is the set of hyper-relation facts \((f_1, f_2,\cdots, f\) n ), where \(f\) j \(\in E\times R\times E\times P(R\times E)\), and \(P\) in the formula represents the power set; the hyper-relation fact \(f\) i \(\in F\) is represented as a tuple where \(h\), \(t\), \(v\) i \(\in E\); \(r\), \(k\) i \(\in R\). In the above tuple, \((h, r, t)\) is the main triple is the additional key-value pair corresponding to the main triple; Hyper-relation knowledge graph completion refers to predicting missing elements from hyper-relation facts. Given an incomplete hyper-relation knowledge graph G, the task objectives include: First, predicting missing entities given some entities and relations. In hyper-relation facts where h, t, v i ∈ E; r, k i ∈ R, the missing entities to be predicted can be h, r, and v i ; Second, predicting the relation r in the missing main triple or the relation v in the attached key-value pair given an entity i .
3. The dual-viewpoint hyper-relation knowledge graph completion method according to claim 1, wherein The above-mentioned geometric space-based representation learning method, the step of performing box embedding to find the intersection of the head entity relation and the additional key-value pairs for each hyper-relation instance is implemented using box embedding, including: The node initial boxing is performed in the following manner: According to the characteristics of the Box, there is: B d (cen, off) = {X ∈ R d |cen(X) - off(X) ≤ x i ≤ cen(X) + off(X), i = 1,... d}, where cen(X) ∈ R d represents the center vector of the Box, and off(X) ∈ R d represents the offset vector of the Box; then each source node x is initialized as a set containing a single entity, modeled using a box with center vector cen(x) and offset vector off(x), and the initial box embedding is (cen(x), 0), where cen(x) ∈ R d is the center vector, and 0 is the all-zero vector of dimension d; for the relation node, there is r = (cen(r), off(r)) ∈ R 2d , where off(r) ≥ 0; The hyper-relation fact box mapping is performed in the following manner: Since each relation \(r\in\mathcal{R}\) is embedded as \(r = (\text{cen}(r),\text{off}(r))\in\mathcal{R}\) 2d , where \(\text{off}(r)\geq0\), given an entity embedding box \(x\), \(r\) serves as a spatial transformation of \(x\): f r : cen(f r (x)) = cen(x) + τ r In the formula, is the Hadamard product, and (τ r ', Δ r ) are relationship-specific parameters; Using the above box mapping method, the head entity and relationship of the hyper-relation fact are mapped into a specific box space, and each set of additional key-value pairs where k i ∈R, v i ∈E, or is mapped to k i for the box space transformation of v i ; The box intersection of the hyper-relation fact is obtained in the following manner: Design differential processing of hierarchical attention mechanism and offset contraction, and perform intersection modeling on a set of box embeddings {p1,...p n} to obtain p inter = (cen(p inter ), off(p inter )) as follows: In the hierarchical attention mechanism, separate attention weights are assigned to the main triple box to ensure its dominance, and shared secondary attention weights are assigned to the additional key-value pair boxes as supplementary constraints. The expression is as follows: cen(p inter ) = α main ·cen(p main ) + α aux ·cen(p aux ) Use two independent MLPs (MLP main , MLP aux ) to process the boxes of the main triples and the attached key-value pairs respectively, and limit the attention weight α main ≥α aux .
4. The method for complementing a dual-view super-relationship knowledge graph according to claim 3, wherein The step of updating the entity representation by constraining the expected answer entity to fall inside the box intersection operation, and enhancing the expression of the main triple through the interaction between the box of the additional key-value pairs and the box of the main triple, and obtaining the hyper-relation fact representation loss includes: The offset of the additional key-value pair is constrained and shrunk by the Sigmoid function: s = DeepSets([off(p main ) ; off(p aux , 1) ;... ; off(p aux , k)]) σ(s) = Sigmoid(s) + γ (γ ∈ [0,1]) where ° is the Hadamard product, and DeepSets(·) is a permutation-invariant deep architecture that obtains the position and size of the box updated with the hyper-relation fact in space through the above two transformation operations: p inter = (cen(p inter ), off(p inter )) Calculate the distance from an entity to a box as follows: For the box after the combined contraction of the main triple and the additional key-value pairs as the query space, check whether the distance from the entity e to the corresponding box conforms to the constraints of the hyper-relation facts. Given the contracted box q = (cen(q), off(q)) ∈ R 2d and the embedding vector of the candidate entity v ∈ R d calculate the distance expression as follows: dist box (v; q) = dist outside (v; q) + α·dist inside (v; q) Where: dist outside = ||max(v - q max , 0) + max(q min - v, 0)||1 dist inside = ||cen(q) - clip(v, q min , q max ) || 1 q max = cen(q) + off(q) q min = cen(q) - off(q) The distance between the external corresponding entity and the nearest corner or edge of the box; the internal distance corresponds to the distance between the center of the box and the corresponding side or corner. If the entity is inside the box, it is the entity itself; The distance inside the box is reduced by using 0 < α < 1; The training objective is set to learn the entity embedding, as well as the geometric projection and intersection operators. Given a training set of missing and incomplete hyper-relation facts and answers, negative sampling loss is used to optimize the distance-based model: where γ represents the hyperparameter of the interval, and v + represents the correct answer entity in the hyper-relation fact q with missing input, and v - is an answer not in the missing answers and is an incorrect negative instance, and k is the number of negative instances.
5. The method for complementing a dual-view super-relation knowledge graph according to claim 1, wherein The step of enriching the semantic information of the ontology corresponding to the hyper-relation instance by initializing the representation through a pre-trained model, then boxing the ontology in the same way, naturally expressing the hierarchical information between ontologies, and realizing the ontology constraint instance by constraining the intersection of the expected prediction answer set and the ontology box, and obtaining the ontology representation loss includes: The ontology box embedding representation is carried out as follows. The hierarchy of the ontology is represented by the position and size of the boxes in space. The semantic information from the ontology layer and the relation r is integrated through the following transformation function: h [CLS] = BERT(O text + 'SEP' + r text ) cen(f r (BOX(O))), off(f r (BOX(O))) = f proj (h [ClS] ) Connect the text content of the information and relationships in the ontology, input a pre-trained model BERT to capture the semantics of the text; obtain the embedding by getting the hidden vector at the [CLS] token, and use a projection function f proj The hidden vector h [CLS] ∈R l Map to R 2d space, where l is the hidden dimension of the BERT encoder and d is the dimension of the box space; implement f proj as a two-layer multi-perceptron with a non-linear activation function, split the hidden vector into two equal-length vectors as the center and offset of the transformed box; for calculating the rationality of the ontology, use the following formula to calculate the approximate conditional probability: Among them, represents the box transformation transformed by the relationship r O ; for the correct ontology triples it is close to 1, and for the incorrect ontology triples it is close to 0; The goal of ontology representation training is to correspond the semantics of the ontology and the boxes. Each ontology concept is represented as a multi-dimensional box, and the corresponding position and size are represented by the center vector and the offset vector. The geometric range of the box directly reflects the semantic coverage range of the concept. Therefore, the loss of ontology representation consists of two parts: the ontology triple rationality loss and the ontology hierarchy inclusion loss; The ontology triple rationality loss is as follows: In the formula, Negative instances constructed using the correct head and tail ontology triples after replacement; The ontology hierarchy inclusion loss is obtained by constraining the size of the ontology hierarchy boxes. The range of the parent ontology box includes the range of the child ontology box. Therefore, the loss function is as follows: Wherein, O sub is the box of the sub-ontology, which is composed of two vector minimum inflection points and maximum inflection points (O sub,min , O sub,max ); O super is the box of the parent-ontology, which is composed of two vector minimum inflection points and maximum inflection points (O super,min , O super,max ). The ontology representation loss obtained is: L ontolgy = L BCE + L contain .
6. The dual-viewpoint hyper-relation knowledge graph completion method according to claim 1, wherein The steps of establishing the semantic connection between the hyper-relation fact embedding and the corresponding ontology embedding, by constraining the distance between the entity and the corresponding ontology vector, defining a metric function to measure the distance between the ontology box and the hyper-relation instance box, and using contrastive learning to minimize the loss function to obtain the dual-view joint modeling loss include: Establish the link distance between the entity and the ontology layer: f d (e, o) = dist out (e, o) + α·dist in (e, o) dist out (e,o) = ||Max(e - μ M , 0) + Max(μ m - e, 0)||1 dist in (e,o) = ||Cen(Box(c)) - Min(μ m , Max(μ m , e))||1 where dist out is the distance from the entity to the nearest corner of the box, and dist in corresponds to the distance between the center of the box and the side of the box or the entity itself. 0 < α < 1 is a fixed scalar to balance the distance between the inside and the outside. If the target entity falls outside the body box, more penalties are imposed, and if the entity is inside the box, the penalty is reduced. The distance from the entity to the box is non - negative; The dual-view joint modeling loss is calculated by the following expression: where γ cross represents a fixed scalar interval.
7. The method for complementing a dual-viewpoint hyper-relation knowledge graph according to claim 1, wherein The calculation expression for the balance of the obtained hyper-relation fact representation loss, ontology representation loss, and dual-view joint modeling loss is as follows: L = L fact + λ1L ontolgy + λ2L cross where, L fact is the loss of hyper-relation fact representation, and L ontolgy is the loss of ontology representation; L cross is the loss of dual-view joint modeling; λ1, λ2 > 0 are two positive hyperparameters to balance the three losses.
8. A dual-viewpoint hyper-relation knowledge graph completion system, characterized in that, Including: The hyper-relation fact representation module is used to, based on the representation learning method in geometric space, perform box embedding on the head entity relation and the additional key-value pairs for each hyper-relation instance respectively and take the intersection. By constraining the expected answer entity to fall inside the box intersection operation to update the representation of the entity, and enhancing the expression of the main triple through the interaction between the box of the additional key-value pair and the main triple box, to obtain the hyper-relation fact representation loss; The ontology representation module is used to, for the ontology corresponding to the hyper-relation instance, initialize the representation through a pre-trained model to enrich the semantic information of the ontology, and then boxify the ontology in the same way to naturally express the hierarchical information between ontologies. By constraining the intersection of the expected prediction answer set and the ontology box, the ontology constraint instance is realized to obtain the ontology representation loss; The dual-view joint modeling module is used to establish the semantic connection between the hyper-relation fact embedding and the corresponding ontology embedding, by constraining the distance between the entity and the corresponding ontology vector, defining a metric function to measure the distance between the ontology box and the hyper-relation instance box, and using contrastive learning to minimize the loss function to obtain the dual-view joint modeling loss; The loss balance module is used to balance the obtained hyper-relation fact representation loss, ontology representation loss, and dual-view joint modeling loss to complete the hyper-relation knowledge graph completion.
9. An electronic device, characterized in that, Including a processor and a memory. The processor is used to execute the computer program stored in the memory to implement the dual-view hyper-relation knowledge graph completion method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by the processor, it implements the dual-view hyper-relation knowledge graph completion method as described in any one of claims 1 to 7.