A knowledge graph triple classification method based on marginal parameter loss

By combining the knowledge representation learning method of isA relation modeling, using projection invariance and partial order constraint modeling of isA relationship transferability and antisymmetry, the problems of data sparsity and isA relationship modeling in the knowledge graph embedding method are solved, and the performance of knowledge graph representation learning and the effect of triple-group classification are improved.

CN116795999BActive Publication Date: 2025-05-16NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310608186.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-25
Publication Date
2025-05-16
Estimated Expiration
2043-05-25

AI Technical Summary

Technical Problem

Due to data sparseness, the existing knowledge graph embedding method cannot effectively encode sparse entities, and cannot fully model the transitiveness and antisymmetry of isA relationship.

Method used

By combining the knowledge representation learning method of isA relation modeling, the transitiveness of isA relationship is modeled using the projection invariance of vectors, the antisymmetry of the isA relationship is introduced, and the hyperparameters of entity-level information are introduced to distinguish entities at different concept levels.

Benefits of technology

It improves the performance of knowledge graph representation learning, can better encode sparse entities, and perform better in triple classification tasks, alleviating the problem of insufficient learning of embedded representations caused by sparse samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116795999B_ABST
    Figure CN116795999B_ABST
Patent Text Reader

Abstract

The present invention discloses a knowledge graph triple classification method based on marginal parameter loss, comprising the following steps: dividing a triple set into two disjoint subsets, namely, an isA triple set and a relationship triple set; for relationship triples, adopting a RotatE model for modeling; modeling the transitivity and antisymmetry of the isA relationship to model the isA relationship chain, and after transitivity modeling, for the case where the same hyponym corresponds to different hypernyms, modeling them as different points on the same straight line; combining the modeling method for transitivity and the modeling for antisymmetry, modeling is performed using the different projection points of each component of the entity embedding representation on the virtual axis of the complex plane after rotation; using a loss function based on marginal parameters as an optimization target for training, and using the trained model to classify the knowledge graph into triples. The vector representation performance learned by the present invention is better than that of other baseline methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine learning, and in particular relates to a knowledge graph triple classification method based on marginal parameter loss. Background Art

[0002] Knowledge graph representation learning has been a hot topic in the field of artificial intelligence in recent years. It aims to map entities and relationships in the graph to a low-dimensional continuous vector space to obtain a dense vector representation to support knowledge calculation and reasoning. Knowledge graph (KG) describes the objects in the objective world and their connections in a directed graph structure, and is represented and stored in the form of a triple (h, r, t), where h, r, and t represent the head entity, relationship, and tail entity, respectively.

[0003] The commonly used knowledge graph embedding methods have poor embedding performance due to data sparsity problems, and cannot encode sparse entities well. Therefore, researchers proposed to improve the performance of representation learning by introducing some external information or prior knowledge. Among them, the isA relationship triple, as an important part of the knowledge graph, contains rich semantic information. By accurately modeling the isA relationship, the performance of embedding representation learning of entities and relationships can be enhanced, alleviating the problem of data sparsity.

[0004] Generally, the isA relationship is a special relationship from a more specific entity to a more abstract entity. That is, given entity c1 and entity c2, if the extension of c2 contains the extension of c1, then c1 and c2 are considered to have an "isA" relationship. From a linguistic point of view, the isA relationship is a hyponym-hypernym relationship between c1 and c2. c2 is called the hypernym of c1, and c1 is the hyponym of c2, which is recorded as isA(c1, c2). For example, in the triple (dog, isA, animal), dog is a more specific animal species, and animal is a more abstract concept. The isA relationship connects the hyponym dog with the hypernym animal. The potential semantic information contained in the isA relationship triple includes the following two aspects: on the one hand, the hypernym provides basic category information for its hyponym, and on the other hand, the hyponym provides specific detail information for its corresponding hypernym, so that they benefit each other when learning the embedded representation, improving the learning effect of the embedded representation. For example, for an animal "sloth" with fewer adjacent samples, since it has the basic category information of its hypernym "animal" like other common animal instances "cat" and "dog", its embedding representation should be related to the embedding representations of the hypernym "animal" and other common animal instances "cat" and "dog". The embedding representation of "sloth" can be roughly determined by observing the representations of other animals and their hypernym "animal" in the embedding space. The embedding representation of the hypernym "animal" can also be roughly determined by observing the representations of multiple known animal instances in the embedding space. For knowledge graphs containing isA relationship triples, the potential semantic information of the isA relationship triples themselves is used to learn their representations, which can alleviate the problem of insufficient learning caused by graph sparsity without increasing the requirements for additional training samples.

[0005] The isA relationship has two important properties. The first is transitivity, that is, from the two triples (X isA Y) and (Y isA Z), we can infer that the triple (X isA Z) is true. From the two triples (Audrey Hepburn isA Actor) and (Actor isAArist), even if the triple (Audrey Hepburn isA Arist) does not exist in the knowledge graph, we can infer that this triple is true through transitivity. The second is antisymmetry. There is a partial order relationship between the head entity and the tail entity in the isA relationship, and this partial order relationship is antisymmetric. The entities "Singer", "Actor", "Arist" and "Person" in the three triples (Singer isA Arist), (Actor isA Arist), and (Arist isA Person) form a tree-like hierarchical structure with antisymmetry, that is, the triples (Arist isA Singer), (Arist isA Actor), and (PersonisA Arist) must not be true.

[0006] In order to integrate the learning of isA relationship triples into the existing knowledge graph representation learning, researchers use the characteristics of the hierarchical relationship between the head and tail entities in the isA relationship triples for modeling. Existing related work can be divided into three categories. The first category is the region-based model. This type of work models the concept as a geometric region in space, and uses the inclusion relationship between regions to represent the hierarchical relationship between the head entity and the tail entity in the isA relationship. However, the geometric bodies modeled by this type of model are usually simple and regular geometric bodies (hyperspheres or hyperrectangles), which cannot accurately model the complex inclusion relationship in the real world. The second category is the method based on partial order relationship modeling. This type of work constrains the head entity and the tail entity with a hierarchical relationship to always remain in order in the positive real number embedding representation space to model antisymmetry and transitivity. However, the model models the embedding space as a positive real number space, which limits the expressiveness of the model. The third category is the method based on linear transformation. This type of work uses linear transformation to model the hierarchical relationship between entities. This type of method cannot model antisymmetry and transitivity. In summary, through the analysis of existing work, it is found that existing work cannot fully model the characteristics of isA relationships. Summary of the invention

[0007] The knowledge representation learning method combined with isA relationship modeling belongs to the knowledge graph representation learning model based on instances and concepts. The existing methods of this type of model include three categories: region-based models, methods based on partial order relationship modeling, and methods based on linear transformation, but none of them can fully model the characteristics of isA relationships. Therefore, the present invention proposes a knowledge representation learning method combined with isA relationship modeling, which improves the knowledge graph representation learning performance by accurately modeling the characteristics of isA relationships. Specifically, the transitivity of the isA relationship is modeled using the projection invariance of the vector, and the partial order constraint is introduced to model the antisymmetry of the isA relationship. After training, the model can model the situation where the same entity corresponds to multiple different hyponyms or hypernyms. In addition, in order to avoid the misjudgment of the existence of isA relationships between different hyponyms or hypernyms at the same level corresponding to the same entity, the model introduces a hyperparameter of entity hierarchy information. The triple classification task was performed on multiple data sets, and the experimental results showed that the vector representation performance learned by the present invention is better than that of other baseline methods.

[0008] The knowledge graph triple classification method based on marginal parameter loss of the present invention comprises the following steps:

[0009] Divide the triple set S into two disjoint subsets, namely the isA triple set S i and relation triple set S R , the isA triple set Where n i YesS i The size of e i 、e j ∈E, whose embedding representations are is an n-dimensional embedding representation space; the relationship triple set Where n r YesS R The size of h, t∈E, r∈R r , whose embeddings are expressed as It is an n-dimensional embedding representation space;

[0010] Different methods are used to analyze the isA triple set S i and relation triple set S RModeling; for the relation triple (h, r, t), the RotatE model is used for modeling; the transitivity and antisymmetry of the isA relationship are modeled to model the isA relationship chain. After transitivity modeling, for the case where the same hyponym corresponds to different hypernyms, these different hypernyms are modeled as different points on the same straight line, that is, these different hyponyms are modeled as different embedding representations. For the case where the same hypernym corresponds to multiple different hyponyms, these different hyponyms are modeled as different embedding representations on the same straight line; combining the modeling method for transitivity and the modeling for antisymmetry, modeling is performed using the different projection points of each component of the entity embedding representation on the imaginary axis of the rotated complex plane;

[0011] In order to improve the distinguishability between positive and negative samples, a loss function based on marginal parameters is used as the optimization objective for training, and the trained model is used to perform triple classification on the knowledge graph.

[0012] Furthermore, the RotatE model maps entities and relations into an n-dimensional complex vector space The relation r is modeled as a vector element-wise rotation from the head entity embedding to the tail entity embedding in the complex vector space, that is:

[0013] e t (k) = rot(e h (k),θ r (k))

[0014] Among them, e h 、e h ,θ r They represent the embedding representations of the head entity h, the tail entity t, and the relation r, respectively, and are all n-dimensional complex vectors. h (k), e t (k),θ r (k) respectively represent e h 、e t ,θ r The k-th dimension component of , which are all one-dimensional complex vectors;

[0015] Re(e h (k)) and Im(e h (k)) represents the complex number e h (k) real and imaginary parts, then e h (k) = Re (e h (k))+i·Im(e h (k)) can also be expressed as: [Re(e h (k)),Im(e h (k))] T ,

[0016] Each component θr The modulus of is constrained to be 1, that is, |θ r (k)|=1, so we can know that θ r It is expressed as:

[0017]

[0018] θ r The corresponding angle about the origin of the complex plane is θ r (k) counterclockwise rotation, rot(e h (k),θ r (k)) is the rotation function, which means that e h (k) Rotation θ r (k) radians;

[0019] According to the coordinate rotation transformation formula, the point p(x,y) on the plane rotates counterclockwise about the origin by θ radians to reach p'(X,Y), and we get:

[0020]

[0021] Therefore, e t (k) = rot(e h (k),θ r (k)) is expressed as:

[0022]

[0023] For a relation triple (h, r, t), the distance function of the RotatE model is defined as:

[0024]

[0025] RotatE cannot handle isA relationships.

[0026] Furthermore, the isA relationship is modeled as follows:

[0027] Define the isA relationship chain (e1, e2..., e m ) represents the set of triples (e1, isA, e2), ..., (e m-2 ,isA,e m-1 ), (e m-1 ,isA,e m ), the tail entity of a triple in the set is the head entity of another triple, where e1,…,e m Its n-dimensional embedding representation vector, e j (k) represents entity e j The k-th dimension vector of (j=1,2,...,m);

[0028] The isA relationship chain is modeled by modeling the transitivity and antisymmetry of the isA relationship. The transitivity modeling method is as follows:

[0029] The isA relation is transitive, that is, if there are triples (A isA B), (B isA B), then there must be a triple (A isA C);

[0030] For a triple (h, r, t), each relation r is modeled as a transformation operation T r , that is, after model training, it will satisfy e t =T r (e h ); for the isA relationship chain (e1, e2, e3), there are (e1, isA, e2), (e2, isA, e3), (e1, isA, e3), e1, e2, e3 are different entities, and the corresponding embedding representations e1, e2, e3 are different; if the transformation operation T can model the transitive relationship, that is, it satisfies

[0031] e2=T(e1), e3=T(e2)=T(T(e1)), e3=T(e1),

[0032] roll out

[0033] T(e1)=T(T(e1)),

[0034] That is, executing T operation multiple times on e1 is equivalent to executing T operation once, which shows that transitivity has the following property: the effect of transitivity acting on an entity multiple times is equivalent to the effect of acting on this entity once. Mathematically, this property is expressed by a linear transformation of a vector - a projection transformation.

[0035] Furthermore, the linear transformation-projection transformation of the vector includes:

[0036] e1(k), e2(k), e3(k) are the k-th components of e1, e2, e3, all of which are complex numbers and are represented by two-dimensional vectors on the complex plane:

[0037] e1(k)=[Re(e1(k)),Im(e1(k))] T

[0038] e2(k)=[Re(e2(k)),Im(e2(k))] T ;

[0039] e3(k)=[Re(e3(k)),Im(e3(k))] T

[0040] Use the projection operation T to act on e1(k) to obtain T(e1(k)), that is, e1(k) is projected onto the real axis once, and the result is:

[0041] T(e1(k))=[Re(e1(k)),0] T

[0042] T(e2(k)) is obtained by projecting e2(k) onto the real axis once:

[0043] T(e2(k))=[Re(e2(k)),0] T

[0044] Since e2(k) = T(e1(k)), T(e2(k)) is the second projection of e1(k) onto the real axis, we get:

[0045] T(e2(k))=T(T(e1(k)))=[Re(e1(k)),0] T

[0046] Therefore, we get

[0047] Re(e2(k))=Re(e1(k))

[0048] T(e3(k)) is the projection of e3(k) onto the real axis, and we get:

[0049] T(e3(k))=[Re(e3(k)),0] T

[0050] Since e3(k) = T(T(e1(k))), T(e3(k)) is the projection of e1(k) onto the real axis three times, we get

[0051] T(e3(k))=T(T(T(e1(k))))=[Re(e1(k)),0] T

[0052] Therefore, we get

[0053] Re(e3(k))=Re(e1(k))

[0054] In summary, we can conclude

[0055] Re(e3(k))=Re(e2(k))=Re(e1(k))

[0056] That is, the projection operation T is applied to e1(k), e2(k), and e3(k) to obtain the same projection point, which is d. Since the imaginary information Im(e3(k)), Im(e2(k)), and Im(e1(k)) of e1(k), e2(k), and e3(k) are different from each other, e1(k), e2(k), and e3(k) are different from each other. Therefore, the projection operation is used to model the transitivity of the isA relationship.

[0057] Furthermore, the projection transformation is equivalent to the idempotent transformation, that is, if p(e) represents the projection transformation of the complex number e, then

[0058]

[0059] It represents the projection transformation result of e on the real axis of the complex plane. Therefore, the imaginary part information of p(e) is 0. That is, if the projection operation is only performed on the real axis, the imaginary part information of the projection points of entities in all different isA relationship chains will be 0, resulting in the loss of one dimension of information.

[0060] In order to obtain a more general projection transformation method, the complex plane coordinate axis is rotated around the origin and the complex plane coordinate axis is rotated counterclockwise by θ p Radians, get the real axis x' and imaginary axis y' after the rotation, project e1(k), e2(k), e3(k) onto the real axis x' after the rotation, and get their projection points in the new coordinate system are all p, that is, they satisfy

[0061] T(e1(k))=T(e2(k))=T(e3(k))

[0062] The entities in different isA relationship chains are projected onto the real axis of the rotated complex plane coordinate axis, and the imaginary part information of each projection point is different. This allows the projection points to retain different imaginary part information while modeling the transitivity of the isA relationship.

[0063] Furthermore, the transitivity modeling process of the isA relationship is described algebraically as follows:

[0064] The rotation vector of the projection coordinate axis is the n-dimensional complex vector θ p , for different transmission chains, θ p The values ​​are the same; θ p Each component θ p The modulus of (k) is constrained to be 1, i.e. |θ p (k)|=1, therefore, θ p (k) is expressed in the form of

[0065]

[0066] θ p(k) means rotating the complex plane coordinate axis counterclockwise about the origin by θ p (k) radians, 0 < θ p (k)≤2π;

[0067] For the isA relationship chain (e1, e2…, e m ) on any entity embedding e j , e j (k) after rotating by θ p (k) The projection point p on the real axis of the complex plane after radians x (e j (k)) is expressed as:

[0068]

[0069] For any entity e in the isA relationship transfer chain j (k)(j=2,3,...,m), all have

[0070] p x (e j (k)) = p x (e1(k))(j=2,3,...,m)

[0071] Combining the above formula, we get:

[0072]

[0073] If the above equation is true, then and only if the following equation is true:

[0074] cosθ p (k)Re(e j (k))+sinθ p (k)Im(e j (k)) = cosθ p (k)Re(e1(k))+sinθ p (k)Im(e1(k))

[0075] If formula (10) is valid for j=2,3,...,m, then:

[0076] cosθ p (k)Re(e j (k))+sinθ p (k)Im(e j (k)) = c k

[0077] Among them, c k Must be a constant, that is:

[0078] c k= cosθ p (k)Re(e1(k))+sinθ p (k)Im(e1(k))

[0079] From the above formula, it is c k and θ p (k) defines a straight line whose slope is -cot(θ p (k)); and all entities e1, e2, ..., e j ,…,e m , all satisfy the above formula, that is, e j (k)(j=1,2,3,...,m) are all generated by c k and θ p (k) is defined as -cot(θ p (k)), and the slope of the real axis of the rotated complex plane is tan(θ p (k)), it can be concluded that the straight line defined by the above formula is perpendicular to the real axis of the rotated complex plane, and the straight lines corresponding to different transfer chains are parallel to each other and perpendicular to the real axis of the rotated complex plane;

[0080] For the isA relation triple (e j ,isA,e j+1 ), the distance function for transitivity modeling is defined as:

[0081]

[0082] where p x For projection operation.

[0083] Furthermore, the antisymmetry of the isA relationship is modeled, including:

[0084] The isA relationship is antisymmetric, that is, if there is a triple (a isA b), then the triple (b isA a) must not exist; for the isA relationship chain (e1, e2…, e m ) in any triple (e j ,isA,e j+1 ), e j , e j+1 The corresponding embedding is denoted as e j 、e j+1 , e j+1 Yes j Hypernym of, or e j+1 E j More abstract, j E j+1 More specifically, based on the quantitative distinction j 、e j+1 Embedding representationj 、e j+1 , to model (e j+1 ,isA,e j ) must not exist;

[0085] If the operation F can model the antisymmetry of the isA relation, that is, if e j+1 Yes j The hypernym of F(e j )≤F(e j+1 ), the transitivity is modeled by using the fact that each component of the entity embedding representation has the same projection point on the real axis of the rotated complex plane coordinate axis, that is, p x (e j (k)) = p x (e j+1 (k));

[0086] Project each component of the entity embedding representation onto the rotated imaginary axis of the complex plane coordinate axis. The smaller the L1 or L2 norm value of the projection vector, the more abstract the entity is, while the larger the value, the more specific the entity is.

[0087] p y (e j (k))、p y (e j+1 (k)) represent entities e j+1 、e j The embedding representation e j 、e j+1 The projection point of the k-th dimension component of the rotated complex plane coordinate axis on the imaginary axis. If e j+1 Yes j A hypernym of iff:

[0088]

[0089] For a vector X = (x(1), x(2)...x(N)), represents each dimension of vector X; that is:

[0090] e j The L1 or L2 norm of the projection point of each component on the imaginary axis of the rotated complex plane coordinate axis is greater than e j+1 The L1 or L2 norm value of the projection point, then p y (e j (k)) is vectorized as:

[0091]

[0092] In order to avoid different hyponyms corresponding to the same hypernym, or different hypernyms corresponding to the same hyponym, being misjudged as having an isA relationship after training, a hyperparameter λ is added to distinguish entities at different conceptual levels. F , that is: if e j+1 Yes j A hypernym of iff:

[0093]

[0094] For the isA relation triple (e j ,isA,e j+1 ), the distance function for antisymmetric modeling is defined as:

[0095]

[0096] where p y (z) is the projection operation; function [z] + Only the positive part of z is retained, and the rest is 0, that is, [z] + =max{0,z};

[0097] Therefore, the isA relationship triple is defined (e j ,isA,e j+1 ) The distance function is the sum of the transitivity modeling distance function and the antisymmetry modeling distance function, that is:

[0098] BRIEF DESCRIPTION OF THE DRAWINGS

[0099] Figure 1 Schematic diagram of the RotatE method;

[0100] Figure 2 Schematic diagram of the projection transformation of vectors e1, e2, and e3 onto the real axis of the coordinate axis;

[0101] Figure 3 Schematic diagram of the projection of the entity embedding in the isA transfer chain onto the real axis of the rotated complex plane. DETAILED DESCRIPTION

[0102] The present invention is further described below in conjunction with the accompanying drawings, but the present invention is not limited in any way. Any changes or substitutions made based on the teachings of the present invention belong to the protection scope of the present invention.

[0103] Judging whether entities A, B, and C satisfy the isA relationship is similar to judging whether A, B, and C are three generations of grandparents and grandchildren in a family, that is, A is the descendant of B, and B is the descendant of C. This relationship also has transitivity and antisymmetry. First, it is necessary to confirm whether A, B, and C belong to the same family with a blood relationship, which can be determined by whether they have a common feature. For example, genetic characteristics, gene sequencing results show that A, B, and C all have the same specific gene, proving that A, B, and C belong to the same family. This is similar to modeling transitivity. Then, other information is needed to determine the generation of A, B, and C, such as age information. The information shows that A is more than 20 years younger than B, and B is more than 20 years younger than C. It can be confirmed that A is a grandchild, B is a child, and C is a father. This is similar to modeling antisymmetry. In summary, it can be judged that A is the descendant of B, and B is the descendant of C. Therefore, the present invention proposes to model the isA relationship triple by modeling the two basic characteristics of the isA relationship-transitivity and antisymmetry. First, we use a universal projection transformation to model transitivity. Specifically, any point on the same straight line in space is projected onto a line perpendicular to the line, and the projection points are the same. This projection transformation has properties consistent with the transitive relationship. Secondly, the present invention introduces partial order constraints to model antisymmetry. Specifically, the projection points of entity embedding representations on the imaginary axis of the rotated complex plane are different. It is believed that the more abstract the entity is, the smaller the modulus value of the projection vector of the entity embedded on the imaginary axis of the rotated complex plane, and vice versa. In addition, a hyperparameter is introduced to model the hierarchical information of the entity. These modeling methods are combined to more completely model the isA relationship.

[0104] The knowledge representation learning method combined with isA relationship modeling proposed in the present invention belongs to a knowledge graph representation learning model based on instances and concepts. It models isA relationship triples by modeling isA relationship characteristics and integrates them into knowledge graph representation learning, thereby improving the effect of knowledge graph representation learning.

[0105] The knowledge graph describes entities and their relationships. In this paper, E, R, and S are used to represent entity sets, relationship sets, and triple sets, respectively. The knowledge graph can be formally described as KG = {E, R, S}. We divide the relationship set into two categories: isA relationship r i and other relation sets R except isA relation r , that is: R = {r i ∪R r In order to fully capture the potential semantic links in each type of triple, the present invention trains each type of triple through different embedding representation learning mechanisms. To this end, the present invention divides the triple set S into two disjoint subsets: isA triple set S iand relational triple set S R :

[0106] 1.isA triple set Where n i YesS i The size of e i 、e j ∈E, whose embeddings are in bold It is an n-dimensional embedding representation space.

[0107] 2. Relation Triple Set Where n r YesS R The size of h, t∈E, r∈R r , whose embeddings are indicated in bold It is an n-dimensional embedding representation space.

[0108] Next, we will follow S = {S i ∪S R}, model and describe two types of triples in the knowledge graph.

[0109] The knowledge graph contains two types of triples: relation triples and isA triples. The present invention will use different methods to model these two types of triples, which are described below respectively.

[0110] RotatE can model various relationship patterns such as symmetry, antisymmetry, combination relationship and reversible relationship. Therefore, for the relationship triple (h, r, t), the present invention adopts the classic RotatE model for modeling. The RotatE model maps entities and relationships to n-dimensional complex vector space. The relation r is modeled as a vector element-wise rotation from the head entity embedding to the tail entity embedding in the complex vector space, that is:

[0111] e t (k) = rot(e h (k),θ r (k)) (1)

[0112] Among them, e h 、e h ,θ r They represent the embedding representations of the head entity h, the tail entity t, and the relation r, respectively, and are all n-dimensional complex vectors. h (k), e t (k),θ r (k) respectively represent e h 、e t ,θr The k-th dimension component of , which are all one-dimensional complex vectors.

[0113] Re(e h (k)) and Im(e h (k)) represents the complex number e h (k) is the real and imaginary part of the x-axis. In the xy coordinate system, replace the x-axis with the real axis and the y-axis with the imaginary axis to obtain the complex plane coordinate axis. Any complex number can be represented by a two-dimensional vector on the complex plane, that is, e h (k) = Re (e h (k))+i·Im(e h (k)) can also be expressed as: [Re(e h (k)),Im(e h (k))] T ,like Figure 1 shown.

[0114] θ r The modulus of each component is constrained to be 1, i.e. |θ r (k)|=1, so we can know that θ r It can be expressed as:

[0115]

[0116] θ r The corresponding angle about the origin of the complex plane is θ r (k) counterclockwise rotation. rot(e h (k),θ r (k)) is the rotation function, which means that e h (k) Rotation θ r (k) Radians.

[0117] According to the coordinate rotation transformation formula, the point p(x,y) on the plane rotates counterclockwise about the origin by θ radians to reach p'(X,Y), and we can get:

[0118]

[0119] Therefore, e t (k) = rot(e h (k),θ r (k)) can be expressed as:

[0120]

[0121] For a relation triple (h, r, t), the distance function of the RotatE model is defined as:

[0122]

[0123] RotatE cannot handle isA relations. The proof is as follows: Assuming that RotatE can handle isA, then for isA relation triples (A, isA, B) and (B, isA, C), we can infer (A, isA, C), where A, B, and C are different entities, and the corresponding embedding representations are e A ,e B ,e C According to the properties of RotatE about the combination relationship, it can be concluded that the argument θ of the relational embedding representation of the three triples is R(A→B) ,θ R(B→C) ,θ R(A→C) Satisfy θ R(A→B) +θ R(B→C) =θ R(A→C) At the same time, since these three relations are all isA relations, θ R(A→B) =θ R(B→C) =θ R(A→C) Therefore, the argument of the isA relation embedding must satisfy θ R(A→B) =θ R(B→C) =θ R(A→C) =2nπ(n=0,1,2...), which will make e A =e B =e C , which is consistent with e A ,e B ,e C The different preconditions are contradictory. Therefore, RotatE cannot handle the isA relationship.

[0124] isA triple modeling:

[0125] Since RotatE cannot process the isA relationship, the present invention adopts the following method to model the isA relationship.

[0126] Define the isA relationship chain (chain of isA)(e1, e2…, e m ) represents the set of triples (e1, isA, e2), ..., (e m-2 ,isA,e m-1 ), (e m-1 ,isA,e m ), the tail entity of a triple in the set is the head entity of another triple, where e1,…,e m Its n-dimensional embedding representation vector, e j (k) represents entity e j The k-th dimensional vector of (j=1,2,...,m).

[0127] The isA relationship chain is modeled by modeling these two properties of the isA relationship - transitivity and antisymmetry. First, model transitivity.

[0128] Transitivity Modeling:

[0129] The isA relation is transitive, that is, if there are triples (A isA B), (B isA B), then there must be a triple (A isA C).

[0130] For a triple (h, r, t), the knowledge graph embedding learning model models each relation r as a transformation operation T r , that is, after model training, it will satisfy e t =T r (e h ). For the isA relationship chain (e1, e2, e3), there are (e1, isA, e2), (e2, isA, e3), (e1, isA, e3), e1, e2, e3 are different entities, and the corresponding embedding representations e1, e2, e3 are different. If the transformation operation T can model the transitive relationship, that is, it satisfies

[0131] e2=T(e1), e3=T(e2)=T(T(e1)), e3=T(e1),

[0132] It can be inferred that

[0133] T(e1)=T(T(e1)),

[0134] This can be understood as: performing T operation on e1 multiple times is equivalent to performing T operation once. This shows that transitivity has the following property: the effect of transitivity acting on an entity multiple times is equivalent to the effect of acting on this entity once. Mathematically, this property can be expressed by a linear transformation of a vector - projection transformation. The projection transformation of a vector is shown as follows Figure 2 As shown, the results of projecting the two-dimensional vector e1 on the x-axis multiple times are the same as projecting it once.

[0135] e1(k), e2(k), e3(k) are the k-th components of e1, e2, e3. They are all complex numbers and can be represented by two-dimensional vectors on the complex plane as

[0136] e1(k)=[Re(e1(k)),Im(e1(k))] T

[0137] e2(k)=[Re(e2(k)),Im(e2(k))] T

[0138] e3(k)=[Re(e3(k)),Im(e3(k))] T

[0139] Using the projection operation T on e1(k) we get T(e1(k)), which can be understood as a projection transformation of e1(k) onto the real axis, and we get

[0140] T(e1(k))=[Re(e1(k)),0] T

[0141] T(e2(k)) can be understood as e2(k) projected onto the real axis once.

[0142] T(e2(k))=[Re(e2(k)),0] T

[0143] Since e2(k) = T(e1(k)), T(e2(k)) can also be understood as the second projection of e1(k) onto the real axis, we get

[0144] T(e2(k))=T(T(e1(k)))=[Re(e1(k)),0] T

[0145] Therefore, we can get

[0146] Re(e2(k))=Re(e1(k))

[0147] T(e3(k)) can be understood as a projection of e3(k) onto the real axis, and we get

[0148] T(e3(k))=[Re(e3(k)),0] T

[0149] Since e3(k) = T(T(e1(k))), T(e3(k)) can also be understood as the projection of e1(k) onto the real axis three times, we get

[0150] T(e3(k))=T(T(T(e1(k))))=[Re(e1(k)),0] T

[0151] Therefore, we can get

[0152] Re(e3(k))=Re(e1(k))

[0153] In summary, it can be concluded that

[0154] Re(e3(k))=Re(e2(k))=Re(e1(k))

[0155] That is, use the projection operation T to act on e1(k), e2(k), and e3(k) to get the same projection point, such as Figure 2As shown, the projection point is d. And because the imaginary information Im(e3(k)), Im(e2(k)), Im(e1(k)) of e1(k), e2(k), and e3(k) are different from each other, e1(k), e2(k), and e3(k) are different from each other, that is, the projection operation can model the transitivity of the isA relationship.

[0156] In linear algebra, a projective transformation is equivalent to an idempotent transformation. That is, if p(e) represents the projective transformation of a complex number e, then

[0157]

[0158] represents the projection transformation result of e on the real axis of the complex plane, so the imaginary part information of p(e) is all 0. That is, if the projection operation is only projected onto the real axis, the imaginary part information of the projection points of all entities in different isA relationship chains will be 0, which will lose one dimension of information. In order to obtain a more general projection transformation method, we can consider rotating the complex plane coordinate axis around the origin. Figure 3 As shown, rotate the complex plane coordinate axis counterclockwise by θ p Radians, get the real axis x' and imaginary axis y' after the rotation, project e1(k), e2(k), e3(k) onto the real axis x' after the rotation, and get their projection points in the new coordinate system are all p, that is, they satisfy

[0159] T(e1(k))=T(e2(k))=T(e3(k))

[0160] The entities in different isA relationship chains are projected onto the real axis of the rotated complex plane coordinate axis, and the imaginary part information of each projection point is different. This can not only allow the projection points to retain different imaginary part information, but also model the transitivity of the isA relationship.

[0161] The modeling process is described algebraically below. The rotation vector of the projection coordinate axis is an n-dimensional complex vector θ p , for different transmission chains, θ p The values ​​are the same. p Each component θ p The modulus of (k) is constrained to be 1, i.e. |θ p (k)|=1, therefore, θ p The form of (k) can be expressed as

[0162]

[0163] θ p (k) means rotating the complex plane coordinate axis counterclockwise about the origin by θ p (k) radians, 0 < θ p(k)≤2π. For the isA relationship chain (e1, e2…, e m ) on any entity embedding e j , e j (k) after rotating by θ p (k) The projection point p on the real axis of the complex plane after radians x (e j (k)) can be expressed as:

[0164]

[0165] For any entity e in the isA relationship transfer chain j (k)(j=2,3,...,m), all have

[0166] p x (e j (k)) = p x (e1(k)) (j=2,3,...,m) (7)

[0167] Combining Formula 6 and Formula 7, we can get:

[0168]

[0169] After simplification, we get:

[0170]

[0171] If formula (9) is true, then and only if formula (10) is true:

[0172] cosθ p (k)Re(e j (k))+sinθ p (k)Im(e j (k)) = cosθ p (k)Re(e1(k))+sinθ p (k)Im(e1(k))(10)

[0173] If formula (10) is valid for j=2,3,...,m, then:

[0174] cosθ p (k)Re(e j (k))+sinθ p (k)Im(e j (k)) = c k (11)

[0175] Among them, c k Must be a constant. That is:

[0176] ck = cosθ p (k)Re(e1(k))+sinθ p (k)Im(e1(k))(12)

[0177] From formula (11), we can see that it is k and θ p (k) defines a straight line whose slope is -cot(θ p (k)). All entities e1, e2, ..., e j ,…,e m , all satisfy formula (11), that is, e j (k)(j=1,2,3,...,m) are all generated by c k and θ p (k) is defined as -cot(θ p (k)). The slope of the real axis of the rotated complex plane is tan(θ p (k)). It can be seen that the straight line defined by formula (11) is perpendicular to the real axis of the rotated complex plane, and the straight lines corresponding to different transfer chains are parallel to each other and perpendicular to the real axis of the rotated complex plane.

[0178] For the isA relation triple (e j ,isA,e j+1 ), the distance function for transitivity modeling is defined as:

[0179]

[0180] where p x is the projection operation, defined by formula (6).

[0181] Through transitive modeling, for the case where the same hyponym corresponds to different hypernyms, these different hypernyms are in the same transitive chain, and the model will model these different hypernyms as different points on the same line. That is, the model can model these different hyponyms as different embedding representations. Similarly, for the case where the same hypernym corresponds to multiple different hyponyms, the model can model these different hyponyms as different embedding representations on the same line. Therefore, the model can model the case where the same hyponym corresponds to different hypernyms and the same hypernym corresponds to different hyponyms.

[0182] Next, the antisymmetry of the isA relationship will be modeled.

[0183] The isA relationship is antisymmetric, that is, if there is a triple (a isA b), then the triple (b isA a) must not exist. m) in any triple (e j ,isA,e j+1 ), e j , e j+1 The corresponding embedding is denoted as e j 、e j+1 .e j+1 Yes j The hypernym of e j+1 E j More abstract, j E j+1 More specifically, we need to quantify the difference between j 、e j+1 Embedding representation j 、e j+1 , in order to model (e j+1 ,isA,e j ) must not exist. If the operation F can model the antisymmetry of the isA relation, that is, if e j+1 Yes j The hypernym of F(e j )≤F(e j+1 ). The transitivity is modeled by using the fact that each component of the entity embedding representation has the same projection point on the real axis of the rotated complex plane coordinate axis, that is, p x (e j (k)) = p x (e j+1 (k)). Combined with the modeling method for transitivity, the modeling of antisymmetry is modeled by using the different projection points of each component of the entity embedding representation on the imaginary axis of the rotated complex plane. Each component of the entity embedding representation is projected on the imaginary axis of the rotated complex plane coordinate axis. The smaller the L1 or L2 norm value of the projection vector, the more abstract the entity is, and the larger the value, the more specific the entity is. y (e j (k))、p y (e j+1 (k)) represent entities e j+1 、e j The embedding representation e j 、e j+1 The projection point of the k-th dimension component of the rotated complex plane coordinate axis on the imaginary axis. If e j+1 Yes j A hypernym of iff:

[0184]

[0185] For a vector X = (x(1), x(2)...x(N)), Represents each dimension of vector X. That is: e jThe L1 or L2 norm of the projection point of each component on the imaginary axis of the rotated complex plane coordinate axis is greater than e j+1 The L1 or L2 norm value of the projection point. y (e j (k)) can be vectorized as:

[0186]

[0187] In order to avoid different hyponyms corresponding to the same hypernym, or different hypernyms corresponding to the same hyponym, being mistakenly judged as having an isA relationship after training. For example, for the two triples (banana, isA, fruit) and (apple, isA, fruit), after training, banana and apple may satisfy equation (14) and be mistakenly judged as having an isA relationship. For this reason, a hyperparameter λ is added to distinguish entities at different conceptual levels. F That is: if e j+1 Yes j A hypernym of iff:

[0188]

[0189] For the isA relation triple (e j ,isA,e j+1 ), the distance function for antisymmetric modeling is defined as:

[0190]

[0191] where p y (z) is the projection operation, defined by formula (15). Function [z] + Only the positive part of z is retained, and the rest is 0, that is, [z] + =max{0,z}.

[0192] Therefore, the isA relationship triple is defined (e j ,isA,e j+1 ) The distance function is the sum of the transitivity modeling distance function and the antisymmetry modeling distance function, that is:

[0193]

[0194] In order to improve the distinguishability between positive and negative samples, the present invention uses a loss function based on marginal parameters as the optimization target for training. During the training process, a self-adversarial negative sampling method, which has been proven to be an effective optimization method for knowledge graph embedding, is used to generate a training negative sample set. This method samples negative triplets according to the current embedding model. The negative sampling sampling distribution is defined as:

[0195]

[0196] Among them, α is the sampling hyperparameter, d r (h k ',t k ') is the kth candidate negative sampling triplet (h k ',r,t k ') corresponds to the model distance function value.

[0197] For the relation triple (h, r, t), the loss function based on the marginal parameter is defined as:

[0198]

[0199] Among them, d r (h, t) represents the distance function of the relation triple (h, r, t), (h k ',r,t k ') is the kth negative sampling triplet (h among the m negative samples of (h,r,t) k ',r,t k '), p(h k ',r,t k ') represents the sampling probability of the negative sampling triplet defined by formula 19, γ r Margin hyperparameter used to represent relation triplets.

[0200] Relation triple set S R The total loss function is:

[0201]

[0202] Similarly, for the isA triple (e j ,isA,e j+1 ) The loss function L defined i for:

[0203]

[0204] Among them, d i (e j ,e j+1 ) represents a triple (e j ,isA,e j+1 ), (e' j+k ,r i ,e' j+1+k ) is (e j ,isA,e j+1 The kth negative sampling triplet (e') among the m negative samples j+k,isA,e' j+1+k ), p(e' j+k ,r i ,e' j+1+k ) represents the sampling probability of the negative sampling triplet defined by Formula 19, γ i Margin hyperparameter used to represent the isA triplet.

[0205] isA triple set S i The total loss function is:

[0206]

[0207] Finally, the overall loss function F is defined as a linear combination of the loss functions of these two triple sets, that is:

[0208] F=F R +β·F i (twenty four)

[0209] Among them, β>0, is F R 、F i Hyperparameters that maintain a balance between

[0210] During model training, in order to avoid overfitting, we enforce the L2 norm of the embedding representation of entities in all triplets to be less than or equal to 1, that is, ||h||2≤1, ||t||2≤1, ||e||2≤1.

[0211] Use N e 、N r Respectively represent the number of entities and relations in the data set, and n represents the dimension of the embedding representation space. For the set of relation triples, its parameter complexity is the parameter complexity of RotatE, which is O(2nN e +nN r ). Since the parameters of the embedding representation of the entity are shared throughout the model, compared with RotatE, the parameters of the isA triplet set in this model add a rotation angle vector θ of the complex plane coordinate axis p , which means the parameter complexity increases by 1. The overall parameter complexity of the model is O(2nN e +nN r +1). Therefore, the parameter complexity of the model is considered to be roughly the same as that of RotatE.

[0212] Next, the effectiveness of the present invention is evaluated on the triple classification task of knowledge graph embedding. First, the evaluation criteria of the task, the specific configuration of the experimental implementation and the corresponding experimental results are introduced, and then the experimental results are analyzed and compared with other methods. All experiments are completed on a Linux server with a hardware configuration of Intel(R) Xeon(R) Gold 5118CPU@2.30GHz processor, 128GB memory and an NVIDIA GeForce GTX 2080GPU.

[0213] The datasets YAGO39K and M-YAGO39K proposed in TransC and YAGO26K and DB111K proposed in JOIE are used to merge the subclassof and instanceof relations in these datasets into the isA relation.

[0214] Triple classification is a test task for the knowledge graph representation learning method, and its main task is to determine whether a given triple is "correct" or "wrong". The triple can be a relation triple or an isA triple. This is a binary classification task, and its evaluation indicators use the accuracy, precision, recall and F1 value commonly used in binary classification tasks. The present invention constructs the negative triples required for the triple classification task test according to the same settings as the NTN model.

[0215] 1) Experimental design. The triple set is divided into training set, validation set and test set, accounting for about 60%, 20% and 20% respectively. For each relation r in the relation triple data set, a threshold δ is set. r For a given test relation triple (h, r, t), calculate its distance function, that is, the value of formula (5). If its score function value is less than δ r , then the label of the triple is predicted to be “correct”, otherwise it is predicted to be “wrong”. Similarly, for the isA triple (x, r i ,c), if its distance function, that is, the score of formula (18) is less than the threshold δ ri , then it is predicted as “correct”, otherwise it is predicted as “wrong”. r and δ ri It is determined by maximizing the classification accuracy on the validation set.

[0216] 2) Comparison method

[0217] The comparison methods are divided into two categories: (1) Basic knowledge graph embedding methods, that is, embedding methods that do not introduce any external information. TransE and RotatE are selected as comparison methods of this category; (2) Knowledge graph representation learning models based on instances and concepts. TransC, CITS, JOIE, etc. are selected as comparison methods. Except for RotatE (there are no experimental results on the YAGO26K and DB111K datasets), the experimental results of all other comparison methods on link prediction are directly obtained from published papers (the experimental results of the two relations subclassof and instanceof in the dataset are merged into the experimental results of the isA relationship).

[0218] 3) Experimental implementation. The best configuration is determined by the accuracy of the validation set. During training, the learning rate λ used in the stochastic gradient descent method is λ = {0.0005, 0.001, 0.01, 0.1}, and the marginal hyperparameter γ F ,γ r ,γ i The value range is {0.1, 0.2, 0.5, 1, 2}, the embedding representation space dimension d ranges from {100, 200, 500, 1000}, and the weight β takes values ​​of {0.5, 1, 2}. The optimal parameters are determined by the Hits@10 of the validation set. The self-adversarial negative sampling parameter α∈{0.5, 1.0}. For each data set, this experiment iterates all training triplets for 1000 rounds.

[0219] Table 1: Classification results of relation triples. The best results are shown in bold.

[0220]

[0221] Compared with TransC and JOIE, the model of the present invention can better model the transitivity of the isA relationship, alleviate the problem of spatial aggregation of instance and concept embedding representations, and enable the model to achieve better results in the triple classification task.

[0222] The present invention proposes a new knowledge graph representation model - a knowledge graph representation model combined with isA relationship modeling. The model uses the isA relationship characteristics - transitivity and antisymmetry for modeling, that is, it uses the projection invariance of vectors to model the transitivity of the isA relationship, introduces partial order constraints to model the antisymmetry of the isA relationship, and introduces a hyperparameter for the entity to distinguish the different conceptual levels of the entity. Experimental results show that our method outperforms the existing baseline model in most cases, indicating that the model can more completely model the isA relationship, thereby alleviating the problem of insufficient embedding representation learning caused by sparse samples.

[0223] The beneficial effects of the present invention are as follows:

[0224] For the first time, a knowledge graph representation method combined with isA relationship modeling is proposed, which learns the embedded representation of isA relationship triples by modeling the two basic properties of isA relationship: transitivity and antisymmetry.

[0225] The model of the present invention can model the situation where the same hypernym corresponds to multiple different hyponyms or the same hyponym corresponds to multiple hypernyms. In addition, by introducing a hyperparameter for modeling entity hierarchy information, it is possible to avoid the misjudgment of the existence of an isA relationship between different hyponyms or hypernyms of the same level corresponding to the same entity.

[0226] Two tasks, link prediction and triple classification, were carried out on two benchmark datasets. The experimental results show that the present invention can alleviate the problem of insufficient embedding representation learning caused by sample sparsity.

[0227] As used herein, the word "preferred" is intended to be used as an example, instance, or illustration. Any aspect or design described herein as "preferred" is not necessarily to be construed as being more advantageous than other aspects or designs. On the contrary, the use of the word "preferred" is intended to present concepts in a specific way. The term "or" as used in this application is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X uses A or B" means any one of the naturally included permutations. That is, if X uses A; X uses B; or X uses both A and B, then "X uses A or B" is satisfied in any of the foregoing examples.

[0228] Moreover, although the present disclosure has been shown and described with respect to one or implementations, those skilled in the art will think of equivalent variations and modifications based on the reading and understanding of this specification and the accompanying drawings. The present disclosure includes all such modifications and variations, and is limited only by the scope of the appended claims. In particular, with respect to the various functions performed by the above-mentioned components (such as elements, etc.), the terms used to describe such components are intended to correspond to any component (unless otherwise indicated) that performs the specified function of the component (such as it is functionally equivalent), even if the structure is not equivalent to the disclosed structure of the function in the exemplary implementation of the present disclosure shown herein. In addition, although the specific features of the present disclosure have been disclosed with respect to only one of several implementations, such features can be combined with one or other features of other implementations that may be desired and advantageous for a given or specific application. Moreover, insofar as the terms "including", "having", "containing" or their variations are used in specific embodiments or claims, such terms are intended to be included in a manner similar to the term "comprising".

[0229] The functional units in the embodiments of the present invention may be integrated into a processing module, or each unit may exist physically separately, or multiple or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc. The above-mentioned devices or systems may execute the storage method in the corresponding method embodiment.

[0230] To sum up, the above embodiment is an implementation mode of the present invention, but the implementation mode of the present invention is not limited by the embodiment. Any other changes, modifications, substitutions, combinations, and simplifications that deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.

Claims

1. A knowledge graph triple classification method based on marginal parameter loss, characterized in that: The following steps are involved: Divide the triple set S into two disjoint subsets, namely the isA triple set S i and relation triple set S R , the isA triple set where n i YesS i The size of e i 、e j ∈E, whose embedding representations are is an n-dimensional embedding representation space; the relationship triple set where n r YesS R The size of h, t∈E, r∈R r , whose embeddings are expressed as Each of the above embedding representations is a complex vector; Different methods are used to analyze the isA triple set S i and relation triple set S R Modeling: For the relation triple (h, r, t), the RotatE model is used for modeling; The isA relationship chain is modeled by modeling the transitivity and antisymmetry of the isA relationship. After transitivity modeling, when the same hyponym corresponds to different hypernyms, these different hypernyms are modeled as different points on the same straight line. Combining the transitivity modeling method with the antisymmetry modeling, the modeling is performed by using the different projection points of each component of the entity embedding representation on the imaginary axis of the rotated complex plane. In order to improve the distinguishability between positive and negative samples, a loss function based on marginal parameters is used as the optimization target for training. The trained model is used to classify triples on the knowledge graph to determine whether a given triple is "correct" or "wrong"; The isA relationship is modeled as follows: Define the isA relationship chain (e1, e2..., e m ) represents the set of triples (e1, isA, e2), ..., (e m-2 ,isA,e m-1 ), (e m-1 ,isA,e m ), the tail entity of a triple in the set is the head entity of another triple, where e1,…,e m Its n-dimensional embedding representation vector, e j (k) represents entity e j The k-th dimension vector of (j=1,2,...,m); The isA relationship chain is modeled by modeling the transitivity and antisymmetry of the isA relationship. The transitivity modeling method is as follows: The isA relation is transitive, that is, if there are triples (AisAB), (BisA C), then there must be a triple (AisAC); For a triple (h, r, t), each relation r is modeled as a transformation operation T r , that is, after model training, it will satisfy e t =T r (e h ); for the isA relationship chain (e1, e2, e3), there are (e1, isA, e2), (e2, isA, e3), (e1, isA, e3), e1, e2, e3 are different entities, and the corresponding embedding representations e1, e2, e3 are different; if the transformation operation T can model the transitive relationship, that is, it satisfies e2=T(e1), e3=T(e2)=T(T(e1)), e3=T(e1), Conclusion T(e1)=T(T(e1)), That is, executing T operation multiple times on e1 is equivalent to executing T operation once, which shows that transitivity has the following property: the effect of transitivity acting on an entity multiple times is equivalent to the effect of acting on this entity once. Mathematically, this property is expressed by the linear transformation-projection transformation of a vector.

2. The knowledge graph triple classification method based on marginal parameter loss according to claim 1, characterized in that: The RotatE model maps entities and relations into an n-dimensional complex vector space. The relation r is modeled as a vector element-wise rotation from the head entity embedding to the tail entity embedding in the complex vector space, that is: et(k)=rot(eh(k),θr(k)) Among them, e h 、e h ,θ r They represent the embedding representations of the head entity h, the tail entity t, and the relation r, respectively, and are all n-dimensional complex vectors. h (k), e t (k),θ r (k) respectively represent e h 、e t ,θ r The k-th dimension component of , which are all one-dimensional complex vectors; Re(e h (k)) and Im(e h (k)) represents the complex number e h (k), then eh(k) = Re(eh(k)) + i·Im(eh(k)) can also be expressed as: [Re(eh(k) h (k)),Im(e h (k))] T , each component θ r The modulus of is constrained to be 1, that is, |θ r (k)|=1, so we can know that θ r It is expressed as: θ r The corresponding angle about the origin of the complex plane is θ r (k) counterclockwise rotation, rot(e h (k),θ r (k)) is the rotation function, which means that e h (k) Rotation θ r (k) radians; According to the coordinate rotation transformation formula, the point p(x,y) on the plane rotates counterclockwise about the origin by θ radians to reach p'(X,Y), and we get: Therefore, e t (k) = rot(e h (k),θ r (k)) is expressed as: For a relation triple (h, r, t), the distance function of the RotatE model is defined as:

3. The knowledge graph triple classification method based on marginal parameter loss according to claim 1, characterized in that: The linear transformation-projection transformation of the vector includes: e1(k), e2(k), e3(k) are the k-th components of e1, e2, e3, all of which are complex numbers and are represented by two-dimensional vectors on the complex plane: e1(k)=[Re(e1(k)),and(e1(k))] T e2(k)=[Re(e2(k)),Im(e2(k))] T ; e3(k)=[Re(e3(k)),Im(e3(k))] T Use the projection operation T to act on e1(k) to obtain T(e1(k)), that is, e1(k) is projected onto the real axis once, and the result is: T(e1(k))=[Re(e1(k)),0] T T(e2(k)) is obtained by projecting e2(k) onto the real axis once: T(e2(k))=[Re(e2(k)),0] T Since e2(k) = T(e1(k)), T(e2(k)) is the second projection of e1(k) onto the real axis, we get: T(e2(k))=T(T(e1(k)))=[Re(e1(k)),0] T Therefore, we get Re(e2(k))=Re(e1(k)) T(e3(k)) is the projection of e3(k) onto the real axis, and we get: T(e3(k))=[Re(e3(k)),0] T Since e3(k)=T(T(e1(k))), T(e3(k)) is the projection of e1(k) onto the real axis three times, we get T(e3(k))=T(T(T(e1(k))))=[Re(e1(k)),0] T Therefore, we get Re(e3(k))=Re(e1(k)) In summary, we can conclude Re(e3(k))=Re(e2(k))=Re(e1(k)) That is, the projection operation T is applied to e1(k), e2(k), and e3(k) to obtain the same projection point, which is d. Since the imaginary information Im(e3(k)), Im(e2(k)), and Im(e1(k)) of e1(k), e2(k), and e3(k) are different from each other, e1(k), e2(k), and e3(k) are different from each other. Therefore, the projection operation is used to model the transitivity of the isA relationship.

4. The knowledge graph triple classification method based on marginal parameter loss according to claim 3 is characterized in that: Projection transformation is equivalent to idempotent transformation, that is, if p(e) represents the projection transformation of the complex number e, then It represents the projection transformation result of e on the real axis of the complex plane. Therefore, the imaginary part information of p(e) is 0. That is, if the projection operation is only performed on the real axis, the imaginary part information of the projection points of entities in all different isA relationship chains will be 0, resulting in the loss of one dimension of information. To obtain a more general projection transformation, rotate the complex plane coordinate axis around the origin and rotate the complex plane coordinate axis counterclockwise by θ p Radians, get the real axis x' and imaginary axis y' after rotation, project e1(k), e2(k), e3(k) onto the real axis x' after rotation, and get their projection points in the new coordinate system are all p, that is, they satisfy T(e1(k))=T(e2(k))=T(e3(k)) The entities in different isA relationship chains are projected onto the real axis of the rotated complex plane coordinate axis, and the imaginary part information of each projection point is different. This allows the projection points to retain different imaginary part information while modeling the transitivity of the isA relationship.

5. The knowledge graph triple classification method based on marginal parameter loss according to claim 4 is characterized in that: The algebraic description of the transitivity modeling process of the isA relationship is as follows: The rotation vector of the projection coordinate axis is the n-dimensional complex vector θ p , for different transmission chains, θ p The values ​​are the same; θ p Each component θ p The modulus of (k) is constrained to be 1, i.e. |θ p (k)|=1, therefore, θ p (k) is expressed in the form of θ p (k) means rotating the complex plane coordinate axis counterclockwise about the origin by θ p (k) radians, 0 < θ p (k)≤2π; For the isA relationship chain (e1, e2…, e m ) on any entity embedding e j , e j (k) after rotating by θ p (k) The projection point p on the real axis of the complex plane after radians x (e j (k)) is expressed as: For any entity e in the isA relationship transfer chain j (k)(j=2,3,...,m), all have pp x (have been j (k))=p x (e1(k))(j=2,3,...,m) Combining the above formula, we get: If the above equation is true, then and only if the following equation is true: cosθp(k)Re(ej(k))+sinθp(k)Im(ej(k))=cosθp(k)Re(e1(k))+sinθp(k)Im(e1(k)) If the above formula is true for j=2,3,...,m, then it is true only if: cosθp(k)Re(ej(k))+sinθp(k)Im(ej(k))=ck Among them, c k Must be a constant, that is: ck=cosθp(k)Re(e1(k))+sinθp(k)Im(e1(k)) From the above formula, c k and θ p The slope of the line defined by (k) is -cot(θ p (k)); and all entities e1, e2, ..., e j ,…,e m , all satisfy the above formula, that is, e j (k)(j=1,2,3,...,m) are all generated by c k and θ p (k) is defined as -cot(θ p (k)), and the slope of the real axis of the rotated complex plane is tan(θ p (k)), it can be concluded that the straight line defined by the above formula is perpendicular to the real axis of the rotated complex plane, and the straight lines corresponding to different transfer chains are parallel to each other and perpendicular to the real axis of the rotated complex plane; For the isA relation triple (e j ,isA,e j+1 ), the distance function for transitivity modeling is defined as: where p x For projection operation.

6. The knowledge graph triple classification method based on marginal parameter loss according to claim 5, characterized in that: Model the antisymmetry of the isA relationship, including: The isA relationship is antisymmetric, that is, if the triple (aisAb) exists, then the triple (bisAa) must not exist; for the isA relationship chain (e1, e2…, e m ) in any triple (e j ,isA,e j+1 ), e j , e j+1 The corresponding embedding is denoted as e j 、e j+1 , e j+1 Yes j Hypernym of, or e j+1 E j More abstract, j E j+1 More specifically, based on the quantitative distinction j 、e j+1 Embedding representation j 、e j+1 , to model (e j+1 ,isA,e j ) must not exist; If the operation F can model the antisymmetry of the isA relation, that is, if e j+1 Yes j The hypernym of F(e j )≤F(e j+1 ), the transitivity is modeled by using the fact that each component of the entity embedding representation has the same projection point on the real axis of the rotated complex plane coordinate axis, that is, p x (e j (k)) = p x (e j+1 (k)); Project each component of the entity embedding representation onto the rotated imaginary axis of the complex plane coordinate axis. The smaller the L1 or L2 norm value of the projection vector, the more abstract the entity is, while the larger the value, the more specific the entity is. p y (e j (k))、p y (e j+1 (k)) represent entities e j+1 、e j The embedding representation e j 、e j+1 The projection point of the k-th dimension component of the rotated complex plane coordinate axis on the imaginary axis. If e j+1 Yes j A hypernym of iff: For a vector X = (x(1), x(2)...x(N)), Represents each dimension of vector X; that is: e j The L1 or L2 norm of the projection point of each component on the imaginary axis of the rotated complex plane coordinate axis is greater than e j+1 The L1 or L2 norm value of the projection point, then p y (e j (k)) is vectorized as: In order to avoid different hyponyms corresponding to the same hypernym, or different hypernyms corresponding to the same hyponym, being misjudged as having an isA relationship after training, a hyperparameter λ is added to distinguish entities at different conceptual levels. F , that is: if e j+1 Yes j A hypernym of iff: For the isA relation triple (e j ,isA,e j+1 ), the distance function for antisymmetric modeling is defined as: where p y (z) is the projection operation; function [z] + Only the positive part of z is retained, and the rest is 0, that is, [z] + =max{0,z}; Therefore, the isA relationship triple is defined (e j ,isA,e j+1 ) The distance function is the sum of the transitivity modeling distance function and the antisymmetry modeling distance function, that is:

7. The knowledge graph triple classification method based on marginal parameter loss according to claim 6, characterized in that: During the training process, a self-adversarial negative sampling method is used to generate a training negative sample set, which samples negative triplets according to the current embedding model; For the relation triple (h, r, t), the loss function based on the marginal parameter is defined as: Among them, d r (h, t) represents the distance function of the relation triple (h, r, t), (h k ',r,t k ') is the kth negative sampling triplet (h among the m negative samples of (h,r,t) k ',r,t k '), p(h k ',r,t k ') represents the sampling probability of the negative sampling triplet, γ r Margin hyperparameter used to represent relation triplets.

8. The knowledge graph triple classification method based on marginal parameter loss according to claim 7, characterized in that: Relation triple set S R The total loss function is: isA triple (e j ,isA,e j+1 ) The loss function L defined i for: Among them, d i (e j ,e j+1 ) represents a triple (e j ,isA,e j+1 ), (e' j+k ,r i ,e' j+1+k ) is (e j ,isA,e j+1 The kth negative sampling triplet (e') among the m negative samples j+k ,isA,e' j+1+k ), p(e' j+k ,r i ,e' j+1+k ) represents the sampling probability of the negative sampling triplet, γ i represents the marginal hyperparameter of the isA triplet; isA triple set S i The total loss function is: The overall loss function F is defined as a linear combination of the loss functions of these two triple sets, namely: F=F R +β·F i Among them, β>0, is F R 、F i Hyperparameters that maintain a balance between 9. The knowledge graph triple classification method based on marginal parameter loss according to claim 8, characterized in that: During model training, in order to avoid overfitting, the L2 norm of the embedding representation of entities in all triples is constrained to be less than or equal to 1, that is, ||h||2≤1, ||t||2≤1, ||e||2≤1.

Citation Information

Patent Citations

  • knowledge graph optimization method based on a fuzzy theory

    CN109840282A

  • Triple classification method based on improved concept and instance

    CN115168602A