Knowledge hypergraph link prediction method fusing external information and internal information

By integrating external and internal information, using technical means such as word vector model, hypergraph convolutional neural network and 3D convolutional neural network, the problem of knowledge hypergraph link prediction is solved, and more accurate and personalized prediction effects are achieved, and the understanding and utilization ability of complex relationship data is improved.

CN119940506APending Publication Date: 2025-05-06LIAONING UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510028958.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively predict and utilize knowledge hypergraph links in complex relationship data, resulting in insufficient understanding and utilization of facts of multiple relationships in the real world.

Method used

A knowledge hypergraph link prediction method that combines external information and internal information is adopted. Through technical means such as word vector model, hypergraph convolutional neural network and 3D convolutional neural network, external information and semantic information of relationships and entities are extracted and fused, and the scores of n-tuples are calculated for link prediction.

Benefits of technology

It realizes more personalized and accurate knowledge hypergraph link prediction, breaks through the performance bottleneck of traditional methods, and improves the understanding and utilization ability of complex relationship data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940506A_ABST
    Figure CN119940506A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge hypergraph link prediction method fusing external information and internal information. The method comprises the following steps of: 1, preprocessing data, embedding a relationship and an entity in a knowledge hypergraph data set by utilizing a word vector model, and converting the relationship and the entity into an initial feature matrix; 2, performing knowledge hypergraph n-tuple external information extraction according to the initial feature matrix and different relation importance degrees; 3, according to the relationship and entity embedding, converting the two-dimensional matrix into a two-dimensional matrix, connecting the two-dimensional matrix into a three-dimensional matrix, and performing semantic information extraction through a 3D convolutional neural network; 4, extracting sequence information of the entities according to the relation and entity embedding; 5, combining the external information and the internal information (semantic information and entity sequence information) as the final feature representation of the n-tuple, and calculating the score of the n-tuple; 6, designing a loss function training model; 7, the entity or the relation is inserted into the missing tuple, and the score of the n tuple is calculated according to the model; and 8, taking the tuple with the highest score as an optimal link prediction n tuple, or selecting several optimal n tuples as candidate tuples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of knowledge hypergraph link prediction, and in particular, relates to a knowledge hypergraph link prediction method that integrates external information and internal information. Background Art

[0002] Knowledge graph is a structured semantic knowledge base used to describe concepts and their relationships in the physical world, making the facts in reality more intuitive. Knowledge graph is mainly represented by triples, which are composed of head entity, relationship and tail entity. As the cornerstone of artificial intelligence, knowledge graph is widely used in search, recommendation system and other fields to assist in completing tasks. In the real world, although many facts can be represented by binary facts to form knowledge graph, knowledge hypergraph is more common in real life than knowledge graph. In the Freebase database, about 60% of the relationships are multi-dimensional relationships. Most relationships cannot be represented by two entities and one relationship. These relationship facts are usually composed of a relationship and multiple entities with sequential relationships in the relationship. Knowledge hypergraph is thus generated.

[0003] Knowledge hypergraphs have become increasingly ubiquitous in diverse domains, primarily due to the widespread existence of n-ary relational facts in the real world. These facts are demonstrated in complex protein structures consisting of multiple interacting elements, intricate social networks involving numerous individuals, and a variety of other multifaceted systems. The inherent complexity and richness of these n-ary relations require advanced representation frameworks, such as knowledge hypergraphs, to efficiently capture and model the underlying relational structures. As a result, the task of knowledge hypergraph link prediction, which attempts to infer and predict the underlying knowledge links in these hypergraphs, has become a daunting yet essential challenge. This problem is of critical importance as it has the potential to significantly improve our understanding and exploitation of complex relational data. The ability to accurately predict links in knowledge hypergraphs could lead to breakthroughs in many fields, including bioinformatics, social network analysis, and others. Therefore, developing robust and efficient knowledge hypergraph link prediction methods remains a critical research area that requires innovative approaches and effective algorithms. Summary of the invention

[0004] In order to solve the problems existing in the above-mentioned prior art, the present invention provides a knowledge hypergraph link prediction method that integrates external information and internal information to provide a more personalized and accurate knowledge hypergraph link prediction method.

[0005] The technical solution adopted by the present invention is as follows:

[0006] A knowledge hypergraph link prediction method integrating external information and internal information, the steps of which are:

[0007] Step 1: Data preprocessing: Use the word vector model word2vec to embed the relations and entities in the n-tuples of the knowledge hypergraph dataset and convert them into an initial feature matrix for extracting external information.

[0008] Step 101: Number each entity and relationship in the data;

[0009] Step 102: Input each tuple into the word vector model for training to obtain the initial vector representation of the relationship and entity;

[0010] Step 103: The initial vector representations of the relationship and entity are combined into an initial feature matrix.

[0011] Step 2: Extract external information of knowledge hypergraph n-tuples based on the initial feature matrix and the importance of different relationships.

[0012] Step 201: Determine the importance of each relationship based on the number of entities connected to it;

[0013] Step 202: taking the first few relations as important relations according to the importance of different relations and the number of relations in the data set;

[0014] Step 203: Initialize virtual nodes for important relationships, and perform hypergraph convolutional neural network deep extraction of relationship information between virtual nodes and relationships;

[0015] Step 204: A hypergraph convolutional neural network is then performed between the relationship and the entity to extract external information of the relationship and the entity in the tuple.

[0016] Step 3: According to the relationship and entity embedding, convert them into a two-dimensional matrix and concatenate them into a three-dimensional matrix, and extract semantic information through a 3D convolutional neural network.

[0017] Step 301: Initially embed the relationship and entity to obtain a one-dimensional feature vector;

[0018] Step 302: linearly transform the one-dimensional embeddings of the relationship and entity to obtain a two-dimensional matrix;

[0019] Step 303: Connect the two-dimensional matrix to obtain a three-dimensional matrix;

[0020] Step 304: define a 3D convolution kernel, and perform 3D convolution on the three-dimensional matrix to obtain tuple semantic information.

[0021] Step 4: Use the sequence model and leverage relation and entity embeddings to extract the sequence information of entities.

[0022] Step 401: Initially embedding the relationship and entity to obtain a one-dimensional feature vector;

[0023] Step 402: embed multiple entities in the tuple for connection;

[0024] Step 403: Input the multi-entity embedding and relations into the Mamba sequence model to extract the positive entity sequence information and integrate the entity potential relationship information;

[0025] Step 404: Reverse the entity embedding order and input it into the Mamba sequence model to further extract reverse sequence information;

[0026] Step 405: Add the forward and reverse sequence information to obtain the final entity overall embedding, and perform 2D convolution using the convolution kernel obtained by the relational linear transformation to obtain entity sequence information.

[0027] Step 5: Combine external information and internal information (semantic information and entity sequence information) as the final feature representation of the n-gram and calculate the n-gram score.

[0028] Step 501: Combine external information, semantic information and entity sequence information to obtain the final n-gram feature representation;

[0029] Step 502: Calculate the score of the n-tuple according to the feature representation of the n-tuple.

[0030] Step 6: Design loss function to train the model.

[0031] Step 601: Design a loss function that is consistent with the knowledge hypergraph score;

[0032] Step 602: Train the model until the loss function value becomes stable and as small as possible.

[0033] Step 7: Insert entities or relations into missing tuples and calculate n-gram scores based on the model.

[0034] Step 701: For the knowledge hypergraph n-tuples with missing entities or relations, we insert each entity or relation into the missing position to form a new n-tuple;

[0035] Step 702: Use the new n-gram as model input and calculate the score of the new n-gram.

[0036] Step 8: Take the tuple with the highest score as the optimal link prediction n-tuple, or select the best n-tuples as candidate tuples.

[0037] Step 801: Arrange in ascending order according to the n-tuple scores;

[0038] Step 802: Take the best score tuple as the best link prediction result or take the TOP-k tuples as the best candidate tuples.

[0039] The beneficial effects of the present invention are:

[0040] 1. This paper proposes a groundbreaking model that can effectively perform knowledge hypergraph link prediction and break through performance bottlenecks;

[0041] 2. This paper proposes a hypergraph convolutional neural network to extract external structural information by introducing virtual nodes, and proposes a customized hypergraph Mamba model to extract sequential relationship information;

[0042] 3. In order to further improve the performance of the model, the present invention proposes two negative sampling methods for knowledge hypergraph link prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a flow chart of the method of the present invention;

[0044] Figure 2 This is a structural diagram of the model recommended by the present invention. DETAILED DESCRIPTION

[0045] The technical solutions in the embodiments of the present invention are described in detail below in conjunction with the accompanying drawings. The following embodiments are applicable to the present invention, but are not intended to limit the scope of the present invention.

[0046] like Figure 1 As shown, the present invention proposes a knowledge hypergraph link prediction method that integrates external information and internal information. The model structure of the invention is as follows Figure 2 As shown, the following steps are included:

[0047] Step 1: Data preprocessing: Use the word vector model word2vec to embed the relations and entities in the n-tuples of the knowledge hypergraph dataset and convert them into an initial feature matrix for extracting external information.

[0048] Step 101: Extract the relations and entities in the facts and represent them as r(e 1 ,e 2 ,...,e k ), where r represents the relationship designed by the fact, e i represents the i-th entity designed in the relationship fact, and k represents the number of entities involved in the relationship fact. Number each entity and relationship in the data to get its corresponding token, which is convenient for embedding and inputting into the model;

[0049] Step 102: Input each tuple into the word vector model CBOW for training to obtain the initial vector representation of the relationship and entity;

[0050] Step 103: The initial vector representations of the relationship and entity are combined into an initial feature matrix H.

[0051] Step 2: Extract external information of knowledge hypergraph n-tuples based on the initial feature matrix and the importance of different relationships.

[0052] Step 201: Since the frequency of occurrence of each relationship in the knowledge hypergraph and the number of entities involved vary greatly, the importance of each relationship is determined based on the number of entities connected to it;

[0053] Step 202: Sort different relations according to their importance and select the first n / ln(n+1) relations as important relations according to the number of relations in different data sets, where n represents the number of relations in the data set;

[0054] Step 203: Initialize virtual nodes for important relationships and initialize virtual node feature matrix V, and perform hypergraph convolutional neural network deep extraction of relationship information between virtual nodes and relationships:

[0055]

[0056] H vr =GCN([V||H r ])

[0057] Among them, GCN() represents graph convolution, D represents degree matrix, A represents adjacency matrix, H represents feature matrix, W represents graph convolution weight, H represents vr The feature matrix after the information aggregation between the relationship and the virtual point is represented;

[0058] Step 204: extract the relationship features and perform a hypergraph convolutional neural network between the relationship and the entity to extract the external information of the relationship and the entity in the tuple:

[0059] H'=GCN(H)

[0060] I=r,e 1 ,e 2 ,...,e n

[0061] [r,e 1 ,e 2 ,...,e n ]=H' I,:

[0062] p 1 =maxpool([r (hgcn) ||e 1(hgcn) ||...||e k(hgcn) ])

[0063] Among them, H' represents the feature matrix after aggregating information between relations and entities, I represents the token sequence of relations and entities, and H' I,:Indicates that the features corresponding to the n-tuples are taken out according to I, p 1 Represents external information.

[0064] Step 3: According to the relationship and entity embedding, convert them into a two-dimensional matrix and concatenate them into a three-dimensional matrix, and extract semantic information through a 3D convolutional neural network.

[0065] Step 301: Initially embed the relationship and entity to obtain a one-dimensional feature vector

[0066] Step 302: Linearly transform the one-dimensional embeddings of relations and entities to obtain a two-dimensional matrix where d 1 ×d 2 =d;

[0067] Step 303: Connect the two-dimensional matrix to obtain a three-dimensional matrix

[0068] Step 304: define the 3D convolution kernel w, perform 3D convolution on the three-dimensional matrix, and perform operations such as maxpool to obtain tuple semantic information:

[0069]

[0070] Among them, p 2 Represents semantic information.

[0071] Step 4: Use the sequence model and leverage relation and entity embeddings to extract the sequence information of entities.

[0072] Step 401: Initially embed the relationship and entity to obtain a one-dimensional feature vector

[0073] Step 402: embed multiple entities in the tuple for connection;

[0074] Step 403: Input the multi-entity feature vector and the relationship feature vector together into the Mamba sequence model to extract the positive entity sequence information and integrate it into the entity potential relationship information:

[0075]

[0076] in, Used to affect the previous hidden state h n-1 , Affects entity embedding, Influence relation embedding, h n Indicates the current hidden state. Used to affect the current hidden state, represents the residual link;

[0077] Step 404: The entity embedding order is then reversed and input into the Mamba sequence model to further extract reverse sequence information:

[0078]

[0079] Step 405: Add the forward and reverse sequence information to obtain the final entity overall embedding, perform 2D convolution using the convolution kernel obtained by linear transformation of the relational embedding, and obtain entity sequence information through maxpool operation:

[0080] y i =y(backward) i +y(forward) i

[0081] w r =vec -1 (r.W 1 )

[0082] m ri =y i *w r =y i *vec -1 (r.W 1 )

[0083] p 3 =maxpool(vec([m r1 ||m r2 ||...||m rk ]))W 2

[0084] Among them, W 1 represents the linear transformation weight matrix, w r represents the convolution kernel after the relational embedding linear transformation, vec represents vectorization, and m ri represents the entity embedding after the convolution to extract the surface relationship information of the previous entity embedding, p 3 Represents entity sequence information.

[0085] Step 5: Combine external information and internal information (semantic information and entity sequence information) as the final feature representation of the n-gram and calculate the n-gram score.

[0086] Step 501: Combine external information, semantic information and entity sequence information to obtain the final n-gram feature representation;

[0087] Step 502: Calculate the score of the tuple according to the n-tuple feature representation:

[0088] score = g(p 1 +p 2 +p3 )

[0089] Step 6: Design loss function to train the model.

[0090] Step 601: Design a loss function that is consistent with the knowledge hypergraph score:

[0091]

[0092] Among them, f() represents the n-tuple score, φ is the regularization term, Represents an n-tuple label, and its values ​​are as follows:

[0093]

[0094] Step 602: Train the model until the loss function value tends to be stable and the evaluation index result of the model on the validation set cannot be improved, and take the model at this time as the optimal model result.

[0095] Step 7: Insert entities or relations into missing tuples and calculate n-gram scores based on the model.

[0096] Step 701: For the knowledge hypergraph n-tuple r(e) of missing entities or relations 1 ,? ,...,e n ) or? (e 1 ,e 2 ,...,e n ), insert each entity or relation into the missing position to form a new n-gram;

[0097] Step 702: Use the new n-gram as model input and calculate the score of the new n-gram.

[0098] Step 8: Take the tuple with the highest score as the optimal link prediction n-tuple, or select the best n-tuples as candidate tuples.

[0099] Step 801: Arrange in ascending order according to the n-tuple scores;

[0100] Step 802: Take the best score tuple as the best link prediction result or take the TOP-k tuples as the best candidate tuples.

Claims

1. A knowledge hypergraph link prediction method integrating external information and internal information, characterized in that: The steps are: Step 1: Data preprocessing: Use the word vector model word2vec to embed the relations and entities in the n-tuples of the knowledge hypergraph dataset and convert them into an initial feature matrix for extracting external information. Step 2: Extract external information of knowledge hypergraph n-tuples based on the initial feature matrix and the importance of different relationships; Step 3: According to the relationship and entity embedding, convert it into a two-dimensional matrix and connect it into a three-dimensional matrix, and extract semantic information through a 3D convolutional neural network; Step 4: Use the sequence model and leverage relation and entity embeddings to extract the sequence information of entities; Step 5: Combine the external information and the internal information as the final feature representation of the n-gram and calculate the n-gram score; Step 6: Design loss function training model; Step 7: Insert the entity or relation into the missing tuple and calculate the n-tuple score based on the model obtained in step 6; Step 8: Take the tuple with the highest score as the optimal link prediction n-tuple, or select the best n-tuples as candidate tuples.

2. A knowledge hypergraph link prediction method integrating external information and internal information according to claim 1, characterized in that: In the step 1, the specific method is: Step 101: Extract the relations and entities in the facts and represent them as r(e1,e2,...,e k ), where r represents the relationship designed by the fact, e i represents the i-th entity designed in the relationship fact, k represents the number of entities involved in the relationship fact, and each entity and relationship in the data is numbered to obtain its corresponding token; Step 102: Input each tuple into the word vector model CBOW for training to obtain the initial vector representation of the relationship and entity; Step 103: The initial vector representations of the relationship and entity are combined into an initial feature matrix H.

3. The method for predicting knowledge hypergraph links by integrating external information and internal information according to claim 1, characterized in that: In the step 2, the specific method is: Step 201: Determine the importance of each relationship based on the number of entities connected to it; Step 202: taking the first few relations as important relations according to the importance of different relations and the number of relations in the data set; Step 203: Initialize virtual nodes for important relationships, and perform hypergraph convolutional neural network deep extraction of relationship information between virtual nodes and relationships; Step 204: A hypergraph convolutional neural network is then performed between the relationship and the entity to extract external information of the relationship and the entity in the tuple.

4. A knowledge hypergraph link prediction method integrating external information and internal information according to claim 3, characterized in that: In the steps 203 and 204, the specific method is: Step 203: Initialize virtual nodes for important relationships and initialize virtual node feature matrix V, and perform hypergraph convolutional neural network deep extraction of relationship information between virtual nodes and relationships: H vr =GCN([V||H r ]) Among them, GCN() represents graph convolution, D represents degree matrix, A represents adjacency matrix, H represents feature matrix, W represents graph convolution weight, H represents vr The feature matrix after the information aggregation between the relationship and the virtual point is represented; Step 204: extract the relationship features and perform a hypergraph convolutional neural network between the relationship and the entity to extract the external information of the relationship and the entity in the tuple: H'=GCN(H) I=r,e1,e 2, ...,have been n [r,e1,e 2, ...,have been n ]=H' I,: p1=maxpool([r (hgcn) ||and 1(hgcn) ||…||and k(hgcn) ]) Among them, H' represents the feature matrix after aggregating information between relations and entities, I represents the token sequence of relations and entities, and H' I,: It means taking out the features corresponding to the n-tuple according to I, and p1 represents external information.

5. The method for predicting knowledge hypergraph links by integrating external information and internal information according to claim 1, characterized in that: In the step 3, the specific method is: Step 301: Initially embed the relationship and entity to obtain a one-dimensional feature vector; Step 302: linearly transform the one-dimensional embeddings of the relationship and entity to obtain a two-dimensional matrix; Step 303: Connect the two-dimensional matrix to obtain a three-dimensional matrix; Step 304: define a 3D convolution kernel, and perform 3D convolution on the three-dimensional matrix to obtain tuple semantic information.

6. The method for predicting knowledge hypergraph links by integrating external information and internal information according to claim 1, characterized in that: In the step 4, the specific method is: Step 401: Initially embedding the relationship and entity to obtain a one-dimensional feature vector; Step 402: embed multiple entities in the tuple for connection; Step 403: Input the multi-entity embedding and relations into the Mamba sequence model to extract the positive entity sequence information and integrate the entity potential relationship information; Step 404: Reverse the entity embedding order and input it into the Mamba sequence model to further extract reverse sequence information; Step 405: Add the forward and reverse sequence information to obtain the final entity overall embedding, and perform 2D convolution using the convolution kernel obtained by the relational linear transformation to obtain entity sequence information.

7. A knowledge hypergraph link prediction method integrating external information and internal information according to claim 6, characterized in that: In the step 4, the specific method is: Step 401: Initially embed the relationship and entity to obtain a one-dimensional feature vector Step 402: embed multiple entities in the tuple for connection; Step 403: Input the multi-entity feature vector and the relationship feature vector together into the Mamba sequence model to extract the positive entity sequence information and integrate it into the entity potential relationship information: in, Used to affect the previous hidden state h n-1 , Affects entity embedding, Influence relation embedding, h n Indicates the current hidden state. Used to affect the current hidden state, represents a residual link. Step 404: The entity embedding order is then reversed and input into the Mamba sequence model to further extract reverse sequence information: Step 405: Add the forward and reverse sequence information to obtain the final entity overall embedding, perform 2D convolution using the convolution kernel obtained by linear transformation of the relational embedding, and obtain entity sequence information through maxpool operation: y i =y(backward) i +y(forward) i w r =vec -1 (r·W1) m ri =y i *w r =y i *vec -1 (r·W1) p3=maxpool(vec([m r1 ||m r2 ||...||m rk ]))W2 Among them, W1 represents the linear transformation weight matrix, w r represents the convolution kernel after the relational embedding linear transformation, vec represents vectorization, and m ri It represents the entity embedding after the convolution operation to extract the surface relationship information of the entity embedding in the previous step, and p3 represents the entity sequence information.

8. The method for predicting knowledge hypergraph links by integrating external information and internal information according to claim 1, characterized in that: In the step 6, the specific method is: Step 601: Design a loss function that is consistent with the knowledge hypergraph score: Among them, f() represents the n-tuple score, φ is the regularization term, Represents an n-tuple label, and its values ​​are as follows: Step 602: Train the model until the loss function value tends to be stable and the evaluation index result of the model on the validation set cannot be improved, and take the model at this time as the optimal model result.

9. The method for predicting knowledge hypergraph links by integrating external information and internal information according to claim 1, characterized in that: In the step 7, the specific method is: Step 701: for the knowledge hypergraph n-tuples with missing entities or relations, insert each entity or relation into the missing position to form a new n-tuple; Step 702: Use the new n-gram as model input and calculate the score of the new n-gram.

10. The method for predicting knowledge hypergraph links by integrating external information and internal information according to claim 1, characterized in that: In the step 8, the specific method is: Step 801: Arrange in ascending order according to the n-tuple scores; Step 802: Take the best score tuple as the best link prediction result or take the TOP-k tuples as the best candidate tuples.