BERT-based neighborhood relation learning knowledge graph completion method and device
By extracting first-order relationship subgraphs from the knowledge graph and using the BERT model for embedded learning, the problem of not being able to fully utilize the neighborhood information of triples and capturing complex semantic information in the prior art is solved, and more accurate knowledge graph completion and relationship prediction effects are achieved.
Patent Information
- Application Number
- CN202510102000.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art cannot fully utilize the neighborhood information of triples in knowledge graph completion to capture complex semantic information, especially in small sample learning scenarios, making it difficult to effectively learn and predict the potential relationship between triples.
The neighborhood relationship learning method based on BERT is adopted to extract the first-order relationship subgraph of the target triplet in the knowledge graph, and the target triplet and its adjacency context relationship are input into the BERT model for embedding learning. Through the mask language model and the next sentence prediction task, the left and right context information of the vocabulary is learned, and the relationship prediction is finally based on the final hidden vector of the BERT model.
The neighborhood relationship information of triples is effectively utilized, the accuracy of knowledge graph completion is improved, the generalization ability in small sample scenarios is improved, and the effect of triple relationship prediction is significantly improved.
Smart Images

Figure CN120046709A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of knowledge graph reasoning, and more specifically, to a knowledge graph completion method and device for neighborhood relationship learning based on BERT. Background Art
[0002] Knowledge graphs play a crucial role in tasks such as question answering systems, semantic reasoning, sentiment analysis, and recommendation systems. However, many existing knowledge graphs are usually incomplete, which makes the knowledge graph completion task (KGC) indispensable in many application scenarios. To address the incompleteness of knowledge graphs, embedding-based transductive completion methods have achieved good results in some KGC benchmarks. The basic idea of such methods is to embed entities and relationships in a knowledge graph into a continuous vector space and use translation models, bilinear models, and neural network-based models for reasoning. This method is mainly applied to static knowledge graphs and has made remarkable progress in many tasks.
[0003] However, transductive completion methods also have obvious limitations. First, such methods usually can only handle known entities and relationships, and for dynamic knowledge graphs with constantly changing, emerging new entities and relationships, the retraining cost is high. In addition, these methods perform poorly in dealing with interpretability, few-shot learning, and transfer learning, and it is difficult to adapt to more complex application scenarios. Inductive completion methods have emerged to address the limitations of transductive methods. Different from transductive methods that rely on the closed-world assumption, inductive methods have stronger generalization ability and can infer unseen entities and relationships. Subgraph-based inductive methods and rule-based methods partially solve the problem of dealing with new entities and relationships by learning inference rules from local structures.
[0004] Although current inductive completion methods show good generalization ability in dealing with unseen entities and relationships, there are still two main problems. First, many methods do not fully utilize the neighborhood information of triples. In fact, in the process of entity and relationship modeling, the neighborhood relationships of triples play a crucial role. The adjacent context relationships of triples provide rich semantic information for entity and relationship modeling, helping to capture the associations between entities more accurately. Specifically, the adjacent context relationships of triples reveal the specific positions and roles of entities in the knowledge graph by examining the surrounding environments of the head and tail entities. For example, when considering a specific head entity and the relationships it emits, the neighborhood relationships can help us understand in what contexts this head entity tends to establish connections with which tail entities. This contextual information is crucial for distinguishing different types of relationships because it can help us identify the unique meanings of relationships in specific contexts. When observing the tail entity and the relationships it receives, the neighborhood relationships are equally important. It reveals the functions of the tail entity in the knowledge graph and how it is related to other entities. This information helps us better understand the attributes of the tail entity and its position in the relationship network.
[0005] Second, traditional encoding methods cannot fully capture complex semantic information, especially in the few-shot learning scenario where they perform limitedly. To address these problems, existing research has attempted to introduce the pre-trained language model BERT into the knowledge graph completion task. For example, the BERTRL model linearizes triples and their local subgraphs into text sequences by leveraging BERT, enabling the model to learn deeper semantic representations from text information. BERT has powerful language understanding capabilities, which can enhance the representations of entities and relationships and improve the generalization ability of the model in the few-shot scenario.
[0006] Generally speaking, current inductive knowledge graph completion models still have deficiencies in fully utilizing triple neighborhood information, capturing complex semantic information, and in the few-shot learning scenario, and are unable to learn and predict potential relationships between triples well. Summary of the Invention
[0007] In view of at least one defect or improvement requirement of the prior art, the present invention provides a BERT-based neighborhood relationship learning knowledge graph completion method and device, which solves the problems of currently being unable to fully utilize triple neighborhood information, capture complex semantic information, and being unable to learn and predict potential relationships between triples well, can effectively utilize neighborhood relationship information, improve the accuracy of knowledge graph completion, and achieves the purpose of enhancing the effect of entity relationship prediction in the knowledge graph.
[0008] To achieve the above object, according to the first aspect of the present invention, a knowledge graph completion method for learning neighborhood relationships based on BERT is provided, and the method includes: extracting a first-order relational subgraph for a target triple in the knowledge graph, where the relational subgraph includes the out-edges and in-edges of the target entity; inputting the target triple and the adjacent context relationships of the head and tail entities corresponding to the target triple into a BERT model for embedding learning, where the BERT model learns the left and right context information of the vocabulary through two pre-training tasks of masked language model and next sentence prediction; setting a prediction function for the target triple, and obtaining a relationship prediction result after performing relationship prediction based on the final hidden vector of the BERT model; and completing the knowledge graph based on the obtained relationship prediction result.
[0009] In an exemplary embodiment, the extracting a first-order adjacent relational subgraph for a target triple in the knowledge graph includes: for each target entity, constructing a first-order relational subgraph of the target entity by using a breadth-first search algorithm, and the first-order relational subgraph is expressed as
[0010]
[0011] where r j represents a certain type of relationship, m represents the number of extracted adjacent relationship contexts, and e is each target entity.
[0012] In an exemplary embodiment, the inputting the target triple and the adjacent context relationships of the head and tail entities corresponding to the target triple into a BERT model for embedding learning includes: determining a first text sequence as the text sequence representing the head entity, the tail entity, and the relationship between the head and tail entities; determining a second text sequence as the text sequence representing the first-order relational subgraph of the head entity; determining a third text sequence as the text sequence representing the first-order relational subgraph of the tail entity; and determining the original input of the BERT model based on the first text sequence, the second text sequence, and the third text sequence.
[0013] In an exemplary embodiment, after determining the original input of the BERT model based on the first text sequence, the second text sequence, and the third text sequence, the method further includes: determining a final input based on the original input, and the final input is expressed as
[0014] x i = E tokens (w i ) + E segments (w i ) + E positions (i)
[0015] where E tokens(w i ) is the token vector of word w i , and E segments (w i ) is the segment vector corresponding to word w i , which is used to distinguish different sentences or sequences; E positions (i) is the position vector of position i, indicating the position of word w i in the sequence.
[0016] In an exemplary embodiment, for the prediction function of setting the target triple, after the relationship prediction is performed based on the final hidden vector of the BERT model, the relationship prediction result includes: using the binary cross-entropy loss function to calculate the difference between the predicted value of the BERT model and the actual label, and the calculation formula is:
[0017]
[0018] where the label set is y τ ∈ {0, 1}, 0 represents a negative example, and 1 represents a positive example; the prediction function is a two-dimensional entity vector, S τ0 , S τ1 ∈ [0, 1], and S τ0 + S τ1 = 1.
[0019] According to the second aspect of the present invention, there is also provided a knowledge graph completion device for neighborhood relationship learning based on BERT, which includes: an extraction unit for extracting a first-order relationship subgraph of a target triple in the knowledge graph, where the relationship subgraph includes the out-edges and in-edges of the target entity; a learning unit for inputting the target triple and the adjacency context relationship of the head and tail entities corresponding to the target triple into the BERT model for embedding learning, where the BERT model learns the left and right context information of the vocabulary through two pre-training tasks of the masked language model and the next sentence prediction; a prediction unit for setting a prediction function for the target triple, and obtaining a relationship prediction result after performing relationship prediction based on the final hidden vector of the BERT model; and a completion unit for completing the knowledge graph based on the obtained relationship prediction result.
[0020] According to the third aspect of the present invention, there is also provided a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the above-mentioned knowledge graph completion method for neighborhood relationship learning based on BERT when running.
[0021] According to the fourth aspect of the present invention, an electronic device is further provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. Wherein, the above-mentioned processor executes the above-mentioned BERT-based neighborhood relationship learning knowledge graph completion method through the computer program.
[0022] Generally speaking, compared with the prior art by the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0023] (1) The present application proposes a BERT-based neighborhood relationship learning knowledge graph completion method. By introducing the extraction of the first-order adjacency relationship subgraph of the target triple and combining the fine-tuning of the pre-trained BERT model, the neighborhood relationship information is fully utilized, and the accuracy of knowledge graph reasoning is improved. Different from the traditional method of realizing reasoning based on a closed subgraph, the neighborhood context relationship information is taken into consideration to construct a first-order relationship subgraph, enabling the full use of limited valuable information resources for reasoning. The method of the present invention not only improves the effect of knowledge graph completion, but also provides a new idea for dealing with new entities and relationships in dynamic knowledge graphs;
[0024] (2) This method first extracts the first-order adjacency relationship subgraph of the given target triple, then inputs the target triple and its adjacency context relationship into the pre-trained BERT model for fine-tuning, and finally predicts the user's interest in the item through a prediction function. The performance of this method on the FB15K-237 and WN18RR data sets is better than that of existing mainstream models. Especially when dealing with sparse subgraphs, it can significantly improve the accuracy of relationship prediction. The experimental results on the benchmark data sets WN18RR and FB15k-237 show that BERTNR is significantly better than existing methods in the knowledge graph completion task. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0026] Figure 1 It is a schematic flowchart of an optional BERT-based neighborhood relationship learning knowledge graph completion method provided by an embodiment of the present application;
[0027] Figure 2 It is a schematic diagram of the extraction of the first-order relationship subgraph of an optional node provided by an embodiment of the present application;
[0028] Figure 3Schematic diagram of an optional relational sub-graph extraction process provided by an embodiment of the present application;
[0029] Figure 4 Schematic diagram of a process for forming entity embedding representations using entities and their first-order neighborhood information in a knowledge graph by means of BERT provided by an embodiment of the present application;
[0030] Figure 5 Schematic diagram of the overall model framework adopted by the present application;
[0031] Figure 6 Schematic diagram of the structure of an optional BERT-based neighborhood relationship learning knowledge graph completion device provided by an embodiment of the present application;
[0032] Figure 7 Schematic diagram of the structure of an optional electronic device provided by an embodiment of the present application. Detailed implementation manners
[0033] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0034] The terms "first", "second", "third", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0035] According to one aspect of an embodiment of the present application, a BERT-based neighborhood relationship learning knowledge graph completion method is provided. The following combines Figure 1 Describe the BERT-based neighborhood relationship learning knowledge graph completion method provided by the embodiments of the present application.
[0036] Figure 1 Is a schematic diagram of a process of an optional BERT-based neighborhood relationship learning knowledge graph completion method provided by an embodiment of the present application. As Figure 1 shown, the process of this method may include the following steps:
[0037] S102, extract the first-order relational subgraph of the target triple in the knowledge graph, where the relational subgraph includes the out-edges and in-edges of the target entity;
[0038] S104, input the target triple and the adjacent context relationships of the head and tail entities corresponding to the target triple into the BERT model for embedding learning, where the BERT model learns the left and right context information of the vocabulary through two pre-training tasks: masked language model and next sentence prediction;
[0039] S106, set the prediction function of the target triple, and the prediction function obtains the relationship prediction result after performing relationship prediction based on the final hidden vector of the BERT model;
[0040] S108, complete the knowledge graph based on the obtained relationship prediction result.
[0041] A knowledge graph completion method based on BERT's neighborhood relationship learning provided by this application is applicable to various downstream task scenarios, including but not limited to scenarios involving recommendation systems, question answering systems, semantic reasoning, sentiment analysis, etc. In this example, taking a book recommendation system as an example, it details how to apply the knowledge graph completion method based on BERT's neighborhood relationship learning of the present invention. For example, when the goal is to recommend books that a user may be interested in, this requires extracting the first-order adjacent relationship subgraph of the user and the book from the knowledge graph and using this information to enrich the embedding representations of the user and the book.
[0042] Exemplarily, for a given knowledge graph, for the target triple, perform subgraph extraction of the head and tail entities and learning of the target relationship extraction, adopt a random negative sample generation strategy to generate negative samples corresponding to the target triple as the positive sample. Use the BERT model to perform embedding learning on the triple and its adjacent context relationships, fully capture the rich information in the knowledge graph, and obtain the final hidden vector [CLS]. Take the obtained final hidden vector [CLS] as the input to the classification layer, and regard this task as a binary classification task to generate the final relationship prediction result.
[0043] Preferably, in the complex network of the knowledge graph, the neighborhood relationships around the nodes often contain rich semantic content, including the attributes, categories of the nodes, and their associations with other nodes. Therefore, extracting the first-order relational subgraph of the nodes is a key step in understanding their context relationships. Specifically, assume that the neighborhood relationships around the target node contain the information required to infer the target relationship. As a first step, a relational subgraph surrounding the target entity is defined, and this relational subgraph is composed of the out-edges and in-edges of the entity. As a first step, define a relational subgraph around the target entity, which consists of the out-edges and in-edges of the entity, first-order and second-order subgraphs. This invention takes the first-order relational subgraph of the main entity as an example for illustration.
[0044] By extracting these subgraphs, the relationships of adjacent contexts can be effectively captured, and the complex connections between nodes can be deeply revealed. The first-order relational subgraphs of entities are obtained using the BFS algorithm These relational subgraphs contain rich adjacent context information. To effectively utilize this information, the BERT model can be selected to embed the target triples and their adjacent context relational subgraphs
[0045] Through the above steps S102 to S108, by extracting the first-order relational subgraphs of the target triples in the knowledge graph, where the relational subgraphs include the out-edges and in-edges of the target entities; inputting the target triples and the adjacent context relationships of the head and tail entities corresponding to the target triples into the BERT model for embedding learning, where the BERT model learns the left and right context information of the vocabulary through two pre-training tasks: masked language model and next sentence prediction; setting the prediction function for the target triples, and obtaining the relationship prediction result after the relationship prediction is based on the final hidden vector of the BERT model; complementing the knowledge graph based on the obtained relationship prediction result, solving the problems that currently the neighborhood information of the triples cannot be fully utilized, complex semantic information cannot be captured, and the potential relationships between the triples cannot be well learned and predicted, being able to effectively utilize the neighborhood relationship information, improving the accuracy of knowledge graph completion, and achieving the purpose of enhancing the effect of entity relationship prediction in the knowledge graph
[0046] In an exemplary embodiment, the extraction of the first-order adjacent relational subgraphs of the target triples in the knowledge graph includes:
[0047] S11. For each target entity, use the breadth-first search algorithm to construct the first-order relational subgraph of the target entity, and the first-order relational subgraph is expressed as
[0048]
[0049] where r j represents a certain type of relationship, m represents the number of extracted adjacent relationship contexts, and e is each target entity
[0050] In this embodiment, since the relational subgraphs of entities cannot be directly obtained and need to be manually constructed, for each entity e that appears in the training set and the test set, the present invention uses the BFS algorithm to traverse the triples in the training set and the test set to obtain the first-order relational subgraphs of the entities (removing the target relationship r during this extraction process), and records them as where r j represents a certain type of relationship, and m represents the number of extracted adjacent relationship contexts. In this way, a relatively rich and accurate information basis can be provided for the relationship prediction between entities
[0051] For example, both the set of users and books are represented as entities of triples \(e = \{e 1 , e 2 , …, e n \}\), and the relationships between entities are represented as \(r = \{r 1 ,..., r m \}\), where \(n\) represents the number of entities and \(m\) represents the number of relationships. We introduce the knowledge graph completion method of the present invention for improvement. For a given target triple, first-order adjacent relationship subgraph extraction is performed. This step includes the out-edges and in-edges of the entities in the user's reading history, as well as the information in the first-order relationship subgraph, whereby the relationships of the adjacent context can be effectively captured and the complex relationships between nodes can be deeply revealed. Use the BFS algorithm to traverse the triples in the knowledge graph, obtain the first-order relationship subgraph of each entity, and record it as where \(r j represents a certain type of relationship and \(m\) represents the number of extracted adjacent relationship contexts. Through this method, a relatively rich and accurate information basis can be provided for the relationship prediction between entities, which is beneficial to the better completion of the knowledge graph required for recommendation.
[0052] See Figure 2 , which defines the relationship subgraph around the target entity, which is composed of the out-edges and in-edges of the entity, the first-order subgraph and the second-order subgraph, and the red part is the first-order relationship subgraph. The first-order relationship subgraph of the entity is obtained by using the BFS algorithm These relationship subgraphs contain rich adjacent context information.
[0053] See Figure 3 , and a specific example is used to more clearly understand the importance of the adjacent context relationship and the extraction process of the subgraph. Suppose the goal is to predict which of the entities "Tang Monk" and "Sand monk" or "Jingu Bang" has the "disciple of" relationship. The first-order relationship subgraph can be constructed for it. The first-order relationship subgraphs corresponding to these two target triples. For "Tang Monk", his relationship subgraph includes relationships such as "fellow", "past life of", "master of", etc. (including out-edges and in-edges); for "Jingu Bang", there are relationships such as "possess", "make by", "locate in", etc. (including out-edges and in-edges); and for "Sand monk", there are relationships such as "fellow disciples", "locate in", "past life of", etc. (including out-edges and in-edges).
[0054] First, the first-order relational subgraph of "Tang Monk" can be focused on. According to the subgraph, "Tang Monk" has relationships such as "fellow", "past life of", and "master of". In particular, "master of" indicates that he is the master of someone. This information implies that "Tang Monk" may have a "disciple of" relationship with a certain entity.
[0055] Then, examine the first-order relational subgraph of "Jingu Bang". There is "fellowdisciples" in the subgraph of "Sand monk", which implies that he is the disciple of a certain master. Combining the information of "fellow disciples" with the "master of" in the "Tang Monk" subgraph, it can be speculated that as a "disciple", "Sand monk" is very likely to have a "disciple of" relationship with "Tang Monk".
[0056] Combined with the first-order relational subgraph of "Jingu Bang" again. Its subgraph includes relationships such as "possess", "make by", and "locate in", all of which indicate that "Jingu Bang" is an item rather than a person. Although it has some associations with "Tang Monk", such as "possess" may imply that "Tang Monk" uses or owns "Jingu Bang", these relationships do not show that it can form a "disciple of" relationship with "Tang Monk".
[0057] Therefore, through the analysis of the adjacent context relationships, it can be clearly concluded that there is a potential "disciple of" relationship between "Tang Monk" and "Sand monk" because the master-disciple relationships such as "master of" and "fellow disciples" are directly involved in their relational subgraphs. And as an item, the association between "Jingu Bang" and "Tang Monk" is mainly reflected in the ownership or manufacturing information and cannot form a master-disciple relationship.
[0058] This kind of analysis fully considers the adjacent context relationships of entities, thus providing a rich information basis for predicting the relationships between entities.
[0059] In an exemplary embodiment, the inputting the target triple and the adjacent context relationships of the head and tail entities corresponding to the target triple into the BERT model for embedding learning includes:
[0060] S21. Determine the text sequence representing the head entity, the tail entity, and the relationship between the head and tail entities as the first text sequence;
[0061] S22. Determine the text sequence representing the first-order relationship subgraph of the head entity as the second text sequence;
[0062] S23. Determine the text sequence representing the first-order relationship subgraph of the tail entity as the third text sequence;
[0063] S24. Determine the original input of the BERT model based on the first text sequence, the second text sequence, and the third text sequence.
[0064] In this embodiment, the text sequences of the head entity, the tail entity, and the relationship between the head and tail entities can be respectively represented as h = {h 1 ,..., h a}, t = {t 1 ,..., t b}, and r = {r 1 ,..., r m}, where a represents the length of the head entity h, b represents the length of the tail entity t, and m represents the length of the relationship between the head and tail entities. The text sequences of the first-order relationship subgraph of the head entity and the first-order relationship subgraph of the tail entity are represented as and where p represents the length of , and q represents the length of . Therefore, the original input of BERT can be represented as:
[0065]
[0066] In an exemplary embodiment, after determining the original input of the BERT model based on the first text sequence, the second text sequence, and the third text sequence, the method further includes:
[0067] S31. Determine the final input based on the original input, and the final input is represented as
[0068] x i = E tokens (w i ) + E segments (w i ) + E positions (i)
[0069] where E tokens (w i ) is the token vector of the word w i , and E segments (w i ) is the corresponding vector for the word w iSegment vectors, used to distinguish different sentences or sequences; E positions (i) is the position vector of position i, representing the word w i 's position in the sequence.
[0070] In this embodiment, they are stored in sequence according to the relationship between the head and tail entities, the first-order relationship subgraph of the head entity and the first-order relationship subgraph sequence of the tail entity, and the order of the head and tail entities. After obtaining the original input, it needs to be processed into a form that the model can understand, and three types of information need to be added respectively: word information (token vector): E tokens , sentence information (segment vector): E segments and position information (position vector): E positions . The final input representation x i of each word w i is the sum of these three components:
[0071] x i = E tokens (w i ) + E segments (w i ) + E positions (i)
[0072] E tokens (w i ) is the token vector of the word w i , E segments (w i ) is the segment vector corresponding to the word w i , used to distinguish different sentences or sequences. E positions (i) is the position vector of position i, representing the word w i 's position in the sequence. For the entire input sequence, it can be represented as:
[0073] X = [x 1 , x 2 ,..., x n
[0074] After generating the embeddings, the next step in the BERT model involves further processing these embeddings through a series of Transformer encoder layers. The processing of each layer can be described by the following formula:
[0075] Multi-head self-attention:
[0076] MultiHead(Q, K, V) = Concat(head 1 , …, head h )W O
[0077]
[0078] Among them, Q, K, and V respectively represent query, key, and value vectors, which are used to calculate attention weights and generate the final multi-head self-attention representation; and W O are parameter matrices learned in the model, and d k is the dimension of the key vector.
[0079] Residual connection and normalization:
[0080] LayerNorm(x + Sublayer(x))
[0081] Among them, Sublayer(x) is the operation of the sublayer itself, which can be a multi-head self-attention network or a position-wise feed-forward network. LayerNorm represents the layer normalization operation, and x is the input to the sublayer.
[0082] Position-wise feed-forward network:
[0083] FFN(x) = max(0, xW 1 + b 1 )W 2 + b 2
[0084] Among them, W 1 、W 2 、b 1 、b 2 are parameters in the network.
[0085] By repeating the above process for each Transformer layer, the embedding representation of the target triple and its adjacent context relationship can be obtained finally. However, for downstream tasks, we select the final hidden vector C corresponding to [CLS] as the aggregated sequence representation for classification calculation, where
[0086] The relationship prediction in the knowledge graph completion task explored in the present invention can be regarded as a typical binary classification problem. In this task, the powerful ability of the BERT model architecture can be utilized, and through its unique bidirectional encoder representation, the final hidden vector C corresponding to the [CLS] token can be extracted. This vector is the key to relationship prediction. This hidden vector C not only contains the semantic information of the input triple but also can provide rich feature representations for subsequent classification tasks.
[0087] In an exemplary embodiment, setting the prediction function of the target triple, the relationship prediction result obtained after the relationship prediction based on the final hidden vector of the BERT model includes:
[0088] S41. Use the binary cross - entropy loss function to calculate the difference between the predicted value of the BERT model and the actual label. The calculation formula is:
[0089]
[0090] where the label set is y τ ∈{0, 1}, 0 represents a negative example, and 1 represents a positive example; the prediction function is a two - dimensional entity vector, S τ0 , S τ1 ∈[0, 1], S τ0 +S τ1 = 1.
[0091] In this embodiment, optionally, during the fine - tuning process of triple classification, a key component, namely the classification layer weights, needs to be introduced. These weights are important parameters in the model learning process. They map the hidden vector C to a two - dimensional output space, where each dimension represents the probability that the triple belongs to a positive or negative example. The design of the prediction function is crucial. We set the prediction function for the triple τ=(h, t, sr) as S τ = f(h, t, sr)= sigmoid(CW t ), is a two - dimensional entity vector, S τ0 , S τ1 ∈[0, 1], S τ0 +S τ1 = 1.
[0092] In the model training stage, two sets of triple sets are prepared: one is the positive triple set, and the other is the negative triple set. For the given positive triple set and the negative triple set use the cross - entropy loss function to calculate the difference between the predicted value of the model and the actual label.
[0093]
[0094] Here, the label set is y τ ∈{0, 1}, where 0 represents a negative example and 1 represents a positive example.
[0095] To generate negative triples, optionally, randomly replace the head entity or the tail entity in the positive triple. In this process, deduplication processing is required to ensure that no negative examples are generated that are duplicates of the positive triple or existing negative triples. This is done to avoid introducing noise during training, which may affect the final performance of the model.
[0096]
[0097] To effectively optimize the model parameters, exemplarily, the Adam optimizer is selected. The Adam optimizer is an efficient adaptive learning rate optimization algorithm. It combines the advantages of the momentum method, which can accumulate historical gradient information to accelerate the learning process. At the same time, it also incorporates the characteristics of the RMSProp algorithm, adjusting the learning rate by calculating the moving average of the squared gradients. The use of this optimizer enables the model to converge more stably and efficiently during training. Especially when dealing with large-scale data and complex model parameters, the Adam optimizer demonstrates its superior performance. Through these methods, more accurate relationship prediction can be achieved in the knowledge graph completion task.
[0098] The rationality of the target triple is predicted through a prediction function, that is, the degree of interest of the user in a certain book is predicted. The binary cross-entropy loss function is used to calculate the difference between the predicted value of the model and the actual label, and negative samples are generated by randomly replacing the user or book entity in the positive sample. The rationality of the target triple is determined by comparing the output of the prediction function with a threshold. If the output value is greater than the threshold, it is considered that the user is interested in the book, otherwise not. The Adam optimizer is selected to optimize the model parameters to achieve stable and efficient convergence of the model during training.
[0099] According to the results predicted by the model, the incomplete knowledge graph will be completed, and a list of books that a user may be interested in will be generated for each user. The recommendation algorithm is adjusted through user feedback to improve the accuracy of the recommendation and user satisfaction.
[0100] See Figure 5 For the given knowledge graph, this model performs subgraph extraction of the head and tail entities of the target triple and extraction of the target relationship to be learned. The strategy of randomly generating negative samples is adopted to generate negative samples corresponding to the target triple as the positive sample. The BERT model is used for embedding learning of the triple and its adjacent context relationships, fully capturing the rich information in the knowledge graph, and obtaining the final hidden vector [CLS] as the input to the classification layer to generate the final relationship prediction result. Since the inductive relationship prediction method is adopted in this paper, as shown in Figure 5 (b), in the test stage, the provided knowledge graphs are all entities not seen in the training part.
[0101] Furthermore, to verify the model effect, tests are carried out on the public datasets FB15k-237 and WN18RR. These two datasets have a total of more than 32,000 entities, which can comprehensively test the rationality and accuracy of the knowledge graph completion scheme proposed by the present invention. The performance of the knowledge graph completion model is evaluated by calculating the metrics AUC-PR and Hits@10, and compared with existing methods.
[0102] Refer to Figure 5, for a given knowledge graph, for the target triple, perform subgraph extraction and learning of the head and tail entities for the target relationship extraction. Adopt the strategy of randomly generating negative samples to generate negative samples corresponding to the target triple as the positive sample. Use the BERT model to perform embedding learning on the triple and its adjacent context relationships, fully capture the rich information in the knowledge graph, and obtain the final hidden vector [CLS] as the input to the classification layer to generate the final relationship prediction result. Since the present invention is an inductive relationship prediction method, therefore, as Figure 5 (b) shows, in the test stage, the provided knowledge graphs are all entities not seen in the training part. This is also beneficial for predicting the relationships between newly added books and entities in the future, and thus for further recommendation. The results are analyzed from two aspects of AUC-PR and Hits@10 respectively below.
[0103] Compared with the existing relatively effective and novel models in terms of the AUC-PR evaluation index on the datasets WN18RR and FB15k-237, the average improvement is 6.4%.
[0104] Compared with the existing relatively effective and novel models in terms of the Hits@10 evaluation index on the datasets WN18RR and FB15k-237, the average improvement is 9.63%.
[0105] It can be seen that the method of the present invention effectively improves the effect of knowledge graph completion. In addition, when it is generalized to the example of book recommendation, after the incomplete knowledge graph is completed by the method of the present invention, the relationships between rich books and reader entities can effectively improve the recommendation performance of the downstream task of book recommendation.
[0106] According to another aspect of the embodiments of the present application, there is also provided a knowledge graph completion device for neighborhood relationship learning based on BERT for implementing the above-mentioned knowledge graph completion method for neighborhood relationship learning based on BERT. Figure 6 is a schematic structural diagram of an optional knowledge graph completion device for neighborhood relationship learning based on BERT according to the embodiments of the present application. As Figure 6 shown, the device may include:
[0107] An extraction unit 602, configured to perform first-order relationship subgraph extraction on a target triple in the knowledge graph, where the relationship subgraph includes the out-edges and in-edges of the target entity;
[0108] A learning unit 604, configured to input the target triple and the adjacent context relationships of the head and tail entities corresponding to the target triple into the BERT model for embedding learning, where the BERT model learns the left and right context information of the vocabulary through two pre-training tasks of masked language model and next sentence prediction;
[0109] A prediction unit 606 for setting a prediction function for the target triple, where the prediction function performs relation prediction based on the final hidden vector of the BERT model to obtain a relation prediction result;
[0110] A completion unit 608 for completing the knowledge graph based on the obtained relation prediction result.
[0111] It should be noted that the extraction unit 602 in this embodiment can be used to execute the above step S102, the learning unit 604 in this embodiment can be used to execute the above step S104, the prediction unit 606 in this embodiment can be used to execute the above step S106, and the completion unit 608 in this embodiment can be used to execute the above step S108.
[0112] Through the above modules, by extracting the first-order relation subgraph of the target triple in the knowledge graph, where the relation subgraph includes the out-edges and in-edges of the target entity; inputting the target triple and the adjacent context relationships of the head and tail entities corresponding to the target triple into the BERT model for embedding learning, where the BERT model learns the left and right context information of the vocabulary through two pre-training tasks of masked language model and next sentence prediction; setting a prediction function for the target triple, and the prediction function performs relation prediction based on the final hidden vector of the BERT model to obtain a relation prediction result; completing the knowledge graph based on the obtained relation prediction result, it solves the problems that currently it is impossible to make full use of the triple neighborhood information, capture complex semantic information, and cannot learn and predict the potential relationships between triples well, can effectively utilize the neighborhood relationship information, improve the accuracy of knowledge graph completion, and achieves the purpose of enhancing the effect of entity relation prediction in the knowledge graph.
[0113] In an exemplary embodiment, the extraction unit includes:
[0114] A construction module for, for each target entity, constructing the first-order relation subgraph of the target entity by using the breadth-first search algorithm, and the first-order relation subgraph is represented as
[0115]
[0116] where r j represents a certain type of relation, m represents the number of extracted adjacent relation contexts, and e is each target entity.
[0117] In an exemplary embodiment, the learning unit includes:
[0118] A first determination module for determining the text sequence representing the head entity, the tail entity, and the relation between the head and tail entities as the first text sequence;
[0119] A second determination module, configured to determine the text sequence representing the first-order relationship sub-graph of the head entity as the second text sequence;
[0120] A third determination module, configured to determine the text sequence representing the first-order relationship sub-graph of the tail entity as the third text sequence;
[0121] A fourth determination module, configured to determine the original input of the BERT model based on the first text sequence, the second text sequence, and the third text sequence.
[0122] In an exemplary embodiment, the apparatus further includes:
[0123] Determining a final input based on the original input, the final input is represented as,
[0124] x i = E tokebs (w i ) + E segments (w i ) + E positiobs (i)
[0125] Wherein, E tokens (w i ) is the token vector of the word w i , E segments (w i ) is the segment vector corresponding to the word w i , used to distinguish different sentences or sequences; E positions (i) is the position vector of position i, indicating the position of the word w i in the sequence.
[0126] In an exemplary embodiment, the prediction unit includes:
[0127] A calculation module, configured to use the binary cross-entropy loss function to calculate the difference between the predicted value of the BERT model and the actual label, and the calculation formula is:
[0128]
[0129] Wherein, the label set is y τ ∈ {0, 1}, 0 represents a negative example, and 1 represents a positive example; the prediction function is a two-dimensional entity vector, S τ0 , S τ1 ∈ [0, 1], S τ0 + S τ1 = 1.
[0130] It should be noted here that the implementation examples and scenarios of the above modules and corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules, as part of the device, can run in a hardware environment, can be implemented by software, or can be implemented by hardware, where the hardware environment includes a network environment.
[0131] According to another aspect of the embodiments of the present application, a storage medium is also provided. Optionally, in this embodiment, the above storage medium can be used to execute the program code of any one of the above BERT-based neighborhood relationship learning knowledge graph completion methods in the embodiments of the present application.
[0132] Optionally, in this embodiment, the storage medium is set to store program code for executing the following steps:
[0133] S1, extract a first-order relational subgraph for a target triple in the knowledge graph, where the relational subgraph includes the out-edges and in-edges of the target entity;
[0134] S2, input the target triple and the adjacent context relationships of the head and tail entities corresponding to the target triple into the BERT model for embedding learning, where the BERT model learns the left and right context information of the vocabulary through two pre-training tasks: masked language model and next sentence prediction;
[0135] S3, set a prediction function for the target triple, and the prediction function performs relationship prediction based on the final hidden vector of the BERT model to obtain a relationship prediction result;
[0136] S4, complete the knowledge graph based on the obtained relationship prediction result.
[0137] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments, and will not be elaborated herein.
[0138] Among them, the computer-readable storage medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs, CD-ROMs, microdrives, and magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, DRAMs, VRAMs, flash memory devices, magnetic cards or optical cards, nanosystems (including molecular memory ICs), or any type of medium or device suitable for storing instructions and / or data.
[0139] According to another aspect of the embodiments of the present application, an electronic device for implementing the above BERT-based neighborhood relationship learning knowledge graph completion method is also provided, and the electronic device can be a server, a terminal, or a combination thereof.
[0140] Figure 7It is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. As Figure 7 shown, it includes a processor 702, a communication interface 704, a memory 706, and a communication bus 708. Among them, the processor 702, the communication interface 704, and the memory 706 complete mutual communication through the communication bus 708. Among them,
[0141] The memory 706 is used to store computer programs;
[0142] The processor 702, when executing the computer program stored on the memory 706, realizes the following steps:
[0143] S1, extract a first-order relational subgraph for the target triple in the knowledge graph, where the relational subgraph includes the outgoing edges and incoming edges of the target entity;
[0144] S2, input the target triple and the adjacent context relationships of the head and tail entities corresponding to the target triple into the BERT model for embedding learning, where the BERT model learns the left and right context information of the vocabulary through two pre-training tasks: masked language model and next sentence prediction;
[0145] S3, set a prediction function for the target triple, and the prediction function obtains a relationship prediction result after performing relationship prediction based on the final hidden vector of the BERT model;
[0146] S4, complete the knowledge graph based on the obtained relationship prediction result.
[0147] Optionally, the communication bus can be a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 7 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used for communication between the above-mentioned electronic device and other devices.
[0148] The memory may include a RAM, and may also include a non-volatile memory, for example, at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0149] As an example, the above-mentioned memory 706 may but is not limited to include the extraction unit 602, the learning unit 604, the prediction unit 606, and the completion unit 608 in the above-mentioned BERT-based neighborhood relationship learning knowledge graph completion device. In addition, it may also include but is not limited to other module units in the above-mentioned BERT-based neighborhood relationship learning knowledge graph completion device, which will not be elaborated in this example.
[0150] The above-mentioned processor may be a general-purpose processor, which may include but is not limited to: CPU (Central Processing Unit, central processor), NP (Network Processor, network processor), etc.; it may also be a DSP (Digital Signal Processing, digital signal processor), ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), FPGA (Field-Programmable Gate Array, field programmable gate array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0151] Optionally, for the specific examples in this embodiment, reference may be made to the examples described in the above-mentioned embodiments, which will not be elaborated herein.
[0152] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps may be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0153] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0154] In the several embodiments provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some service interfaces. The indirect coupling or communication connection of the device or unit may be in an electrical or other form.
[0155] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0156] In addition, each functional unit in various embodiments of the present application may be integrated into a processing unit, may exist independently as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0157] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned memory includes: USB flash drives, read-only memory (ROM), random access memory (RAM), mobile hard disks, magnetic disks, or optical discs, etc., which can store program codes.
[0158] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable memory. The memory can include: flash drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc.
[0159] The above are only exemplary embodiments of the present disclosure and should not be used to limit the scope of the present disclosure. That is, all equivalent changes and modifications made according to the teachings of the present disclosure still fall within the scope covered by the present disclosure. After considering the specification and practicing the present disclosure, those skilled in the art will easily think of other implementation manners of the present disclosure. The present application aims to cover any variations, uses, or adaptive changes of the present disclosure. These variations, uses, or adaptive changes follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not described in the present disclosure. The specification and embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.
[0160] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered that the scope described in this specification is covered.
[0161] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included within the protection scope of the present invention.
Claims
1. A knowledge graph completion method based on BERT neighborhood relationship learning, characterized in that: include: Extracting a first-order relation subgraph for the target triple in the knowledge graph, wherein the relation subgraph includes outgoing edges and incoming edges of the target entity; Inputting the target triple and the adjacency contextual relationship of the head and tail entities corresponding to the target triple into the BERT model for embedding learning, wherein the BERT model learns the left and right contextual information of the vocabulary through two pre-training tasks: a masked language model and next sentence prediction; Setting a prediction function for the target triplet, wherein the prediction function performs relationship prediction based on the final hidden vector of the BERT model to obtain a relationship prediction result; Complete the knowledge graph based on the acquired relationship prediction results.
2. The BERT-based neighborhood relationship learning knowledge graph completion method as claimed in claim 1, characterized in that: The first-order adjacency subgraph extraction of the target triple in the knowledge graph includes: For each target entity, a breadth-first search algorithm is used to construct a first-order relationship subgraph of the target entity. The first-order relationship subgraph is expressed as: Among them, r j represents a certain type of relationship, m represents the number of extracted adjacency contexts, and e is each target entity.
3. The BERT-based neighborhood relationship learning knowledge graph completion method as claimed in claim 1, characterized in that: The step of inputting the target triple and the adjacency contextual relationship of the head and tail entities corresponding to the target triple into the BERT model for embedding learning includes: Determine a text sequence representing a head entity, a tail entity, and a relationship between the head and tail entities as a first text sequence; Determine a text sequence representing a first-order relationship subgraph of a head entity as a second text sequence; Determine a text sequence expressing a first-order relationship subgraph of the tail entity as a third text sequence; An original input of the BERT model is determined based on the first text sequence, the second text sequence, and the third text sequence.
4. The BERT-based neighborhood relationship learning knowledge graph completion method as claimed in claim 3, characterized in that: After determining the original input of the BERT model based on the first text sequence, the second text sequence, and the third text sequence, the method further includes: The final input is determined based on the original input, and the final input is expressed as, x i =E tokens (w i )+E segments (w i )+E positions (i) Among them, E tokens (w i ) is the word w i The label vector, E segments (w i ) corresponds to word w i segment vectors, used to distinguish different sentences or sequences; positions (i) is the position vector of position i, representing word w i Position in the sequence.
5. The BERT-based neighborhood relationship learning knowledge graph completion method as claimed in claim 1, characterized in that: The setting of the prediction function of the target triplet, wherein the prediction function performs relationship prediction based on the final hidden vector of the BERT model to obtain a relationship prediction result, comprises: The binary cross entropy loss function is used to calculate the difference between the predicted value of the BERT model and the actual label. The calculation formula is: Among them, the label set is y τ ∈{0, 1}, 0 represents a negative example, 1 represents a positive example; prediction function is a two-dimensional solid vector, S τ0 , S τ1 ∈[0, 1], S τ0 +S τ1 =1.
6. A knowledge graph completion device for neighborhood relationship learning based on BERT, characterized in that: include: An extraction unit, configured to extract a first-order relation subgraph for a target triple in a knowledge graph, wherein the relation subgraph includes outgoing edges and incoming edges of the target entity; A learning unit, used for inputting the target triple and the adjacency context relationship of the head and tail entities corresponding to the target triple into a BERT model for embedding learning, wherein the BERT model learns the left and right context information of the vocabulary through two pre-training tasks of a masked language model and next sentence prediction; A prediction unit, used for setting a prediction function for the target triple, wherein the prediction function performs relationship prediction based on the final hidden vector of the BERT model to obtain a relationship prediction result; A completion unit is used to complete the knowledge graph based on the acquired relationship prediction results.
7. The BERT-based neighborhood relationship learning knowledge graph completion device according to claim 6, characterized in that: The extraction unit comprises: A construction module is used to construct a first-order relationship subgraph of each target entity using a breadth-first search algorithm, wherein the first-order relationship subgraph is represented as: Among them, r j represents a certain type of relationship, m represents the number of extracted adjacency contexts, and e is each target entity.
8. The BERT-based neighborhood relationship learning knowledge graph completion device as claimed in claim 6, characterized in that: The learning units include: A first determination module, configured to determine a text sequence representing a head entity, a tail entity, and a relationship between the head and tail entities as a first text sequence; A second determination module, configured to determine a text sequence representing a first-order relationship subgraph of a head entity as a second text sequence; A third determination module, used to determine the text sequence expressing the first-order relationship subgraph of the tail entity as a third text sequence; The fourth determination module is used to determine the original input of the BERT model based on the first text sequence, the second text sequence and the third text sequence.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method according to any one of claims 1 to 5 when executed.
10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 5 through the computer program.
Citation Information
Cited By
Knowledge extraction and mining method for smart home multi-mode dialogue system
CN120850235A
Knowledge extraction and mining method for smart home multi-modal dialogue system
CN120850235B