A knowledge graph embedding method and apparatus based on a pre-trained model

By introducing a super network and an additional parameter layer, and optimizing the parameter offset based on cross-entropy and KL divergence loss, the problem of inefficient updating of knowledge graph embedding in existing technologies is solved, enabling fast, local updates and error correction, and improving the adaptability and efficiency of knowledge graphs.

CN116975307BActive Publication Date: 2025-10-28ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310782122.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-29
Publication Date
2025-10-28
Estimated Expiration
2043-06-29

AI Technical Summary

Technical Problem

Existing knowledge graph embedding methods based on pre-trained models cannot efficiently perform local updates when faced with changes or errors in the knowledge graph, resulting in the model being unable to adapt to new facts or correct errors, and retraining the model consumes a lot of computational resources.

Method used

By introducing a supernetwork and additional parameter layers, the system utilizes mask prediction tasks to filter out erroneous text knowledge, constructs cross-entropy loss and KL divergence loss to optimize parameter offsets, and achieves fast, local updates of knowledge graph embeddings.

Benefits of technology

It enables rapid and efficient updating of knowledge graph embeddings without retraining the model, improving flexibility and practicality while reducing time and resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116975307B_ABST
    Figure CN116975307B_ABST
Patent Text Reader

Abstract

This invention discloses a knowledge graph embedding method and apparatus based on a pre-trained model. The method involves training a pre-trained model to obtain a knowledge graph embedding model, then introducing a hypernetwork to predict parameter offsets based on sample data. An additional parameter layer is added to the knowledge graph embedding model to form an editing model. This additional parameter layer has the same structure as the feedforward neural network in the knowledge graph embedding model, with parameters consisting of randomly initialized parameters superimposed with the parameter offsets predicted by the hypernetwork. A loss function is constructed to optimize the hypernetwork parameters. For the text knowledge to be edited, the parameter-optimized hypernetwork predicts parameter offsets based on the text knowledge, and the editing model is updated using these parameter offsets. The updated editing model is then used to embed the text knowledge into a representation. This method and apparatus achieve fast, efficient, and local updates to the embedding vectors without requiring significant computational overhead to re-optimize the model parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of natural language processing and knowledge graph in computer science, and specifically relates to a knowledge graph embedding method and apparatus based on a pre-trained model. Background Technology

[0002] A knowledge graph is a structured semantic knowledge base based on symbolic representation. It is a multi-relation graph containing a large number of symbolic facts, used to describe the relationships between entities and concepts, providing backend support for a wide range of knowledge-intensive tasks, including information retrieval, question answering, and recommendation systems. Knowledge graph embedding is a method of knowledge representation learning. To better utilize the symbolic knowledge in knowledge graphs for machine learning models, many knowledge graph embedding methods strive to represent knowledge graphs in a low-dimensional vector space. Traditional knowledge graph embedding models, such as TransE and RotatE, are naturally classified as structure-based methods. This is supervised machine learning that preserves the inherent structure of the knowledge graph by optimizing the target object using a scoring function. Recent trends in knowledge graph embedding have shifted away from explicitly modeling such structures. Instead, they apply textual descriptions and expressive black-box models, such as pre-trained language models, hoping that the model can capture the structure without explicit instructions.

[0003] A pre-trained model is a deep learning model that has been pre-trained using a large-scale corpus. Its performance can be effectively improved by fine-tuning it on downstream tasks. Pre-trained models have demonstrated remarkable performance in many fields, including natural language processing, computer vision, and speech artificial intelligence.

[0004] Framework-based knowledge graph embedding using language models is an increasingly promising technique with proven success. When pre-trained models are applied to knowledge graph embedding, their advantage lies in leveraging the powerful expressive capabilities and rich semantic information of the pre-trained model, along with the structural information within the knowledge graph, to obtain more accurate and robust knowledge representations, especially for long-tail and emerging entities. However, pre-trained model-based knowledge graph embedding typically fixes the parameters during the training phase, making it impossible to modify or update the embeddings during deployment. This means they struggle to adapt to changes in the knowledge graph, such as newly added facts or corrected errors, without retraining.

[0005] Current work has explored editing knowledge graph embeddings based on pre-trained models by modifying the model architecture or introducing hypernetworks. However, these methods have limitations, such as excessively long training times and limited editing capabilities.

[0006] When learning knowledge graph embeddings using pre-trained models, the expressive power of the pre-trained models is strong enough, and the semantic information is rich enough. However, there is no efficient way to deal with problems when the semantic information is compromised. Existing methods either modify the model architecture or introduce hypernetworks for editing, but they fail to efficiently edit knowledge graph embeddings. To cope with changes (e.g., newly emerging facts) or corrections to facts in existing knowledge graphs, the ability to flexibly update knowledge to knowledge graph embeddings is desirable. Summary of the Invention

[0007] In view of the above, the purpose of this invention is to provide a knowledge graph embedding method and apparatus based on a pre-trained model, which can achieve fast, efficient and local updates of the embedding vector without consuming a lot of computational overhead to re-optimize the model parameters, while minimizing the impact on the performance of other parts, so as to adapt to changes or errors in the knowledge graph.

[0008] To achieve the above-mentioned objectives, an embodiment provides a knowledge graph embedding method based on a pre-trained model, comprising the following steps:

[0009] Text knowledge extracted from text corpora and knowledge graphs is used to train the pre-trained model for a mask-based prediction task, resulting in a knowledge graph embedding model. Then, text knowledge that is predicted incorrectly is selected as sample data based on the knowledge graph embedding model.

[0010] A supernetwork is introduced to predict parameter offsets based on sample data;

[0011] An additional parameter layer is added to the knowledge graph embedding model to form an editing model. The structure of this additional parameter layer is the same as that of the feedforward neural network in the knowledge graph embedding model. The parameters are randomly initialized parameters superimposed with the parameter offsets predicted by the hypernetwork.

[0012] The sample data is divided into data that needs to be edited and data that does not need to be edited. Based on the data that needs to be edited, a cross-entropy loss is constructed between the prediction results of the editing model and the target entity. Based on the data that does not need to be edited, a KL divergence loss is constructed between the prediction results of the knowledge graph embedding model and the editing model respectively. The parameters of the hypernetwork are optimized based on the cross-entropy loss and the KL divergence loss.

[0013] For the text knowledge to be edited, a parameter-optimized hypernetwork is used to predict parameter offsets based on the text knowledge to be edited. The parameter offsets are then used to update the editing model, and the updated editing model is used to embed the text knowledge to be edited.

[0014] Preferably, the extracted text knowledge is represented by triples (h,r,t), where h and t represent the head entity and tail entity, and t represents the relation. When training the pre-trained model for a mask-based prediction task based on triples (h,r,t), either h or t in the triples (h,r,t) is masked and then input into the pre-trained model to obtain an embedding representation. A classifier is then introduced to predict and classify the masked entities based on the embedding representation. The parameters of the pre-trained model are optimized based on the cross-entropy loss between the predicted classification results and the real entities. The pre-trained model with optimized parameters serves as a knowledge graph embedding model.

[0015] The text knowledge that is incorrectly predicted based on the knowledge graph embedding model is represented as (x,y,a), where x represents the model input, y represents the incorrect or outdated entity, i.e. the model output, and a represents the target entity.

[0016] Preferably, the hypernetwork includes a hidden state extraction module and a mapping module, wherein the input text knowledge is processed by the hidden state extraction module to extract the hidden state, and the hidden state is then mapped and predicted by the mapping module to obtain vectors α,β∈R. m ,γ,δ∈R n And a scalar η∈R, where m×n represents the size of the random initialization parameters of the extra parameter layer, and R represents a real number;

[0017] The parameter offset is calculated based on the vectors α, β, γ, δ, the scalar η, and the cross-entropy loss gradient of the input text knowledge in the knowledge graph embedding model.

[0018]

[0019]

[0020]

[0021] in, σ represents the Softmax function, and σ represents the Sigmoid function. The input text knowledge x represents the input parameters. The gradient of the cross-entropy loss between the knowledge graph embedded in the model and the target entity a. ⊙ represents the parameter offset, and ⊙ represents the dot product.

[0022] Preferably, the hidden state extraction module uses a bidirectional LSTM network, and the mapping module uses a fully connected network.

[0023] Preferably, after adding an extra parameter layer to the knowledge graph embedding model, the feedforward neural network in the edit model is represented as follows:

[0024]

[0025] Where H represents the output of the attention layer in the Transformer of the knowledge graph embedding model. This represents the original parameters in the knowledge graph embedding model. This represents the randomized parameters generated by the additional parameter layer. This represents the parameter offset predicted by the supernetwork, and ΔFFN represents the additional parameter layer added to achieve parameter offset. Implementation Let denot be the parameter matrix of the pre-trained language model, GELU be the activation function, and FFN be the original feedforward neural network. ′ This represents a feedforward neural network with added parameter layers.

[0026] Preferably, the cross-entropy loss is expressed as:

[0027]

[0028] Where, x i This represents the text knowledge that needs to be edited, including head entities and relations, or relations and tail entities, a i Represents the target entity. The parameter is The editing model predicts the probability of the target entity. This represents the cross-entropy loss.

[0029] Preferably, the KL divergence loss is expressed as:

[0030]

[0031] in, This represents the KL divergence loss. This represents the text knowledge x that does not need to be edited within the same batch. ′ The set that is formed This represents the set of erroneous entities y. x represents ′ Input parameters are The parameterized representation of the discrete distribution of incorrect entities predicted by the knowledge graph embedding model in the sample space. x represents ′ Input parameters are The edit model predicts the parameterized representation of the discrete distribution of erroneous entities in the sample space.

[0032] To achieve the above-mentioned objectives, the present invention also provides a knowledge graph embedding device based on a pre-trained model, comprising a pre-training unit, a hypernetwork introduction unit, an additional parameter layer addition unit, a hypernetwork training unit, and an embedding representation unit.

[0033] The pre-training unit is used to train the pre-trained model for a mask-based prediction task using text knowledge extracted from a text corpus and a knowledge graph, to obtain a knowledge graph embedding model, and to filter out text knowledge that is predicted incorrectly as sample data based on the knowledge graph embedding model.

[0034] The hypernetwork introduction unit is used to introduce a hypernetwork for predicting parameter offsets based on sample data.

[0035] The additional parameter layer addition unit is used to add an additional parameter layer to the knowledge graph embedding model to form an editing model. The structure of the additional parameter layer is the same as that of the feedforward neural network in the knowledge graph embedding model, and the parameters are randomly initialized parameters superimposed with the parameter offset predicted by the hypernetwork.

[0036] The hypernetwork training unit is used to divide the sample data into data that needs to be edited and data that does not need to be edited. Based on the data that needs to be edited, a cross-entropy loss is constructed between the prediction results of the editing model and the target entity. Based on the data that does not need to be edited, a KL divergence loss is constructed between the prediction results of the knowledge graph embedding model and the editing model, respectively. The parameters of the hypernetwork are optimized based on the cross-entropy loss and the KL divergence loss.

[0037] The embedding representation unit is used to predict parameter offsets based on the knowledge of the text to be edited using a parameter-optimized hypernetwork, update the editing model using the parameter offsets, and then use the updated editing model to perform the embedding representation of the knowledge of the text to be edited.

[0038] To achieve the above-mentioned objectives, the present invention also provides a knowledge graph embedding device based on a pre-trained model, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the knowledge graph embedding method based on the pre-trained model described above.

[0039] Compared with the prior art, the beneficial effects of the present invention include at least the following:

[0040] By introducing hypernetworks and additional parameter layers, knowledge graph embeddings can be updated quickly, effectively, and locally without retraining the entire model to adapt to changes or errors in the knowledge graph, thereby improving the flexibility and practicality of knowledge graph embeddings. Compared to traditional model editing techniques, it eliminates the need to adjust too many parameters, making the model easier to train and optimize. In addition, it significantly improves editing efficiency, greatly reducing the consumption of time and resources. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 This is a flowchart of the knowledge graph embedding method based on a pre-trained model provided in the embodiment;

[0043] Figure 2 This is a flowchart at the model level provided in the embodiment;

[0044] Figure 3 This is a schematic diagram of the structure of the knowledge graph embedding device based on a pre-trained model provided in the embodiment. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and do not limit the scope of protection of this invention.

[0046] The inventive concept of this invention is to address the technical problem that existing pre-trained models are not accurate enough in embedding new textual knowledge, and that retraining the model incurs high computational costs. The embodiment provides a knowledge graph embedding method based on a pre-trained model. This method introduces a hypernetwork and additional parameter layers to add facts for editing the knowledge graph embedding. Specifically, it can update the knowledge graph. Through knowledge graph embedding editing, the embedding representation can be updated quickly without retraining the entire model, thus ensuring the timeliness and accuracy of the knowledge. It can also be used in personalized recommendation systems, where knowledge graph embedding can be edited based on users' personal preferences and behavioral history to provide personalized knowledge graph services.

[0047] Figure 1 This is a flowchart of the knowledge graph embedding method based on a pre-trained model provided in the embodiment, such as... Figure 1 As shown in the embodiment, the knowledge graph embedding method based on a pre-trained model includes the following steps:

[0048] Step 1: Use text knowledge extracted from text corpora and knowledge graphs to train the pre-trained model for a mask-based prediction task to obtain a knowledge graph embedding model, and use the knowledge graph embedding model to select text knowledge that is predicted incorrectly as sample data.

[0049] In this embodiment, textual knowledge extracted from the text corpus and knowledge graph is represented by triples (h,r,t), where h and t represent the head and tail entities, and t represents the relation. When training the pre-trained model for a mask-based prediction task based on the triples (h,r,t), either h or t in the triples (h,r,t) is masked and input into the pre-trained model to obtain an embedding representation. A classifier is then introduced to predict and classify the masked entities based on the embedding representation. The parameters of the pre-trained model are optimized based on the cross-entropy loss between the predicted classification results and the real entities. The pre-trained model with optimized parameters serves as a knowledge graph embedding model, which can simultaneously represent the embedding vectors of entities and relations. Textual knowledge that is incorrectly predicted based on the knowledge graph embedding model is represented as (x,y,a), where x represents the model input, i.e., (h,r) or (r,t), y represents the incorrect or outdated entity, i.e., the model output, and a represents the target entity. The constructed editing model aims to change the model's prediction result from y to a through editing.

[0050] Specifically, a knowledge graph embedding model was constructed based on a pre-trained language model (hereinafter referred to as the pre-trained model). The text corpus and knowledge graph triples used were the FB15k-237 and WN18RR datasets, and two sub-task datasets for editing and adding were constructed based on these two datasets respectively. Recent research has shown that pre-trained language models can acquire a large amount of factual knowledge and have a strong ability to predict the correct tail entities, which may hinder the proper evaluation of the editing task. Therefore, relatively simple triples that can be inferred from the knowledge stored within the language model were excluded to ensure accurate evaluation. In addition, some long-tail triples need to be extracted to address the potential skewed distribution of the data. Specifically, data was extracted based on the link prediction results (rank values ​​greater than 2,500). Taking the FB15k-237 dataset as an example, this resulted in a total of 6,174 triples in the FB15k-237 dataset. Due to the high cost of constructing the dataset, creating an accurate dataset that requires modification is not feasible. Therefore, a random replacement method was chosen to generate the dataset, and the results were then manually validated.

[0051] In this embodiment, both the editing and adding subtasks belong to the mask prediction task. For the editing subtask, entities in the FB15k-237 training set are replaced with the Top-1 facts predicted by the original pre-trained model to construct the pre-training dataset (original pre-training data), and standard facts are considered as the target editing data. Note that for the editing subtask, the goal is to evaluate the ability to change knowledge; therefore, a training and test set are constructed. Specifically, the model is initially guided to learn editing capabilities using the training set, and then evaluated using different test sets. The specific dataset construction process for the editing subtask is as follows:

[0052] (1) Randomly destroy the existing knowledge graph to generate a corrupted dataset;

[0053] (2) Use this dataset to fine-tune a pre-trained model to obtain a knowledge graph embedding model with erroneous knowledge that needs to be edited;

[0054] (3) Filtering: Considering the knowledge stored in the pre-trained model, it is not feasible to directly edit the damaged triples. Therefore, it is necessary to use the knowledge graph embedding model to re-evaluate the data, and to assign the correctly labeled data to the test locality dataset, while assigning the incorrect data to a designated dataset for editing.

[0055] For the addition subtask, a pre-training dataset (original pre-training data) was constructed using the original training set of FB15k-237, and data from the standard inference setting was used because they were not previously seen. Unlike the editing subtask, since the new knowledge was not present during the training process for the addition subtask, the training set can be directly used to evaluate the ability to absorb new knowledge. Simultaneously, the same strategy was used to construct datasets for both subtasks based on WN18RR. To evaluate the locality of edited knowledge, a test locality dataset based on link prediction performance was further created to assess the ability to locally update knowledge graph embeddings.

[0056] Step 2: Introduce a supernetwork to predict parameter offsets based on sample data.

[0057] In this embodiment, the introduced hypernetwork includes a hidden state extraction module and a mapping module. The hidden state extraction module can employ a bidirectional long short-term memory (LSTM) network, and the mapping module employs a fully connected network. After the input text knowledge is processed by the hidden state extraction module to extract the hidden state h, the hidden state h is mapped and predicted by the mapping module to obtain vectors α, β ∈ R. m ,γ,δ∈R n And a scalar η∈R, where m×n represents the size of the random initialization parameters of the extra parameter layer, and R represents a real number;

[0058] The parameter offset is calculated based on the vectors α, β, γ, δ, the scalar η, and the cross-entropy loss gradient of the input text knowledge in the knowledge graph embedding model.

[0059]

[0060]

[0061]

[0062] in, σ represents the Softmax function, and σ represents the Sigmoid function. The input text knowledge x represents the input parameters. The gradient of the cross-entropy loss between the knowledge graph embedding model and the target entity 'a', which includes how the weights are accessed. The knowledge in ⊙ represents the parameter offset, and ⊙ represents the dot product.

[0063] Step 3: Add an extra parameter layer to the knowledge graph embedding model to form an editing model. The structure of this extra parameter layer is the same as that of the feedforward neural network in the knowledge graph embedding model. The parameters are randomly initialized parameters superimposed with the parameter offsets predicted by the supernetwork.

[0064] In this embodiment, an additional parameter layer is added to the knowledge graph embedding model, also using a feedforward neural network. The structure is the same as the feedforward neural network in the knowledge graph embedding model, and the parameters are randomly initialized parameters superimposed with parameter offsets predicted by the hypernetwork. This knowledge graph embedding model with the added parameter layer serves as the editing model. The feedforward neural network in the editing model is represented as follows:

[0065]

[0066] Where H represents the output of the attention layer in the Transformer of the knowledge graph embedding model. This represents the original parameters in the knowledge graph embedding model. This represents the randomized parameters generated by the additional parameter layer. This represents the parameter offset predicted by the supernetwork, and ΔFFN represents the additional parameter layer added to achieve parameter offset. Implementation Let denot be the parameter matrix of the pre-trained language model, GELU be the activation function, and FFN be the original feedforward neural network. ′ This represents a feedforward neural network with added parameter layers.

[0067] Step 4: Divide the sample data into data that needs to be edited and data that does not need to be edited. Based on the data that needs to be edited, construct the cross-entropy loss between the prediction results of the editing model and the target entity. Based on the data that does not need to be edited, construct the KL divergence loss between the prediction results of the knowledge graph embedding model and the editing model respectively. Optimize the parameters of the hypernetwork based on the cross-entropy loss and the KL divergence loss.

[0068] In this embodiment, the sample data is divided into data that needs editing and data that does not need editing. Based on the prediction results of the editing model and the target entity of the data that needs editing, a cross-entropy loss is constructed, which is expressed as:

[0069]

[0070] Where, xi This represents the text knowledge that needs to be edited, including head entities and relations, or relations and tail entities, a i Represents the target entity. The parameter is The editing model predicts the probability of the target entity. This represents the cross-entropy loss. The smaller the cross-entropy loss value, the higher the prediction probability of the editing model for the target entity, which means that the reliability of knowledge is improved.

[0071] In this embodiment, the KL divergence loss is constructed based on the prediction results of the knowledge graph embedding model and the editing model, respectively, for data that does not require editing. It is expressed as:

[0072]

[0073] in, This represents the KL divergence loss. This represents the text knowledge x that does not need to be edited within the same batch. ′ The set that is formed This represents the set of erroneous entities y. x represents ′ Input parameters are The parameterized representation of the discrete distribution of incorrect entities predicted by the knowledge graph embedding model in the sample space. x represents ′ Input parameters are The edit model predicts the parameterized representation of the discrete distribution of erroneous entities in the sample space.

[0074] Based on the cross-entropy loss and KL divergence loss mentioned above, the total loss is constructed as follows: for:

[0075]

[0076] Where λ represents the adjustment ratio, and the total loss is used The parameters of the hypernetwork are optimized to reduce the difference between the target entity and the edited prediction results, while keeping the prediction results of the unedited entity unchanged before and after the model update. Specifically, the first part uses the cross-entropy loss function, commonly used in multi-label classification tasks, to measure the difference between the real entity label and the predicted entity result, thereby optimizing the hypernetwork parameters; the second part uses KL divergence loss to encourage the editing model to keep its predicted output distribution as consistent as possible with the output distribution of the knowledge graph embedding model for the unedited entity.

[0077] Step 5: For the text knowledge to be edited, use a parameter-optimized hypernetwork to predict parameter offsets based on the text knowledge to be edited, update the editing model using the parameter offsets, and use the updated editing model to embed the text knowledge to be edited.

[0078] In this embodiment, a knowledge graph to be edited is received, textual knowledge such as entities and relationships is extracted from the knowledge graph, and an input sequence is constructed according to a specified format. A parameter-optimized hypernetwork is then used to predict the parameter changes of additional parameter layers. These parameter changes are superimposed on the editing model to obtain a new editing model. The new editing model is then used to predict the relevant knowledge graph embeddings to obtain the edited embedding prediction results, thereby achieving the editing of the knowledge graph embeddings.

[0079] Based on the same inventive concept, the embodiment also provides a knowledge graph embedding device based on a pre-trained model, including a pre-training unit, a hypernetwork introduction unit, an additional parameter layer addition unit, a hypernetwork training unit, and an embedding representation unit.

[0080] The pre-training unit trains the pre-trained model using text knowledge extracted from a text corpus and knowledge graph to perform a mask-based prediction task, resulting in a knowledge graph embedding model. It then filters out incorrectly predicted text knowledge as sample data based on this model. The hypernetwork introduction unit introduces a hypernetwork to predict parameter shifts based on the sample data. The additional parameter layer addition unit adds an extra parameter layer to the knowledge graph embedding model to form an editing model. This extra parameter layer has the same structure as the feedforward neural network in the knowledge graph embedding model, and its parameters are randomly initialized parameters superimposed with the parameter shifts predicted by the hypernetwork. The hypernetwork training unit is used to divide the sample data into data that needs to be edited and data that does not need to be edited. Based on the data that needs to be edited, it constructs a cross-entropy loss between the prediction results of the editing model and the target entity. Based on the data that does not need to be edited, it constructs a KL divergence loss between the prediction results of the knowledge graph embedding model and the editing model, respectively. The parameters of the hypernetwork are optimized based on the cross-entropy loss and the KL divergence loss. The embedding representation unit is used to predict the parameter offset based on the text knowledge to be edited using the parameter-optimized hypernetwork. The parameter offset is used to update the editing model. The updated editing model is used to embed the text knowledge to be edited.

[0081] It should be noted that the knowledge graph embedding device based on the pre-trained model provided in the above embodiments should be illustrated using the above-described division of functional units as an example when performing knowledge graph embedding based on the pre-trained model. The functions can be assigned to different functional units as needed, that is, the internal structure of the terminal or server can be divided into different functional units to complete all or part of the functions described above. Furthermore, the knowledge graph embedding device based on the pre-trained model provided in the above embodiments and the knowledge graph embedding method based on the pre-trained model belong to the same concept. For details of its implementation, please refer to the knowledge graph embedding method embodiment based on the pre-trained model, which will not be repeated here.

[0082] Based on the same inventive concept, the embodiment also provides a knowledge graph embedding device based on a pre-trained model, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-described knowledge graph embedding method based on a pre-trained model, including the following steps:

[0083] Step 1: Use text knowledge extracted from text corpora and knowledge graphs to train the pre-trained model for a mask-based prediction task to obtain a knowledge graph embedding model, and use the knowledge graph embedding model to filter out text knowledge that is predicted incorrectly as sample data.

[0084] Step 2: Introduce a supernetwork to predict parameter offsets based on sample data;

[0085] Step 3: Add an extra parameter layer to the knowledge graph embedding model to form an editing model. The structure of this extra parameter layer is the same as that of the feedforward neural network in the knowledge graph embedding model. The parameters are randomly initialized parameters superimposed with the parameter offsets predicted by the super network.

[0086] Step 4: Divide the sample data into data that needs to be edited and data that does not need to be edited. Based on the data that needs to be edited, construct the cross-entropy loss between the prediction results of the editing model and the target entity. Based on the data that does not need to be edited, construct the KL divergence loss between the prediction results of the knowledge graph embedding model and the editing model respectively. Optimize the parameters of the hypernetwork based on the cross-entropy loss and the KL divergence loss.

[0087] Step 5: For the text knowledge to be edited, use a parameter-optimized hypernetwork to predict parameter offsets based on the text knowledge to be edited, update the editing model using the parameter offsets, and use the updated editing model to embed the text knowledge to be edited.

[0088] The knowledge graph embedding method and apparatus based on the pre-trained model provided in the above embodiments can quickly, effectively and locally update the knowledge graph embedding without retraining the entire model, so as to adapt to changes or errors in the knowledge graph. Compared with traditional model editing techniques, it is no longer necessary to adjust too many parameters, making the model easier to train and optimize. In addition, it significantly improves editing efficiency and greatly reduces the consumption of time and resources.

[0089] The specific embodiments described above illustrate the technical solution and beneficial effects of the present invention in detail. It should be understood that the above description is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, additions, and equivalent substitutions made within the scope of the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A knowledge graph embedding method based on a pre-trained model, characterized in that, Includes the following steps: Text knowledge extracted from text corpora and knowledge graphs is used to train the pre-trained model for a mask-based prediction task, resulting in a knowledge graph embedding model. Then, text knowledge that is predicted incorrectly is selected as sample data based on the knowledge graph embedding model. A supernetwork is introduced to predict parameter offsets based on sample data; An additional parameter layer is added to the knowledge graph embedding model to form an editing model. The structure of this additional parameter layer is the same as that of the feedforward neural network in the knowledge graph embedding model. The parameters are randomly initialized parameters superimposed with the parameter offsets predicted by the hypernetwork. The sample data is divided into data that needs to be edited and data that does not need to be edited. Based on the data that needs to be edited, a cross-entropy loss is constructed between the prediction results of the editing model and the target entity. Based on the data that does not need to be edited, a KL divergence loss is constructed between the prediction results of the knowledge graph embedding model and the editing model respectively. The parameters of the hypernetwork are optimized based on the cross-entropy loss and the KL divergence loss. For the text knowledge to be edited, a parameter-optimized hypernetwork is used to predict parameter offsets based on the text knowledge to be edited. The parameter offsets are then used to update the editing model, and the updated editing model is used to embed the text knowledge to be edited.

2. The knowledge graph embedding method based on a pre-trained model according to claim 1, characterized in that, The extracted text knowledge is represented by triples (h,r,t), where h and t represent the head entity and tail entity, and r represents the relation. When training the pre-trained model for a mask-based prediction task based on triples (h,r,t), either h or t in the triples (h,r,t) is masked and then input into the pre-trained model to obtain an embedding representation. A classifier is then introduced to predict and classify the masked entities based on the embedding representation. The parameters of the pre-trained model are optimized based on the cross-entropy loss between the predicted classification results and the real entities. The pre-trained model with optimized parameters serves as a knowledge graph embedding model. The text knowledge that is incorrectly predicted based on the knowledge graph embedding model is represented as (x,y,a), where x represents the model input, y represents the incorrect or outdated entity, i.e. the model output, and a represents the target entity.

3. The knowledge graph embedding method based on a pre-trained model according to claim 1, characterized in that, The hypernetwork includes a hidden state extraction module and a mapping module. The input text knowledge is processed by the hidden state extraction module to extract the hidden state, and this hidden state is then mapped and predicted by the mapping module to obtain vectors α, β ∈ R. m ,γ,δ∈R n And a scalar η∈R, where m×n represents the size of the random initialization parameters of the extra parameter layer, and R represents a real number; The parameter offset is calculated based on the vectors α, β, γ, δ, the scalar η, and the cross-entropy loss gradient of the input text knowledge in the knowledge graph embedding model. in, σ represents the Softmax function, and σ represents the Sigmoid function. The input text knowledge x represents the input parameters. The gradient of the cross-entropy loss between the knowledge graph embedded in the model and the target entity a. ⊙ represents the parameter offset, and ⊙ represents the dot product.

4. The knowledge graph embedding method based on a pre-trained model according to claim 3, characterized in that, The hidden state extraction module uses a bidirectional LSTM network, and the mapping module uses a fully connected network.

5. The knowledge graph embedding method based on a pre-trained model according to claim 1, characterized in that, After adding an extra parameter layer to the knowledge graph embedding model, the feedforward neural network in the edited model is represented as follows: Where H represents the output of the attention layer in the Transformer of the knowledge graph embedding model. This represents the original parameters in the knowledge graph embedding model. This represents the randomized parameters generated by the additional parameter layer. This represents the parameter offset predicted by the supernetwork, and ΔFFN represents the additional parameter layer added to achieve parameter offset. Implementation Let denot be the parameter matrix of the pre-trained language model, GELU be the activation function, and FFN be the original feedforward neural network. ′ This represents a feedforward neural network with added parameter layers.

6. The knowledge graph embedding method based on a pre-trained model according to claim 1, characterized in that, The cross-entropy loss is expressed as: Where, x i This represents the text knowledge that needs to be edited, including head entities and relations, or relations and tail entities, a i Represents the target entity. The parameter is The editing model predicts the probability of the target entity. This represents the cross-entropy loss.

7. The knowledge graph embedding method based on a pre-trained model according to claim 1, characterized in that, The KL divergence loss is expressed as: in, This represents the KL divergence loss. This represents the text knowledge x that does not need to be edited within the same batch. ′ The set that is formed This represents the set of erroneous entities y. x represents ′ Input parameters are The parameterized representation of the discrete distribution of incorrect entities predicted by the knowledge graph embedding model in the sample space. x represents ′ Input parameters are The edit model predicts the parameterized representation of the discrete distribution of erroneous entities in the sample space.

8. A knowledge graph embedding device based on a pre-trained model, characterized in that, It includes pre-training units, hypernetwork introduction units, extra parameter layer addition units, hypernetwork training units, and embedding representation units. The pre-training unit is used to train the pre-trained model for a mask-based prediction task using text knowledge extracted from a text corpus and a knowledge graph, to obtain a knowledge graph embedding model, and to filter out text knowledge that is predicted incorrectly as sample data based on the knowledge graph embedding model. The hypernetwork introduction unit is used to introduce a hypernetwork for predicting parameter offsets based on sample data. The additional parameter layer addition unit is used to add an additional parameter layer to the knowledge graph embedding model to form an editing model. The structure of the additional parameter layer is the same as that of the feedforward neural network in the knowledge graph embedding model, and the parameters are randomly initialized parameters superimposed with the parameter offset predicted by the hypernetwork. The hypernetwork training unit is used to divide the sample data into data that needs to be edited and data that does not need to be edited. Based on the data that needs to be edited, a cross-entropy loss is constructed between the prediction results of the editing model and the target entity. Based on the data that does not need to be edited, a KL divergence loss is constructed between the prediction results of the knowledge graph embedding model and the editing model, respectively. The parameters of the hypernetwork are optimized based on the cross-entropy loss and the KL divergence loss. The embedding representation unit is used to predict parameter offsets based on the knowledge of the text to be edited using a parameter-optimized hypernetwork, update the editing model using the parameter offsets, and then use the updated editing model to perform the embedding representation of the knowledge of the text to be edited.

9. A knowledge graph embedding device based on a pre-trained model, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the knowledge graph embedding method based on the pre-trained model as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Video production method and system based on machine learning algorithm

    CN114915841A

  • Question generation method based on social text and related device

    CN115700513A