Knowledge graph completion method based on pre-training language model

Through the combination of multi-perspective generation and re-evaluation models, the problem of insufficient modeling of knowledge graph structure information in the existing technology is solved, and a transparent and interpretable knowledge graph completion method is realized, which improves prediction accuracy and credibility.

CN120578769APending Publication Date: 2025-09-02BEIJING INST OF TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510489027.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The existing knowledge graph completion method based on pre-trained language models relies too much on triple text information, lacks structural information modeling of knowledge graphs, resulting in error accumulation and a lack of visual inference process, and is unable to fully utilize the ability of language model generation.

Method used

A multi-view generation paradigm is constructed, and relevant triplets are obtained through neighbor selection models and inserted text sequences are inserted. The pre-trained language model is used to generate target entities at the triple and path level. Combined with the re-evaluation model to make consistency and difficult prediction judgments, and a heuristic fusion sorting strategy is used to improve prediction accuracy.

Benefits of technology

Effective modeling of local triplets and far-reaching path structures is achieved, transparent and interpretable inference processes are provided, and the performance and credibility of knowledge graph completion is improved, especially on WN18RR and FB15K-237 datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120578769A_ABST
    Figure CN120578769A_ABST
Patent Text Reader

Abstract

The invention provides a pre-training language model-based knowledge graph completion method, which comprises the following steps of: S1, giving an input query pair, and constructing a text sequence of the query pair; s2, obtaining a plurality of neighbor triples most related to the query pair through a neighbor selection model, and inserting the neighbor triples into the tail of the text sequence of the query pair to form a new text sequence; s3, adding an instruction label for the new text sequence, inputting the instruction label into a pre-training language model, and generating a target entity sequence at a triple level and a path level; and S4, judging two groups of target entities generated by the triple level and the path level, and if the judgment result is consistency prediction, taking the prediction of the path level as final prediction output. According to the method, the pre-training language model can be subjected to multi-view generation, the knowledge graph completion capability of the pre-training language model is effectively improved, and the interpretability of a prediction result is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing and knowledge graph technology, and in particular to a knowledge graph completion method based on a pre-trained language model. Background Art

[0002] The knowledge graph aims to provide a good data organization form for the storage, management, query and retrieval of massive information data on the Internet, and plays an important supporting role in many downstream fields. However, although the number and scale of current knowledge graphs are constantly growing, such as DBpedia, Freebase, NELL, YAGO, etc., they are often incomplete due to their own knowledge capacity. Therefore, it is of great significance to use appropriate automated knowledge graph completion algorithms to mine potential ternary relationships in knowledge graphs. Among them, link prediction for entities and relationships is one of the core technologies of knowledge graph completion, which aims to use a given query pair, that is, an entity and relationship pair. , combined with existing knowledge in the knowledge graph to predict possible tail entities , in order to realize the potential ternary relationship in the knowledge graph of excavation.

[0003] Traditional knowledge graph completion methods mainly focus on embedding the symbolic rules or structures of knowledge graphs, constructing a logical system or representation space, and using this system or space to mine potential missing triple relationships in the knowledge graph. However, they lack the representation and understanding of the semantic level of knowledge, resulting in a relative gap between knowledge graph representation learning and natural language processing tasks. With the in-depth development of language models, the combination of knowledge graph tasks and language models has emerged, and has demonstrated strong generalization performance and higher prediction accuracy. For example, reference [1] inputs the text representation of the test triple into the encoder-based pre-trained language model BERT, and predicts the true value of the candidate triple from the perspective of text representation. On this basis, reference [2] constructs a dual-tower encoding model for query pairs and candidate entities, and introduces the entity description text. By sharing the encoding of the candidate entity, the computational space consumption is greatly reduced, and the prediction accuracy is effectively improved. Different from the above ideas, the literature [3] constructed an encoder-decoder prediction generation framework, which used the text and description information of the query triple as encoding information, and generated the predicted target tail entity through the decoder side of the model, effectively utilizing the generation ability of the language model.

[0004] However, existing knowledge graph completion methods based on pre-trained language models rely too much on the textual information of triples and do not adequately model the structural information of knowledge graphs. This problem becomes more acute after the introduction of entity descriptions. Although some literature [4] distinguishes textual knowledge from knowledge graph knowledge through soft hint engineering, or independently models the embedding representation of entities and relationships, and integrates or distills this embedding rich in structural information into the language model, their indirect modeling methods not only bring about error accumulation problems, but also restrict the language model from directly modeling and understanding the structural information of the knowledge graph. Moreover, they do not have the ability to model the global and multi-hop structure of the knowledge graph, which restricts the further improvement of the performance of such methods. In addition, the prediction results of such methods are closed black boxes, which do not give full play to the generation ability of the language model and lack visual reasoning process or human-readable explanation, which undermines the credibility of the prediction results.

[0005] References:

[0006] [1]Yao L, Mao C, Luo Y. KG-BERT: BERT for knowledge graph completion[J]. arXiv preprint arXiv:1909.03193, 2019. DOI: 10.48550 / arXiv.1909.03193.

[0007] [2]Wang L, Zhao W, Wei Z, et al. SimKGC: Simple Contrastive KnowledgeGraph Completion with Pre-trained Language Models[C] / / Proceedings of the 60thAnnual Meeting of the Association for Computational Linguistics (Volume 1:Long Papers). 2022: 4281-4294. DOI: 10.18653 / v1 / 2022.acl-long.295.

[0008] [3]Saxena A, Kochsiek A, Gemulla R. Sequence-to-Sequence KnowledgeGraph Completion and Question Answering[C] / / Proceedings of the 60th AnnualMeeting of the Association for Computational Linguistics (Volume 1: LongPapers). 2022: 2814-2828. DOI: 10.18653 / v1 / 2022.acl-long.201.

[0009] [4]Lv 10.18653 / v1 / 2022.findings-acl.282. Summary of the Invention

[0010] In order to solve the above problems, the present invention proposes a knowledge graph completion method based on a pre-trained language model, which includes:

[0011] S1. Given an input query pair , construct the text sequence of query pairs;

[0012] S2. Obtain several neighbor triplets that are most relevant to the query pair through the neighbor selection model and insert them into the end of the text sequence of the query pair to form a new text sequence;

[0013] S3. Add instruction labels to the new text sequence and input it into the pre-trained language model. Use the multi-perspective generation paradigm to generate target entity sequences at the triple level and path level.

[0014] S4. Determine the two groups of target entities generated at the triple level and the path level. If the determination result is a consistent prediction, take the path level prediction as the final prediction output.

[0015] Further, in step S1, the text sequence includes the query pair Entity name, entity description text, relationship Corresponding knowledge soft prompt vector row , relationship name, where the knowledge soft prompt matrix corresponding to the knowledge soft prompt vector row is updated when training the pre-trained language model.

[0016] Furthermore, in step S2, the neighbor selection model includes an encoder and a decoder, and the training method includes:

[0017] S21. Select the triples adjacent to the query pair in the training set in the knowledge graph to form an adjacent set ;

[0018] S22. The triples in the adjacent set that contain the target entity corresponding to the query pair are regarded as positively correlated samples, and the other triples are regarded as negatively correlated samples; and the neighbor selection model is trained based on the positive and negative correlation samples.

[0019] Furthermore, in step S22, the contrast loss function of the training is:

[0020] ,in, represents the concatenated text vector of the input, is the interval factor, represents the set of decoding scores of negative samples, represents the decoding score set of positive samples, Represents the decoding score of the neighbor selection model for the current triplet prediction .

[0021] Furthermore, in S3, the training method of the pre-trained language model is:

[0022] The triplets in the knowledge graph and the triplets with paths are used as training data to train the pre-trained language model. The loss functions in both perspectives are: ,in, represents the length of the decoder prediction sequence, is the input text sequence, for Step 1 generates characters; then adjust the parameters of the pre-trained language model through the back propagation algorithm, including the knowledge soft prompt matrix .

[0023] Furthermore, in S4, if the entity ranked first in the prediction set from the triple perspective and the entity ranked first in the path-level prediction are the same entity, the prediction is called a consistent prediction.

[0024] Furthermore, in S4, if the number of entities predicted from the triple perspective , or the number of identical entities predicted by the two views is less than , then the prediction is called difficult to predict; if the result is difficult to predict, the two sets of target entity sets are re-evaluated, then merged and sorted by confidence score to obtain the final prediction result.

[0025] Furthermore, in S4, re-evaluation is performed by a re-evaluation model, which is composed of a dual-tower model based on an encoder. The training method of the re-evaluation model includes:

[0026] S41. Consider all triples in the knowledge graph as positive training examples and construct false triples as negative examples.

[0027] S42, all triples in the training set Split into query pairs With the target entity Two parts, input to the reassessment model for encoding;

[0028] S43. Calculate the cosine similarity between the two sets of coding features and use it as a triplet Confidence score for being true;

[0029] S44. Use contrastive learning to train a re-evaluation model to minimize the cosine distance between the query pair of the true triple and the target entity.

[0030] Furthermore, in S44, the training goal of the re-evaluation model is to minimize the cosine distance between the query pairs and the target entity of all positive sample triplets. The contrastive learning loss function used in the training process is:

[0031] ,in, is the spacing factor to ensure that the model increases the score of positive samples, is the temperature coefficient, which is used to adjust the relative importance of negative samples.

[0032] Furthermore, in S4, if the evaluation result is neither a consistent prediction nor a difficult prediction, the path-level prediction is taken as the final prediction output.

[0033] The beneficial effects of the present invention are as follows:

[0034] (1) This paper proposes a new multi-view generation paradigm that can effectively capture local triple-level knowledge and profound path structure knowledge, and can simultaneously model text and structural information, achieving a transparent and explainable reasoning process.

[0035] (2) This paper proposes a re-evaluation model and a heuristic fusion ranking strategy for the triples generated from multiple perspectives and the two sets of reasoning results on the path. By efficiently cross-validating and fusing the generated results from multiple perspectives, the reasoning performance is further improved.

[0036] (3) The present invention achieves excellent performance on two standard datasets, WN18RR and FB15K-237, and enables the knowledge graph completion results to have a human-readable reasoning process. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0038] Figure 1 Schematic diagram of a flow chart of a knowledge graph completion method based on a pre-trained language model according to an embodiment of the present invention;

[0039] Figure 2 A schematic diagram of a training process of a neighbor selection model according to an embodiment of the present invention;

[0040] Figure 3 A schematic diagram of a process for generating predictions from multiple perspectives according to an embodiment of the present invention;

[0041] Figure 4 A schematic diagram of a flow chart of a heuristic fusion strategy according to an embodiment of the present invention;

[0042] Figure 5 2 is a schematic diagram of the training process of the re-evaluation model according to one embodiment of the present invention. DETAILED DESCRIPTION

[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be reviewed and fully described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0044] The present invention provides a method for completing a knowledge graph based on a pre-trained language model. The pre-trained language model used can be a sequence-to-sequence model that conforms to an encoder-decoder structure, such as T5 or BART. Figure 1 As shown, the method includes the following steps:

[0045] S1. Given an input query pair , construct the text sequence of query pairs;

[0046] S2. Obtain several neighbor triplets that are most relevant to the query pair through the neighbor selection model and insert them into the end of the text sequence of the query pair to form a new text sequence;

[0047] S3. Add instruction labels to the new text sequence and input it into the pre-trained language model for encoding. Using a multi-perspective generation paradigm, the decoder alternately generates target entity sequences at the triple level and the path level.

[0048] S4. Determine the two sets of target entities generated at the triple level and the path level. If the determination result is a consistent prediction, take the path level prediction as the final prediction output;

[0049] If the result is difficult to predict, the two sets of target entities are re-evaluated and then merged and sorted to obtain the final prediction result;

[0050] If the evaluation result is a general prediction, the path-level prediction is taken as the final prediction output.

[0051] In step S1, is the entity of the query pair, To distinguish knowledge graph knowledge from natural language text knowledge and emphasize the structural knowledge of knowledge graph, this paper introduces the knowledge soft prompt matrix ,in represents the number of relations in the knowledge graph, Represents the dimension of the pre-trained language model. As part of the parameters of the pre-trained language model, learning updates are performed synchronously during the training process of the pre-trained language model. When constructing the input, the query Entity name, entity description text, relationship Corresponding knowledge soft prompt vector row , the relation names are concatenated in order to form an input text sequence. For example, given the query pair "(Diamond Head Mountain, located in)", the corresponding text sequence is "Diamond Head Mountain [Diamond Head Volcano is an extinct volcano, and its top is like a Victoria leaf.] < > Located <mask>",in" <mask>" represents the target entity that needs to be predicted Covering.

[0052] In step S2, in order to fully mine and utilize the structural knowledge of the knowledge graph and the contextual information of the query pair, the present invention designs a neighbor selection model based on an encoder-decoder to score and filter the correlation between the query pair and its adjacent triples. Figure 2 As shown, the neighbor selection model training method includes:

[0053] S21. Select the triples adjacent to the query pair in the training set in the knowledge graph to form an adjacent set ;

[0054] S22. The triples in the adjacent set that contain the target entity corresponding to the query pair are regarded as positively correlated samples, and the other triples are regarded as negatively correlated samples; and the neighbor selection model is trained based on the positive and negative correlation samples.

[0055] In step S21, select the entity of the query pair in the knowledge graph surrounding The triples in the order adjacent subgraph constitute the adjacency set :

[0056]

[0057] In the above formula, and Representing the Entities or relationships of the same order.

[0058] In step S22, the adjacency set The triples containing the target entity corresponding to the query pair are regarded as positively correlated samples, and the triples not containing the target entity are regarded as negatively correlated samples with the query pair, and the two sets of sample sets are used as training sample sets. The text of the triplet in the adjacent set is concatenated and sent to the encoder of the neighbor selection model for joint encoding. The correlation between the two is judged at the decoder. The judgment target is , Indicates that the triple is irrelevant to the query pair. Indicates that the triple is related to the query pair. Then, based on the positive and negative labels of the adjacent triples, the neighbor selection model is trained using the following contrastive loss function:

[0059]

[0060] In the above formula, represents the concatenated text vector of the input, is the interval factor, represents the set of decoding scores of negative samples, represents the decoding score set of positive samples, Represents the decoding score of the neighbor selection model for the current triplet prediction , which is the correlation probability predicted by the decoder of the neighbor selection model.

[0061] After the neighbor selection model is trained, it is responsible for calculating the correlation between all adjacent triples and the query pair during the pre-trained language model prediction process, and selecting the top triples with the highest correlation. triples form a related set , and are concatenated to the end of the query text sequence to form a new input text sequence.

[0062] In step S3, in order to distinguish the generated perspectives of the current round, the present invention defines instruction tags and , namely: "Predict tail / head entity:" and "Predict tail / head entity with path:", as instructions to guide the decoder of the current pre-trained language model to generate triples or paths, and insert them into the head of the new input text sequence to form a new input text sequence . Then, as Figure 3 As shown, in order to realize the pre-trained language model to understand the local knowledge of triples and the profound knowledge of multi-hop at the same time, especially through the reasoning of multi-hop paths, the reasoning process can be visualized. The input is encoded into the encoder of the pre-trained language model, and the decoder generates the target entity alternately from the triple perspective and the path perspective The prediction sequence of Figure 3 As shown, the methods for generating predictions include:

[0063] S31. Encode input text sequence , directly generate all possible target entities and form a set of predicted target entities from the perspective of triples ;

[0064] In this step, the decoder generates the name of the target entity and the description of the target entity to form a target entity set. For example, given the query pair "(Diamond Head Mountain, located in)", the generated sequence of triple perspectives is "Hawaii [is an archipelago state consisting of 132 islands in the central Pacific]", where the predicted target entity For "Hawaii".

[0065] S32. Encode input text sequence , generate a multi-hop reasoning path starting from the query entity and the target entity name and description, forming a set of predicted target entities from the path perspective .

[0066] In this step, the decoder simultaneously generates a multi-hop reasoning path from the head entity to the tail entity, the name of the target entity, and the text description of the target entity to form the target entity set. For example: given the query pair "(Diamond Head, located)", the path-level generation sequence is "Diamond Head  located  Oahu  located in Honolulu  located in  Hawaii [is an archipelago consisting of 132 islands in the central Pacific]", where the predicted target entity For "Hawaii".

[0067] The training method of the pre-trained language model is to use the triples and paths in the knowledge graph as training data to train the pre-trained language model. The loss function from both perspectives can be expressed as: ,in, represents the length of the decoder prediction sequence, is the input text sequence, for Step 1 generates characters; then adjust the parameters of the pre-trained language model through the back propagation algorithm, including the knowledge soft prompt matrix .

[0068] In step S4, Figure 4 As shown, the present invention proposes a heuristic fusion sorting strategy for cross-validating and merging the two sets of results generated from the above different perspectives. That is, the two sets of results are scored fairly with confidence under a unified framework, and then the two sets of results are merged and sorted according to the scores. This can further improve the prediction accuracy and richness of the pre-trained language model through cross-validation of the two sets of prediction results. The fusion sorting strategy includes:

[0069] (1) If the entity ranked first in the prediction set of the triple perspective , if the entity ranked first with path level prediction If the predictions are the same entity, then the prediction is called a consistent prediction. Consistent prediction means that the target entity lists predicted from the two perspectives are relatively consistent, which is a relatively simple prediction task. At this time, you can directly select the prediction results at the path level. Output as the final prediction result without re-evaluation operation. The final output is expressed as:

[0070]

[0071] (2) If the number of entities predicted from the triple perspective , or the number of identical entities predicted by the two views is less than , then the prediction is called difficult prediction. Difficult prediction means that there is a big difference between the predictions of the two perspectives, which is a more difficult prediction. Therefore, it is necessary to generate the two sets of target entity sets. and The re-evaluation model is used to perform a unified confidence score, which is then merged and sorted by confidence score to perform cross-validation from two perspectives. The final output is a set of entities that are the result of merging the two sets of predictions and sorting them by score:

[0072]

[0073] in, and is a pre-defined threshold value. In the following experimental verification The value is 5. The value is 10.

[0074] (3) The results that are neither consistent nor difficult to predict can be regarded as general predictions. In order to simplify the calculation, the prediction results at the path level are directly used. The final output is: .

[0075] In strategy (2), the re-evaluation operation of the re-evaluation model requires additional computational overhead, so it is necessary to identify which predictions do not need to be re-evaluated in order to reduce the number of re-evaluation triplets and unnecessary computation. The re-evaluation model is composed of a two-tower model based on the encoder. and Perform fair confidence scoring under a unified framework, such as Figure 5 As shown, the training method of the re-evaluation model includes:

[0076] S41. Consider all triples in the knowledge graph as positive training examples and construct false triples as negative examples.

[0077] S42, all triples in the training set Split into query pairs With the target entity Two parts, input to the reassessment model for encoding;

[0078] S43. Calculate the cosine similarity between the two sets of coding features and use it as a triplet Confidence score for being true;

[0079] S44. Use contrastive learning to train a re-evaluation model to minimize the cosine distance between the query pair of the true triple and the target entity.

[0080] In step S41, the triples that already exist in the knowledge graph are regarded as positive examples, and false triples that do not exist in the knowledge graph are constructed. And regarded as negative examples, positive and negative samples together constitute the training sample set, where is a randomly selected entity in the knowledge graph, but it is consistent with the query They cannot be combined into triples that actually exist in the knowledge graph.

[0081] In step S42, the query pair is re-evaluated. Encode the text and description information, and perform maximum pooling on the output features to obtain the encoded feature vector of the query pair , for the target entity The text and description information are encoded to obtain the encoding feature vector of the target entity .

[0082] In step S43, the cosine similarity between any set of query pairs and the target entity is calculated by cosine distance, and the formula is as follows:

[0083]

[0084] In step S44, the training objective of the re-evaluation model is to minimize the cosine distance between the query pair and the target entity for all positive triples. Therefore, the following contrastive learning loss function is used to train the re-evaluation model:

[0085]

[0086] in, is the spacing factor to ensure that the model increases the score of positive samples, is the temperature coefficient, which is used to adjust the relative importance of negative samples.

[0087] The trained re-evaluation model can be used in strategy (2) to re-evaluate and sort the unpredictable results. Specifically, take the entity set All target entities and their corresponding query pairs The two entities are input into the re-evaluation model together to obtain the cosine similarity between them, which is used as the evaluation score. Then all entities in the set are sorted according to the evaluation score to obtain the final ordered set. as the final output.

[0088] Finally, the set of predicted target entities output by the pre-trained language model The higher score Entity and input query pairs Forming triples , which is inserted into the original knowledge graph as a newly discovered triple, and the multi-hop path generated by the path perspective can serve as a visual evidence of the triple, thereby achieving transparent and explainable completion of the knowledge graph.

[0089] Experimental verification

[0090] In one embodiment, the datasets used are the WN18RR dataset and the FB15K-237 dataset. The WN18RR dataset is a more challenging dataset obtained by removing multiple test leak source relationships and triples from the original WN18 dataset. The FB15K-237 dataset is a more realistic dataset obtained by removing additional interdependent relationships from the FB15K dataset.

[0091] This paper effectively enhances the performance of pre-trained language models on the WN18RR and FB15K-237 datasets. Specifically, the present invention uses T5 as the base model for experiments, namely, S1, given an input query pair, constructing the query pair's text sequence; S2, the neighbor selection model obtains several neighbor triplets that are most relevant to the query pair and inserts them into the end of the query pair's text sequence; S3, the text sequence is input into the encoder for encoding, and a multi-perspective generation paradigm is used to decode and generate answer entities at the triple level and path level; S4, the re-evaluation model evaluates, verifies, and merges and sorts the two sets of answer entities generated at the triple level and path level, and then outputs the final prediction result.

[0092] This paper experimentally validates the results on two standard datasets, using KG-BERT, StAR, GenKGC, KGT5, CoLE, SimKGC, KG-S2S, CSProm-KG, PDKGC, BMKGC, COSIGN, and PEMLM as baseline models for comparison on four metrics: MRR, Hits@1, Hits@3, and Hits@10. The experimental results are shown in Tables 1 and 2, respectively.

[0093] The best result is in bold and the second best result is underlined.

[0094] Table 1 Results on the WN18RR dataset

[0095]

[0096] Table 2 Results on the FB15k-237 dataset

[0097]

[0098] Those skilled in the art will understand that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art will understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.< / mask> < / mask>

Claims

1. A knowledge graph completion method based on a pre-trained language model, characterized in that: The method comprises: S1. Given an input query pair, construct a text sequence of the query pair; S2. Obtain several neighbor triplets that are most relevant to the query pair through the neighbor selection model and insert them into the end of the text sequence of the query pair to form a new text sequence; S3. Add instruction labels to the new text sequence and input it into the pre-trained language model. Use the multi-perspective generation paradigm to generate target entity sequences at the triple level and path level. S4. Determine the two groups of target entities generated at the triple level and the path level. If the determination result is a consistent prediction, take the path level prediction as the final prediction output.

2. The knowledge graph completion method according to claim 1, characterized in that: In step S1, the text sequence includes the entity name of the query pair, the entity description text, the knowledge soft prompt vector row corresponding to the relationship, and the relationship name, wherein the knowledge soft prompt matrix corresponding to the knowledge soft prompt vector row is updated when training the pre-trained language model.

3. The knowledge graph completion method according to claim 1, characterized in that: In step S2, the neighbor selection model includes an encoder and a decoder, and the training method includes: S21. Select the triples adjacent to the query pair in the training set in the knowledge graph to form an adjacent set ; S22. The triples in the adjacent set that contain the target entity corresponding to the query pair are regarded as positively correlated samples, and the other triples are regarded as negatively correlated samples; and the neighbor selection model is trained based on the positive and negative correlation samples.

4. The knowledge graph completion method according to claim 3, characterized in that: In step S22, the contrast loss function of the training is: ,in, represents the concatenated text vector of the input, is the interval factor, represents the set of decoding scores of negative samples, represents the decoding score set of positive samples, Represents the decoding score of the neighbor selection model for the current triplet prediction .

5. The knowledge graph completion method according to claim 1, characterized in that: In S3, the training method of the pre-trained language model is: The triplets in the knowledge graph and the triplets with paths are used as training data to train the pre-trained language model. The loss functions in both perspectives are: ,in, represents the length of the decoder prediction sequence, is the input text sequence, for The characters are generated in steps; then the parameters of the pre-trained language model, including the knowledge soft prompt matrix, are adjusted through the back-propagation algorithm.

6. The knowledge graph completion method according to claim 1, characterized in that: In S4, if the entity ranked first in the prediction set from the triple perspective is the same as the entity ranked first in the path-level prediction, the prediction is called a consistent prediction.

7. The knowledge graph completion method according to claim 1, characterized in that: In S4, if the number of entities predicted from the triple perspective , or the number of identical entities predicted by the two views is less than , then the prediction is called difficult to predict; If the result is difficult to predict, the two sets of target entities are re-evaluated, then merged and sorted by confidence score to obtain the final prediction result.

8. The knowledge graph completion method according to claim 7, characterized in that: In S4, re-evaluation is performed through a re-evaluation model, which is composed of a two-tower model based on an encoder. The training method of the re-evaluation model includes: S41. Consider all triples in the knowledge graph as positive training examples and construct false triples as negative examples. S42, all triples in the training set Split into query pairs With the target entity Two parts, input to the reassessment model for encoding; S43. Calculate the cosine similarity between the two sets of coding features and use it as a triplet Confidence score for being true; S44. Use contrastive learning to train a re-evaluation model to minimize the cosine distance between the query pair of the true triple and the target entity.

9. The knowledge graph completion method according to claim 8, characterized in that: In S44, the training goal of the re-evaluation model is to minimize the cosine distance between the query pairs and the target entity for all positive sample triplets. The contrastive learning loss function used in the training process is: ,in, is the spacing factor to ensure that the model increases the score of positive samples, is the temperature coefficient, which is used to adjust the relative importance of negative samples.

10. The knowledge graph completion method according to claim 1, characterized in that: In S4, if the evaluation result is neither a consistent prediction nor a difficult prediction, the path-level prediction is taken as the final prediction output.

Citation Information

Patent Citations

  • Knowledge graph completion method based on entity description and relationship path

    CN111026875A

  • Single-sample knowledge graph completion method based on path enhancement and entity metric cooperation

    CN117689015A

  • Pre-training language model knowledge graph completion method in combination with comparative learning

    CN117829280A

  • Open knowledge graph completion method and device based on pre-training language model prompt fine tuning

    CN117892807A

  • Knowledge graph completion method and system based on entity description and soft prompt enhancement

    CN118536588A