Knowledge graph completion method combining embedding method and large language model

By combining embedding methods and large language models, structured embeddings and semantic embeddings are generated, and a comparison learning framework is used to integrate them, the existing knowledge graph completion method has solved the problems of insufficient learning of graph structure information, limited generalization ability and poor adaptability of specific knowledge graphs, achieving a more efficient and accurate knowledge graph completion effect.

CN120218199APending Publication Date: 2025-06-27FUDAN UNIVERSITY
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510178899.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing knowledge graph completion method has defects in insufficient learning of graph structure information, limited generalization ability, and poor adaptability of specific knowledge graphs, resulting in poor performance in complex multi-relationship graphs or specific domain graphs.

Method used

Combining the embedding method and large language model, the knowledge graph completion is achieved by generating structured embedding and semantic embedding, and fusion using a contrast learning framework. This method not only utilizes the structural information of the knowledge graph, but also integrates deep semantic features, enhancing the adaptability and generalization capabilities of the model.

Benefits of technology

It significantly improves the ability to learn graph structure information, enhances the generalization ability of the model, can better adapt to the knowledge graph characteristics of specific fields, and improves the accuracy and efficiency of knowledge graph completion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218199A_ABST
    Figure CN120218199A_ABST
Patent Text Reader

Abstract

The invention provides a knowledge graph completion method combining an embedding method and a large language model, and belongs to the technical field of knowledge graph processing. According to the knowledge graph complementing method, the model based on embedding and the large language model are combined, the graphic characteristics of the structured data and the semantic depth of the text data are fully utilized, and the accuracy and efficiency of knowledge graph complementing are effectively improved. Compared with the traditional knowledge graph completion technology, the method has the advantages that the large language model is introduced to process the text description, and the text description is fused with the embedding generated based on the embedding model, so that the performance of the completion task is greatly improved. According to the method, a comparative learning framework is introduced to optimize prediction and enhance the generalization ability of a model (completion method) on various knowledge maps.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of knowledge graph processing, and particularly relates to a knowledge graph completion method combining an embedding method and a large language model. Background Art

[0002] A knowledge graph is an information network constructed by structurally representing entities and their relationships in the real world. Its purpose is to simulate the human knowledge system to support various intelligent applications, such as semantic search, personalized recommendation, natural language understanding, and intelligent question-answering systems. The core value of a knowledge graph lies in its ability to provide rich structured knowledge to help machines understand and process a large amount of unstructured data.

[0003] Although the concept of a knowledge graph has been widely promoted and applied in scientific research and industrial applications, due to the dynamic and evolving nature of knowledge, existing knowledge graphs often face the problem of incomplete knowledge. This incompleteness limits the application potential of the knowledge graph because the missing links may be crucial for achieving high-quality knowledge reasoning and decision-making. For example, in a knowledge graph in the field of healthcare, there may be a lack of associations between certain diseases and symptoms, which directly affects the accuracy and reliability of diagnosis and treatment recommendations.

[0004] To address this issue, the Knowledge Graph Completion (KGC) task emerged. The goal of this task is to automatically infer and fill in the missing relationships in the knowledge graph to make the graph more complete, thereby improving its performance in various applications. Early methods were mainly structured learning methods, such as TransE [1] and RotatE [2] , which predict the missing links between entities by learning the low-dimensional vector representations of entities and relationships. While maintaining the entity relationship structure, these methods attempt to infer the potential connections between entities through distances or angles in the vector space. For example, the RotatE model captures features such as symmetry, antisymmetry, and transitivity of relationships by treating relationships as rotations in the complex space.

[0005] Text-based knowledge graph completion methods improve the accuracy of knowledge graph completion by integrating the text descriptions of entities and relationships and leveraging the semantic information in these descriptions. For example, the DKRL model first introduced text descriptions into entity embeddings and used convolutional neural networks to process text data, thereby enhancing the semantic information of entity representations. Subsequent studies such as KG-BERT [3] , StAR [4]Et cetera, all use pre-trained language models to more deeply analyze text descriptions and extract rich semantic features of entities and relationships. Recent methods use descriptions of entities and relationships in triples as prompts for large language models and utilize the model outputs for relationship prediction. [5] 。

[0006] The existing technologies have the following problems:

[0007] (1) Insufficient learning of graph structure information: Existing text-based knowledge graph completion methods can effectively utilize semantic relationships to predict entities and relationships, but they often ignore the structural information in the knowledge graph. Traditional embedding methods only rely on the semantic information of entities and relationships and fail to fully capture the global structural characteristics of the graph. Therefore, these methods do not perform well when dealing with complex multi-relational graphs or knowledge graphs with sparse structures.

[0008] (2) Limited generalization ability: Most existing knowledge graph completion methods rely on specific datasets for fine-tuning and training, and thus have weak generalization ability. When applied to different knowledge graphs, these methods usually require a large amount of model adjustment and cannot flexibly adapt to knowledge graph completion in different domains. Especially in scenarios with long-tail entities and class imbalance, the prediction ability of existing methods significantly decreases, and overfitting or misjudgment is likely to occur.

[0009] (3) Poor adaptability to specific knowledge graphs: Although existing methods can handle some standardized knowledge graphs, they often have difficulty adapting to the unique semantics and structures of graphs in specific domains. Existing methods lack a flexible parameter adjustment mechanism and cannot quickly adapt to knowledge graphs in different domains, which greatly limits their efficiency and effectiveness in practical applications.

[0010] [1] Bordes A, Usunier N, Garcia-Duran A, et al. Translating embeddings for modeling multi-relational data[J]. Advances in neural information processing systems, 2013, 26.

[0011] [2] Xie R, Liu Z, Jia J, et al. Representation learning of knowledge graphs with entity descriptions[C] / / Proceedings of the AAAI conference on artificial intelligence. 2016, 30(1).

[0012] [3]Yao L,Mao C,Luo Y.KG-BERT:BERT for knowledge graph completion[J].arXiv preprint arXiv:1909.03193,2019.

[0013] [4]Wang B,Shen T,Long G,et al.Structure-augmented text representation learning for efficient knowledge graph completion[C] / / Proceedings of the Web Conference 2021.2021:1737-1748.

[0014] [5]YaoL,PengJ,Mao C,Luo Y.Exploring Large Language Models for Knowledge Graph Completion.https: / / arxiv.org / pdf / 2308.13916v3,2023. Summary of the Invention

[0015] This invention is made to solve the above problems, and aims to provide a knowledge graph completion method that combines embedding methods and large language models.

[0016] This invention provides a knowledge graph completion method that combines embedding methods and large language models, with the following features, including the following steps: S10, loading the basic data of the knowledge graph, where the basic data includes the initial structured information of entities and relationships and the text descriptions of entities and relationships; S20, using an embedding-based model and generating embedding vectors for each entity and relationship according to the initial structured information to obtain structured embeddings; S30, using a large language model to process the text descriptions to extract the deep semantic features of entities and relationships to obtain semantic embeddings; S40, using the semantic embeddings as positive samples, and randomly selected other entities and relationships as negative samples, and through a hybrid learning strategy, aligning and optimizing the structured embeddings and semantic embeddings to achieve the fusion of the two and obtain fused embeddings; S50, through natural language prompts, using the fused embeddings to predict the missing entities or relationships in the knowledge graph in a contrastive learning framework and output the results, finally realizing the completion of the knowledge graph.

[0017] In the knowledge graph completion method combining the embedding method and the large language model provided by the present invention, it may further have the following features: Among them, in step S20, the embedding-based model includes the RotatE model, and in step S30, the large language model includes llama.

[0018] In the knowledge graph completion method combining the embedding method and the large language model provided by the present invention, it may further have the following features: Among them, step S20 includes the following sub-steps: S21, using the embedding-based model, training the corresponding model with the known triples in the knowledge graph as the training set, and generating the embedding vector of the entity as the entity embedding; S22, using the embedding-based model, training the corresponding model with the known triples in the knowledge graph as the training set, and generating the embedding vector of the relationship as the relationship embedding. The entity embedding and the relationship embedding together serve as the structured embedding.

[0019] In the knowledge graph completion method combining the embedding method and the large language model provided by the present invention, it may further have the following features: Among them, step S30 includes the following sub-steps: S31, combining the name of the entity and its text description and passing them through the large language model to obtain the semantic embedding vector of the entity; S32, combining the name of the relationship and its text description and passing them through the large language model to obtain the semantic embedding vector of the relationship. The semantic embedding vectors of the entity and the relationship together serve as the semantic embedding.

[0020] In the knowledge graph completion method combining the embedding method and the large language model provided by the present invention, it may further have the following features: Among them, step S40 includes the following sub-steps: S41, constructing a contrastive learning framework, using the semantic embedding vector of the entity obtained in step S31 as the positive sample, and randomly extracting the remaining entities as negative samples, so as to obtain the projection of the entity embedding on the word vector dimension of the large language model using a multi-layer perceptron, denoted as the entity embedding projection; S42, constructing a contrastive learning framework, using the semantic embedding vector of the relationship obtained in step S32 as the positive sample, and randomly extracting the remaining relationships as negative samples, so as to obtain the projection of the relationship embedding on the word vector dimension of the large language model using a multi-layer perceptron, denoted as the relationship embedding projection. The entity embedding projection and the relationship embedding projection together serve as the fused embedding.

[0021] In the knowledge graph completion method combining the embedding method and the large language model provided by the present invention, it may further have the following features: Among them, in step S41, the multi-layer perceptron makes the distance between the entity embedding projection and the positive sample in the word vector space as close as possible, and the distance from the negative samples of other entities as far as possible. In step S42, the multi-layer perceptron makes the distance between the relationship embedding projection and the positive sample in the word vector space as close as possible, and the distance from the negative samples of other relationships as far as possible.

[0022] In the knowledge graph completion method combining the embedding method and the large language model provided by the present invention, it may further have the following features: Among them, step S50 includes the following sub-steps: S51, after constructing a natural language prompt, passing the natural language prompt through the large language model to obtain a prompt embedding; S52, using the concatenation of the prompt embedding and the fusion embedding as input to fine-tune the large language model; S53, in combination with the fine-tuned large language model, inputting the names and text descriptions of the entities and the names and text descriptions of the relationships of the triples in the knowledge graph, and predicting and outputting the results for unknown entities or relationships through the corresponding structured embeddings obtained by the embedding-based model, finally realizing knowledge graph completion.

[0023] In the knowledge graph completion method combining the embedding method and the large language model provided by the present invention, it may further have the following features: Among them, in step S52, a linear layer is added to the last layer of the large language model, corresponding to the number of entities, to perform a classification task, and the large language model is fine-tuned by calculating the classification loss.

[0024] In the knowledge graph completion method combining the embedding method and the large language model provided by the present invention, it may further have the following features: Among them, in step S52, the optimizer used includes the Adam optimizer, and the learning rate is 1e-5.

[0025] In the knowledge graph completion method combining the embedding method and the large language model provided by the present invention, it may further have the following features: Among them, in step S50, the output result is a list of possible entities or relationships, and the one with the highest score is the entity or relationship predicted for knowledge graph completion.

[0026] Functions and effects of the invention

[0027] The present invention improves the ability to learn graph structure information: Current text-based methods mainly use semantic relationships to enable the model to understand entities and relationships in the knowledge graph, but this process often ignores learning the structure information in the knowledge graph. To solve this problem, this patent combines the structured embedding method and the large language model to semantically embed the knowledge graph entities and relationships. This dual embedding strategy significantly improves the model's prediction ability for complex relationships, especially when dealing with entities and relationships with rich text descriptions, and can more accurately capture and utilize the deep semantic information in these descriptions.

[0028] The present invention improves the generalization ability of the model: The model generalization ability of existing methods is insufficient. Especially in knowledge graphs with long-tail entities and class imbalance, the weak generalization ability will lead to difficult improvement in the performance of such relationship prediction. This patent introduces a contrastive learning framework, which optimizes the generation and adjustment process of embedding vectors. By comparing the embedding vectors of positive and negative samples, the model can more effectively learn the features to distinguish correct and incorrect relationships, reducing misjudgment. This progress not only enhances the prediction accuracy of the model, but also enables the model to be applicable to a wider range of knowledge graph application scenarios, including those with sparse data or class imbalance.

[0029] The present invention improves the adaptability to specific knowledge graphs: There is rich domain knowledge and structural features in knowledge graphs in specific domains. However, existing methods do not process domain knowledge and structural information deeply enough, resulting in insufficient adaptability to such knowledge graphs. The present invention uses the LoRA technology to adjust the parameters of the large language model to better adapt to the features of specific knowledge graphs. The introduction of this adaptive adjustment technology enables the model to more accurately understand and process the semantic structure and relationships related to specific knowledge graphs while maintaining the original semantic processing ability. In addition, this technology supports fast fine-tuning and deployment, greatly improving the flexibility and efficiency of the model in practical applications. Especially when facing large-scale knowledge graphs, it can achieve fast model adaptation and update. Brief Description of the Drawings

[0030] Figure 1 is the overall flowchart of a knowledge graph completion method combining an embedding method and a large language model in Embodiment 1 of the present invention;

[0031] Figure 2 is the model training process of the first stage of a knowledge graph completion method combining an embedding method and a large language model in Embodiment 1 of the present invention;

[0032] Figure 3 is the model training process of the second stage of a knowledge graph completion method combining an embedding method and a large language model in Embodiment 1 of the present invention. Detailed Embodiment

[0033] In order to make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the following embodiments will specifically elaborate on a knowledge graph completion method combining an embedding method and a large language model of the present invention in conjunction with the drawings.

[0034] Before further elaborating on the technical details of this embodiment, it is necessary to explain several key concepts to better understand the technical background and implementation mechanism of this embodiment. These concepts include semantic enhancement, embedding space regularization, and cross-modal learning. The following is a detailed introduction to these three concepts:

[0035] Definition 1: Semantic Augmentation: Semantic Augmentation is the process of improving the quality and richness of data representation by introducing additional semantic information. In knowledge graph completion, by introducing external knowledge (such as Wikipedia, professional databases, etc.) or context information generated by natural language processing into the descriptions of entities and relationships, the model's ability to understand the dynamics of entity attributes and relationships can be significantly enhanced.

[0036] Definition 2: Embedding Space Regularization: In the embedding learning of knowledge graphs, regularization techniques are used to avoid overfitting and improve the model's generalization ability on unseen data. By imposing regularization constraints (such as L1 or L2 regularization) in the embedding space, the model can be helped to learn smoother and more consistent vector representations of entities and relationships, thus achieving more stable and reliable performance in complex knowledge graph completion tasks.

[0037] Definition 3: Cross-modal Learning: Cross-modal Learning involves extracting and integrating information from multiple data sources (such as text, images, sounds) to enhance the performance of learning models. In knowledge graph applications, combining data from different modalities can not only provide more comprehensive entity descriptions but also improve the accuracy and depth of relationship inference through multiple perceptual dimensions. Although this embodiment does not directly use traditional multi-modal data, the process of combining structured entity embeddings and natural language refers to common processing methods of cross-modal learning.

[0038] <Example 1>

[0039] Figure 1 It is the overall flowchart of a knowledge graph completion method combining an embedding method and a large language model in Example 1 of the present invention.

[0040] As Figure 1 shown, this embodiment provides a knowledge graph completion method combining an embedding method and a large language model, including the following steps:

[0041] S10, Initialize the knowledge graph:

[0042] Load the basic data of the knowledge graph, where the basic data includes the initial structured information of entities and relationships and the text descriptions of entities and relationships.

[0043] Among them, the text description includes definitions and additional information extracted from relevant databases and literature.

[0044] S20, Generate structured embeddings, including the following sub-steps S21 - S22:

[0045] S21. Use the embedding-based model (specifically, the RotatE model in this embodiment), and after training the corresponding model using the known triples in the knowledge graph as the training set, generate the embedding vectors of entities as entity embeddings.

[0046] S22. Use the RotatE model, and after training the corresponding model using the known triples in the knowledge graph as the training set, generate the embedding vectors of relations as relation embeddings, and use the entity embeddings and relation embeddings together as structured embeddings.

[0047] In the above steps S21 - S22, the embedding vectors can capture and express various relationships between entities, including symmetry and / or antisymmetry.

[0048] Figure 2 This is the first-stage model training process of a knowledge graph completion method that combines embedding methods and large language models in Embodiment 1 of the present invention.

[0049] As Figure 2 shown, the first stage of the completion method in this embodiment is a contrastive learning framework, including the following steps S30 - S40:

[0050] S30. The LLM extracts semantic features, including the following sub-steps S31 - S32:

[0051] S31. Combine the name of the entity and its text description and pass them through the large language model (specifically, llama in this embodiment) to obtain the semantic embedding vector of the entity.

[0052] S32. Combine the name of the relation and its text description and pass them through the large language model (llama) to obtain the semantic embedding vector of the relation, and use the semantic embedding vectors of the entity and the relation together as semantic embeddings, so as to extract the deep semantic features of the entity and the relation.

[0053] S40. Embedding fusion, including the following sub-steps S41 - S42:

[0054] S41. Construct a contrastive learning framework, use the semantic embedding vector of the entity obtained in step S31 as the positive sample, and randomly select the remaining entities as negative samples, so as to obtain the projection of the entity embedding on the word vector dimension of the large language model (llama), denoted as the entity embedding projection.

[0055] Among them, the multi-layer perceptron makes the distance between the entity embedding projection and the positive sample in the word vector space as close as possible, and the distance from other entity negative samples as far as possible.

[0056] S42. Construct a contrastive learning framework. Use the semantic embedding vectors of the relationships obtained in step S32 as positive samples, and randomly select the remaining relationships as negative samples. Then, use a multi-layer perceptron to obtain the projection of the relationship embedding on the word vector dimension of the large language model (Llama), denoted as the relationship embedding projection. Use the entity embedding projection and the relationship embedding projection together as the fused embedding.

[0057] Among them, the multi-layer perceptron makes the distance between the relationship embedding projection and the positive samples in the word vector space as close as possible, and the distance from other relationship negative samples as far as possible.

[0058] In the above steps S41 - S42, through a hybrid learning strategy, fuse the structured embeddings generated by the RotatE model with the semantic embeddings extracted from Llama, which involves the alignment and optimization of the embeddings to ensure that the two embeddings can effectively express the information of the knowledge graph in the same vector space.

[0059] Figure 3 This is the second-stage model training process of a knowledge graph completion method combining an embedding method and a large language model in Embodiment 1 of the present invention.

[0060] As Figure 3 shown, the second stage of the completion method in this embodiment fine-tunes the Llama model, including the following step S50:

[0061] S50. Link prediction to achieve knowledge graph completion, including the following sub-steps S51 - S53:

[0062] S51. After constructing a natural language prompt, pass the natural language prompt through the large language model (Llama) to obtain a prompt embedding.

[0063] S52. Use the concatenated prompt embedding and fused embedding as the input, add a linear layer to the last layer of the large language model (Llama), corresponding to the number of entities, and perform a classification task. Fine-tune the large language model (Llama) by calculating the classification loss.

[0064] S53. Combine the fine-tuned large language model in step S52, input the names and text descriptions of the entities and the names and text descriptions of the relationships in the triples of the knowledge graph, and after obtaining the corresponding structured embeddings by the RotatE model, predict the unknown entities or relationships and output the results (the output results are a list of possible entities or relationships, and the one with the highest score is the entity or relationship predicted for knowledge graph completion), finally achieving knowledge graph completion.

[0065] <Embodiment 2>

[0066] Based on Embodiment 1, this embodiment provides a knowledge graph completion method that combines an embedding method and a large language model, including the following steps:

[0067] Step S10, initialize the graph data: Load the basic data of the knowledge graph, where the basic data includes the initial structured information of entities and relationships, as well as the text descriptions of entities and relationships. Among them, the text descriptions include definitions and additional information extracted from relevant databases and literature.

[0068] Given a triple " / m / 02jx1 / location / location / contains / m / 013t85". Among them, " / m / 02jx1" is the head entity serial number, " / location / location / contains" is the relationship name, and " / m / 013t85" is the tail entity serial number. Here, taking the tail entity completion task as an example, that is, the tail entity is unknown to the model.

[0069] Find the name "England" and the text description "England is a country that is part of the United Kingdom. It shares land borders with Scotland to the north and Wales to the west." corresponding to the head entity " / m / 02jx1" in the dataset, and find the text description "A scenario where one geographical entity encompasses another." corresponding to the relationship " / location / location / contains".

[0070] S20, generate structured embeddings, including the following sub-steps S21 - S22:

[0071] S21, use the RotatE model, utilize the known triples in the knowledge graph as the training set, train the RotatE model corresponding to the dataset, and through this model, obtain the embedding vector corresponding to the head entity " / m / 02jx1" as the entity embedding (head_embedding).

[0072] S22. Use the RotatE model. After training the corresponding model with the known triples in the knowledge graph as the training set, generate the embedding vector of the relation " / location / location / contains" as the relation embedding, and use the entity embedding of the head entity and the relation embedding together as the structured embedding.

[0073] S30. The LLM extracts semantic features, including the following sub-steps S31 - S32:

[0074] S31. Combine the name "England" of the head entity " / m / 02jx1" and its text description "England is a country that is part of the United Kingdom. It shares land borders with Scotland to the north and Wales to the west." into "England: England is a country that is part of the United Kingdom. It shares land borders with Scotland to the north and Wales to the west." Pass it through llama to obtain the semantic embedding vector of the head entity.

[0075] S32. Pass the text description of the relation " / location / location / contains" through llama to obtain the semantic embedding vector of the relation. Use the semantic embedding vector of the head entity and the semantic embedding vector of the relation together as the semantic embedding, so as to extract the deep semantic features of the entity and the relation.

[0076] S40. Embedding fusion, including the following sub-steps S41 - S42:

[0077] S41. Build a contrastive learning framework. Use the semantic embedding vector of the head entity obtained in step S31 as the positive sample (positive_token_embedding). Similarly, randomly select the remaining entities as negative samples (negative_token_embedding). Use a multi-layer perceptron to obtain the projection of the entity embedding (head_embedding) on the word vector dimension of the large language model (llama), denoted as the entity embedding projection (head_projection_embedding).

[0078] Using a multi-layer perceptron (MLP) (two layers, with the activation function being GELU, and the hidden layer dimension being 4096→4096), project the entity embedding (head_embedding) into the semantic space of the large language model (llama), making it close to the positive sample (positive_token_embedding) and far from the negative sample (negative_token_embedding), to obtain the entity embedding projection (head_projection_embedding).

[0079] S42. Obtain the projection of the relation embedding (relation_embedding) on the word vector dimension of the large language model (llama) in a manner corresponding to step S41, denoted as the relation embedding projection (relation_projection_embedding). Use the entity embedding projection (head_projection_embedding) and the relation embedding projection (relation_embedding) together as the fused embedding.

[0080] In the above steps S41 - S42, through a hybrid learning strategy, fuse the structured embeddings (relation embeddings and entity embeddings) generated by the RotatE model in step S20 with the semantic embeddings (the semantic embedding vectors of the head entity and the semantic embedding vectors of the relations) extracted from llama in step S30. This involves the alignment and optimization of the embeddings to ensure that the two types of embeddings can effectively represent the information of the knowledge graph within the same vector space.

[0081] S50. Link prediction to achieve knowledge graph completion, including the following sub-steps S51 - S53:

[0082] S51, Write a natural language prompt, such as "Please predict the relationship. head entity: {name: England, description:, England is a country that is part of the United Kingdom. It shares land borders with Scotland to the north and Wales to the west.}, relation: {name: / location / location / contains, description: A scenario where one geographical entity encompasses another.}", and obtain the prompt embedding through llama for the natural language prompt.

[0083] S52, Concatenate the prompt embedding with the fused embedding (the entity embedding projection and relation embedding projection obtained in step S40) as the input (input_embedding). Add a linear layer to the last layer of llama, corresponding to the number of entities, and perform a classification task. Fine-tune llama by calculating the classification loss (using the Adam optimizer with a learning rate of 1e-5). Fine-tune for 3 epochs.

[0084] S53, Combine the fine-tuned llama in step S52, input the head entity and its text description, and the relation and its text description of the triples in the knowledge graph, and through the corresponding structured embeddings (relation embedding and entity embedding) obtained by the RotatE model, the unknown tail entity can be predicted and the result can be output (the output result is a list of possible entities or relations, and the one with the highest score is the entity or relation predicted for knowledge graph completion. In this embodiment, the predicted tail entity is " / m / 013t85"), and finally, knowledge graph completion is achieved.

[0085] Output the completed knowledge graph and conduct an evaluation to verify the completion effect and the efficiency of the completion method.

[0086] <Test Example>

[0087] This test example uses a standard knowledge graph dataset (FB15k-237). After conducting actual tests using a knowledge graph completion method that combines an embedding method and a large language model provided in Example 2, the performance of the completion method in Example 2 is compared with that of multiple existing knowledge graph completion models. To comprehensively evaluate its performance, this test example uses the following commonly used evaluation metrics:

[0088] (1) MRR (Mean Reciprocal Rank): This metric is used to evaluate the average ranking performance of the model / method's predictions, reflecting whether the model / method can accurately predict the target entity. A higher MRR value indicates that the model has better ranking ability during prediction.

[0089] (2) Hits@1: This metric indicates whether the model / method can accurately hit the target entity in the first prediction result. The Hits metric can intuitively reflect the prediction accuracy of the model / method, especially its ability to predict the correct answer in multiple-choice scenarios.

[0090] The test results are shown in Table 1 below:

[0091] Table 1 (Comparison of the performance of the completion method in Example 2 with multiple existing knowledge graph completion models)

[0092]

[0093] As shown in Table 1 above, the knowledge graph completion method that combines the embedding method and the large language model outperforms multiple existing knowledge graph completion models in terms of both the MRR and Hits@1 metrics on the FB15k-237 dataset. Among them:

[0094] MRR = 0.362, showing a certain improvement compared to the traditional RotatE (0.338) and HAKE (0.346), indicating that this method is superior in overall ranking performance.

[0095] Hits@1 = 0.289, which is more than 3.6% higher than models such as MEM-KGC (0.253) and HAKE (0.250), indicating that this method has an advantage in the accuracy of the first prediction result.

[0096] Functions and effects of the example

[0097] (1) Making full use of graph structure and semantic information to improve prediction ability: This example combines traditional embedding methods and large language models, not only utilizing the structural information in the knowledge graph but also integrating semantic features extracted from text. Through this dual embedding strategy, the model can perform more excellently in capturing complex relationships and inferring sparse structure graphs, overcoming the deficiencies of existing technologies in learning graph structure information and greatly improving the accuracy and generality of predictions.

[0098] (2) Enhanced generalization ability and efficiency: In this embodiment, by introducing a contrastive learning framework, the generation and adjustment process of the embedding vectors are effectively optimized, thereby enhancing the generalization ability of the model. Compared with the prior art, this embodiment can quickly adapt to the data characteristics of different knowledge graphs without large-scale fine-tuning. Especially when dealing with long-tail entities and imbalanced data sets, it can still maintain a high prediction accuracy and significantly reduce the occurrence of misjudgments.

[0099] (3) Adaptability and flexibility to specific knowledge graphs: This embodiment uses the LoRA technique to fine-tune the parameters of the large language model, enabling it to better adapt to the characteristics of knowledge graphs in specific domains. This technique not only greatly improves the adaptability while maintaining the semantic processing ability, but also supports the rapid deployment and update of the model. Especially when facing large-scale knowledge graphs, it can quickly adjust the model, enhancing the flexibility and efficiency in practical applications.

[0100] Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only to illustrate the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A knowledge graph completion method combining an embedding method and a large language model, characterized in that: The following steps are involved: S10, loading basic data of the knowledge graph, wherein the basic data includes initial structured information of entities and relationships and text descriptions of entities and relationships; S20, using an embedding-based model and generating an embedding vector for each of the entities and relationships according to the initial structured information to obtain a structured embedding; S30, using a large language model to process the text description to extract deep semantic features of the entities and relations to obtain semantic embedding; S40, using the semantic embedding as a positive sample and the remaining randomly selected entities and relations as negative samples, aligning and optimizing the structured embedding with the semantic embedding through a hybrid learning strategy to achieve fusion of the two, thereby obtaining a fused embedding; S50, through natural language prompts, using the fusion embedding under the contrastive learning framework to predict missing entities or relationships in the knowledge graph and output the results, ultimately completing the knowledge graph.

2. The knowledge graph completion method combining the embedding method and the large language model according to claim 1, characterized in that: in, In step S20, the embedding-based model includes a RotatE model, In step S30, the large language model includes llama.

3. The knowledge graph completion method combining the embedding method and the large language model according to claim 1 or 2, Features: in, Step S20 includes the following sub-steps: S21, using the embedding-based model, using the known triples in the knowledge graph as a training set to train the corresponding model and then generate an embedding vector of the entity as an entity embedding; S22, using the embedding-based model, using the known triples in the knowledge graph as a training set to train the corresponding model and then generate an embedding vector of the relationship as the relationship embedding, The entity embedding and the relation embedding together serve as the structured embedding.

4. The knowledge graph completion method combining the embedding method and the large language model according to claim 3, characterized in that: in, Step S30 includes the following sub-steps: S31, combining the name of the entity and its text description through the large language model to obtain a semantic embedding vector of the entity; S32, combining the name of the relationship and its text description through the large language model to obtain a semantic embedding vector of the relationship, The semantic embedding vectors of the entity and the relationship are collectively used as the semantic embedding.

5. The knowledge graph completion method combining the embedding method and the large language model according to claim 4, characterized in that: in, Step S40 includes the following sub-steps: S41, constructing a contrastive learning framework, taking the semantic embedding vector of the entity obtained in step S31 as a positive sample, and randomly extracting the remaining entities as negative samples, thereby using a multi-layer perceptron to obtain a projection of the entity embedding on the word vector dimension of the large language model, recorded as entity embedding projection; S42, constructing a contrastive learning framework, taking the semantic embedding vector of the relationship obtained in step S32 as a positive sample, randomly extracting the remaining relationships as negative samples, and using a multi-layer perceptron to obtain a projection of the relationship embedding on the dimension of the large language model word vector, recorded as the relationship embedding projection, The entity embedding projection and the relationship embedding projection are used together as the fused embedding.

6. The knowledge graph completion method combining the embedding method and the large language model according to claim 5, characterized in that: in, In step S41, the multi-layer perceptron makes the distance between the entity embedding projection and the positive sample in the word vector space as close as possible, and the distance from other entity negative samples as far as possible. In step S42, the multilayer perceptron makes the distance between the relation embedding projection and the positive sample in the word vector space as close as possible and the distance between the relation embedding projection and the negative sample as far as possible.

7. The knowledge graph completion method combining the embedding method and the large language model according to claim 1 or 2, Features: in, Step S50 includes the following sub-steps: S51, after constructing the natural language prompt, the natural language prompt is passed through the large language model to obtain a prompt embedding; S52, concatenating the prompt embedding and the fused embedding as input, thereby fine-tuning the large language model; S53, combined with the fine-tuned large language model, input the name of the entity and its text description and the name of the relationship and its text description of the triple in the knowledge graph, and predict the unknown entity or relationship after obtaining the corresponding structured embedding through the embedding-based model and output the result, finally realizing the knowledge graph completion.

8. The knowledge graph completion method combining the embedding method and the large language model according to claim 7, characterized in that: in, In step S52, a linear layer is added to the last layer of the large language model, corresponding to the number of entities, so as to perform the classification task, and the large language model is fine-tuned by calculating the classification loss.

9. The knowledge graph completion method combining the embedding method and the large language model according to claim 8, characterized in that: in, In step S52, the optimizer used includes the Adam optimizer, and the learning rate is 1e-5.

10. The knowledge graph completion method combining the embedding method and the large language model according to claim 1, characterized in that: in, In step S50, the output result is a list of possible entities or relationships, among which the one with the highest score is the predicted entity or relationship to be used for knowledge graph completion.

Citation Information

Cited By

  • End-to-end fusion method, device and equipment for atlas and large model, medium and product

    CN120744078A

  • Knowledge graph entity completion method

    CN120851176A

  • Heterogeneous double-model-based large model output reliability enhancement method

    CN122264113A

  • A large model output reliability enhancement method based on a heterogeneous double model

    CN122264113B