Knowledge graph completion method based on semantic similarity model
Through the knowledge graph completion method based on semantic similarity model, the problems of incompleteness and lag in updates are solved, and the accurate completion and efficient update of the knowledge graph are achieved.
Patent Information
- Application Number
- CN202510270123.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, the knowledge graph has a large number of unconnected or undefined entities and relationships due to the limitations of data acquisition, lag in information updates or incomplete original data, which affects application effects and service quality.
Using a knowledge graph completion method based on semantic similarity model, a new triple is generated by extracting the ontology and entity networks of the existing knowledge graph, and a new triple with similarity is evaluated through the semantic similarity model, and a new triple with similarity below the threshold is filtered and added.
It effectively reduces the introduction of irrelevant or redundant information, ensures the accuracy of complementary knowledge, significantly improves the quality of acquisition of new knowledge, reduces the computational complexity, and speeds up the problem solving speed.
Smart Images

Figure CN120197680A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of knowledge graph completion, and particularly to a knowledge graph completion method based on a semantic similarity model. Background Art
[0002] A knowledge graph is a structured way of representing knowledge, which expresses entities in the real world and their relationships in the form of a graph. This graph consists of nodes (representing entities or concepts) and edges (representing relationships between entities). A knowledge graph can be regarded as a specific implementation of a semantic network, which not only contains the data itself, but also the relevance between the data and the context information of these data.
[0003] Large Language Models (LLMs) are advanced models in the field of Natural Language Processing (NLP). They learn the complex structure and semantics of language through pre-training on large-scale text data. These models can understand and generate natural language, and perform various tasks such as text classification, sentiment analysis, machine translation, and question answering systems. With the progress of deep learning technology, the scale and capabilities of large language models are constantly growing, and they perform excellently in understanding and generating natural language.
[0004] A semantic similarity model is a technology for evaluating the semantic similarity between two text fragments, entities, or concepts. This model works by converting text into points in a vector space and then calculating the distance or similarity metric (such as cosine similarity) between these points. Semantic similarity models have a wide range of applications in fields such as information retrieval, question answering systems, and text classification. They can help the system understand the intention of the user's query, recommend relevant information, or classify the text.
[0005] Knowledge graph completion is a technical operation performed on a constructed knowledge graph, aiming to discover and fill in the missing entities and relationships in the knowledge graph. In practical applications, due to limitations in data acquisition, lag in information update, or the incompleteness of the original data itself, there are often a large number of unconnected or undefined entities and relationships in the knowledge graph, which will affect the application effect and service quality of the knowledge graph. Summary of the Invention
[0006] Aiming at the above deficiencies in the prior art, the present invention provides a knowledge graph completion method based on a semantic similarity model, which solves the problems of incomplete knowledge graph and lagging update in the prior art.
[0007] To achieve the above invention purpose, the technical solution adopted by the present invention is: A knowledge graph completion method based on a semantic similarity model, comprising:
[0008] S1. Extract the ontology network and entity network from the existing knowledge graph, and perform parsing to obtain the triples of the existing knowledge graph;
[0009] S2. Input the unstructured text data into the large language model, and generate new triples in combination with Schema constraints;
[0010] S3. Encode the new triples and the triples of the existing knowledge graph through the semantic similarity model, and convert them into semantic vectors of fixed length;
[0011] S4. According to the semantic vectors of fixed length, evaluate the similarity between the new triples and the triples in the existing entity network;
[0012] S5. According to the similarity evaluation results, filter out the new triples whose similarity to the triples in the existing entity network is lower than the similarity threshold, and add them to the existing knowledge graph.
[0013] Furthermore: In S1, the extraction of the ontology network and entity network from the existing knowledge graph includes the extraction of entities, attributes, relationships and their hierarchical structures.
[0014] Furthermore: In S2, the large language model is a pre-trained natural language processing model, and by introducing the constraints of the ontology hierarchy and concept relationships, new triples that conform to the knowledge graph architecture are generated.
[0015] Furthermore: S3 includes:
[0016] S31. Encode the new triple A' and the triple B' of the existing knowledge graph through the semantic similarity model to obtain the low-dimensional vector representations of triple A and triple B;
[0017] Among them, both triple A and triple B are represented as (h, r, t), where h represents the head entity, r represents the relationship, and t represents the tail entity;
[0018] S32. Perform embedding representation on the encoded triples through a shared Transformer encoder to obtain the concatenated embedding vector;
[0019] S33. Train the semantic similarity model through the contrastive learning loss function;
[0020] S34. Fine-tune the semantic similarity model through a triple dataset containing similarity matching annotations, and convert the concatenated embedding vector into a semantic vector of fixed length.
[0021] Furthermore: In S33, during the training process, a positive and negative sample pair is set for each triple, and the training objective is: minimize the distance between positive examples and maximize the distance between negative examples.
[0022] Furthermore, S5 includes:
[0023] S51. Adjust the similarity threshold according to the actual application scenario of the knowledge graph;
[0024] S52. According to the similarity evaluation results, filter out the new triples whose similarity to the triples in the existing entity network is lower than the similarity threshold;
[0025] S53. Conduct a consistency check on the filtered new triples, and add the triples that pass the consistency check to the existing knowledge graph.
[0026] Furthermore, in S53, the method of adding the triples that pass the consistency check to the existing knowledge graph is as follows:
[0027] S531. Update the entity nodes and relationship edges in the knowledge graph according to the triples that pass the consistency check;
[0028] S532. Check whether the triples that pass the consistency check introduce new entities or relationships:
[0029] If so, adjust the overall structure of the knowledge graph and complete the update of the knowledge graph;
[0030] If not, complete the update of the knowledge graph.
[0031] The beneficial effects of the present invention are as follows:
[0032] 1. The present invention discriminates the similarity between new triples and existing triples through a semantic similarity model, ensuring that only new triples with a large difference from the existing knowledge are added to the knowledge graph, effectively reducing the introduction of irrelevant or redundant information and ensuring the accuracy of the supplemented knowledge;
[0033] 2. Using the ontology network structure of the knowledge graph as a constraint to guide the large language model to extract new triples that conform to the knowledge graph structure from unstructured text, significantly improving the quality of new knowledge acquisition;
[0034] 3. By optimizing the calculation of semantic vectors and similarity evaluation, the computational complexity of adding new knowledge to the knowledge graph is effectively reduced, and the problem-solving speed is accelerated, especially showing higher efficiency and accuracy when dealing with large-scale complex knowledge graphs. Description of the Drawings
[0035] Figure 1 It is a flowchart of the knowledge graph completion method based on the semantic similarity model.
[0036] Figure 2 It is an analysis diagram of the knowledge graph completion method based on the semantic similarity model. Detailed Embodiments
[0037] The specific embodiments of the present invention will be described below to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
[0038] As Figure 1 and Figure 2 shown, in an embodiment of the present invention, a knowledge graph completion method based on a semantic similarity model is provided, including:
[0039] S1. Extract the ontology network and entity network in the existing knowledge graph, and perform parsing to obtain the triples of the existing knowledge graph;
[0040] S2. Input the unstructured text data into a large language model, and combine with Schema constraints to generate new triples;
[0041] S3. Encode the new triples and the triples of the existing knowledge graph through a semantic similarity model, and convert them into semantic vectors of a fixed length;
[0042] S4. According to the semantic vectors of the fixed length, perform a similarity evaluation on the new triples and the triples in the existing entity network;
[0043] S5. According to the similarity evaluation results, filter out the new triples whose similarity to the triples in the existing entity network is lower than the similarity threshold, and add them to the existing knowledge graph.
[0044] In S1, the extraction of the ontology network and entity network in the existing knowledge graph includes the extraction of entities, attributes, relationships, and their hierarchical structures.
[0045] In S2, the large language model is a pre-trained natural language processing model, and by introducing the constraints of the ontology hierarchy and concept relationships, new triples that conform to the knowledge graph architecture are generated.
[0046] S3 includes:
[0047] S31. Encode the new triple A' and the triple B' of the existing knowledge graph through a semantic similarity model to obtain the low-dimensional vector representations of the triple A and the triple B;
[0048] Among them, both the triple A and the triple B are represented as (h, r, t), where h represents the head entity, r represents the relationship, and t represents the tail entity;
[0049] S32. Perform an embedding representation on the encoded triples through a shared Transformer encoder to obtain a concatenated embedding vector;
[0050] S33. Train the semantic similarity model using a contrastive learning loss function;
[0051] S34. Fine-tune the semantic similarity model using a triple dataset containing similarity matching annotations, and convert the concatenated embedding vectors into fixed-length semantic vectors.
[0052] In S33, during the training process, set a positive and negative sample pair for each triple, and the training objective is to minimize the distance between positive examples and maximize the distance between negative examples.
[0053] S5 includes:
[0054] S51. Adjust the similarity threshold according to the actual application scenario of the knowledge graph;
[0055] S52. Filter out new triples whose similarity to the triples in the existing entity network is lower than the similarity threshold according to the similarity evaluation results;
[0056] S53. Perform a consistency check on the filtered new triples, and add the triples that pass the consistency check to the existing knowledge graph.
[0057] In S53, the method of adding the triples that pass the consistency check to the existing knowledge graph is as follows:
[0058] S531. Update the entity nodes and relationship edges in the knowledge graph according to the triples that pass the consistency check;
[0059] S532. Check whether the triples that pass the consistency check introduce new entities or relationships:
[0060] If so, adjust the overall structure of the knowledge graph and complete the update of the knowledge graph;
[0061] If not, complete the update of the knowledge graph.
[0062] In an embodiment of the present invention, taking the knowledge graph of superalloy materials as an example, data including 205 superalloy grades, 50 ontology network triples and their corresponding related network structure features were obtained, which are: physical and chemical parameters, shear modulus, microstructure, preparation process, misfit degree, performance parameters, element composition, creep temperature, etc., and 7338 entity network triples.
[0063] It should be noted that this embodiment only takes the knowledge graph of superalloy materials as an example, and does not mean that the knowledge graph completion method based on the semantic similarity model provided by this application is limited to the knowledge graph of superalloy materials. In the application scenarios of different types of knowledge graphs, the similarity threshold can be dynamically adjusted, and the setting is adjusted according to different fields, data sparsity, and task requirements to meet the requirements of precision and recall for graph completion in different scenarios and ensure the flexibility and adaptability of the method.
[0064] In this embodiment, S1: Extract the ontology network and entity network in the existing knowledge graph, and perform parsing to obtain the triples of the existing knowledge graph;
[0065] S2: Input the unstructured text data into the large language model, and generate new triples in combination with Schema constraints;
[0066] S3: Encode the new triples and the triples of the existing knowledge graph through the semantic similarity model, and convert them into semantic vectors of fixed length;
[0067] S4: According to the semantic vectors of fixed length, perform similarity evaluation on the new triples and the triples in the existing entity network;
[0068] S5: According to the similarity evaluation results, screen out the new triples with similarity lower than the similarity threshold with the triples in the existing entity network, and add them to the existing knowledge graph.
[0069] Specifically, the specific method of S1 is: Connect to the neo4j graph database, and use Cypher query to extract the ontology network and entity network in the knowledge graph of superalloy materials;
[0070] The ontology network includes the basic attributes and conceptual relationships of materials, such as composition, microstructure, preparation process, and mechanical properties, etc. All ontology relationships are represented in the form of triples (Head-Relation-Tail).
[0071] The entity network includes specific superalloy material grades and related attributes. All entity relationships are represented in the form of triples (Head-Relation-Tail).
[0072] In S2, the large language model is a pre-trained natural language processing model. By introducing the constraints of the semantic hierarchy and conceptual relationships of the ontology network, new triples that conform to the knowledge graph architecture are generated. The specific steps are as follows:
[0073] S21: Collect unstructured data in the superalloy field to form a dedicated corpus;
[0074] Data sources include scientific research papers, technical manuals (such as "China Superalloy Handbook"), patent literature, experimental simulations, etc.
[0075] S22: According to the ontology network, extract semantic levels and conceptual relationships, define entity types and relationship types, and design prompts for large language models (such as GPT-4) as constraints;
[0076] The prompts will limit the large language model to only extract specified types of entities and relationships, ensuring that the triples extracted by the large language model match the pattern of the superalloy material knowledge graph;
[0077] S23: Input the unstructured text and prompts into the large language model for entity-relationship triple extraction, and the new triples output by the large language model will be used in the subsequent knowledge graph completion process.
[0078] S3 specifically includes:
[0079] S31: For text input, the triples will be converted into a format similar to natural language, such as "Entity 1 is related to Entity 2 through a relationship", and then this text will be input into the Transformer model of the semantic similarity model Triple-BERT for encoding;
[0080] For embedding input, if the embedding vectors of existing entities and relationships in the knowledge graph are available, these vectors can be directly input into the semantic similarity model Triple-BERT for further semantic encoding;
[0081] Whether it is text input or embedding input, the semantic similarity model Triple-BERT will uniformly map the triples into the same semantic vector space, ensuring the representational consistency of new and old triples;
[0082] S32: Use the semantic similarity model Triple-BERT to encode the concatenated sentences;
[0083] The semantic similarity model Triple-BERT uses a shared Transformer encoder to process each triple, generating a fixed-length semantic vector. Through the siamese network structure, the model can encode two triples simultaneously, ensuring encoding consistency.
[0084] S33: Train the semantic similarity model Triple-BERT through contrastive learning. The positive samples are triples with similar semantics, and the negative samples are triples with different semantics; by minimizing the distance between positive samples and maximizing the distance between negative samples, the model can distinguish between semantically similar and dissimilar triples;
[0085] S34: Fine-tune the semantic similarity model Triple-BERT using the labeled triple similarity dataset to ensure that the similarity between new and old triples can be calculated in the same semantic space.
[0086] S5 includes: setting the similarity threshold to 0.9;
[0087] If the similarity of the new triple is higher than the threshold, it is considered semantically similar to the existing triple and this triple is ignored; if the similarity is lower than the threshold, it is considered semantically different and further processing is required;
[0088] S52: Compare the entities in the new triples with low similarity with the entities in the existing entity network;
[0089] If the new entity does not exist in the knowledge graph, add it as a new entity node to the superalloy material knowledge graph.
[0090] S53: Calculate the semantic similarity of the relationship types in the new triples;
[0091] If it is found that the new relationship is similar to the existing relationship type, align them and merge them into the existing relationship; if the new relationship type is different from the existing one, add it as a new relationship type to the superalloy material knowledge graph and connect it with the relevant entity nodes.
[0092] S54: Add the new triples to the knowledge graph in the form of nodes and edges;
[0093] Update the Neo4j graph database through Cypher statements to realize the visual display of the knowledge graph, ensuring dynamic growth and the integrity of the graph structure.
[0094] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. A knowledge graph completion method based on a semantic similarity model, characterized in that: include: S1. Extract the ontology network and entity network in the existing knowledge graph and parse them to obtain the triples of the existing knowledge graph; S2, input the unstructured text data into the large language model, combine the Schema constraints, and generate new triples; S3, encode the new triples and the triples of the existing knowledge graph through the semantic similarity model and convert them into fixed-length semantic vectors; S4, based on the fixed-length semantic vector, the similarity between the new triples and the triples in the existing entity network is evaluated; S5. Based on the similarity evaluation results, new triples whose similarity to triples in the existing entity network is lower than the similarity threshold are screened out and added to the existing knowledge graph.
2. The knowledge graph completion method based on the semantic similarity model according to claim 1 is characterized in that: In S1, the extraction of ontology network and entity network in the existing knowledge graph includes the extraction of entities, attributes, relationships and their hierarchical structures.
3. The knowledge graph completion method based on the semantic similarity model according to claim 1 is characterized in that: In S2, the large language model is a pre-trained natural language processing model that generates new triples that conform to the knowledge graph architecture by introducing constraints on ontology hierarchy and concept relationships.
4. The knowledge graph completion method based on semantic similarity model according to claim 1, characterized in that S3 include: S31, encoding the new triple A' and the triple B' of the existing knowledge graph through a semantic similarity model to obtain a low-dimensional vector representation of triple A and triple B; Among them, triple A and triple B are both expressed as (h, r, t), h represents the head entity, r represents the relationship, and t represents the tail entity; S32, embedding the encoded triples through the shared Transformer encoder to obtain a concatenated embedding vector; S33, training the semantic similarity model by contrastive learning loss function; S34. Fine-tune the semantic similarity model using a triple dataset containing similar matching annotations to convert the concatenated embedding vector into a fixed-length semantic vector.
5. The knowledge graph completion method based on the semantic similarity model according to claim 4 is characterized in that: In S33, a positive and negative sample pair is set for each triplet during the training process, and the training goal is to minimize the distance between positive examples and maximize the distance between negative examples.
6. The knowledge graph completion method based on the semantic similarity model according to claim 1, characterized in that S5 include: S51. Adjust the similarity threshold according to the actual application scenario of the knowledge graph; S52, based on the similarity evaluation result, screening out new triples whose similarity to triples in the existing entity network is lower than a similarity threshold; S53. Perform a consistency check on the screened new triples, and add the triples that pass the consistency check to the existing knowledge graph.
7. The knowledge graph completion method based on semantic similarity model according to claim 6, characterized in that: In S53, the method of adding the triples that pass the consistency check to the existing knowledge graph is: S531, updating entity nodes and relationship edges in the knowledge graph according to the triples that pass the consistency check; S532: Check whether the triples that pass the consistency check introduce new entities or relationships: If yes, adjust the overall structure of the knowledge graph and complete the knowledge graph update; If not, the knowledge graph update is completed.
Citation Information
Cited By
Inference method, system and equipment of wide constraint large language model and medium
CN121146067A