Knowledge graph completion method, system and device
By sampling entities and relationships in the knowledge graph, combining graph neural networks and pre-trained language models to integrate structure and semantic information, the problem of difficult to deal with complex semantic relationships and neglecting structural characteristics in the existing technology is solved, and the efficiency and accuracy of knowledge graph completion is achieved.
Patent Information
- Application Number
- CN202510190056.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-06
AI Technical Summary
The existing knowledge graph completion method has shortcomings in processing structure and semantic information. Embedded methods are difficult to deal with complex semantic relationships, while pre-trained language models ignore the structural characteristics of the graph.
By sampling entities and relationships in the knowledge graph to form a sub-graph, combining graph neural networks and neural network adapters to aggregate structural information, and splicing them with auxiliary text prompt information, inputting pre-trained language models to fuse structural and semantic information, and finally predicting and completing them through convolutional neural networks.
This method can effectively combine the structure and semantic information of the knowledge graph to improve the effectiveness and practical value of graph completion, and is particularly suitable for aggregating multi-hop relationship information in the knowledge graph.
Smart Images

Figure CN120106194A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of knowledge engineering technology, and in particular to a knowledge graph completion method, system and device. Background Art
[0002] Knowledge graphs have been widely used in many fields, especially in semantic search, recommendation systems and other fields. Knowledge graphs provide support for different information services by organizing a large amount of factual information (including head entities, relationships and tail entities). However, in the process of building knowledge graphs, due to the limitations of data sources and the imperfections of automated extraction algorithms, knowledge graphs often face the problem of incomplete data. This problem of incomplete data may lead to the missing of key links or entities in the graph, thus affecting its overall quality and application effect.
[0003] To solve this problem, the Knowledge Graph Completion (KGC) method came into being. Traditional KGC methods mainly focus on graph embedding-based models, which retain the structural information of the knowledge graph by mapping entities and relationships into a low-dimensional vector space. However, such methods rely too much on the structural characteristics of the knowledge graph and are limited by the problem of data sparsity. In addition, with the development of pre-trained language models (PLMs) such as BERT, T5, and GPT-3, researchers have begun to explore the use of PLMs to fuse semantic information in the graph to improve the accuracy and generalization ability of the model. Nevertheless, these methods often focus too much on textual information and fail to make full use of the structural information in the graph.
[0004] In summary, existing KGC methods have many deficiencies in processing the structure and semantic information of knowledge graphs: embedding-based methods, although they can capture structural information, are not good at processing complex semantic relationships; PLM-based methods, although they have advantages in text semantic processing, ignore the structural characteristics of graphs. Therefore, a new method is urgently needed to effectively integrate these two aspects of information to comprehensively improve the completion ability and practical value of knowledge graphs. Summary of the invention
[0005] The purpose of the present invention is to provide a knowledge graph completion method, system and device to solve at least one of the above-mentioned technical problems existing in the prior art.
[0006] In a first aspect, in order to solve the above technical problems, the present invention provides a knowledge graph completion method, comprising the following steps: Step 1: By querying entities, entities and relationships are sampled in the knowledge graph (KG) to form a subgraph; based on the query entity, corresponding auxiliary text prompt information (A2P) is constructed; Step 2: embed the subgraph and perform structural information aggregation (SIA) through a graph neural network (GNN) and a neural network adapter; then concatenate it with the auxiliary text prompt information to obtain a concatenated vector, input it into a pre-trained language model (PLMs), and output a vector representation that integrates structural information and semantic information; the auxiliary text prompt information includes query triples, entity descriptions, and entity types in the benchmark dataset; Step 3: Input the vector representation into a convolutional neural network, output the prediction result and add it to the knowledge graph, so as to complete the query entity related information in the knowledge graph; the convolutional neural network is used to reorder the subgraphs according to the weights of the candidate entities, and the candidate entities closer to the head entity in the vector representation have greater weights; Through the above method, we can construct an innovative framework for graph completion (SSH-KGC), which can cleverly combine the structure and semantic information of the knowledge graph to improve the efficiency of graph completion; by combining graph neural network and knowledge graph embedding (KGE) technology, it surpasses the traditional single-hop method (also called one hop, one hop in the knowledge graph is one relationship), and is particularly suitable for aggregating multi-hop relationship information in the knowledge graph.
[0007] In a feasible implementation manner, the query entity in step 1 is a query triplet that is missing a head entity or a tail entity.
[0008] In a feasible implementation manner, the specific method for constructing the auxiliary text prompt information in step 1 includes: Step 11: splice the query triples in the benchmark dataset according to the input format of the pre-trained language model to obtain a coherent text sequence, and add it to the preset position of the preset prompt template of the auxiliary text prompt information; this helps the query triples to be better aligned with the subsequent pre-trained language model, so that the latter can more easily understand the words in the benchmark dataset; Step 12: Add the entity description to the preset position of the preset prompt template of the auxiliary text prompt information, so as to help the subsequent pre-trained language model grasp the semantic nuances of the query, so as to more effectively understand the context of the current completion; Step 13: All triples containing the query entity in the benchmark data set are spliced into prompt information, input into the large model (LLM), and after identifying the entity type of the query entity, the entity type is added to the preset position of the preset prompt template of the auxiliary text prompt information; this makes it easier for the subsequent pre-trained language model to understand the specific semantics of the same query entity in different query triples.
[0009] In a feasible implementation manner, the benchmark dataset includes the FB15k-237 dataset and / or the WN18RR dataset.
[0010] In a feasible implementation manner, the specific method of aggregating the structural information in step 2 is: Step 21: Treat the subgraph as an undirected subgraph , calculate the head entity The undirected aggregation embedding of is expressed as follows: ; in, express The embedding vector of ; Representing relationships The embedding vector of Represents the tail entity The embedding vector of express middle Undirected aggregation embedding of; Represents an aggregation operation, which is specifically defined as: ; in, express The matrix of neighboring entity embeddings in ; express Normalized adjacency matrix; express The degree matrix of Represents the weight matrix of the GNN layer; represents a nonlinear activation function; Step 22: Treat the relationship as a node that is equally important as the entity, and include the relationship and the head and tail entities related to the relationship within the one-hop neighborhood into the relationship subgraph. ; Compute the relation-aware aggregation embedding, the specific expression is: ; in, express Middle head entity Relation-aware aggregation embedding of Relationship subgraph The specific expression is: ; in, Indicates the header entity associated with a one-hop neighbor and tail entity ; represents sub-graph sampling; Step 23: and Concatenate and input into the neural network adapter to get the final structure embedding , the specific expression is: ; in, Indicates an adapter.
[0011] In a feasible implementation, the neural network adapter includes a trainable two-layer neural network.
[0012] In a feasible implementation, the pre-trained language model is a BERT model.
[0013] In a feasible implementation, the loss function of the convolutional neural network in step 3 includes: Negative Sampling Loss , the specific formula can be: ; in, represents a hyperparameter, i.e., a score boundary value or interval (margin); represents the scoring function; Represents the head entity; To express a relationship; Represents the tail entity; Indicates auxiliary text prompt information; Indicates the number of negative samples; Represents the embedding of negative samples; Contains smoothing coefficient The cross entropy loss , the specific formula is: ; in, represents the predicted probability of the target entity given the query entity, relation, and auxiliary text hint information; represents the size of the label space, that is, the number of all possible values of the tail entity t; represents the predicted probability of negative samples given the query entity, relation, and auxiliary text hint information; Total loss function , the specific formula is: ; in, Represents a triple; represents the weight hyperparameter; Through the total loss function, the convolutional neural network can strike a balance between the two tasks of predicting the target entity and distinguishing negative samples.
[0014] In a feasible implementation, the evaluation index of the prediction result in step 3 includes the average reciprocal ranking and index; Average last ranking The specific formula is: ; in, Representing a collection of entities the number of The value range is from 0 to 1. The higher the value, the better the performance of the knowledge graph. When the value is 1, it means that for each query, the most relevant item is always ranked first in the query results. The specific formula of the indicator is: ; in, Indicators used to measure Entities Ranking list The proportion of queries where at least one relevant item appears in the first K positions; Represents an indicator function. If the condition is true, the function value is 1, otherwise the function value is 0; Indicates the number of digits.
[0015] In the second aspect, based on the same inventive concept, the present application also provides a knowledge graph completion method system, including a data receiving module, a data processing module and a result generating module; The data receiving module is used to receive the knowledge graph and the query entity; The data processing module includes a sub-graph unit, an auxiliary text prompt information unit, a structural information aggregation unit, a pre-trained language model unit and a convolutional neural network unit; The subgraph unit samples entities and relationships in the knowledge graph through the query entity to form a subgraph; The auxiliary text prompt information unit constructs corresponding auxiliary text prompt information based on the query entity; the auxiliary text prompt information includes the query triple, entity description and entity type in the benchmark data set; The structural information aggregation unit embeds the subgraph and aggregates the structural information through the graph neural network and the neural network adapter; and then splices it with the auxiliary text prompt information to obtain a splicing vector; The pre-trained language model unit stores a pre-trained language model, inputs a concatenated vector, and outputs a vector representation of the fusion structure information and semantic information; The convolutional neural network unit stores a convolutional neural network, inputs the vector representation, outputs a prediction result and adds it to the knowledge graph; the convolutional neural network is used to reorder the subgraphs according to the weights of the candidate entities, and the candidate entities closer to the head entity in the vector representation have greater weights; The result generation module is used to send the knowledge graph externally.
[0016] On the third aspect, based on the same inventive concept, the present application also provides a knowledge graph completion device, including a processor, a memory and a bus, wherein the memory stores instructions and data read by the processor, the processor is used to call the instructions and data in the memory to execute the knowledge graph completion method as described above, and the bus connects the functional components for transmitting information.
[0017] By adopting the above technical solution, the present invention has the following beneficial effects: The present invention provides a knowledge graph completion method, system and device, constructing an innovative framework for graph completion (SSH-KGC), which can cleverly combine the structure and semantic information of the knowledge graph to improve the efficiency and practical value of graph completion; by combining graph neural network and knowledge graph embedding technology, it surpasses the traditional single-hop method and is particularly suitable for aggregating multi-hop relationship information in the knowledge graph; this solution also effectively assembles text data from the knowledge graph to provide appropriate input for the pre-trained language model, thereby improving the computational efficiency and effect of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0019] Figure 1 A diagram illustrating the principle of a knowledge graph completion method is provided for an embodiment of the present invention; Figure 2 A flow chart of a knowledge graph completion method is provided for an embodiment of the present invention; Figure 3 for Figure 2 Flow chart of the specific construction method of the auxiliary text prompt information in step 1; Figure 4 for Figure 2 Flowchart of the specific method for aggregating structural information in step 2; Figure 5 A knowledge graph completion system diagram provided for an embodiment of the present invention. DETAILED DESCRIPTION
[0020] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0021] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.
[0022] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0023] The present invention is further explained below in conjunction with specific implementation modes.
[0024] It should also be noted that the following specific embodiments or specific implementations are a series of optimized settings listed in the present invention to further explain the specific content of the invention, and these settings can be used in combination or in association with each other.
[0025] Embodiment 1: like Figure 1-2 As shown, a knowledge graph completion method provided in this embodiment includes the following steps: Step 1: By querying entities, entities and relationships are sampled in the knowledge graph (KG) to form a subgraph; based on the query entity, corresponding auxiliary text prompt information (A2P) is constructed; Step 2: embed the subgraph and perform structural information aggregation (SIA) through a graph neural network (GNN) and a neural network adapter; then concatenate it with the auxiliary text prompt information to obtain a concatenated vector, input it into a pre-trained language model (PLMs), and output a vector representation that integrates structural information and semantic information; the auxiliary text prompt information includes query triples, entity descriptions, and entity types in the benchmark dataset; Step 3: Input the vector representation into a convolutional neural network, output the prediction result and add it to the knowledge graph, so as to complete the query entity related information in the knowledge graph; the convolutional neural network is used to reorder the subgraphs according to the weights of the candidate entities, and the candidate entities closer to the head entity in the vector representation have greater weights; Through the above method, we can construct an innovative framework for graph completion (SSH-KGC), which can cleverly combine the structure and semantic information of the knowledge graph to improve the efficiency of graph completion; by combining graph neural network and knowledge graph embedding (KGE) technology, it surpasses the traditional single-hop method (also called one hop, one hop in the knowledge graph is one relationship), and is particularly suitable for aggregating multi-hop relationship information in the knowledge graph.
[0026] Furthermore, the query entity in step 1 is a query triplet that is missing a head entity or a tail entity.
[0027] Furthermore, if Figure 3 As shown, the specific method for constructing the auxiliary text prompt information in step 1 includes: Step 11: splice the query triples in the benchmark dataset according to the input format of the pre-trained language model to obtain a coherent text sequence, and add it to the preset position of the preset prompt template of the auxiliary text prompt information; this helps the query triples to be better aligned with the subsequent pre-trained language model, so that the latter can more easily understand the words in the benchmark dataset; Step 12: Add the entity description to the preset position of the preset prompt template of the auxiliary text prompt information, so as to help the subsequent pre-trained language model grasp the semantic nuances of the query, so as to more effectively understand the context of the current completion; Step 13: All triples containing the query entity in the benchmark data set are spliced into prompt information, input into the large model (LLM), and after identifying the entity type of the query entity (i.e., 3-5 nouns expressing the information type), the entity type is added to the preset position of the preset prompt template of the auxiliary text prompt information; this facilitates the subsequent pre-trained language model to understand the specific semantics of the same query entity in different query triples; For example: in "Bill Gates founded Microsoft", the entity type of the query entity "Bill Gates" is "entrepreneur"; in "Bill Gates married Melinda Gates", the entity type of the query entity "Bill Gates" is "husband".
[0028] Further, the benchmark dataset includes the FB15k-237 dataset and / or the WN18RR dataset; The FB15k-237 dataset is an extension of the Freebase knowledge graph and contains 237 relationships, each of which consists of multiple words connected by underscores. The WN18RR dataset is derived from the WordNet knowledge graph and contains triplets in the form of (subject, relation, object) that express the relationship between entities. For statistical information about the dataset, see Table 1.
[0029]
[0030] Furthermore, if Figure 4 As shown, the specific method of aggregating the structural information in step 2 is: Step 21: Treat the subgraph as an undirected subgraph , calculate the head entity The undirected aggregation embedding of is expressed as follows: ; in, express The embedding vector of ; Representing relationships The embedding vector of Represents the tail entity The embedding vector of express middle Undirected aggregation embedding of; Represents an aggregation operation, which is specifically defined as: ; in, express The matrix of neighboring entity embeddings in ; express Normalized adjacency matrix; express The degree matrix of Represents the weight matrix of the GNN layer; represents a nonlinear activation function (such as ReLU); Step 22: Treat the relationship as a node that is equally important as the entity, and include the relationship and the head and tail entities related to the relationship within the one-hop neighborhood into the relationship subgraph. ; Compute the relation-aware aggregation embedding, the specific expression is: ; in, express Middle head entity Relation-aware aggregation embedding of Relationship subgraph The specific expression is: ; in, Indicates the header entity associated with a one-hop neighbor and tail entity ; represents sub-graph sampling; Step 23: and Concatenate and input into the neural network adapter to get the final structure embedding , the specific expression is: ; in, Indicates an adapter.
[0031] Furthermore, the neural network adapter includes a trainable two-layer neural network.
[0032] Furthermore, the pre-trained language model is a BERT model.
[0033] Furthermore, the loss function of the convolutional neural network in step 3 includes: Negative Sampling Loss , the specific formula can be: ; in, represents a hyperparameter, i.e., a score boundary value or interval (margin); represents the scoring function; Represents the head entity; To express a relationship; Represents the tail entity; Indicates auxiliary text prompt information; Indicates the number of negative samples; Represents the embedding of negative samples; Contains smoothing coefficient The cross entropy loss , the specific formula is: ; in, represents the predicted probability of the target entity given the query entity, relation, and auxiliary text hint information; represents the size of the label space, that is, the number of all possible values of the tail entity t; represents the predicted probability of negative samples given the query entity, relation, and auxiliary text hint information; Total loss function , the specific formula is: ; in, Represents a triple; represents the weight hyperparameter; Through the total loss function, the convolutional neural network can strike a balance between the two tasks of predicting the target entity and distinguishing negative samples.
[0034] Furthermore, the evaluation indicators of the prediction results in step 3 include the average reciprocal ranking and Indicators (K = 1, 3, 10); Average last ranking The specific formula is: ; in, Representing a collection of entities the number of The value range is from 0 to 1. The higher the value, the better the performance of the knowledge graph. When the value is 1, it means that for each query, the most relevant item is always ranked first in the query results. The specific formula of the indicator is: ; in, Indicators used to measure Entities Ranking list The proportion of queries where at least one relevant item appears in the first K (K = 1, 3, 10) positions; Represents an indicator function. If the condition is true, the function value is 1, otherwise the function value is 0; Indicates the number of digits, and can be 1, 3, or 10.
[0035] Embodiment 2: like Figure 5 As shown, this embodiment provides a knowledge graph completion method system, including a data receiving module, a data processing module and a result generating module; The data receiving module is used to receive the knowledge graph and the query entity; The data processing module includes a sub-graph unit, an auxiliary text prompt information unit, a structural information aggregation unit, a pre-trained language model unit and a convolutional neural network unit; The subgraph unit samples entities and relationships in the knowledge graph through the query entity to form a subgraph; The auxiliary text prompt information unit constructs corresponding auxiliary text prompt information based on the query entity; the auxiliary text prompt information includes the query triple, entity description and entity type in the benchmark data set; The structural information aggregation unit embeds the subgraph and aggregates the structural information through the graph neural network and the neural network adapter; and then splices it with the auxiliary text prompt information to obtain a splicing vector; The pre-trained language model unit stores a pre-trained language model, inputs a concatenated vector, and outputs a vector representation of the fusion structure information and semantic information; The convolutional neural network unit stores a convolutional neural network, inputs the vector representation, outputs a prediction result and adds it to the knowledge graph; the convolutional neural network is used to reorder the subgraphs according to the weights of the candidate entities, and the candidate entities closer to the head entity in the vector representation have greater weights; The result generation module is used to send the knowledge graph externally.
[0036] Embodiment three: This embodiment provides a knowledge graph completion device, including a processor, a memory and a bus, wherein the memory stores instructions and data read by the processor, the processor is used to call the instructions and data in the memory to execute the knowledge graph completion method as described above, and the bus connects the functional components for transmitting information.
[0037] In another embodiment, the present solution can be implemented by an integrated device, which may include corresponding modules for performing each or several steps in the above-mentioned embodiments. The module may be one or more hardware modules specially configured to perform the corresponding steps, or implemented by a processor configured to perform the corresponding steps, or stored in a computer-readable medium for implementation by a processor, or implemented by some combination.
[0038] The processor performs the various methods and processes described above. For example, the method implementation in the present solution can be implemented as a software program, which is tangibly contained in a machine-readable medium, such as a memory. In some embodiments, part or all of the software program can be loaded and / or installed via a memory and / or a communication interface. When the software program is loaded into the memory and executed by the processor, one or more steps in the method described above can be executed. Alternatively, in other embodiments, the processor can be configured to perform one of the above methods by any other appropriate means (for example, by means of firmware).
[0039] The device can be implemented using a bus architecture. The bus architecture can include any number of interconnecting buses and bridges, depending on the specific application and overall design constraints of the hardware. The bus connects various circuits including one or more processors, memories, and / or hardware modules together. The bus can also connect various other circuits such as peripherals, voltage regulators, power management circuits, external antennas, etc.
[0040] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Component (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc.
[0041] Embodiment 4: The method in Example 1 (SSH-KGC for short) was compared with the existing knowledge graph completion model for performance testing. The specific test conditions include: Comparison models: TransE, DistMult, ComplEx, ConvE and KG-BERT; Test hardware: NVIDIA RTX A5000 GPU; NVIDIA GeForce RTX 3090; Model optimizer: Adam; Learning rate: 1e-3; Batch size: 128 When exploring hyperparameters, input the subgraph embedding dimension of the pre-trained language model: 196.
[0042] For the comparison of knowledge graph completion effects on the WN18RR dataset, see Table 2;
[0043] For the comparison of knowledge graph completion effects on the FB15K-237 dataset, see Table 3 for details;
[0044] According to the above comparison results, it can be seen that on the benchmark dataset, this scheme shows significant performance improvement compared with the traditional knowledge graph completion method.
[0045] The comparison results prove that: this scheme uses graph neural networks to aggregate the contextual information of query entities in the knowledge graph, and uses adapters to align the vector space, which can solve the spatial inconsistency problem between the words in the text in the heterogeneous embedding space and the word embedding vectors in the knowledge graph; the auxiliary text prompt information prompt can supplement the learning of semantic information; the various modules in this scheme are combined together to jointly improve the knowledge graph completion performance.
[0046] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A knowledge graph completion method, characterized in that: include: Step 1: By querying entities, sample entities and relationships in the knowledge graph to form a subgraph; Based on the query entity, construct corresponding auxiliary text prompt information; Step 2: embed the subgraph and aggregate the structural information through the graph neural network and the neural network adapter; then concatenate it with the auxiliary text prompt information to obtain a concatenated vector, input it into the pre-trained language model, and output a vector representation that integrates the structural information and semantic information; the auxiliary text prompt information includes the query triple, entity description and entity type in the benchmark data set; Step 3: Input the vector representation into a convolutional neural network, output the prediction result and add it to the knowledge graph; the convolutional neural network is used to reorder the subgraphs according to the weights of the candidate entities, and the candidate entities that are closer to the head entity in the vector representation have greater weights.
2. The method according to claim 1, characterized in that The query entity in step 1 is a query triplet that is missing a head entity or a tail entity.
3. The method according to claim 2, characterized in that The specific method for constructing the auxiliary text prompt information in step 1 includes: Step 11: concatenate the query triples in the benchmark data set according to the input format of the pre-trained language model to obtain a coherent text sequence, and add it to the preset position of the preset prompt template of the auxiliary text prompt information; Step 12: Add the entity description to a preset position of a preset prompt template of the auxiliary text prompt information; Step 13: All triples containing the query entity in the benchmark data set are spliced into prompt information, input into the large model, and after identifying the entity type of the query entity, the entity type is added to the preset position of the preset prompt template of the auxiliary text prompt information.
4. The method according to claim 1, characterized in that The benchmark datasets include the FB15k-237 dataset and / or the WN18RR dataset.
5. The method according to claim 3, characterized in that: The specific method of structural information aggregation in step 2 is: Step 21: Treat the subgraph as an undirected subgraph , calculate the head entity The undirected aggregation embedding of is expressed as follows: ; in, express The embedding vector of ; Representing relationships The embedding vector of Represents the tail entity The embedding vector of express middle Undirected aggregation embedding of; Represents an aggregation operation, which is specifically defined as: ; in, express The matrix of neighboring entity embeddings in ; express Normalized adjacency matrix; express The degree matrix of Represents the weight matrix of the GNN layer; represents a nonlinear activation function; Step 22: Treat the relationship as a node that is equally important as the entity, and include the relationship and the head and tail entities related to the relationship within the one-hop neighborhood into the relationship subgraph. ; Calculate the relation-aware aggregate embedding. The specific expression is: ; in, express Middle head entity Relation-aware aggregation embedding of Relationship subgraph The specific expression is: ; in, Indicates the header entity associated with a one-hop neighbor and tail entity ; represents sub-graph sampling; Step 23: and Concatenate and input into the neural network adapter to get the final structure embedding , the specific expression is: ; in, Indicates an adapter.
6. The method according to claim 1, characterized in that The pre-trained language model is a BERT model.
7. The method according to claim 1, characterized in that The loss function of the convolutional neural network in step 3 includes: Negative Sampling Loss , the specific formula is: ; in, represents a hyperparameter, i.e., a boundary value or interval of a score; represents the scoring function; Represents the head entity; To express a relationship; Represents the tail entity; Indicates auxiliary text prompt information; Indicates the number of negative samples; Represents the embedding of negative samples; Contains smoothing coefficient The cross entropy loss , the specific formula is: ; in, represents the predicted probability of the target entity given the query entity, relation, and auxiliary text hint information; Indicates the size of the label space; represents the predicted probability of negative samples given the query entity, relation, and auxiliary text hint information; Total loss function , the specific formula is: ; in, Represents a triple; represents the weight hyperparameter.
8. The method according to claim 1, characterized in that The evaluation indicators of the prediction results in step 3 include the average reciprocal ranking and index; Average last ranking The specific formula is: ; in, Representing a collection of entities the number of The value range is from 0 to 1. The higher the value, the better the performance of the knowledge graph. When the value is 1, it means that for each query, the most relevant item is always ranked first in the query results. Indicators used to measure Entities Ranking list The proportion of queries where at least one relevant item appears in the first K positions is: ; in, Represents an indicator function. If the condition is true, the function value is 1, otherwise the function value is 0; Indicates the number of digits.
9. A knowledge graph completion method system, characterized in that: It includes a data receiving module, a data processing module and a result generating module; The data receiving module is used to receive the knowledge graph and the query entity; The data processing module includes a sub-graph unit, an auxiliary text prompt information unit, a structural information aggregation unit, a pre-trained language model unit and a convolutional neural network unit; The subgraph unit samples entities and relationships in the knowledge graph through the query entity to form a subgraph; The auxiliary text prompt information unit constructs corresponding auxiliary text prompt information based on the query entity; the auxiliary text prompt information includes the query triple, entity description and entity type in the benchmark data set; The structural information aggregation unit embeds the subgraph and aggregates the structural information through the graph neural network and the neural network adapter; and then splices it with the auxiliary text prompt information to obtain a splicing vector; The pre-trained language model unit stores a pre-trained language model, inputs a concatenated vector, and outputs a vector representation of the fusion structure information and semantic information; The convolutional neural network unit stores a convolutional neural network, inputs the vector representation, outputs a prediction result and adds it to the knowledge graph; the convolutional neural network is used to reorder the subgraphs according to the weights of the candidate entities, and the candidate entities closer to the head entity in the vector representation have greater weights; The result generation module is used to send the knowledge graph externally.
10. A knowledge graph completion device, characterized in that: It includes a processor, a memory and a bus, wherein the memory stores instructions and data read by the processor, the processor is used to call the instructions and data in the memory to execute any method as claimed in claim 1-8, and the bus connects the functional components for transmitting information.