Knowledge Graph Link Prediction Method, Apparatus, Device, Medium and Program Product
By generating a collection of candidate tail entities and using the language analysis model and embedded feature representation model for comprehensive scoring, the problem of inability to effectively explore the potential features of triples in the prior art is solved, and the accuracy of knowledge graph link prediction is improved.
Patent Information
- Application Number
- CN202510457301.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing technology is unable to effectively explore potential features in triplets, resulting in low accuracy in knowledge graph link prediction.
By generating a collection of candidate tail entities, using the language analysis model for semantic analysis, combining the embedded feature representation model and comprehensive score, the target tail entity elements are selected for link prediction.
Effectively tapping potential features in triplets improves the accuracy and purity of entity predictions, and improves the accuracy of knowledge graph link prediction.
Smart Images

Figure CN119990283B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge graphs, and in particular, to a method, device, equipment, medium and program product for knowledge graph link prediction. Background Art
[0002] As an important technology in the field of artificial intelligence, knowledge graphs have also been widely used in many scenarios such as knowledge answering, financial investment consulting, and medical assistance in recent years. Knowledge graph link prediction is a core task in knowledge graph research, aiming to complete the missing entities or relationships in the knowledge graph through reasoning technology to improve its integrity and practicality. The existing knowledge graph link prediction methods are mainly triple-based methods. However, the current triple-based knowledge graph link prediction cannot effectively mine the potential features in triples, so it cannot accurately complete the missing entities in triples, resulting in a low accuracy rate of knowledge graph link prediction and unable to meet the actual needs of knowledge graph link prediction in the application field. Summary of the Invention
[0003] The main purpose of the present invention is to provide a method, device, equipment, medium and program product for knowledge graph link prediction, aiming to solve the technical problem that in the prior art, the potential features in triples cannot be effectively mined, so the missing entities in triples cannot be accurately completed, resulting in a low accuracy rate of knowledge graph link prediction.
[0004] To achieve the above object, the present invention provides a method for knowledge graph link prediction, and the method includes the following steps:
[0005] Generating a candidate tail entity set based on the head entity element and the relationship element of the triple to be queried in the knowledge graph, where the candidate tail entity set includes multiple candidate tail entity elements;
[0006] Performing semantic analysis on the triple to be queried through a language analysis model to obtain the semantic information of the triple to be queried;
[0007] Filtering the candidate tail entity set based on the semantic information, and generating a tail entity set to be scored according to the filtering result;
[0008] Performing a comprehensive score on each tail entity element to be scored in the tail entity set to be scored, and screening out the target tail entity element from the tail entity set to be scored according to the comprehensive score result;
[0009] Performing link prediction on the knowledge graph based on the target tail entity element and the triple to be queried.
[0010] Optionally, the generating a candidate tail entity set based on the head entity element and the relationship element of the triple to be queried in the knowledge graph includes:
[0011] Quantize the embedding features of the head entity element and the relationship element of the triple to be queried in the knowledge graph through the embedding feature representation model to obtain the head entity embedding feature vector and the relationship embedding feature vector;
[0012] Generate a plurality of original tail entity elements based on the head entity embedding feature vector and the relationship embedding feature vector;
[0013] Score the original tail entity elements through the embedding feature representation model to obtain a first scoring result;
[0014] Filter the original tail entity elements based on the first scoring result to obtain a first tail entity set;
[0015] Perform semantic relevance analysis on the head entity element and the relationship element through the language analysis model to obtain semantic relevant information;
[0016] Generate a second tail entity set according to the semantic relevant information;
[0017] Aggregate the first tail entity set and the second tail entity set to obtain an initial tail entity set;
[0018] Score each initial tail entity element in the initial tail entity set to obtain a second scoring result;
[0019] Filter out a plurality of candidate tail entity elements from the initial tail entity elements based on the second scoring result, and generate a candidate tail entity set based on the plurality of candidate tail entity elements.
[0020] Optionally, the scoring each initial tail entity element in the initial tail entity set to obtain a second scoring result includes:
[0021] Quantize the embedding features of each initial tail entity element in the initial tail entity set through the embedding feature representation model to obtain an initial tail entity embedding feature vector;
[0022] Evaluate the correlation degree between each initial tail entity element and the triple to be queried based on the initial tail entity embedding feature vector, the head entity embedding feature vector and the relationship embedding feature vector to obtain a correlation scoring result;
[0023] Perform a diversity scoring on each initial tail entity element according to the cosine similarity between each initial tail entity embedding feature vector to obtain a diversity scoring result;
[0024] Score each initial tail entity element in the initial tail entity set based on the correlation scoring result and the diversity scoring result to obtain a second scoring result:
[0025]
[0026] Among them, represents the second scoring result of the initial tail entity element ; represents the scoring importance weight used to control the importance weights of the relevance scoring result and the diversity scoring result represents the initial tail entity set represents the initial tail entity element 's relevance scoring result represents the initial tail entity set in the initial tail entity elements 's diversity scoring result.
[0027] Optionally, the semantic information includes: the context information of the head entity element and the semantic description information and adversarial description information of the triple to be queried;
[0028] The semantic analysis of the triple to be queried by the language analysis model to obtain the semantic information of the triple to be queried includes:
[0029] Inputting the head entity element and the relationship element into the language analysis model for semantic expansion analysis to obtain the context information of the head entity element:
[0030]
[0031] Among them, represents the context information of the head entity element ; represents the language analysis model represents the language prompt template of the language analysis model;
[0032] Based on the head entity element and the relationship element, performing semantic description analysis on the triple to be queried to obtain the semantic description information of the triple to be queried:
[0033]
[0034] Among them, represents the semantic description information of the triple to be queried represents the head entity element represents the relationship element;
[0035] Screen out the adversarial analysis tail entity elements that meet the preset adversarial conditions from the candidate tail entity set according to the second scoring results of the respective candidate tail entity elements;
[0036] Perform adversarial analysis on the triple to be queried based on the adversarial analysis tail entity element, the head entity element, and the relationship element, and obtain the adversarial description information of the triple to be queried:
[0037]
[0038] Among them, represents the adversarial description information of the triple to be queried, represents the adversarial analysis tail entity element.
[0039] Optionally, the set of candidate tail entities to be scored includes a third set of tail entities, a fourth set of tail entities, and a fifth set of tail entities; the screening of the set of candidate tail entities based on the semantic information and generating the set of candidate tail entities to be scored according to the screening results includes:
[0040] Generate entity type constraint conditions based on the relationship element, and perform entity type analysis on each candidate tail entity element according to the entity type constraint conditions to obtain the entity type analysis result:
[0041]
[0042] Among them, represents the entity type analysis result, represents the expected tail entity type of the head entity element under the relationship element , represents the tail entity type that conforms to the entity type constraint conditions;
[0043] Perform entity type screening on the set of candidate tail entities according to the entity type constraint conditions to obtain a third set of tail entities;
[0044] Perform semantic relevance analysis on each candidate tail entity element in the set of candidate tail entities according to the embedding feature vector of the head entity element to obtain the semantic relevance analysis result between the head entity element and each candidate tail entity element:
[0045]
[0046] Among them, represents the semantic relevance analysis result, represents the embedding feature vector of the head entity element, represents the embedding feature vector of the candidate tail entity element;
[0047] Perform semantic relevance screening on the set of candidate tail entities based on the semantic relevance analysis result to obtain a fourth set of tail entities;
[0048] Perform context consistency analysis on each candidate tail entity element in the candidate tail entity set according to the semantic information, and obtain the context consistency analysis results between each candidate tail entity element and the knowledge graph:
[0049]
[0050] Among them, represents the context consistency analysis result, represents the relationship element and the semantic similarity between relationship elements, represents the knowledge graph, represents the indicator function, used to judge whether the triple already exists in the knowledge graph ;
[0051] Perform context consistency screening on the candidate tail entity set based on the context consistency analysis results to obtain the fifth tail entity set.
[0052] Optionally, the comprehensive scoring of each to-be-scored tail entity element in the to-be-scored tail entity set, and screening out the target tail entity element from the to-be-scored tail entity set according to the comprehensive scoring result, includes:
[0053] Perform weight ratio matching based on the entity type analysis result, the semantic relevance analysis result, and the context consistency analysis result to obtain weight parameters;
[0054] Perform comprehensive scoring on each to-be-scored tail entity element in the to-be-scored tail entity set according to the weight parameters to obtain the comprehensive scoring result:
[0055]
[0056] Among them, is the entity type weight parameter, is the semantic relevance weight parameter, is the context consistency weight parameter, represents the entity type analysis result, represents the semantic relevance analysis result, represents the context consistency analysis result, represents the comprehensive scoring result;
[0057] Screen out the target tail entity element from the to-be-scored tail entity set according to the comprehensive scoring result.
[0058] In addition, to achieve the above object, the present invention also proposes a knowledge graph link prediction device, and the knowledge graph link prediction device includes:
[0059] A tail entity generation module, configured to generate a candidate tail entity set based on a head entity element and a relationship element of a triple to be queried in a knowledge graph, where the candidate tail entity set includes a plurality of candidate tail entity elements;
[0060] A semantic analysis module, configured to perform semantic analysis on the triple to be queried through a language analysis model to obtain semantic information of the triple to be queried;
[0061] A tail entity screening module, configured to screen the candidate tail entity set based on the semantic information and generate a tail entity set to be scored according to the screening result;
[0062] A comprehensive scoring module, configured to comprehensively score each tail entity element to be scored in the tail entity set to be scored, and screen out a target tail entity element from the tail entity set to be scored according to the comprehensive scoring result;
[0063] A knowledge graph prediction module, configured to perform link prediction on the knowledge graph based on the target tail entity element and the triple to be queried.
[0064] In addition, to achieve the above object, the present application further provides a knowledge graph link prediction device, where the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the knowledge graph link prediction method as described above.
[0065] In addition, to achieve the above object, the present application further provides a computer-readable storage medium, where a computer program is stored on the computer-readable storage medium, and the computer program, when executed by a processor, implements the steps of the knowledge graph link prediction method as described above.
[0066] In addition, to achieve the above object, the present application further provides a computer program product, where the computer program product includes a computer program, and the computer program, when executed by a processor, implements the steps of the knowledge graph link prediction method as described above.
[0067] The present invention generates a candidate tail entity set based on the head entity element and the relationship element of the triple to be queried in the knowledge graph. The candidate tail entity set includes multiple candidate tail entity elements. The semantic analysis of the triple to be queried is performed through a language analysis model to obtain the semantic information of the triple to be queried. The candidate tail entity set is filtered based on the semantic information, and a tail entity set to be scored is generated according to the filtering result. The comprehensive scores of each tail entity element in the tail entity set to be scored are calculated, and the target tail entity element is selected from the tail entity set to be scored according to the comprehensive score result. Link prediction is performed on the knowledge graph based on the target tail entity element and the triple to be queried. Since the present invention performs semantic analysis on the triple to be queried through a language analysis model, the potential features in the triple can be effectively mined, and the candidate tail entity set is filtered based on the semantic information, thus effectively avoiding the problems of entity type mismatch and context semantic mismatch between the predicted tail entity element and the triple to be queried, improving the accuracy of entity prediction. The comprehensive scores of the filtered tail entity set to be scored are calculated, and the tail entity is filtered again based on the score result to obtain the target tail entity element. By filtering the predicted tail entity element in multiple stages, the entity prediction noise is reduced, the purity of the tail entity set is effectively improved, and the accurate complement of the missing triple entity is realized, thereby greatly improving the accuracy of knowledge graph link prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following briefly introduces the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0069] Figure 1 It is a schematic structural diagram of a knowledge graph link prediction device for the hardware operating environment involved in the embodiment of the present invention;
[0070] Figure 2 It is a schematic flowchart of the first embodiment of the knowledge graph link prediction method of the present invention;
[0071] Figure 3 It is a schematic flowchart of the second embodiment of the knowledge graph link prediction method of the present invention;
[0072] Figure 4 It is a schematic flowchart of the third embodiment of the knowledge graph link prediction method of the present invention;
[0073] Figure 5 It is a schematic block diagram of the first embodiment of the knowledge graph link prediction device of the present invention.
[0074] The implementation, functional features, and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0075] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0076] Refer to Figure 1 , Figure 1 which is a schematic structural diagram of a knowledge graph link prediction device for the hardware operating environment involved in the embodiment solution of the present invention.
[0077] As Figure 1 shown, the knowledge graph link prediction device may include: a processor 1001, such as a Central Processing Unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a Wireless-Fidelity (WI-FI) interface). The memory 1005 may be a high-speed Random Access Memory (RAM) or a stable Non-Volatile Memory (NVM), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0078] Those skilled in the art can understand that Figure 1 the structure shown in
[0079] does not constitute a limitation on the knowledge graph link prediction device, and it may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. Figure 1 As
[0080] shown in Figure 1In the knowledge graph link prediction device shown, the network interface 1004 is mainly used for data communication with a network server; the user interface 1003 is mainly used for data interaction with a user; the processor 1001 and the memory 1005 in the knowledge graph link prediction device of the present invention can be arranged in the knowledge graph link prediction device. The knowledge graph link prediction device calls the knowledge graph link prediction program stored in the memory 1005 through the processor 1001 and executes the knowledge graph link prediction method provided by the embodiments of the present invention.
[0081] Embodiments of the present invention provide a knowledge graph link prediction method. Refer to Figure 2 , Figure 2 which is a schematic flowchart of the first embodiment of the knowledge graph link prediction method of the present invention.
[0082] In this embodiment, the knowledge graph link prediction method includes the following steps:
[0083] Step S10: Generate a candidate tail entity set based on the head entity element and the relationship element of the triple to be queried in the knowledge graph.
[0084] It should be understood that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or a terminal electronic device capable of implementing the above functions. Hereinafter, a knowledge graph link prediction device (prediction device) will be used as an example to illustrate this embodiment and the following embodiments.
[0085] It should be noted that the candidate tail entity set includes multiple candidate tail entity elements. The triple to be queried can be a triple that needs to complete a missing tail entity element. For example, the triple to be queried is , where " " is the head entity element, " " is the relationship element, and " " is the missing tail entity element.
[0086] In some embodiments, the prediction device can generate multiple candidate tail entities by means of embedded feature learning, or can also generate multiple candidate tail entities based on a pre-trained large language model, and construct a candidate tail entity set based on the candidate tail entities.
[0087] In some embodiments, the prediction device performs embedded feature learning on the query triple to obtain candidate tail entities based on an embedding model; generates candidate tail entities based on the large language model based on the query head entity and the relationship ; and obtains a candidate tail entity set based on the screening of the relevance and diversity of the candidate entities .
[0088] Further, in order to improve the diversity of the candidate entity set and thus improve the prediction accuracy, in some embodiments, step S10 described above may include:
[0089] Step S101: Quantize the embedding features of the head entity element and the relationship element of the triple to be queried in the knowledge graph through an embedding feature representation model, and obtain a head entity embedding feature vector and a relationship embedding feature vector;
[0090] Step S102: Generate a plurality of original tail entity elements based on the head entity embedding feature vector and the relationship embedding feature vector;
[0091] Step S103: Score the original tail entity elements through the embedding feature representation model to obtain a first scoring result;
[0092] Step S104: Screen the original tail entity elements based on the first scoring result to obtain a first tail entity set;
[0093] Step S105: Perform semantic relevance analysis on the head entity element and the relationship element through a language analysis model to obtain semantic relevant information;
[0094] Step S106: Generate a second tail entity set according to the semantic relevant information;
[0095] Step S107: Aggregate the first tail entity set and the second tail entity set to obtain an initial tail entity set;
[0096] Step S108: Score each initial tail entity element in the initial tail entity set to obtain a second scoring result;
[0097] Step S109: Screen out a plurality of candidate tail entity elements from the initial tail entity elements based on the second scoring result, and generate a candidate tail entity set based on the plurality of candidate tail entity elements.
[0098] It should be noted that the embedding feature representation model may be the SAttLE model. The language analysis model may be a large language model (LLM). The above first scoring result may be the scoring result corresponding to each original tail entity element generated based on the embedding feature dimension. The above second scoring result may be the scoring result of each initial tail entity element in the initial tail entity set. The above initial tail entity set may be a set obtained by aggregating the first tail entity set generated based on the embedding feature dimension and the second tail entity set generated based on the semantic feature dimension.
[0099] It should be understood that in this embodiment, candidate tail entity elements can be generated from the embedding feature dimension and the semantic feature dimension respectively:
[0100] The prediction device can select the SAttLE model, an embedding feature representation model, to perform embedding feature representation learning on the triple, generate multiple original tail entity elements, score the original tail entity elements through the embedding feature representation model, and select the top N original tail entity elements with higher scores to form the first tail entity set. For example, select the top 200 entities with the highest scores to form the first tail entity set , and the scoring function refers to the following formula:
[0101]
[0102] where, represents the scoring result of the embedding representation model for the triple , is the original tail entity element.
[0103] Input the head entity element and the relationship element in the query vector into the language analysis model for semantic analysis, and generate the second tail entity set based on the semantic relevant information in the semantic analysis result. This set has potential semantic relevance. For example, the head entity element and the relationship element in the query vector are = ("A Museum", "location"), and the language analysis model generates relevant tail entity elements based on the head entity element and the relationship element, such as "the city where A Museum is located", "the country where A Museum is located", etc. The process of the language analysis model generating the second tail entity set can be expressed by the following formula:
[0104]
[0105] where, represents the second tail entity set, represents the language analysis model.
[0106] It can be understood that the prediction device will merge the first tail entity set generated based on the embedding feature dimension and the second tail entity set generated based on the semantic feature dimension, and remove the duplicate tail entities to obtain the initial tail entity set, referring to the following formula:
[0107]
[0108] where, represents the initial tail entity set.
[0109] It should be noted that, in order to improve the purity of the tail entity set, in this embodiment, the initial tail entity set can be scored and screened again. For each initial tail entity element in the initial tail entity set, it is scored through the embedding feature representation model to evaluate the relevance between the initial tail entity element and the triple to be queried. Entities with high scores in the embedding model usually have a closer relevance to the head entity and the relationship. The initial tail entity elements with high scores in the embedding feature representation model are used as candidate tail entity elements.
[0110] Further, in order to evaluate the relevance and diversity of the initial tail entity elements and thus accurately screen out the candidate tail entity elements, step S108 above may include:
[0111] Step S1081: Quantify the embedding features of each initial tail entity element in the initial tail entity set through the embedding feature representation model to obtain an initial tail entity embedding feature vector;
[0112] Step S1082: Evaluate the degree of correlation between each initial tail entity element and the triple to be queried based on the initial tail entity embedding feature vector, the head entity embedding feature vector, and the relationship embedding feature vector to obtain a correlation scoring result;
[0113] Step S1083: Perform a diversity score on each initial tail entity element according to the cosine similarity between the initial tail entity embedding feature vectors to obtain a diversity scoring result;
[0114] Step S1084: Score each initial tail entity element in the initial tail entity set based on the correlation scoring result and the diversity scoring result to obtain a second scoring result.
[0115] It should be noted that for each initial tail entity element, based on the score of the embedding feature representation model evaluate its relevance to the triple to be queried. Entities with high scores in the embedding model usually have a closer relevance to the head entity and the relationship.
[0116] To avoid insufficient semantic coverage caused by overly similar candidate entities, a diversity function is defined , and the diversity score is calculated based on the cosine similarity between the initial tail entity embedding feature vectors, so as to ensure that the candidate entity set is in different semantic categories. A scoring function is defined for correlation and diversity scoring, which is specifically expressed as follows:
[0117]
[0118] Among them, represents the second scoring result of the initial tail entity element , Indicates the scoring importance weight, which is used to control the importance weights of the relevance scoring result and the diversity scoring result, Indicates the initial set of tail entities, Indicates an initial tail entity element of the relevance scoring result, Indicates the initial set of tail entities in the initial tail entity elements of the diversity scoring result.
[0119] The prediction device can select the top scoring candidate tail entity elements from the initial set of tail entities by maximizing the score, to form a candidate tail entity set that has both high relevance and diversity, with reference to the following formula:
[0120]
[0121]
[0122] Among them, Indicates the initial set of tail entities, Indicates the candidate tail entity set, can be the number of candidate tail entity elements, for example can be 50.
[0123] Step S20: Perform semantic analysis on the to-be-query triple through a language analysis model to obtain the semantic information of the to-be-query triple.
[0124] It should be noted that the semantic information can be semantic-related information of the to-be-query triple. For example, the semantic information can include the semantic information of the head entity of the to-be-query triple, the semantic information of the relationship, the context semantic information, the triple description semantic information, and the triple adversarial description semantic information.
[0125] In some embodiments, the prediction device can design a large language model prompt template from three aspects: entity enhancement, triple description enhancement, and adversarial description enhancement based on the head entity element and the relationship element, and generate the semantic information of the to-be-query triple based on the large language model prompt template.
[0126] Step S30: Screen the candidate tail entity set based on the semantic information, and generate a to-be-scored tail entity set according to the screening result.
[0127] It should be noted that the tail entity set to be scored can be an entity set composed of tail entity elements selected from the candidate tail entity set. For example, the candidate tail entity set includes 200 candidate tail entity elements, and 50 candidate tail entity elements that meet the constraint conditions of the semantic information are selected from the candidate tail entity set based on the semantic information. These 50 candidate tail entity elements are used as the tail entity elements to be scored, so as to construct the tail entity set to be scored.
[0128] It can be understood that the prediction device can generate semantic constraint conditions based on the semantic information, screen the candidate tail entity set based on the constraint conditions, eliminate the tail entities that do not meet the constraint conditions, select the candidate tail entity elements that meet the constraint conditions, and generate the tail entity set to be scored based on the selected candidate tail entity elements that meet the constraint conditions. Among them, the above-mentioned constraint conditions can include entity semantic constraint conditions, entity type constraint conditions, and semantic context matching constraint conditions.
[0129] Step S40: Perform a comprehensive score on each tail entity element to be scored in the tail entity set to be scored, and screen out the target tail entity elements from the tail entity set to be scored according to the comprehensive score result.
[0130] It should be noted that the comprehensive score can score each tail entity element to be scored in the tail entity set to be scored from multiple dimensions.
[0131] In some embodiments, the prediction device can evaluate each tail entity element to be scored from three dimensions of entity type matching degree, semantic matching degree, and context consistency respectively, and then perform a weighted average on the evaluations of the three dimensions to obtain the comprehensive score result.
[0132] Step S50: Perform link prediction on the knowledge graph based on the target tail entity element and the triple to be queried.
[0133] It should be noted that the prediction device selects the target tail entity element with the highest evaluation / score from the tail entity set to be scored based on the comprehensive score result as the tail entity prediction result to complete the triple to be queried, and realizes link prediction on the knowledge graph based on the completed triple.
[0134] In this embodiment, a candidate tail entity set is generated based on the head entity element and the relationship element of the triple to be queried in the knowledge graph. The candidate tail entity set includes multiple candidate tail entity elements. The triple to be queried is semantically analyzed through a language analysis model to obtain the semantic information of the triple to be queried. The candidate tail entity set is filtered based on the semantic information, and a tail entity set to be scored is generated according to the filtering result. Comprehensive scoring is performed on each tail entity element in the tail entity set to be scored, and the target tail entity element is filtered out from the tail entity set to be scored according to the comprehensive scoring result. Link prediction is performed on the knowledge graph based on the target tail entity element and the triple to be queried. Since this embodiment semantically analyzes the triple to be queried through a language analysis model, the potential features in the triple are effectively mined, and the candidate tail entity set is filtered based on the semantic information, thus effectively avoiding the problems of entity type mismatch and context semantic mismatch between the predicted tail entity element and the triple to be queried, improving the accuracy of entity prediction. Comprehensive scoring is performed on the filtered tail entity set to be scored, and tail entity filtering is performed again based on the scoring result to obtain the target tail entity element. By filtering the predicted tail entity element in multiple stages, the entity prediction noise is reduced, the purity of the tail entity set is effectively improved, and the accurate completion of the missing triple entity is realized, thereby greatly improving the accuracy of knowledge graph link prediction.
[0135] Reference Figure 3 , Figure 3 is a schematic flowchart of the second embodiment of the knowledge graph link prediction method of the present invention.
[0136] Based on the above first embodiment, in this embodiment, the semantic information includes: the context information of the head entity element, as well as the semantic description information and adversarial description information of the triple to be queried;
[0137] In this embodiment, the above step S20 further includes:
[0138] Step S21: Input the head entity element and the relationship element into a language analysis model for semantic expansion analysis to obtain the context information of the head entity element.
[0139] It should be noted that in this embodiment, the prediction device can adopt an entity enhancement strategy to semantically expand the head entity element based on the triple to be queried , specifically, the large language model prompt template of the language analysis model is used to generate the context information of the head entity. For example, the head entity in the triple to be queried = "City B" and the relationship = "capital", the head entity enhancement strategy will provide more background information about "City B", such as it being the capital of "Country B" and a major economic and cultural center. The specific representation is as follows:
[0140]
[0141] Among them, represents the context information of the head entity element of represents the language analysis model, represents the language prompt template of the language analysis model.
[0142] Step S22: Based on the head entity element and the relationship element, perform semantic description analysis on the triple to be queried, and obtain the semantic description information of the triple to be queried.
[0143] It should be noted that in this embodiment, the prediction device can adopt triple description enhancement: for a given head entity and relationship , use the large language model prompt template to generate the description of the triple. The triple description enhancement enables the model to effectively infer candidate tail entities that conform to the triple logic by providing coherent context information. For example, given the triple ("City b", "capital", "?"), the generated description can be "City b is an important city and it is the capital of a certain country". It is expressed as follows:
[0144]
[0145] Among them, represents the semantic description information of the triple to be queried, represents the head entity element, represents the relationship element.
[0146] Step S23: Screen out the adversarial analysis tail entity elements that meet the preset adversarial conditions from the candidate tail entity set according to the second scoring results of each candidate tail entity element.
[0147] It should be noted that the adversarial analysis tail entity element can be a candidate entity with a relatively low ranking in the second scoring result. The adversarial description enhancement highlights the candidates that are less likely to be the target entity, prompting the model to be able to distinguish the rationality of the candidate entities.
[0148] Step S24: Based on the adversarial analysis tail entity element, the head entity element and the relationship element, perform description adversarial analysis on the triple to be queried, and obtain the adversarial description information of the triple to be queried.
[0149] It can be understood that in this embodiment, the prediction device can adopt adversarial description enhancement: for a given head entity and relationship , adopt the prompting template of the large language model to generate adversarial descriptions of triples. The adversarial descriptions enhance and highlight candidates that are less likely to be target entities, prompting the model to be able to distinguish the rationality of candidate entities. By analyzing the candidates ranked at the bottom, reasons for these entities not being suitable as tail entities are generated, thereby improving the prediction accuracy. It is expressed as follows:
[0150]
[0151] Among them, represents the adversarial description information of the triple to be queried, represents the adversarial analysis tail entity element.
[0152] In this embodiment, by inputting the head entity element and the relationship element into a language analysis model for semantic expansion analysis, the context information of the head entity element is obtained. Based on the head entity element and the relationship element, semantic description analysis of the triple to be queried is performed to obtain the semantic description information of the triple to be queried. According to the second scoring results of each candidate tail entity element, adversarial analysis tail entity elements that meet the preset adversarial conditions are screened out from the candidate tail entity set. Based on the adversarial analysis tail entity elements, the head entity element, and the relationship element, description adversarial analysis of the triple to be queried is performed to obtain the adversarial description information of the triple to be queried, thereby realizing the analysis of the semantic information of the triple to be queried from multiple dimensions, respectively realizing entity semantic enhancement, triple description enhancement, and adversarial description enhancement of the triple to be queried, effectively mining the potential semantic features in the triple, thereby improving the prediction accuracy, and effectively avoiding the problem of semantic irrelevance between the predicted entity and the triple.
[0153] Refer to Figure 4 , Figure 4 which is the flowchart of the third embodiment of the knowledge graph link prediction method of the present invention.
[0154] Based on the above second embodiment, in this embodiment, the set of tail entities to be scored includes the third tail entity set, the fourth tail entity set, and the fifth tail entity set;
[0155] In this embodiment, the above step S30 further includes:
[0156] Step S301: Generate entity type constraint conditions based on the relationship element, and perform entity type analysis on each candidate tail entity element according to the entity type constraint conditions to obtain an entity type analysis result.
[0157] It should be noted that in this embodiment, multi-round question-and-answer thought chain reasoning can be designed: through simulating the user's step-by-step analysis and elimination thought process, the thought chain reasoning can provide multi-level verification for the knowledge graph link prediction task, thereby improving the prediction accuracy of the model for candidate entities. For the knowledge graph link prediction problem, multi-round question-and-answer thought chain reasoning is designed, and explicit logical reasoning based on context hint learning is carried out from three aspects: entity type matching, semantic relevance analysis, and context consistency check, and the candidate tail entity set is gradually screened and optimized. The designed multi-round question-and-answer template is shown in Table 1 below:
[0158]
[0159] It can be understood that in this embodiment, by defining an entity type matching function , the tail entity elements that meet the entity type constraint conditions are screened out from the candidate tail entity elements, with reference to the following formula:
[0160]
[0161] where represents the entity type analysis result, represents the expected tail entity type of the head entity element under the relationship element , represents the tail entity type that meets the entity type constraint conditions. If represents that the candidate tail entity does not meet the entity type constraint conditions.
[0162] Step S302: Perform entity type screening on the candidate tail entity set according to the entity type constraint conditions to obtain a third tail entity set.
[0163] It should be noted that the third tail entity set contains candidate tail entity elements that meet the entity type constraint conditions.
[0164] Step S303: Perform semantic relevance analysis on each candidate tail entity element in the candidate tail entity set according to the embedding feature vector of the head entity element to obtain the semantic relevance analysis result between the head entity element and each candidate tail entity element.
[0165] It should be noted that in this embodiment, by defining a semantic matching function for screening the semantic relevance between the head entity and the candidate tail entity:
[0166]
[0167] where represents the semantic relevance analysis result, represents the embedding feature vector of the head entity element, Represents the embedded feature vector of the candidate tail entity element, If the value of is less than the specified threshold, the tail entity is considered The semantic information of is not sufficient to support it as a reasonable tail entity, and it is excluded.
[0168] Step S304: Based on the semantic relevance analysis result, perform semantic relevance screening on the candidate tail entity set to obtain the fourth tail entity set.
[0169] It should be noted that the fourth tail entity set contains candidate tail entity elements with strong semantic information relevance.
[0170] Step S305: Perform context consistency analysis on each candidate tail entity element in the candidate tail entity set according to the semantic information to obtain the context consistency analysis result between each candidate tail entity element and the knowledge graph.
[0171] It should be noted that in this embodiment, by defining the context consistency matching function , the context consistency between the candidate tail entity element and the knowledge graph is analyzed, referring to the following formula:
[0172]
[0173] Among them, Represents the context consistency analysis result, Represents the relationship element And The semantic similarity between relationship elements, Represents the knowledge graph, Represents the indicator function, Used to judge whether the triple Already exists in the knowledge graph If the triple Exists in the knowledge graph Returns 1, otherwise 0. The context consistency check reflects whether there is supporting evidence for the candidate entity in the knowledge graph, and can further screen out candidate entities that are semantically reasonable and consistent with the background.
[0174] Step S306: Based on the context consistency analysis result, perform context consistency screening on the candidate tail entity set to obtain the fifth tail entity set.
[0175] It should be noted that the fifth tail entity set contains candidate tail entity elements that are context consistent with the knowledge graph.
[0176] It can be understood that in this embodiment, by screening the candidate tail entity set from the entity type constraint dimension, semantic relevance constraint dimension, and context consistency constraint dimension respectively, the tail entity sets obtained by screening in the three dimensions (the third tail entity set, the fourth tail entity set, and the fifth tail entity set) are obtained. By comprehensively scoring each candidate tail entity element in the tail entity sets of the three dimensions, it is avoided to score the candidate tail entity elements that do not meet the constraint conditions of all dimensions. Based on the comprehensive scoring results, the candidate tail entity element with the highest score is selected as the target tail entity element, thereby ensuring that the target tail entity element is optimal in the entity type constraint dimension, semantic relevance constraint dimension, and context consistency constraint dimension.
[0177] Further, based on the above embodiment, in order to accurately screen out the target tail entity element, in some embodiments, step S40 described above may include:
[0178] Step S401: Perform weight allocation based on the entity type analysis result, the semantic relevance analysis result, and the context consistency analysis result to obtain a weight parameter;
[0179] Step S402: Comprehensively score each to-be-scored tail entity element in the to-be-scored tail entity set according to the weight parameter to obtain a comprehensive scoring result;
[0180] Step S403: Screen out the target tail entity element from the to-be-scored tail entity set according to the comprehensive scoring result.
[0181] It should be noted that the prediction device can comprehensively score the tail entity sets obtained by screening in the three dimensions (the third tail entity set, the fourth tail entity set, and the fifth tail entity set) based on the analysis results of the three dimensions, and use the candidate tail entity element with the highest comprehensive score as the target tail entity element:
[0182]
[0183] Among them, is the entity type weight parameter, is the semantic relevance weight parameter, is the context consistency weight parameter, represents the entity type analysis result, represents the semantic relevance analysis result, represents the context consistency analysis result, represents the comprehensive scoring result.
[0184] Based on the relationship elements, this embodiment generates entity type constraint conditions, performs entity type analysis on each candidate tail entity element according to the entity type constraint conditions to obtain an entity type analysis result, performs entity type screening on the candidate tail entity set according to the entity type constraint conditions to obtain a third tail entity set, performs semantic relevance analysis on each candidate tail entity element in the candidate tail entity set according to the embedding feature vector of the head entity element to obtain a semantic relevance analysis result between the head entity element and each candidate tail entity element, performs semantic relevance screening on the candidate tail entity set based on the semantic relevance analysis result to obtain a fourth tail entity set, performs context consistency analysis on each candidate tail entity element in the candidate tail entity set according to the semantic information to obtain a context consistency analysis result between each candidate tail entity element and the knowledge graph, and performs context consistency screening on the candidate tail entity set based on the context consistency analysis result to obtain a fifth tail entity set. Thus, the purity of the candidate tail entity set is effectively improved. Based on multiple screenings, through the complementarity of screening objectives in different stages, the scoring complexity and analysis noise are effectively reduced, and the scoring accuracy and scoring quality are greatly improved, thereby effectively improving the prediction accuracy.
[0185] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which a knowledge graph link prediction program is stored. When the knowledge graph link prediction program is executed by a processor, the steps of the knowledge graph link prediction method as described above are implemented.
[0186] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0187] The above computer-readable storage medium can be included in the knowledge graph link prediction device; it can also exist independently without being assembled into the knowledge graph link prediction device.
[0188] In addition, an embodiment of the present invention also provides a computer program product, including a knowledge graph link prediction program, and when the knowledge graph link prediction program is executed by a processor, it implements the steps of the knowledge graph link prediction method described above.
[0189] The specific implementation manner of the computer program product of the present invention is basically the same as that of the above embodiments of the knowledge graph link prediction method, and will not be elaborated here.
[0190] Refer to Figure 5 , Figure 5 which is the structural block diagram of the first embodiment of the knowledge graph link prediction device of the present invention.
[0191] As Figure 5 shown, the knowledge graph link prediction device proposed by the embodiment of the present invention includes:
[0192] A tail entity generation module 10, configured to generate a candidate tail entity set based on the head entity element and the relationship element of the triple to be queried in the knowledge graph, and the candidate tail entity set includes multiple candidate tail entity elements;
[0193] The semantic analysis module 20 is configured to perform semantic analysis on the to-be-query triple through a language analysis model to obtain the semantic information of the to-be-query triple;
[0194] The tail entity screening module 30 is configured to screen the candidate tail entity set based on the semantic information and generate a to-be-scored tail entity set according to the screening result;
[0195] The comprehensive scoring module 40 is configured to comprehensively score each to-be-scored tail entity element in the to-be-scored tail entity set, and screen out the target tail entity element from the to-be-scored tail entity set according to the comprehensive scoring result;
[0196] The knowledge graph prediction module 50 is configured to perform link prediction on the knowledge graph based on the target tail entity element and the to-be-query triple.
[0197] In this embodiment, a candidate tail entity set is generated based on the head entity element and the relationship element of the to-be-query triple in the knowledge graph. The candidate tail entity set includes multiple candidate tail entity elements. The language analysis model is used to perform semantic analysis on the to-be-query triple to obtain the semantic information of the to-be-query triple. The candidate tail entity set is screened based on the semantic information, and a to-be-scored tail entity set is generated according to the screening result. Each to-be-scored tail entity element in the to-be-scored tail entity set is comprehensively scored, and the target tail entity element is screened out from the to-be-scored tail entity set according to the comprehensive scoring result. Link prediction is performed on the knowledge graph based on the target tail entity element and the to-be-query triple. Since the semantic analysis of the to-be-query triple is performed through the language analysis model in this embodiment, the potential features in the triple are effectively mined. The candidate tail entity set is screened based on the semantic information, so that the problems of entity type mismatch and context semantic mismatch between the predicted tail entity element and the to-be-query triple are effectively avoided, and the accuracy of entity prediction is improved. The to-be-scored tail entity set obtained by screening is comprehensively scored, and the tail entity is screened again based on the scoring result to obtain the target tail entity element. By screening the predicted tail entity element in multiple stages, the entity prediction noise is reduced, the purity of the tail entity set is effectively improved, and the accurate complement of the missing triple entity is realized, thereby greatly improving the accuracy of the knowledge graph link prediction.
[0198] The knowledge graph link prediction device provided by the present application adopts the knowledge graph link prediction method in the above embodiment and can solve the technical problem of knowledge graph link prediction. Compared with the prior art, the beneficial effects of the knowledge graph link prediction device provided by the present application are the same as those of the knowledge graph link prediction method provided by the above embodiment, and other technical features in the knowledge graph link prediction device are the same as the features disclosed in the above embodiment method, which will not be elaborated here.
[0199] It should be understood that the above is only an example and does not constitute any limitation to the technical solution of the present invention. In specific applications, those skilled in the art can set according to needs, and the present invention does not make any restrictions in this regard.
[0200] It should be noted that the above-described workflow is only illustrative and does not constitute a limitation to the protection scope of the present invention. In actual applications, those skilled in the art can select some or all of them according to actual needs to achieve the purpose of the solution of this embodiment, and no restrictions are made here.
[0201] In addition, for the technical details not described in detail in this embodiment, reference can be made to the knowledge graph link prediction method provided in any embodiment of the present invention, and details will not be repeated here.
[0202] It should be noted that in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, article or system. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or system including that element.
[0203] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.
[0204] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as a read-only memory / random access memory, magnetic disk, optical disk), and includes several instructions to enable a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in various embodiments of the present invention.
[0205] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, shall be included in the patent protection scope of the present invention by the same token.
Claims
1. A method for knowledge graph link prediction, characterized in that, The knowledge graph link prediction method is applied to knowledge answering, and the knowledge graph link prediction method includes: Generating a candidate tail entity set based on the head entity element and the relationship element of the triple to be queried in the knowledge graph, where the candidate tail entity set includes multiple candidate tail entity elements; Performing semantic analysis on the triple to be queried through a language analysis model to obtain the semantic information of the triple to be queried; Filtering the candidate tail entity set based on the semantic information, and generating a tail entity set to be scored according to the filtering result; Comprehensively scoring each tail entity element to be scored in the tail entity set to be scored, and screening out the target tail entity element from the tail entity set to be scored according to the comprehensive scoring result; Performing link prediction on the knowledge graph based on the target tail entity element and the triple to be queried; The generating a candidate tail entity set based on the head entity element and the relationship element of the triple to be queried in the knowledge graph includes: Performing embedded feature quantization on the head entity element and the relationship element of the triple to be queried in the knowledge graph through an embedded feature representation model to obtain a head entity embedded feature vector and a relationship embedded feature vector; Generating multiple original tail entity elements based on the head entity embedded feature vector and the relationship embedded feature vector; Scoring the original tail entity elements through the embedded feature representation model to obtain a first scoring result; Filtering the original tail entity elements based on the first scoring result to obtain a first tail entity set; Performing semantic relevance analysis on the head entity element and the relationship element through a language analysis model to obtain semantic relevant information; Generating a second tail entity set according to the semantic relevant information; Aggregating the first tail entity set and the second tail entity set to obtain an initial tail entity set; Scoring each initial tail entity element in the initial tail entity set to obtain a second scoring result; Filtering out multiple candidate tail entity elements from the initial tail entity elements based on the second scoring result, and generating a candidate tail entity set based on the multiple candidate tail entity elements.
2. The knowledge graph link prediction method according to claim 1, wherein The scoring each initial tail entity element in the initial tail entity set to obtain a second scoring result includes: Performing embedded feature quantization on each initial tail entity element in the initial tail entity set through the embedded feature representation model to obtain an initial tail entity embedded feature vector; Evaluating the correlation degree between each initial tail entity element and the triple to be queried based on the initial tail entity embedded feature vector, the head entity embedded feature vector and the relationship embedded feature vector to obtain a correlation scoring result; Performing diversity scoring on each initial tail entity element according to the cosine similarity between the initial tail entity embedded feature vectors to obtain a diversity scoring result; Scoring each initial tail entity element in the initial tail entity set based on the correlation scoring result and the diversity scoring result to obtain a second scoring result: Among them, represents the second scoring result of the initial tail entity element , represents the scoring importance weight used to control the importance weights of the relevance scoring result and the diversity scoring result represents the initial tail entity set represents the initial tail entity element 's relevance scoring result represents the initial tail entity set in the initial tail entity element 's diversity scoring result.
3. The knowledge graph link prediction method according to any one of claims 1 or 2, characterized in that The semantic information includes: the context information of the head entity element, and the semantic description information and adversarial description information of the triple to be queried; Performing semantic analysis on the to-be-query triple by means of a language analysis model to obtain semantic information of the to-be-query triple, including: Inputting the head entity element and the relationship element into the language analysis model for semantic augmentation analysis to obtain context information of the head entity element: Among them, represents the context information of the head entity element , represents the language analysis model, represents the language prompt template of the language analysis model; Performing semantic description analysis on the to-be-query triple based on the head entity element and the relationship element to obtain semantic description information of the to-be-query triple: Among them, represents the semantic description information of the triple to be queried, represents the head entity element, represents the relationship element; Filtering out adversarial analysis tail entity elements that meet a preset adversarial condition from the candidate tail entity set according to the second scoring results of each candidate tail entity element; Performing descriptive adversarial analysis on the to-be-query triple based on the adversarial analysis tail entity element, the head entity element, and the relationship element to obtain adversarial description information of the to-be-query triple: Among them, represents the adversarial description information of the triple to be queried, represents the adversarial analysis tail entity element.
4. The knowledge graph link prediction method according to claim 3, wherein The to-be-scored tail entity set includes a third tail entity set, a fourth tail entity set, and a fifth tail entity set; Filtering the candidate tail entity set based on the semantic information and generating a to-be-scored tail entity set according to the filtering results, including: Generating an entity type constraint condition based on the relationship element and performing entity type analysis on each candidate tail entity element according to the entity type constraint condition to obtain an entity type analysis result: Among them, represents the result of entity type analysis, represents under the relationship element the expected tail entity type of the head entity element ; represents the tail entity type that meets the entity type constraint conditions. Performing entity type filtering on the candidate tail entity set according to the entity type constraint condition to obtain a third tail entity set; Performing semantic relevance analysis on each candidate tail entity element in the candidate tail entity set according to the embedding feature vector of the head entity element to obtain a semantic relevance analysis result between the head entity element and each candidate tail entity element: Among them, represents the result of semantic relevance analysis, represents the embedded feature vector of the head entity element, represents the embedded feature vector of the candidate tail entity element; Performing semantic relevance filtering on the candidate tail entity set based on the semantic relevance analysis result to obtain a fourth tail entity set; Performing context consistency analysis on each candidate tail entity element in the candidate tail entity set according to the semantic information to obtain a context consistency analysis result between each candidate tail entity element and the knowledge graph: Among them, represents the result of context consistency analysis, represents the relationship element and the semantic similarity between relationship elements, represents the knowledge graph, represents the indicator function, which is used to judge whether the triple already exists in the knowledge graph ; Performing context consistency filtering on the candidate tail entity set based on the context consistency analysis result to obtain a fifth tail entity set.
5. The knowledge graph link prediction method according to claim 4, wherein Performing comprehensive scoring on each to-be-scored tail entity element in the to-be-scored tail entity set and filtering out a target tail entity element from the to-be-scored tail entity set according to the comprehensive scoring result, including: Performing weight ratio matching based on the entity type analysis result, the semantic relevance analysis result, and the context consistency analysis result to obtain weight parameters; Performing comprehensive scoring on each to-be-scored tail entity element in the to-be-scored tail entity set according to the weight parameters to obtain a comprehensive scoring result: Among them, is the entity type weight parameter, is the semantic relevance weight parameter, is the context consistency weight parameter, represents the entity type analysis result, represents the semantic relevance analysis result, represents the context consistency analysis result, represents the comprehensive score result; Filtering out a target tail entity element from the to-be-scored tail entity set according to the comprehensive scoring result.
6. A knowledge graph link prediction device, characterized in that The knowledge graph link prediction device is applied to knowledge answering, and the knowledge graph link prediction device includes: A tail entity generation module, configured to generate a candidate tail entity set based on a head entity element and a relationship element of a to-be-query triple in a knowledge graph, where the candidate tail entity set includes a plurality of candidate tail entity elements; A semantic analysis module, configured to perform semantic analysis on the to-be-query triple through a language analysis model to obtain the semantic information of the to-be-query triple; A tail entity screening module, configured to screen the candidate tail entity set based on the semantic information and generate a to-be-scored tail entity set according to the screening result; A comprehensive scoring module, configured to comprehensively score each to-be-scored tail entity element in the to-be-scored tail entity set, and screen out the target tail entity element from the to-be-scored tail entity set according to the comprehensive scoring result; A knowledge graph link prediction module, configured to perform link prediction on the knowledge graph based on the target tail entity element and the to-be-query triple; The tail entity generation module is further configured to perform embedded feature quantization on the head entity element and the relationship element of the to-be-query triple in the knowledge graph through an embedded feature representation model to obtain a head entity embedded feature vector and a relationship embedded feature vector; generate a plurality of original tail entity elements based on the head entity embedded feature vector and the relationship embedded feature vector; score the original tail entity elements through the embedded feature representation model to obtain a first scoring result; screen the original tail entity elements based on the first scoring result to obtain a first tail entity set; perform semantic relevance analysis on the head entity element and the relationship element through a language analysis model to obtain semantic relevant information; generate a second tail entity set according to the semantic relevant information; aggregate the first tail entity set and the second tail entity set to obtain an initial tail entity set; score each initial tail entity element in the initial tail entity set to obtain a second scoring result; screen out a plurality of candidate tail entity elements from the initial tail entity elements based on the second scoring result, and generate a candidate tail entity set based on the plurality of candidate tail entity elements.
7. A knowledge graph link prediction device, characterized in that, The knowledge graph link prediction device includes: a memory, a processor, and a knowledge graph link prediction program stored on the memory and executable on the processor, and the knowledge graph link prediction program is configured to implement the knowledge graph link prediction method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, A knowledge graph link prediction program is stored on the computer-readable storage medium, and when the knowledge graph link prediction program is executed by a processor, it implements the knowledge graph link prediction method according to any one of claims 1 to 5.
9. A computer program product, characterized in that, The computer program product includes a knowledge graph link prediction program, and when the knowledge graph link prediction program is executed by a processor, it implements the steps of the knowledge graph link prediction method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Knowledge graph completion method and device based on vector embedding and transfer learning fusion
CN116701647A