An iterative relevance search method on a knowledge graph

Through the iterative related search method on the knowledge graph, the maximum weighted point coverage algorithm and positive and negative case markings in user interaction are used to solve the relevant search problems in case of no-sample entity, and the fast and accurate search result return is achieved.

CN114860849BActive Publication Date: 2025-06-10NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110158137.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-04
Publication Date
2025-06-10
Estimated Expiration
2041-02-04

AI Technical Summary

Technical Problem

Existing relevant search technologies cannot complete the search without the user providing sample entity pairs.

Method used

A iterative related search method on the knowledge graph is designed to generate a list of diverse results through a greedy algorithm covered by the maximum weighted point, and the metapath weight is adjusted using positive and negative example marks in user interaction to gradually explore user intentions.

Benefits of technology

Without a positive example, you can quickly return the diversity result list and continuously adjust the search results through iterative interaction until the user needs are met, reducing the difficulty of users in related searches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114860849B_ABST
    Figure CN114860849B_ABST
Patent Text Reader

Abstract

An iterative relevant search method on a knowledge graph, given a knowledge graph, a query radius, and the length of the result list, outputs a query entity, obtains the corresponding meta-paths within the radius, and then obtains a diverse result list with the largest possible weight covering the meta-paths. The relevance of each meta-path is obtained based on the positive and negative examples annotated by the user, and the most relevant relevant result list is obtained according to the meta-path relevance. The relevant result lists are merged to obtain the final result list; the user interacts with the final result list, selects and annotates positive and negative examples, and returns a new round of results according to the user's interaction information; if the user does not want to continue the interaction, this search is completed. The present invention solves the problem that results cannot be obtained in the case of relevant search when the user does not specify a positive example, that is, the "cold start" problem, and gradually determines the search results through interaction when the positive and negative examples are not clear, ensuring the normal operation of the search system and the return of results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer technology, relates to the related search of graph data, and is an iterative related search method on a knowledge graph. Technical Background

[0002] Graph data has good expressiveness. For example, Resource Description Framework (RDF) data can intuitively display the complex relationships between entities, i.e., the points on the graph, and is applied in more and more fields. In particular, related search, as a classic research topic in the field of data mining, hopes to find the answer entity most relevant to the query entity in graph data. Generally, for the definition of related search, it is considered that related means that the query entity and the answer entity satisfy certain specific meta-path combinations.

[0003] Graph data is a finite directed graph, also known as an entity association graph, denoted as G = <E, A, R, l>. Among them, E is a set of entities, represented as vertices in the graph; A is a set of arcs, and the direction of each arc a ∈ A is from the tail node t(a) to the head node h(a), where t(a) ∈ E and h(a) ∈ E; R is a set of relationships; l: A → R represents that l labels the arc a (a ∈ A), and the relationship l(a) ∈ R. For a given query entity q and answer entity r, there can be several paths (e 1 , a 1 , e 2 , a 2 , …, e n ) between them. Among them if a i+1 exists; at the same time, we call (a 1 , …, a n-1 ) the meta-path corresponding to this path.

[0004] When the prior art completes the related search task, it often requires the user to give example entity pairs {(s 1 , t 1 ), …, (s m , t m )} for the given query entity q, that is, m groups of examples, where s m , t m represent the start and end entities in an entity pair, or require other format data that can be converted into example entity pairs. In other words, the user needs to tell the search system some examples to represent their intentions, so that the system can complete the search task for the query entity after understanding the intentions. However, this brings a problem, that is, when the user cannot provide examples, the prior art cannot complete the related search task. Therefore, the present invention proposes an iterative related search method on a knowledge graph. Summary of the Invention

[0005] The problem to be solved by the present invention is that existing related search technical solutions cannot complete the search without the user providing sample entity pairs. In view of the situation where samples may not exist during related search on graph data, the present invention designs an interactive method process to mine the user's intention and return the results to the user.

[0006] The technical solution of the present invention is: an iterative related search method on a knowledge graph. In the knowledge graph G, with the entity input by the user as the query entity, the maximum distance from the query entity to the candidate entity as the query radius. For the given query entity q, query radius constraint L, and result list length K, obtain the meta-path, assign an initial weight to all meta-paths, and use the greedy algorithm of maximum weighted vertex cover to obtain a diverse result list E with as large a covered meta-path weight as possible. D Perform the following steps:

[0007] 1) If there are positive examples marked by the user, use a linear classifier to obtain the correlation weight of each meta-path, obtain the relevance of each entity according to the correlation weight of the meta-path, and select the K entities with the highest relevance as the relevant result list E. R From the diverse result list E D and the relevant result list E R Sample to form the final result list and return it to the user;

[0008] 2) If there are no positive examples marked by the user, then use the diverse result list E D as the final result list and return it to the user;

[0009] 3) For the returned final result list, confirm whether the user intends to continue interacting. If the user chooses to continue interacting, mark the entities in the final result list as positive or negative examples according to the user's choice, adjust the weights of the meta-paths of the entities according to the marking results, and use the greedy algorithm to update and calculate the diverse result list E D Then perform a new round of iteration of steps 1) and 2). If the user chooses to end the interaction, end the search.

[0010] As a preferred method, after obtaining the relevant result list E R Determine the number of entities sampled from the diverse result list E D according to the set sampling ratio parameter ∈ and the number of entities sampled from the relevant result list E R ​Sample the entities with the highest rankings from two result lists in turn to form the final result list. If the number of entities sampled from a certain result list has reached the requirement, continue sampling from the other result list until the number of entities in the final result list reaches K, and then return the final result list to the user.

[0011] Further, when continuing the interaction, the marking methods selected by the user include: marking several positive and negative examples, only marking positive examples, only marking negative examples, and not performing positive and negative example annotation. For the cases of only marking positive examples and not annotating, the first entity in the final result list that is not marked as a positive example is regarded as a negative example.

[0012] As a preferred method, the diverse result list E D is obtained as follows:

[0013] (1) Initially, according to the query entity q and the given query radius constraint L, find all the meta-path sets MP in the knowledge graph, and set an initial diversity weight Div for each meta-path mp mp = log(Freq(mp) + 1) * ρ(len(mp)), where Freq is the number of instance occurrences of the meta-path mp in the subgraph centered on q with L as the radius, len(mp) is the length of the meta-path mp, and ρ is a density-based normalization coefficient related to the length;

[0014] (2) Based on the meta-path set MP, the weight Div of each meta-path mp and the coverage relationship between entities and each meta-path, model the problem as a maximum weighted vertex cover problem, and use the greedy algorithm to obtain the corresponding diverse result list E D :

[0015] (2.1) Initialize E D as an empty set, and the covered meta-path set P as an empty set;

[0016] (2.2) Visit all the entities in the candidate entity set that are not included in E D . For each entity, calculate the diversity weight corresponding to the meta-paths that it covers and are not in P, and sum the calculated diversity weights as the diversity score of this entity;

[0017] (2.3) Select the entity with the highest diversity score and add it to E D , and at the same time add all the meta-paths it covers to P;

[0018] (2.4) If the number of entities in E D has reached K, return the final E D ; otherwise, go to (2.2).

[0019] As a preferred method, the relevant result list E RThe acquisition is as follows:

[0020] (3.1) Determine the positive example set E according to user annotation + and the negative example set E _ . If there are no user-annotated negative examples, randomly sample from the candidate entity set to obtain them; define the feature space as the union of the meta-paths covered by all entities in E + ; the feature value of an entity is defined as the number of instances covering the corresponding meta-path;

[0021] (3.2) For each pair of positive example entities e + selected from E + and negative example entities e _ selected from E _ , construct two training samples (e + - e - , 1) and (e - - e + , -1); where e + - e - means that for each dimension of the feature, the value of this sample is the difference between the value corresponding to e + and the value corresponding to e - ; e - - e + is the same;

[0022] (3.3) Use all the training samples constructed in (3.2) to train a linear classifier, and regard the weight of each dimension of the classifier as the relevance weight of the corresponding meta-path;

[0023] (3.4) For each entity in the candidate entity set, its relevance score is the sum of the products of the relevance weights of all the meta-paths it covers and the number of covering instances;

[0024] (3.5) Select K entities with the highest relevance scores to form the relevant result list E R Return.

[0025] As a preferred method, the specific adjustment of the weights of the meta-paths of the positive and negative example entities marked by the user is as follows:

[0026] For all the meta-path sets MP + covered by the entities marked as positive examples E + , divide the diversity weight of each meta-path by the attenuation factor β; for all the meta-path sets MP _ covered by the entities marked as negative examples E - , multiply the diversity weight of each meta-path by the attenuation factor β; at the same time, according to the number of positive example entities contributed by the diversity result list and the relevant result list respectively, while ensuring ∈ min ≤∈′≤∈ maxAdjust and update the sampling ratio coefficient according to the formula in the case of:

[0027]

[0028] ∈ min and ∈ max are the upper and lower limits of the set sampling ratio coefficient, and λ is the update rate.

[0029] The iterative relevance search of the present invention utilizes the interaction with the user to obtain the user's intention, thereby greatly reducing the difficulty of the user performing relevance search by avoiding asking the user to directly provide examples. The present invention proposes a diversity-based module that can, in the case where the user does not provide example entity pairs, i.e., "cold start", as quickly as possible through interaction, enable the user to obtain and mark answers, so that the user can obtain satisfactory relevance search results.

[0030] The beneficial effects of the present invention are as follows: For the situation where the user cannot provide positive examples during relevance search on graph data, the existing methods cannot complete the search task. The present invention takes the lead in proposing a solution to the relevance search task in the case of no positive examples, returns the search results to the user as a diversity result list in the case of no positive and negative examples, and then further utilizes the iterative user interaction process to mine the user's intention and complete the search task. The present invention can start the search without the user providing example entity pairs. As long as the user annotates the returned results to achieve interactive correction, the search can continuously approach the user's search purpose. At the same time, the method of the present invention innovatively proposes the concept of diversity ranking in the relevance search task, proposes a diverse result list, and can detect the user's intention faster by covering as many meta-paths as possible. On the one hand, it realizes the "cold start" problem of search in the case of no example entity pairs. In addition, due to detecting the user's intention, it can effectively reduce the cost that the user needs to spend to obtain the desired information. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is the overall processing flow chart of the present invention.

[0032] Figure 2 is the flow chart for obtaining the diverse result list of the present invention.

[0033] Figure 3 is the flow chart for obtaining the relevant result list of the present invention.

[0034] Figure 4 is an example diagram of an association query under the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0035] For relevant searches in the case where no positive examples are included (i.e., "cold start"), the present invention proposes an iterative relevant search method. Through the combined action of the diversity method and the relevance method, the user's intention is captured during the interaction with the user and the results are returned to the user. The present invention assumes that the user's intention can be represented by a combination of several meta-paths on the knowledge graph, and there is a specific meta-path combination between the query entity and the result entity. The entity input by the user is the query entity, and the maximum distance from the query entity to the candidate entity is the query radius. For a given knowledge graph, query radius constraint, query entity, and the length of the result list, the final result list is calculated and returned to the user; then, based on the user's feedback, the returned final result list is improved, or the interaction is ended according to the user's selection. First, the diversity module uses the greedy algorithm of maximum weighted vertex cover according to the meta-path weights to obtain a diverse result list that covers as large a meta-path weight as possible. If there are no user-labeled positive examples at this time, the diverse result list is used as the final result list; if there are user-labeled positive examples at this time, the relevance weights of each meta-path are obtained using a linear classifier based on the labeled positive and negative examples, and the entity with the largest sum of the relevance weights of the covered meta-paths is selected as the relevant result list. Then, the final result list is sampled from the diverse result list and the relevant result list according to a dynamically adjusted proportionality coefficient and returned to the user. The user can choose:

[0036] (1) Mark positive and negative examples according to the intention for the final result list to represent the user's intention. The marked positive and negative example data returns to the search system, and the system makes corresponding adjustments to the diverse result list according to the intention and obtains the next-round results.

[0037] (2) Be satisfied with the current results or not want to perform the next-round interaction, and choose to end this search.

[0038] Such as Figure 1 、 Figure 2 and Figure 3 , the complete flowchart of the present invention is as shown in Figure 1 shown, the process of obtaining the diverse result list is as shown in Figure 2 shown, and the process of obtaining the relevant result list is as shown in Figure 3 shown. The specific implementation of the present invention is as follows.

[0039] (1) Initially, according to the query entity q and the given query radius constraint L, find all meta-path sets MP, and set an initial diversity weight Div for each meta-path mp in it mp = log(Freq(mp)+1)*ρ(len(mp)), where Freq is the number of instance occurrences of the meta-path mp in the sub-graph centered on q with L as the radius, len(mp) is the length of the meta-path mp, and ρ is a density-based normalization coefficient related to the length.

[0040] (2) Based on the meta-path set MP and the weight Div of each meta-path mp and the coverage relationship between entities and each meta-path, the problem is modeled as a maximum weighted vertex cover problem, and the corresponding diverse result list E is obtained using a greedy algorithm D , the list is an ordered list, and the entities are ranked 1 to K in order:

[0041] (2.1) Initialize E D as an empty set, and the covered meta-path set P as an empty set;

[0042] (2.2) Visit all entities in the graph candidate entity set that are not included in E D , for each entity, calculate the diversity weight corresponding to the meta-paths that it covers and are not in P, and sum the calculated diversity weights as the diversity score of this entity;

[0043] (2.3) Select the entity with the highest diversity score and add it to E D , and at the same time add all the meta-paths it covers to P;

[0044] (2.4) If the number of entities in E D has reached K, return E D ; otherwise, go to (2.2).

[0045] (3) If there are positive examples marked by the user, according to the positive examples marked by the user and the negative examples, use a linear classifier to obtain the relevance weight of each meta-path, obtain the relevance degree of each entity according to the relevance weight of the meta-path, and select the K entities with the highest relevance degree as the relevant result list E R :

[0046] (3.1) For the set of positive examples E + marked by the user, and the set of negative examples E _ marked by the user, if there are no negative examples marked by the user, randomly sample from the candidate entity set; the feature space is defined as the union of the meta-paths covered by all entities in E + ; the feature value of an entity is defined as the number of instances covering the corresponding meta-path;

[0047] (3.2) For each pair of positive example entity e + selected from E + and negative example entity e _ selected from E _ , construct two training samples (e + - e - , 1) and (e - - e + , -1); where e + - e -Indicates that for each dimension feature, the value of the sample is e + The corresponding value is - The value of a sample refers to the eigenvalue of the sample corresponding to a certain dimension of the feature, such as e + The corresponding value is the feature value of the positive entity on this dimension feature; - -e + Similarly;

[0048] (3.3) Use all the training samples constructed in (3.2) to train the linear classifier, and regard the weight of each dimension of the classifier as the correlation weight of the corresponding meta-path;

[0049] (3.4) For each entity in the candidate entity set, its relevance score is the sum of the relevance weights of all covered meta-paths multiplied by the number of covered instances;

[0050] (3.5) Select K entities with the highest relevance scores to form the relevant result list E R , E R It is also an ordered list, and the entities are ranked 1 to K in order.

[0051] (4) Get the relevant result list E R After that, the sampling ratio parameter ∈ is used to determine the number of samples from the diverse result list E D The number of entities sampled and from the relevant result list E R The number of entities sampled ∈ is initialized to 0.5 and then adjusted according to the iteration. The top ranked and unselected entities are sampled from the two result lists in turn to form the final result list. If the number of entities sampled from a result list has reached the requirement during sampling, continue sampling from the other result list until the number of entities in the final result list is K and return it to the user, then go to (6).

[0052] (5) If there is no positive example marked by the user, that is, the relevant result list is empty, the diverse result list is returned to the user as the final result, and go to (6).

[0053] (6) If the user wishes to continue the interaction, the entities in the final result list are marked as positive and negative examples according to the user's choice. The positive and negative examples are used to indicate whether the entity meets the search requirements. Go to (7). The marking method selected by the user includes: marking a number of positive and negative examples, marking only positive examples, marking only negative examples, and not marking positive and negative examples. For the cases of marking only positive examples and not marking, the first entity in the final result list that is not marked as a positive example is marked as a negative example. If the user does not wish to continue the interaction, go to (9).

[0054] (7) For all entities E marked as positive examples + The set of covered metapaths MP+ , divide the diversity weight of each meta - path therein by the decay factor β; for all entities E marked as negative examples _ the set of meta - paths MP covered - , multiply the diversity weight of each meta - path therein by the decay factor β. If there are no user - labeled negative examples, only process the weights of the meta - paths corresponding to positive examples; at the same time, according to the number of positive - example entities contributed by the diversity result list and the relevant result list respectively, adjust the sampling ratio coefficient according to the formula:

[0055]

[0056] where ∈ min ≤∈′≤∈ max , ∈ min and ∈ max are the upper and lower limits of the set sampling ratio coefficient, and λ is the update rate, used to control the update degree of the sampling ratio coefficient.

[0057] (8) According to the adjusted meta - path weights and sampling ratio coefficient, go to (2) to update the diversity result list E D and the relevant result list E R , and complete a new round of search by iteration.

[0058] (9) End this search.

[0059] The method of the present invention can start the search without example entity pairs, first calculate the diversity result list and return it to the user as the final result list. Then the user can make independent adjustments according to the returned results. Through the annotation of positive and negative examples in the iterative search, the weight of the meta - path corresponding to the positive - example entity is increased, and the search results are continuously adjusted to the user's search requirements. Even if the user does not make annotations or only annotates negative examples during the interaction of the iterative search, that is, the user expresses that the search results do not meet the requirements, because the negative examples will also affect the diversity weight, the results of the next round can still be updated, so that the search results are continuously adjusted to explore the user's intention. Compared with the prior art that requires the user to provide data, the present invention only requires the user to make selections based on the existing results, and can even update itself without selection, greatly reducing the requirements for search conditions, and the search target can continuously approach the user's intention.

[0060] Figure 4 is an embodiment of the present invention and an example of an entity association graph. The rounded - rectangle boxes represent entities, and the arrow lines represent the relationships between two entities. Among them, the query entity is marked in black characters on a white background, that is, Don Lusk, and the other entities are some entities within L = 2 steps around the query entity. For the relationship director, that is, the director of the movie, use director -1 to represent its inverse relationship, that is, the movies produced by this director, use <director -1,starring> is used to represent the meta - path, that is, the actors in the movies produced by this director, narrator represents the narrator, and composer represents the dubbing. The following will take Figure 4 as an example to specifically explain the process of the present invention, where L = 2 and K = 2.

[0061] First, analyze all the meta - paths <director -1 >, <director -1 ,starring>, <director -1 ,narrator>, <director -1 ,composer>. Assume that the normalization coefficient ρ(l)=1. We can obtain the weights of all the meta - paths: Div(<director -1 ) = 1.61, Div(<director -1 ,starring>) = 1.39, Div(<director -1 ,narrator>) = 1.10, Div(<director -1 ,composer>) = 0.69. Therefore, the two entities with the largest diversity weights we obtain are Michael Rye = 2.49, Midnight Patrol: Adventures in the Dream Zone = The Jetsons Meet the Flintstones = Yogi and the Invasion of the Space Bears = The Greatest Adventure: Stories from the Bible = 1.61. It should be particularly noted that when Michael Rye is selected in step 2, the meta - paths <director -1 ,starring>, <director -1 ,narrator> have been covered, and at this time, the weights of Nick Turnbull and Patric Zimmerman become 0. Assume that the diversity result list at this time is (Michael Rye, The Jetsons Meet the Flintstones).

[0062] Since the user does not provide positive examples, in the first - round case, the diversity result list is the final result list, and we mark the result with the superscript 1 .

[0063] Assume that the user's intention is the meta - path <director-1 , narrator>, at this time, the user hopes to continue the relevant search process and marks Michael Rye as a positive example (marked with a subscript + for marking), and marks The Jetsons Meet the Flintstones as a negative example (marked with a subscript - for marking).

[0064] Enter the second round of interaction. In the diverse result list section, let's set β = 0.1. Based on the known positive and negative example situations, we adjust the weights to Div(<director -1 >) = 0.16, Div(<director -1 , starring>) = 13.9, Div(<director -1 , narrator>) = 11.0, Div(<director -1 , composer>) = 0.69. Therefore, the entities selected for the second-round diverse result list should be (Michael Rye, Hoyt Curtin).

[0065] In the relevant result list section, considering that the meta-paths corresponding to the feature space are <director -1 , starring> and <director -1 , narrator>, and currently they have the same weight, so the relevant result list is (Michael Rye, Patric Zimmerman).

[0066] At this time, the sampling coefficient ∈ = 0.5, which means that each of the diverse result list and the relevant result list contributes one result during sampling. Let's assume that the finally returned result is (Michael Rye, Patric Zimmerman).

[0067] At this time, the user can continue to mark Patric Zimmerman as a negative example. In the subsequent relevant modules, the system will find that there are samples (Michael Rye - Patric Zimmerman, 1) and (Patric Zimmerman - Michael Rye, -1). Through the linear classifier, the system will find that the user's intention is the meta-path <director -1 , narrator>, and return all the results that meet this intention to the user. The user can then end this relevant search process, obtain the required results, and complete the search task.

Claims

1. An iterative relevant search method on a knowledge graph, characterized in that In the knowledge graph G, with the entity input by the user as the query entity, the maximum distance from the query entity to the candidate entity is the query radius. For the given query entity q, query radius constraint L, and result list length K, obtain the meta-path, assign an initial weight to all meta-paths, and use the greedy algorithm of maximum weighted vertex cover to obtain a diverse result list E that covers as large a meta-path weight as possible D , perform the following steps: 1) If there are positive examples marked by users, use a linear classifier to obtain the correlation weights of each meta-path, get the relevance degrees of each entity according to the correlation weights of the meta-paths, and select the top K entities with the highest relevance degrees as the relevant result list E R , sample from the diverse result list E D and the relevant result list E R to form the final result list and return it to the user; 2) If there is no positive example marked by the user, then the diverse result list E D is returned to the user as the final result list; 3) For the returned list of final results, confirm the user's intention to continue the interaction. If the user chooses to continue the interaction, mark the positive and negative examples of the entities in the final result list according to the user's selection, adjust the weights of the entity's meta-paths according to the marking results, and use the greedy algorithm to update and calculate the diverse result list E D , and then perform a new round of iteration of steps 1) and 2). If the user chooses to end the interaction, end the search; the weights of the meta-paths of the positive and negative example entities marked by the user are adjusted specifically as follows: For all entities E marked as positive examples + The set of meta-paths MP covered + , divide the diversity weight of each meta-path therein by the decay factor β; for all entities E marked as negative examples _ The set of meta-paths MP covered - , multiply the diversity weight of each meta-path therein by the decay factor β; at the same time, according to the number of positive example entities contributed by the diversity result list and the relevant result list respectively, while ensuring ∈ min ≤ ∈′ ≤ ∈ max , adjust and update the sampling ratio coefficient according to the formula under the condition of ∈ min and ∈ max are the upper and lower limits of the set sampling ratio coefficient, ∈ is the set sampling ratio parameter, ∈′ is the updated sampling ratio coefficient, and λ is the update rate.

2. The iterative relevant search method on a knowledge graph according to claim 1, characterized in that Obtain the relevant result list E R After that, determine the number of entities sampled from the diverse result list E according to the set sampling ratio parameter ∈ D The number of entities sampled And the number of entities sampled from the relevant result list E R The number of entities sampled Take turns sampling the entities with the highest ranking from the two result lists to form the final result list. If the number of entities sampled from a certain result list has reached the requirement, continue sampling from the other result list until the number of entities in the final result list reaches K, and return the final result list to the user.

3. The iterative relevant search method on a knowledge graph according to claim 1 or 2, characterized in that When continuing the interaction, the marking methods selected by the user include: marking a number of positive and negative examples, only marking positive examples, only marking negative examples, and not performing positive and negative example annotation. For the cases of only marking positive examples and not annotating, the first entity in the final result list that is not marked as a positive example is regarded as a negative example.

4. The iterative relevant search method on a knowledge graph according to claim 1 or 2, characterized in that Diverse result list E D is obtained as follows: (1) Initially, according to the query entity q and the given query radius constraint L, find all the meta-path sets MP in the knowledge graph, and set the initial diversity weight Div for each meta-path mp among them mp = log(Freq(mp)+1)*ρ(len(mp)), where Freq is the number of instance occurrences of the meta-path mp in the subgraph centered on q with L as the radius, len(mp) is the length of the meta-path mp, and ρ is a density-based normalization coefficient related to the length; (2) Based on the meta-path set MP and the weight Div of each meta-path mp and the coverage relationship between entities and each meta-path, the problem is modeled as a maximum weighted vertex cover problem, and the corresponding diverse result list E is obtained using a greedy algorithm D : (2.1) Initialize E D is an empty set, and the covered meta-path set P is an empty set; (2.2) Access all entities not included in E in the candidate entity set. For each entity, calculate the diversity weights corresponding to the meta-paths covered by it that are not in P, and sum up the calculated diversity weights as the diversity score of this entity; D ​ (2.3) Select the entity with the highest diversity score and add it to E D and add all the meta-paths it covers to P at the same time; (2.4) If the number of entities in E D has reached K, return the final E D ; otherwise, go to (2.2).

5. The iterative relevant search method on a knowledge graph according to claim 1 or 2, characterized in that List of relevant results E R is obtained as follows: (3.1) Determine the positive example set E according to user annotation + and the negative example set E _ . If there is no user-annotated negative example, randomly sample from the candidate entity set; define the feature space as E + the union of the meta-paths covered by all entities; the feature value of an entity is defined as the number of instances covering the corresponding meta-path (3.2) For each pair of positive example entities e + selected from E + and negative example entities e _ selected from E _ , construct two training samples (e + -e - , 1) and (e - -e + , -1); where e + -e - means that for each dimension of features, the value of this sample is the difference between the value corresponding to e + and the value corresponding to e - ; e - -e + is the same; (3.3) Use all the training samples constructed in (3.2) to train a linear classifier, and regard the weight of each dimension of the classifier as the correlation weight of the corresponding meta-path; (3.4) For each entity in the candidate entity set, its correlation score is the sum of the products of the correlation weights of all the covered meta-paths and the number of covered instances; (3.5) Select K entities with the highest relevance scores to form the relevant result list E R Return.

Citation Information

Patent Citations

  • Knowledge map representation learning method based on multi-step relational path

    CN108959472A

  • Using activation paths to cluster proximity query results

    US20080183695A1