A search recall method and system based on knowledge graph representation learning

Through the method based on knowledge graph representation learning, the problem of difficulty in using deep semantic information and poor interpretability in document search recall in professional fields is solved, rapid response and effective cold start are achieved, and the semantic recall ability and interpretability of search recall are improved.

CN115618113BActive Publication Date: 2025-06-20NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211368701.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-03
Publication Date
2025-06-20
Estimated Expiration
2042-11-03

AI Technical Summary

Technical Problem

In the search and recall scenarios of professional documents, the prior art has problems such as difficulty in using deep semantic information, poor interpretability, and difficulty in taking into account the system's response speed and effective cold start.

Method used

A search and recall method based on knowledge graph representation learning is adopted, and the domain knowledge graph is represented and learned through offline stages, and a vector representation of the entity is obtained, and a word-vector library, vector index library and inverted index library are established. In the online stage, based on the input query statement, the accurate vector and candidate similarity vector are obtained through the index library query, vector similarity calculation is performed, semantic similarity words are obtained, and the candidate document list is finally returned.

Benefits of technology

The semantic recall capability of search recall is improved, the cold start of the search recall process is realized, the dependence of the training model on corpus is reduced, the interpretability of the recall process is improved, and the response time is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115618113B_ABST
    Figure CN115618113B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a search recall method and system based on knowledge graph representation learning. The method uses representation learning based on the knowledge graph to obtain vectors corresponding to entities in the graph, establishes an index library corresponding to the target document according to the vectors, and then through entity linking, query and retrieval of vectors, fine calculation of vector similarity, and document retrieval, realizes the process from the user input query statement to obtaining the retrieval result, improving the semantic recall ability of search recall and achieving the effect of cold start. Through representation learning, vectors corresponding to entities can be obtained, and entities in the query statement can be accurately matched through entity linking. Compared with the prior art, it can solve the problem of inaccurate recall caused by word segmentation in the prior art, realize the effective recall of documents in the scenarios of "one word with multiple meanings" and "multiple words with the same meaning", and at the same time reduce the dependence of the training model on the corpus during the recall process, improving the interpretability while ensuring the response time of the recall process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a search and recall method and system based on knowledge graph representation learning. Background Art

[0002] The two most important stages in the search process are the recall stage and the ranking stage. Among them, the recall stage is the basis of the ranking stage. In the recall stage, resources that the user is interested in are screened out from a vast resource pool as the candidate resource set for the ranking stage. If there are too many recall results, it will lead to a large computational pressure and increased computational difficulty in the subsequent refined ranking process, resulting in failure to achieve accurate retrieval according to the user's input intention; while if there are too few recall results, search results may be missed, resulting in missing content that the user sees. Therefore, the quality of the recall results basically determines the quality of the entire search results.

[0003] Currently, common search and recall methods include traditional search and recall methods based on literal (word or keyword) matching and semantic recall methods based on vectors. The search method based on literal matching is sometimes also called the word-based search and recall method or the keyword matching-based search and recall method. The underlying implementation of the search and recall based on literal matching is based on an inverted index. By building an inverted index for the search target documents, when performing a search, the keywords input by the user are segmented, and then the keywords are retrieved and matched in the inverted index. Through calculation using scoring functions such as BM25, the documents with scores higher than a specified threshold are returned.

[0004] The advantages of the search method based on literal matching are simple and easy to understand, but its disadvantages are also very obvious, including: 1) Incorrect or inaccurate recall results caused by segmentation errors. For example, when a user searches for "Nanjing Yangtze River Bridge", it is difficult to determine whether it is to search for "the mayor named Jiang Daqiao in Nanjing" or "the bridge on the Yangtze River in Nanjing"; 2) Unable to distinguish the situation of polysemy. For example, when a user retrieves and inputs "apple", the computer cannot determine whether it refers to the "apple" in fruits or the famous American company "Apple" and its series of products; 3) Unable to implement the retrieval of synonyms. For example, when a user searches for "tomato", documents containing "tomato" cannot be retrieved.

[0005] The core idea of the vector-based semantic recall method is to map the search input and the target document into a semantic space with distributed vector representations in the same dimension, and then calculate the vector similarity to obtain the documents that match the search input. The vector-based semantic recall method can solve the problems in the search recall method based on literal matching to a certain extent. There are two implementation ideas: One is to train an end-to-end similarity model based on deep neural network technology, and use the large-scale corpus generated based on historical click behavior data to train and learn a hidden layer similarity semantic model, thus realizing retrieval and recall. The main problems of this type of method are: 1) It requires a large amount of training corpus support and cannot work during the cold start of the system; 2) The system interpretability during the recall process is poor. The other is to calculate the relevance between the search input and the target document based on a pre-trained language model. By introducing the pairwise approach mode, the optimization goal is changed to the ranking position (matching degree) between two candidate documents in the search recall scenario. This method can reduce the dependence on the training corpus, but it still cannot improve the interpretability of the recall process. At the same time, the response speed of the system will be reduced due to the introduction of a large model. In addition, it is difficult for the vector-based semantic recall method to capture the deep semantic information between the search input and the document.

[0006] In the search recall scenario of professional field documents, there are problems such as difficulty in using deep semantic information, poor interpretability, and it is difficult to balance the response speed of the system and effective cold start. Therefore, new technologies need to be introduced to optimize the search recall of professional field documents. Summary of the Invention

[0007] The present invention provides a search recall method and system based on knowledge graph representation learning to solve the problems existing in the existing search recall methods in the search recall scenario for professional field documents, such as difficulty in using deep semantic information, poor interpretability, and difficulty in balancing the response speed of the system and effective cold start.

[0008] The purpose of this application is to provide the following aspects:

[0009] In the first aspect, this application provides a search recall method based on knowledge graph representation learning, and the method includes:

[0010] Obtain a domain knowledge graph and a target document, where the domain knowledge graph and the target document belong to the same professional field;

[0011] Offline stage: Through the representation learning based on the knowledge graph, obtain the vectors corresponding to the entities in the domain knowledge graph, and establish an index library corresponding to the target document according to the vectors. The index library includes a word-vector library, a vector index library, and an inverted index library;

[0012] The word-vector library is used to store the feature vector representations corresponding to the entities in the domain knowledge graph and store the mapping relation table between the entities and the feature vector representations. The vector index library is used to provide indexes for the feature vector representations corresponding to the entities in the domain knowledge graph. The inverted index library is used to store all the entities in the domain knowledge graph;

[0013] Online stage: Match the entities in the domain knowledge graph according to the input query statement;

[0014] Query through the index library according to the entities obtained by matching to obtain accurate vectors and candidate similar vectors;

[0015] Calculate the vector similarity according to the accurate vectors and candidate similar vectors to obtain a list of similar vectors;

[0016] Query vectors in the vector index library according to the list of similar vectors to obtain semantic similar words corresponding to the entities of the query statement;

[0017] Obtain a list of candidate documents through the inverted index library according to the semantic similar words and return the search results.

[0018] In the present invention, only the knowledge graph representation learning in the offline stage requires large-scale training and computing work. The online stage completes the whole process from the user input query statement to obtaining the retrieval result. This process is vector retrieval and calculation based on the offline stage. Therefore, compared with the method based on the pre-trained language model, the response speed of this application during search is faster. In addition, compared with the vector-based method, this method does not require a large-scale training corpus to train the model and can effectively perform cold start; and the intermediate results of the semantic generalization process in the online stage are all understandable by people, thus enhancing the interpretability of the recall process.

[0019] In addition, in the online stage of the present invention, obtaining the semantic similar words means obtaining a candidate set of semantic similar words. The number of the semantic similar words can be determined by those skilled in the art according to experience. Usually, the number can be selected from 5 to 50. Specifically, it is necessary to determine the actual selected value of the number of semantic similar words in the candidate set according to the density of data features. If the density of data features is high, fewer semantic similar words can be selected without considering the situation that too few semantic similar words will miss search results; on the contrary, if the density of data features is low, more semantic similar words need to be selected to avoid missing search results.

[0020] Combined with the first aspect, in one implementation, the offline stage includes:

[0021] Input the domain knowledge graph into a graph convolutional network, perform feature selection through the graph convolutional network, and output the first intermediate feature vector representation of the entity;

[0022] According to the first intermediate feature vector representation, through activation function transformation and filtering (dropout), obtain the second intermediate feature vector representation of the entity;

[0023] Perform multiple graph convolutions, activation function transformations, and filtering on the second intermediate feature vector representation to obtain the final feature vector representation of the entity;

[0024] Save the final feature vector representation of the entity in the word-vector library and the vector index library.

[0025] In the present invention, after inputting the domain knowledge graph into a graph convolutional network to obtain the intermediate feature vector representation, perform convolution on the intermediate feature vector representation, that is, realize operations such as continuous transformation, screening, and filtering of the intermediate feature vector representation to obtain the final feature vector representation; due to the high generality of the graph convolutional network, good results can be achieved in all aspects.

[0026] Combined with the first aspect, in one implementation, the establishing the index library corresponding to the target document according to the vector includes:

[0027] Perform entity linking between the target document and the domain knowledge graph, that is, match the entities in the target document with the entities in the domain knowledge graph to obtain the entity list linked by each target document;

[0028] Construct an index for each target document and the entity list linked by the target document through the inverted index technology, that is, establish the inverted index library;

[0029] Establish the vector index library according to the vectors corresponding to the entities linked by all the target documents.

[0030] In the present invention, performing entity linking between the target document and the domain knowledge graph is a prerequisite for accurately matching the entities in the query statement in the online stage. Using entity linking can effectively solve the semantic matching problem in the scenarios of "one word with multiple meanings" and "multiple words with the same meaning", and improve the accuracy and coverage rate of the subsequent search results matching the query statement.

[0031] Combined with the first aspect, in one implementation, the matching the entities in the domain knowledge graph according to the input query statement includes:

[0032] Obtain the input query statement;

[0033] Entity-link the query statement with the domain knowledge graph, that is, obtain the entities in the query statement, and match the entities in the query statement with the entities in the domain knowledge graph.

[0034] In this step, entity-link the query statement with the domain knowledge graph. In the offline phase, entity-linking has been performed on the target documents and the domain knowledge graph, which can clarify the semantics of each entity (word) in the query statement, and finally effectively solve the semantic matching problem in the scenarios of "one word with multiple meanings" and "multiple words with the same meaning". Specifically, in the process of entity-linking, on the one hand, semantic judgment is made based on the literal meaning of the word itself, and at the same time, semantic supplementation is carried out based on the relevant context of the word; after representing the word and its context with the current popular pre-trained language model for disambiguation, the problem of inaccurate recall results caused by "one word with multiple meanings" and "multiple words with the same meaning" can be solved, achieving effective recall and improving the accuracy and coverage of recall.

[0035] Combined with the first aspect, in one implementation, querying through the index library according to the entities obtained by matching to obtain accurate vectors and candidate similar vectors includes:[[]]

[0036] Query through the word-vector library in the index library according to the entities obtained by matching to obtain accurate vectors;

[0037] Query through the vector index library in the index library according to the entities obtained by matching to obtain candidate similar vectors, and the candidate similar vectors are a set of similar vectors. Specifically, in this step, a vector search algorithm can be used to quickly find similar vectors to the input query vector from a large number of vectors, greatly improving the search speed.

[0038] Combined with the first aspect, in one implementation, calculating the vector similarity according to the accurate vectors and the candidate similar vectors to obtain a list of similar vectors includes:[[]]

[0039] Determine the number of the accurate vectors;

[0040] If the number of the accurate vectors is one, calculate the similarity between each candidate similar vector and the accurate vector through the cosine similarity algorithm;

[0041] If the number of the accurate vectors is two or more, use weighted average to calculate the similarity between each candidate similar vector and each accurate vector; among them, the weight of each accurate vector adopts the weight of the entity corresponding to the accurate vector;

[0042] Arrange the list of similar vectors in descending order according to the similarity scores and return the list of similar vectors as the search result.

[0043] Second aspect, an embodiment of the present invention provides a search and recall system based on knowledge graph representation learning, and the system includes:

[0044] A preparation module, configured to obtain a domain knowledge graph, where the domain knowledge graph belongs to the same professional field as the target document;

[0045] An offline module, configured to obtain vectors corresponding to entities in the domain knowledge graph through representation learning based on the knowledge graph, and establish an index library corresponding to the target document according to the vectors, where the index library includes a word-vector library, a vector index library, and an inverted index library;

[0046] The word-vector library is used to store feature vectors corresponding to entities in the domain knowledge graph and store a mapping relationship table between entities and feature vectors. The vector index library is used to provide indexes of feature vectors corresponding to entities in the domain knowledge graph. The inverted index library is used to store all entities in the domain knowledge graph;

[0047] An online module, configured to match entities in the domain knowledge graph according to an input query statement;

[0048] Query through the index library according to the entities obtained by matching to obtain accurate vectors and candidate similar vectors;

[0049] Calculate vector similarity according to the accurate vectors and candidate similar vectors to obtain a list of similar vectors;

[0050] Query vectors in the vector index library according to the list of similar vectors to obtain semantic similar words corresponding to the entities of the query statement;

[0051] Obtain a list of candidate documents according to the semantic similar words through the inverted index library and return a search result.

[0052] Combined with the second aspect, in one implementation, the offline module includes:

[0053] A final feature vector representation obtaining unit, configured to input the domain knowledge graph into a graph convolutional network, perform feature selection through the graph convolutional network, and output a first intermediate feature vector representation of the entity; according to the first intermediate feature vector representation, perform transformation and filtering through an activation function to obtain a second intermediate feature vector representation of the entity; perform multiple graph convolutions, activation function transformation and filtering on the second intermediate feature vector representation to obtain a final feature vector representation of the entity;

[0054] A storage unit, configured to save the final feature vector representation of the entity in the word-vector library and the vector index library.

[0055] In a third aspect, the present application further provides a search recall program based on knowledge graph representation learning. The program, when executed, implements the steps of the search recall method based on knowledge graph representation learning described in the first aspect above.

[0056] In a fourth aspect, a computer-readable storage medium stores computer instructions thereon. When the instructions are executed by a processor, the steps of the search recall method based on knowledge graph representation learning described in the first aspect above are implemented.

[0057] In addition, the present invention further provides a search recall processing device. The search recall processing device includes: at least one processor; and a memory communicatively connected to the at least one processor. Wherein, the memory stores instructions executable by the one processor. When the instructions are executed by the at least one processor, the at least one processor executes the search recall method based on knowledge graph representation learning described in the first aspect above.

[0058] As can be seen from the above technical solutions, the embodiments of the present invention provide a search recall method and system based on knowledge graph representation learning. The method includes: obtaining a domain knowledge graph and a target document, where the domain knowledge graph and the target document belong to the same professional field;

[0059] Offline stage: Through representation learning based on the knowledge graph, vectors corresponding to entities in the domain knowledge graph are obtained, and an index library corresponding to the target document is established according to the vectors. The index library includes a word-vector library, a vector index library, and an inverted index library. The word-vector library is used to store the feature vector representations corresponding to entities in the domain knowledge graph and store the mapping relationship table between the entities and the feature vector representations. The vector index library is used to provide indexes of the feature vector representations corresponding to entities in the domain knowledge graph. The inverted index library is used to store all entities in the domain knowledge graph.

[0060] Online stage: According to the input query statement, entities in the domain knowledge graph are matched. According to the obtained entities, queries are made through the index library to obtain accurate vectors and candidate similar vectors. Vector similarity calculation is performed based on the accurate vectors and the candidate similar vectors to obtain a list of similar vectors. According to the list of similar vectors, vector queries are made in the vector index library to obtain semantic similar words corresponding to the entities of the query statement. According to the semantic similar words, a candidate document list is obtained through the inverted index library and the search results are returned.

[0061] In the prior art, traditional recall methods based on literal (word or keyword) matching and semantic recall methods based on vectors have difficulties in using deep semantic information, poor interpretability, and it is difficult to balance the response speed of the system and effective cold start in the document search and recall scenarios for professional fields. By using the foregoing methods, through representation learning based on a knowledge graph in the offline stage, vectors corresponding to entities in the domain knowledge graph are obtained, and an index library corresponding to the target document is established according to the vectors. In the online stage, through entity linking, query and retrieval of vectors, fine calculation of vector similarity, and document retrieval, the whole process from the user inputting a query statement to obtaining a retrieval result is realized, achieving the effects of improving the semantic recall ability of search and recall and realizing cold start in the search and recall process. Through representation learning, vectors corresponding to entities can be obtained, and entities in the query statement can be accurately matched through entity linking. Therefore, compared with the prior art, the present invention can solve the problem of inaccurate recall caused by word segmentation in the literal matching method, realize effective recall of documents in the scenarios of "one word with multiple meanings" and "multiple words with the same meaning", and at the same time reduce the dependence of the training model on the corpus during the recall process, and improve the interpretability during the recall process on the basis of ensuring the response time of recall. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the technical solutions of the present invention, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0063] Figure 1 is a schematic diagram of the workflow of a search and recall method based on knowledge graph representation learning provided by the present application;

[0064] Figure 2 is a schematic diagram of the workflow in the online stage of a search and recall method based on knowledge graph representation learning provided by the present application;

[0065] Figure 3 is a schematic diagram of the offline stage of a search and recall method based on knowledge graph representation learning provided by the present application;

[0066] Figure 4 is a schematic diagram of the workflow of establishing an index library corresponding to the target document in a search and recall method based on knowledge graph representation learning provided by the present application;

[0067] Figure 5 is a schematic diagram of the workflow of matching entities in the domain knowledge graph in a search and recall method based on knowledge graph representation learning provided by the present application;

[0068] Figure 6It is a schematic diagram of the workflow for obtaining accurate vectors and candidate similar vectors in a search recall method based on knowledge graph representation learning provided by this application;

[0069] Figure 7 It is a schematic diagram of the workflow for calculating vector similarity in a search recall method based on knowledge graph representation learning provided by this application;

[0070] Figure 8 It is a schematic diagram of the application process of a search recall method based on knowledge graph representation learning provided by this application;

[0071] Figure 9 It is a schematic diagram of the application process of a search recall method based on knowledge graph representation learning provided by this application in the field of power failure maintenance;

[0072] Figure 10 It is a block diagram of the structure of a search recall system based on knowledge graph representation learning provided by this application. Specific Embodiments

[0073] Here, exemplary embodiments will be described in detail, and their examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of methods consistent with some aspects of the present invention as detailed in the appended claims.

[0074] The following elaborates in detail on a search recall method and system based on knowledge graph representation learning provided by this application through specific embodiments.

[0075] An embodiment of the present invention discloses a search recall method and system based on knowledge graph representation learning, and this solution is applied to the retrieval of professional field documents. Specifically, different from ordinary documents, the professional field documents described in this application are documents related to a certain professional field, which contain proprietary nouns and definitions in this professional field, and the nouns and definitions involved in documents in different professional fields are often different. For example, professional fields such as the medical field, the power field, the financial field, and the legal field can all be distinguished as different professional field documents.

[0076] For example, referring to Figure 1 , it shows a schematic diagram of the workflow of a search recall based on knowledge graph representation learning provided by an embodiment of the present invention. Combining Figure 1 it can be seen that in the process of search recall based on knowledge graph representation learning, it can include an offline stage and an online stage. Among them, before the offline stage, it includes: S101, obtaining a domain knowledge graph and target documents.

[0077] The domain knowledge graph and the target document belong to the same professional field. It should be noted that before retrieving for different professional fields, a relatively complete domain knowledge graph needs to be obtained. The domain knowledge graph is the same as the professional field of the content to be queried and the target document. This application does not provide specific methods for building the domain knowledge graph, that is, it does not limit the method of obtaining the domain knowledge graph, but only takes obtaining a complete domain knowledge graph as a prerequisite for this application. In this domain knowledge graph, necessary contents such as entities, attributes, and relationships corresponding to the professional field are included. In addition, the target document is also the basis for realizing the retrieval. When the user inputs a query statement, a search recall method based on knowledge graph representation learning provided by this application is adopted to screen and obtain search results in the target document.

[0078] S102, in the offline stage, through representation learning based on the knowledge graph, obtain the vectors corresponding to the entities in the domain knowledge graph, and establish an index library corresponding to the target document according to the vectors. The index library includes a word-vector library, a vector index library, and an inverted index library;

[0079] The word-vector library is used to store the feature vector representations corresponding to the entities in the domain knowledge graph and store the mapping relationship table between the entities and the feature vector representations. The vector index library is used to provide indexes of the feature vector representations corresponding to the entities in the domain knowledge graph. The inverted index library is used to store all entities in the domain knowledge graph;

[0080] In this embodiment, obtaining the corresponding vector representation through the knowledge graph representation learning method can effectively enhance the capture of knowledge features. Specifically, the knowledge graph-based representation learning can adopt: the Trans series models based on translation, tensor decomposition models, neural network-based models, and graph neural network models. In the offline step, any knowledge graph-based representation learning method can be used to obtain the vectors corresponding to the entities in the domain knowledge graph. In the specific implementation process, the knowledge graph-based representation learning method can be selected based on the data characteristics through experiments.

[0081] S103, in the online stage, complete the user input query statement and obtain the query result; See Figure 2 , which shows the workflow schematic diagram of the online stage in a search recall method based on knowledge graph representation learning provided by an embodiment of the present invention, including:

[0082] S201, according to the input query statement, match the entities in the domain knowledge graph;

[0083] S202. Query through the index library based on the obtained entity by matching to obtain the accurate vector and candidate similar vectors.

[0084] S203. Calculate the vector similarity based on the accurate vector and candidate similar vectors to obtain a list of similar vectors.

[0085] S204. Query vectors in the vector index library according to the list of similar vectors to obtain semantic similar words of the entity corresponding to the query statement.

[0086] S205. Obtain a list of candidate documents through the inverted index library according to the semantic similar words and return the search results.

[0087] See Figure 3 , which shows a schematic diagram of the workflow in the offline stage of a search and recall method based on knowledge graph representation learning provided by an embodiment of the present invention. The offline stage includes:

[0088] S301. Input the domain knowledge graph into a graph convolutional network, perform feature selection through the graph convolutional network, and output the first intermediate feature vector representation of the entity.

[0089] S302. According to the first intermediate feature vector representation, obtain the second intermediate feature vector representation of the entity through activation function transformation and filtering (dropout).

[0090] S303. Perform multiple graph convolutions, activation function transformations and filtering on the second intermediate feature vector representation to obtain the final feature vector representation of the entity.

[0091] S304. Save the final feature vector representation of the entity in the word-vector library and the vector index library.

[0092] See Figure 4 , which shows a schematic diagram of the workflow for establishing an index library corresponding to a target document based on vectors in a search and recall method based on knowledge graph representation learning provided by an embodiment of the present invention, including:

[0093] S401. Perform entity linking between the target document and the domain knowledge graph, that is, match the entities in the target document with the entities in the domain knowledge graph to obtain a list of entities linked to each target document.

[0094] S402. Construct an index for each target document and the list of entities linked to the target document through the inverted index technology, that is, establish the inverted index library.

[0095] S403. Establish the vector index library according to the vectors corresponding to the entities linked to all the target documents.

[0096] Specifically, the entity linking (EL) in this embodiment refers to the process of unambiguously and correctly pointing the identified entities in the target document to the target entities in the domain knowledge graph. That is, in the domain knowledge graph, according to the meaning of the entities in the target document, the most suitable target items are found. If there are corresponding entities in the domain knowledge graph, those entities are returned; if there are no such entities in the domain knowledge graph, the entities are marked as the null value NIL. By performing entity linking between the target document and the domain knowledge graph, it is a prerequisite for accurately matching the entities in the query statement in the online stage. In the subsequent online process, the entities in the query statement can be accurately matched, effectively realizing the semantic matching problems in the scenarios of "one word with multiple meanings" and "multiple words with the same meaning", and improving the accuracy and coverage rate of the matching between the search results and the query statement.

[0097] In this embodiment, the word-vector library, inverted index library, and vector index library can be established in the offline stage, enabling querying for corresponding vectors according to entities, or querying for corresponding entities through vectors, and also enabling querying for the corresponding target document through the entities linked to the target document.

[0098] See Figure 5 , which shows a schematic workflow diagram of matching entities in the domain knowledge graph in a search recall method based on knowledge graph representation learning provided by an embodiment of the present invention, including:

[0099] S501, obtain the input query statement;

[0100] S502, perform entity linking between the query statement and the domain knowledge graph, that is, obtain the entities in the query statement, and match the entities in the query statement with the entities in the domain knowledge graph. In this embodiment, similar to the entity linking step in the offline stage, after performing entity linking between the query statement and the domain knowledge graph on the basis of the entity linking between the target document and the domain knowledge graph in the offline stage, the semantic matching problems in the scenarios of "one word with multiple meanings" and "multiple words with the same meaning" are finally effectively realized, and the accuracy and coverage rate of the matching between the search results and the query statement in the online stage are improved.

[0101] See Figure 6 , which shows a schematic workflow diagram of obtaining accurate vectors and candidate similar vectors in a search recall method based on knowledge graph representation learning provided by an embodiment of the present invention, including:

[0102] S601, query for accurate vectors through the word-vector library in the index library according to the entities obtained by the matching;

[0103] S602. Query for candidate similar vectors through the vector index library in the index library according to the entities obtained by the matching. The candidate similar vectors are a set of similar vectors.

[0104] See Figure 7 , which shows a schematic workflow diagram of vector similarity calculation in a search and recall method based on knowledge graph representation learning provided by an embodiment of the present invention. The calculation includes:

[0105] S701. Determine the number of accurate vectors.

[0106] S702. If the number of accurate vectors is one, calculate the similarity between each candidate similar vector and the accurate vector through the cosine similarity algorithm.

[0107] S703. If the number of accurate vectors is two or more, use weighted average to calculate the similarity between each candidate similar vector and each accurate vector; wherein, the weight of each accurate vector adopts the weight of the entity corresponding to the accurate vector.

[0108] In this step, the weight of the entity corresponding to the accurate vector can be obtained by iteratively calculating using the PageRank algorithm in the domain knowledge graph until the calculation result converges; the PageRank algorithm can obtain the importance degree of each node (entity) in the domain knowledge graph. In the actual calculation process of this step, a comparison result between the current iteration and the previous iteration, that is, the change rate, will be obtained. When the change rate drops to a certain extent, it can converge; specifically, the advisable empirical value of the number of iterations is 10; in addition, in addition to the PageRank algorithm, the centrality analysis algorithm can also be used for calculation.

[0109] S704. Arrange the similar vector list in descending order according to the similarity score and return the similar vector list as the search result.

[0110] See Figure 8 , which is a schematic application process diagram of a search and recall method based on knowledge graph representation learning provided by this application. See Figure 9 , which is a schematic application process diagram of a search and recall method based on knowledge graph representation learning provided by this application in the power field. Combining Figure 8 and Figure 9 , this application provides an embodiment of using this method to search and recall professional documents in the power fault repair field:

[0111] First, obtain the power fault repair knowledge graph as the domain knowledge graph and obtain the power fault document as the target document.

[0112] In the offline stage, the implementation steps are as follows:

[0113] 1) Based on the power failure repair knowledge graph, perform representation learning of the entities therein based on the knowledge graph to obtain a vector library. The representation learning method based on the knowledge graph can select the LINE algorithm based on the idea of random walk, or the graph convolutional network can also be used;

[0114] 2) Perform entity linking between the fault report document and the power failure repair knowledge graph. Specifically, the end-to-end learning method can be selected for entity linking; then combine the linked entities with the document to construct an inverted index of the document and build an inverted index library; assume that the document is related to "the film of the pressure relief valve is damaged", then "pressure relief valve" may be described as "gray scale relief valve" or "gray scale vacuum relief valve" in the document, and both can be recognized through entity linking, thus solving the problem of "multiple words with one meaning" in the traditional method; at the same time, "film" has many meanings, it can be a part of a device or component, or a material for fault repair. Here, according to the context, it can be judged that it is the first meaning, so the problem of "one word with multiple meanings" in the traditional method can be solved.

[0115] 3) Establish a vector index library according to the vectors corresponding to all the entities linked in all the fault report documents.

[0116] Next, take a query process as an example to show the implementation process of the online stage. The overall process is as Figure 9 shown, where the italic part is the actual query statement:

[0117] 1) Entity linking: For the input query statement "the film of the pressure relief valve is damaged", the result of entity linking is "pressure relief valve, film damage". Compared with the method of literal recall, the usual word segmentation method will divide the input into "pressure, release, film, damage", which obviously loses the reference meaning of a device deployment in the original input;

[0118] 2) Vector query and retrieval: Search the vector index library with the vectors corresponding to "pressure relief valve" and "film damage" to obtain their respective corresponding candidate vector lists;

[0119] 3) Fine calculation of vector similarity: Calculate the two candidate vector lists with "pressure relief valve" and "film damage" respectively. Here, the weights of the two words are the same, so the scores can be directly added. Finally, merge the two score lists, and finally select a certain number of vectors with the highest similarity as similar vectors;

[0120] 4) Document retrieval: Perform a vector query on the final list of similar vectors to obtain corresponding semantically similar words. The first five of these similar words are respectively "pressure relief valve, ash bunker relief valve, ash bunker vacuum relief valve, film breakage, film damage, film impairment". By observing this result, it can be found that the problems of "multiple words with the same meaning" and "one word with multiple meanings" solved by entity linking in the offline stage are effectively resolved. At the same time, the final query result can also be explained. Finally, search for documents in the document inverted index library using the list of semantically similar words to obtain a candidate document list as the search result for return.

[0121] In the prior art, traditional recall methods based on literal (word or keyword) matching and semantic recall methods based on vectors have difficulties in using deep semantic information, poor interpretability, and it is difficult to balance the response speed of the system and effective cold start in the document search and recall scenarios for professional fields. By using the foregoing method, through representation learning based on a knowledge graph in the offline stage, vectors corresponding to entities in the domain knowledge graph are obtained, and an index library corresponding to the target document is established based on these vectors. In the online stage, through entity linking, query and retrieval of vectors, fine calculation of vector similarity, and document retrieval, the entire process from the user inputting a query statement to obtaining a retrieval result is realized, achieving the effect of improving the semantic recall ability of search and recall and realizing cold start in the search and recall process. Through representation learning, vectors corresponding to entities can be obtained, and entities in the query statement can be accurately matched through entity linking. Therefore, compared with the prior art, the present invention can solve the problem of inaccurate recall caused by word segmentation in the literal matching method, realize effective recall of documents in the scenarios of "one word with multiple meanings" and "multiple words with the same meaning", and at the same time reduce the dependence of the training model on the corpus during the recall process, improving the interpretability during the recall process on the basis of ensuring the response time of the recall.

[0122] Each method embodiment described in this article can be an independent solution or can be combined according to internal logic, and these solutions all fall within the protection scope of this application.

[0123] It can be understood that in the above-mentioned various method embodiments, the methods and operations implemented by the network device can also be implemented by components (such as chips or circuits) available for the network device.

[0124] The above embodiments have introduced the search and recall method based on knowledge graph representation learning provided by the present application. It can be understood that in order for a network device to implement the above functions, it includes the corresponding hardware structure and / or software module for each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0125] The embodiments of the present application can divide the functional modules of the network device according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation.

[0126] Above, in combination with Figures 1 to 9 The method provided by the embodiments of the present application has been described in detail. Below, in combination with Figure 10 The system provided by the embodiments of the present application will be described in detail. It should be understood that the description of the system embodiments corresponds to the description of the method embodiments. Therefore, the content not described in detail can refer to the above method embodiments. For the sake of brevity, it will not be repeated here.

[0127] This embodiment also provides a search and recall system 900 based on knowledge graph representation learning. The system includes:

[0128] A preparation module 901, configured to obtain a domain knowledge graph, where the domain knowledge graph belongs to the same professional field as the target document;

[0129] An offline module 902, configured to obtain vectors corresponding to entities in the domain knowledge graph through representation learning based on the knowledge graph, and establish an index library corresponding to the target document according to the vectors. The index library includes a word-vector library, a vector index library, and an inverted index library;

[0130] The word-vector library is used to store the feature vectors corresponding to entities in the domain knowledge graph and store the mapping relationship table between entities and feature vectors. The vector index library is used to provide indexes of the feature vectors corresponding to entities in the domain knowledge graph. The inverted index library is used to store all entities in the domain knowledge graph;

[0131] An online module 903 for matching entities in the domain knowledge graph according to an input query statement;

[0132] Query through the index library according to the entities obtained by matching to obtain an accurate vector and candidate similar vectors;

[0133] Calculate the vector similarity according to the accurate vector and candidate similar vectors to obtain a list of similar vectors;

[0134] Query vectors in the vector index library according to the list of similar vectors to obtain semantic similar words of the entity corresponding to the query statement;

[0135] Obtain a list of candidate documents according to the semantic similar words through the inverted index library and return the search results.

[0136] This embodiment also provides a search recall program based on knowledge graph representation learning. The program used is used to implement the steps of the search recall method based on knowledge graph representation learning when executed.

[0137] This embodiment also provides a computer-readable storage medium, on which computer instructions are stored. When the instructions are executed by a processor, the steps of the search recall method based on knowledge graph representation learning are implemented.

[0138] Those of ordinary skill in the art can realize that the various illustrative logical blocks and steps described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0139] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described in detail here.

[0140] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0141] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0142] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0143] The search and recall system based on knowledge graph representation learning provided by the embodiments of the present application is used to execute the methods provided above. Therefore, the beneficial effects it can achieve can refer to the beneficial effects corresponding to the methods provided above, and will not be elaborated here.

[0144] It should be understood that in the various embodiments of the present application, the execution order of each step should be determined according to its function and internal logic. The size of each step number does not mean the sequence of execution, and does not limit the implementation process of the embodiment.

[0145] Each part of this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the search and recall system based on knowledge graph representation learning, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the description in the method embodiments for the relevant parts.

[0146] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present application.

[0147] The embodiments of the present application described above do not constitute a limitation on the protection scope of the present application.

Claims

1. A search recall method based on knowledge graph representation learning, characterized in that, Including: Obtain a domain knowledge graph and a target document, where the domain knowledge graph and the target document belong to the same professional field; Offline stage: Through representation learning based on the knowledge graph, obtain the vectors corresponding to the entities in the domain knowledge graph, and establish an index library corresponding to the target document according to the vectors. The index library includes a word-vector library, a vector index library, and an inverted index library; The word-vector library is used to store the feature vector representations corresponding to the entities in the domain knowledge graph and store the mapping relationship table between the entities and the feature vector representations. The vector index library is used to provide indexes of the feature vector representations corresponding to the entities in the domain knowledge graph. The inverted index library is used to store all the entities in the domain knowledge graph; Online stage: According to the input query statement, match the entities in the domain knowledge graph; According to the entities obtained by matching, query through the index library to obtain accurate vectors and candidate similar vectors; Calculate the vector similarity according to the accurate vectors and the candidate similar vectors to obtain a list of similar vectors; According to the list of similar vectors, perform vector query in the vector index library to obtain the semantic similar words of the entity corresponding to the query statement; According to the semantic similar words, obtain a list of candidate documents through the inverted index library and return the search results.

2. The search recall method based on knowledge graph representation learning according to claim 1, characterized in that, The offline stage includes: Input the domain knowledge graph into a graph convolutional network, perform feature selection through the graph convolutional network, and output the first intermediate feature vector representation of the entity; According to the first intermediate feature vector representation, perform transformation and filtering through an activation function to obtain the second intermediate feature vector representation of the entity; Perform multiple graph convolutions, activation function transformation and filtering on the second intermediate feature vector representation to obtain the final feature vector representation of the entity; Save the final feature vector representation of the entity in the word-vector library and the vector index library.

3. The search recall method based on knowledge graph representation learning according to claim 2, characterized in that, The establishment of the index library corresponding to the target document according to the vectors includes: Perform entity linking between the target document and the domain knowledge graph, that is, match the entities in the target document with the entities in the domain knowledge graph to obtain a list of entities linked by each target document; Construct an index for each target document and the list of entities linked by the target document through the inverted index technology, that is, establish the inverted index library; Establish the vector index library according to the vectors corresponding to the entities linked by all the target documents.

4. The search recall method based on knowledge graph representation learning according to claim 1, characterized in that, The matching of the entities in the domain knowledge graph according to the input query statement includes: Obtain the input query statement; Perform entity linking between the query statement and the domain knowledge graph, that is, obtain the entities in the query statement, and match the entities in the query statement with the entities in the domain knowledge graph.

5. The search recall method based on knowledge graph representation learning according to claim 1, characterized in that, The query through the index library to obtain accurate vectors and candidate similar vectors according to the entities obtained by matching includes: According to the entities obtained by matching, query through the word-vector library in the index library to obtain accurate vectors; Entities obtained according to the matching are queried through the vector index library in the index library to obtain candidate similar vectors, and the candidate similar vectors are a set of similar vectors.

6. The search recall method based on knowledge graph representation learning according to claim 1, characterized in that, Performing vector similarity calculation based on the accurate vector and the candidate similar vectors to obtain a list of similar vectors, including: Determine the number of the accurate vectors; If the number of the accurate vectors is one, calculate the similarity between each candidate similar vector and the accurate vector through the cosine similarity algorithm; If the number of the accurate vectors is two or more, use weighted average to calculate the similarity between each candidate similar vector and each accurate vector; wherein, the weight of each accurate vector adopts the weight of the entity corresponding to the accurate vector; Arrange the list of similar vectors in descending order of the similarity scores, and return the list of similar vectors as the search result.

7. A search and recall system based on knowledge graph representation learning, characterized in that, A preparation module for obtaining a domain knowledge graph, where the domain knowledge graph belongs to the same professional field as the target document; An offline module for obtaining vectors corresponding to entities in the domain knowledge graph through representation learning based on the knowledge graph, and establishing an index library corresponding to the target document according to the vectors, where the index library includes a word-vector library, a vector index library, and an inverted index library; The word-vector library is used to store the feature vectors corresponding to entities in the domain knowledge graph and store the mapping relationship table between entities and feature vectors. The vector index library is used to provide indexes of the feature vectors corresponding to entities in the domain knowledge graph. The inverted index library is used to store all entities in the domain knowledge graph; An online module for matching entities in the domain knowledge graph according to the input query statement; According to the entities obtained by the matching, query through the index library to obtain accurate vectors and candidate similar vectors; Performing vector similarity calculation based on the accurate vectors and the candidate similar vectors to obtain a list of similar vectors; According to the list of similar vectors, perform vector query in the vector index library to obtain semantic similar words of the entity corresponding to the query statement; According to the semantic similar words, obtain a list of candidate documents through the inverted index library and return the search result.

8. A search and recall program based on knowledge graph representation learning, and the program is used to implement the steps of a search and recall method based on knowledge graph representation learning described in any one of claims 1 to 6 when executed.

9. A computer-readable storage medium, on which computer instructions are stored, and the instructions implement the steps of a search and recall method based on knowledge graph representation learning described in any one of claims 1 to 6 when executed by a processor.

10. A search and recall processing device, characterized in that, The search recall processing device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a search recall method based on knowledge graph representation learning according to any one of claims 1 to 6 above.

Citation Information

Patent Citations

  • Mapping knowledge domain questioning and answering system and method based on template matching technique

    CN105868313A

  • Question and answer knowledge retrieval method, device and equipment based on graph embedding

    CN113947084A