Knowledge Graph-based Document Retrieval Method and Related Devices

Through the document search method based on the knowledge graph, the problem of low document search efficiency in the existing technology is solved, efficient and accurate document search is achieved, and machine learning costs are reduced.

CN114780746BActive Publication Date: 2025-06-17华润数字科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210428422.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-22
Publication Date
2025-06-17
Estimated Expiration
2042-04-22

AI Technical Summary

Technical Problem

The prior art is less efficient in document retrieval, traditional keyword retrieval is difficult to provide accurate results, while neural network-based methods are more expensive to train samples.

Method used

The document search method based on knowledge graph is adopted, and word segmentation is processed by obtaining the collection of documents to be retrieved, the target knowledge graph is constructed, the semantic distance is calculated when the search keyword is received, the maximum semantic subgraph is constructed, and the feature vector is extracted using graph convolution neural network, and the similarity of the topic embedded vector is calculated to determine the target search document.

Benefits of technology

It improves the efficiency and accuracy of document retrieval, reduces machine learning costs, and reduces the interference of irrelevant documents on search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114780746B_ABST
    Figure CN114780746B_ABST
Patent Text Reader

Abstract

The embodiments of the present application belong to the field of artificial intelligence, and relate to a document retrieval method based on a knowledge graph and related devices, including performing word segmentation on documents in a document set to be retrieved to obtain a document word segmentation set, and constructing a target knowledge graph; calculating the semantic distance between retrieval keywords based on the target knowledge graph, and determining the retrieval keyword corresponding to the maximum semantic distance as the central keyword; constructing a first subgraph and a second subgraph based on the central keyword, calculating the number of nodes in the first subgraph and the second subgraph, and selecting the maximum semantic subgraph according to the number of nodes; performing feature extraction on the maximum semantic subgraph based on a graph convolutional neural network to obtain a feature vector; extracting the topic words of the document and calculating the topic embedding vector of the topic words; calculating the vector similarity between the topic embedding vector and the feature vector, and determining the document corresponding to the topic embedding vector with the vector similarity greater than or equal to a preset similarity threshold as the target retrieval document. The present application realizes the efficient screening of the target retrieval document.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a document retrieval method based on a knowledge graph and related devices. Background Art

[0002] In archive management, it is often necessary to retrieve effective information from a vast number of documents. Traditional document retrieval is keyword-based, but even when retrieving by keywords, it is difficult for users to provide the most accurate keywords. Currently, information retrieval methods mainly include the TF-IDF retrieval method based on traditional search engines and neural network models based on supervised learning.

[0003] However, the retrieval method that combines traditional search engines with TF-IDF often has the problem of low recall rate due to an excessive number of candidate answers obtained, and other technologies are needed to further screen the answers. The retrieval method of the neural network model based on supervised learning locates the starting position of the answer in the document by learning the semantic association between the question and the answer, and the cost of collecting and organizing the training samples of this type of method is relatively high. Therefore, how to improve the document retrieval efficiency while reducing the machine learning cost is an urgent problem to be solved. Summary of the Invention

[0004] The purpose of the embodiments of this application is to propose a document retrieval method based on a knowledge graph and related devices to solve the technical problem of low document retrieval efficiency.

[0005] To solve the above technical problem, the embodiments of this application provide a document retrieval method based on a knowledge graph, and adopt the following technical solutions:

[0006] Obtain a set of documents to be retrieved, perform word segmentation processing on the documents in the set of documents to be retrieved to obtain a set of document word segments, and construct a target knowledge graph based on the set of document word segments;

[0007] When receiving multiple retrieval keywords, calculate the semantic distance between the retrieval keywords based on the target knowledge graph, and determine the two retrieval keywords corresponding to the maximum semantic distance as two central keywords;

[0008] Construct a first subgraph and a second subgraph based on the two central keywords respectively, calculate the number of nodes in the first subgraph and the second subgraph respectively, and select the maximum semantic subgraph from the first subgraph and the second subgraph according to the number of nodes;

[0009] Obtain a preset graph convolutional neural network, and perform feature extraction on the maximum semantic subgraph based on the graph convolutional neural network to obtain a feature vector;

[0010] Extract the subject terms of each document in the to-be-retrieved document set, and calculate the subject embedding vectors of the subject terms;

[0011] Calculate the vector similarity between each of the subject embedding vectors and the feature vector, determine the subject embedding vectors with the vector similarity greater than or equal to the preset similarity threshold as the target embedding vectors, and use the documents corresponding to the target embedding vectors as the target retrieval documents.

[0012] Further, the step of calculating the semantic distance between the retrieval keywords based on the target knowledge graph includes:

[0013] Obtain the reference knowledge graph corresponding to the target knowledge graph, and determine the distance weights between all the retrieval keywords according to the reference knowledge graph;

[0014] Calculate the embedding similarity between the retrieval keywords and the sum of the embedding vectors of the corresponding edge nodes of each retrieval keyword in the target knowledge graph, and calculate the semantic distance between the retrieval keywords according to the distance weights, the embedding similarity, and the sum of the embedding vectors.

[0015] Further, the step of determining the distance weights between the retrieval keywords according to the reference knowledge graph includes:

[0016] Obtain the category attributes and levels of the retrieval keywords in the reference knowledge graph, and determine the distance weights between the retrieval keywords according to the category attributes and the levels.

[0017] Further, the step of determining the distance weights between the retrieval keywords according to the category attributes and the levels includes:

[0018] Judge whether the category attributes between each pair of the retrieval keywords are the same, and whether the levels between each pair of the retrieval keywords are the same. When the category attributes of the retrieval keywords are the same and the levels are the same, determine the distance weight between the retrieval keywords as the preset weight;

[0019] When the category attributes are different, or the category attributes are the same but the levels are different, obtain the common superior entity of the retrieval keywords in the reference knowledge graph, calculate the hierarchical distance between the superior entity and the retrieval keywords, and calculate the distance weight between the retrieval keywords according to the hierarchical distance.

[0020] Further, the step of extracting the subject terms of each document in the to-be-retrieved document set includes:

[0021] Obtain the number of words of each document, sort the documents in ascending order according to the number of words to obtain a document queue;

[0022] Obtain the number of topic words corresponding to the document with the lowest order in the document queue, use the number of topic words corresponding to the document with the lowest order as the lowest threshold, and based on the lowest threshold, incrementally process the number of topic words of other documents in the document queue in the arrangement order of the document queue until the number of topic words reaches a preset maximum threshold;

[0023] Extract the topic words of the documents in the document queue in sequence according to the order and quantity from the lowest threshold to the maximum threshold.

[0024] Further, the step of extracting feature vectors from the maximum semantic subgraph based on the graph convolutional neural network includes:

[0025] Calculate the adjacency matrix and in-out degree matrix of the maximum semantic subgraph;

[0026] Obtain a preset weight matrix, and calculate the feature vectors through the graph convolutional neural network according to the weight matrix, the adjacency matrix, and the in-out degree matrix.

[0027] Further, before the step of calculating the semantic distance between the retrieval keywords based on the target knowledge graph, it further includes:

[0028] Retrieve the target knowledge graph to determine whether the retrieval keywords exist in the target knowledge graph;

[0029] When the retrieval keywords do not exist in the target knowledge graph, obtain a preset pre-trained language model, input the retrieval keywords and the word segments in the document word segment set into the pre-trained language model respectively, and calculate to obtain a first representation vector and a second representation vector;

[0030] Calculate the word similarity between the retrieval keywords and the word segments according to the first representation vector and the second representation vector, determine the word segments with the word similarity greater than or equal to the preset similarity as candidate keywords, and replace the retrieval keywords with the candidate keywords.

[0031] To solve the above technical problems, an embodiment of the present application further provides a document retrieval device based on a knowledge graph, which adopts the following technical solutions:

[0032] A construction module, configured to obtain a set of documents to be retrieved, perform word segmentation processing on the documents in the set of documents to be retrieved to obtain a document word segment set, and construct a target knowledge graph based on the document word segment set;

[0033] The first calculation module is configured to, when receiving multiple retrieval keywords, calculate the semantic distance between the retrieval keywords based on the target knowledge graph, and determine two retrieval keywords corresponding to the maximum semantic distance as two central keywords;

[0034] The selection module is configured to respectively construct a first sub-graph and a second sub-graph based on the two central keywords, respectively calculate the number of nodes in the first sub-graph and the second sub-graph, and select the maximum semantic sub-graph from the first sub-graph and the second sub-graph according to the number of nodes;

[0035] The second calculation module is configured to obtain a preset graph convolutional neural network, and perform feature extraction on the maximum semantic sub-graph based on the graph convolutional neural network to obtain a feature vector;

[0036] The extraction module is configured to extract the topic words of each document in the to-be-retrieved document set, and calculate the topic embedding vectors of the topic words;

[0037] The confirmation module is configured to calculate the vector similarity between each topic embedding vector and the feature vector, determine the topic embedding vectors whose vector similarity is greater than or equal to a preset similarity threshold as target embedding vectors, and use the documents corresponding to the target embedding vectors as target retrieval documents.

[0038] To solve the above technical problems, an embodiment of the present application further provides a computer device, which adopts the following technical solutions:

[0039] Obtain a to-be-retrieved document set, perform word segmentation processing on the documents in the to-be-retrieved document set to obtain a document word segmentation set, and construct a target knowledge graph based on the document word segmentation set;

[0040] When receiving multiple retrieval keywords, calculate the semantic distance between the retrieval keywords based on the target knowledge graph, and determine two retrieval keywords corresponding to the maximum semantic distance as two central keywords;

[0041] Respectively construct a first sub-graph and a second sub-graph based on the two central keywords, respectively calculate the number of nodes in the first sub-graph and the second sub-graph, and select the maximum semantic sub-graph from the first sub-graph and the second sub-graph according to the number of nodes;

[0042] Obtain a preset graph convolutional neural network, and perform feature extraction on the maximum semantic sub-graph based on the graph convolutional neural network to obtain a feature vector;

[0043] Extract the topic words of each document in the to-be-retrieved document set, and calculate the topic embedding vectors of the topic words;

[0044] Calculate the vector similarity between each of the subject embedding vectors and the feature vector, determine the subject embedding vectors with the vector similarity greater than or equal to a preset similarity threshold as the target embedding vectors, and use the documents corresponding to the target embedding vectors as the target retrieval documents.

[0045] To solve the above technical problems, an embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solutions:

[0046] Obtain a set of documents to be retrieved, perform word segmentation processing on the documents in the set of documents to be retrieved to obtain a set of document word segments, and construct a target knowledge graph based on the set of document word segments;

[0047] When receiving multiple retrieval keywords, calculate the semantic distance between the retrieval keywords based on the target knowledge graph, and determine the two retrieval keywords corresponding to the maximum semantic distance as the two central keywords;

[0048] Construct a first subgraph and a second subgraph respectively based on the two central keywords, calculate the number of nodes in the first subgraph and the second subgraph respectively, and select the maximum semantic subgraph from the first subgraph and the second subgraph according to the number of nodes;

[0049] Obtain a preset graph convolutional neural network, and perform feature extraction on the maximum semantic subgraph based on the graph convolutional neural network to obtain a feature vector;

[0050] Extract the subject words of each document in the set of documents to be retrieved, and calculate the subject embedding vectors of the subject words;

[0051] Calculate the vector similarity between each of the subject embedding vectors and the feature vector, determine the subject embedding vectors with the vector similarity greater than or equal to a preset similarity threshold as the target embedding vectors, and use the documents corresponding to the target embedding vectors as the target retrieval documents.

[0052] In this application, by obtaining a collection of documents to be retrieved, performing word segmentation on the documents in the collection of documents to be retrieved to obtain a collection of document word segments, constructing a target knowledge graph based on the collection of document word segments, the documents can be efficiently sorted and retrieved through this target knowledge graph; then, when receiving multiple retrieval keywords, calculating the semantic distance between the retrieval keywords based on the target knowledge graph, and determining the two retrieval keywords corresponding to the maximum semantic distance as two central keywords; constructing a first sub-graph and a second sub-graph respectively based on the two central keywords, calculating the number of nodes in the first sub-graph and the second sub-graph respectively, and selecting the maximum semantic sub-graph from the first sub-graph and the second sub-graph according to the number of nodes, thereby screening out the semantic sub-graph including more retrieval keywords, avoiding the interference of irrelevant retrieval keywords, and further improving the accuracy and efficiency of document retrieval; then, obtaining a preset graph convolutional neural network, performing feature extraction on the maximum semantic sub-graph based on the graph convolutional neural network to obtain a feature vector, thereby improving the accuracy of feature extraction; finally, extracting the topic words of each document in the collection of documents to be retrieved, and calculating the topic embedding vectors of the topic words; calculating the vector similarity between each topic embedding vector and the feature vector, and determining the topic embedding vectors with vector similarity greater than or equal to the preset similarity threshold as target embedding vectors, and taking the documents corresponding to the target embedding vectors as target retrieval documents, ultimately achieving the efficient screening of target retrieval documents, reducing the interference of irrelevant documents on the retrieval results, saving the machine learning cost, and improving the efficiency and accuracy of target document retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the solutions in this application, the following will briefly introduce the drawings required for the description of the embodiments of this application. Obviously, the drawings in the following description are some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0054] Figure 1 is an exemplary system architecture diagram to which this application can be applied;

[0055] Figure 2 is a flowchart of an embodiment of the method for document retrieval based on a knowledge graph according to this application;

[0056] Figure 3 is a schematic diagram of the maximum semantic sub-graph in an embodiment according to this application;

[0057] Figure 4 is a schematic structural diagram of an embodiment of the apparatus for document retrieval based on a knowledge graph according to this application;

[0058] Figure 5 is a schematic structural diagram of an embodiment of a computer device according to this application.

[0059] Reference numerals: Knowledge graph-based document retrieval device 400, construction module 401, first calculation module 402, selection module 403, second calculation module 404, extraction module 405, and confirmation module 406. Detailed implementation

[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.

[0061] Reference to "embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0062] In order to enable those skilled in the art of this technology to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0063] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0064] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0065] The terminal devices 101, 102, and 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers, desktop computers, and so on.

[0066] The server 105 can be a server that provides various services, such as a background server that supports the pages displayed on the terminal devices 101, 102, and 103.

[0067] It should be noted that the method for document retrieval based on a knowledge graph provided in the embodiments of the present application is generally executed by a server / terminal device. Correspondingly, the device for document retrieval based on a knowledge graph is generally set in the server / terminal device.

[0068] It should be understood that Figure 1 the numbers of the terminal devices, network, and server in

[0069] Continue to refer to Figure 2 , which shows a flowchart of an embodiment of the method for document retrieval based on a knowledge graph according to the present application. The method for document retrieval based on a knowledge graph includes the following steps:

[0070] Step S201: Obtain a set of documents to be retrieved, perform word segmentation on the documents in the set of documents to be retrieved to obtain a set of document word segments, and construct a target knowledge graph based on the set of document word segments.

[0071] In this embodiment, the set of documents to be retrieved is a set of pre-collected documents. Obtain the set of documents to be retrieved and perform word segmentation on the documents in the set of documents to be retrieved to obtain a set of document word segments. Specifically, a preset word segmentation tool can be used to perform word segmentation on the content of each document to obtain the word segments corresponding to different documents. The word segments corresponding to all the documents in the set of documents to be retrieved are aggregated to obtain a set of document word segments. When the set of document word segments is obtained, named entity recognition is used to extract entities and relationships from the word segments in the set of document word segments to obtain the entities corresponding to each word segment and the relationships between different entities; then, triples in the format of (entity 1, relationship, entity 2) are constructed according to the entities and relationships, and the triples are input into a graph database (such as neo4j), and a target knowledge graph is output based on the graph database.

[0072] Step S202: When multiple retrieval keywords are received, calculate the semantic distance between the retrieval keywords based on the target knowledge graph, and the two retrieval keywords corresponding to the maximum semantic distance are two central keywords.

[0073] In this embodiment, the retrieval keywords are the input document retrieval keywords, which can be obtained by segmenting the retrieval statement input by the user. When multiple retrieval keywords are received, calculate the semantic distance between different retrieval keywords based on the target knowledge graph, and screen the retrieval keywords based on this semantic distance. Specifically, this semantic distance is the semantic distance between two different retrieval keywords. When the retrieval keywords are obtained, match the retrieval keywords with the entity nodes in the target knowledge graph to determine the entity nodes matching the retrieval keywords; then, calculate the semantic distance between different entity nodes, and use the semantic distance between different entity nodes in the target knowledge graph as the semantic distance between the matching retrieval keywords. When determining the entity nodes matching the retrieval keywords, calculate the embedding vector of the entity nodes, which can be obtained by performing one-hot encoding on the retrieval keywords through the bag of words (BOW); then, calculate the cosine similarity between different embedding vectors according to the calculation formula of cosine similarity, and this cosine similarity is the semantic distance between the entity nodes. When the semantic distances between all retrieval keywords are obtained, the two retrieval keywords corresponding to the maximum semantic distance are used as the two central keywords.

[0074] Step S203: Based on the two central keywords, construct a first subgraph and a second subgraph respectively, calculate the number of nodes in the first subgraph and the second subgraph respectively, and select the maximum semantic subgraph from the first subgraph and the second subgraph according to the number of nodes.

[0075] In this embodiment, when determining the central keywords, construct a first subgraph and a second subgraph based on the central keywords. Among them, the central keywords are the two optimal retrieval keywords selected by the semantic distance. One central keyword constructs one subgraph, and the first subgraph and the second subgraph can be respectively constructed according to the two central keywords. Specifically, arbitrarily select one of the central keywords as the center of the circle, and use the maximum semantic distance as the radius to circle the first subgraph from the target knowledge graph; use the remaining central keyword as the center of the circle, and use the maximum semantic distance as the radius to circle the second subgraph from the target knowledge graph. The first subgraph and the second subgraph include different numbers of entity nodes. Obtain the number of nodes in the first subgraph and the second subgraph, and use the subgraph with the largest number of nodes in the first subgraph and the second subgraph as the maximum semantic subgraph. As Figure 3 shown, Figure 3Schematic diagram of the maximum semantic subgraph, where the semantic distance between kw1 and kw3 is the longest, and kw1 and kw3 are the two central keywords. Based on kw1, the first subgraph (i.e., sub Figure 1 ) is obtained, and based on kw2, the second subgraph (i.e., sub Figure 2 ) is obtained; the number of nodes in the second subgraph is greater than that of the first subgraph, and the second subgraph is the maximum semantic subgraph.

[0076] Step S204: Obtain a preset graph convolutional neural network, and based on the graph convolutional neural network, perform feature extraction on the maximum semantic subgraph to obtain a feature vector.

[0077] In this embodiment, when obtaining the maximum semantic subgraph, a preset graph convolutional neural network is obtained, and based on the graph convolutional neural network, feature extraction is performed on the maximum semantic subgraph to obtain a feature vector. Specifically, the graph convolutional neural network (GCN, Graph Convolutional Network) is a convolutional neural network used to process graph structures. Through this graph convolutional neural network, feature extraction can be performed on the input maximum semantic subgraph to obtain a feature vector. Specifically, when obtaining the maximum semantic subgraph, calculate the adjacency matrix and in-degree / out-degree matrix of the maximum semantic subgraph; then, according to the graph convolutional neural network, calculate the adjacency matrix, in-degree / out-degree matrix, and the embedding vectors of the nodes in the maximum semantic subgraph to obtain the Laplacian matrix eigenvector, and this Laplacian matrix eigenvector is the feature vector of the maximum semantic subgraph.

[0078] Step S205: Extract the topic words of each document in the to-be-retrieved document set, and calculate the topic embedding vectors of the topic words.

[0079] In this embodiment, the topic word is the core word of each document. The topic word can be a word whose occurrence times in the document are greater than or equal to the times threshold, or a word whose semantic weight is greater than or equal to the weight threshold, or a word whose total weight obtained by calculating the occurrence times and semantic weight in the document is greater than or equal to the total weight threshold. Extract the topic words of each document in the to-be-retrieved document set, and calculate the embedding vectors of the topic words of each document. The embedding vector of the topic word is the topic embedding vector. The topic embedding vector can be calculated through a preset pre-trained language model (bert). Specifically, when obtaining the topic word of the document, input the topic word into the pre-trained language model, and based on the embedding encoding of the pre-trained language model, obtain the embedding vector of each topic word.

[0080] Step S206: Calculate the vector similarity between each topic embedding vector and the feature vector, determine the topic embedding vectors whose vector similarity is greater than or equal to the preset similarity threshold as the target embedding vectors, and use the documents corresponding to the target embedding vectors as the target retrieval documents.

[0081] In this embodiment, based on the calculation formula of cosine similarity, the vector similarity between the topic embedding vector and the feature vector is calculated to determine whether the vector similarity is greater than or equal to a preset similarity threshold. If the vector similarity is greater than or equal to the preset similarity threshold, the topic embedding vector corresponding to the vector similarity is determined as the target embedding vector. If the vector similarity is less than the preset similarity threshold, the topic embedding vector corresponding to the vector similarity is determined as a non-target embedding vector. When the target embedding vector is obtained, the document corresponding to the target embedding vector is used as the target retrieval document.

[0082] This embodiment realizes the efficient screening of target retrieval documents, reduces the interference of irrelevant documents on the retrieval results, saves the machine learning cost, and improves the efficiency and accuracy of target document retrieval.

[0083] In some alternative implementation manners of this embodiment, the step of calculating the semantic distance between the retrieval keywords based on the target knowledge graph includes:

[0084] Obtain the reference knowledge graph corresponding to the target knowledge graph, and determine the distance weights between all the retrieval keywords according to the reference knowledge graph;

[0085] Calculate the embedding similarity between the retrieval keywords and the sum of the embedding vectors of the corresponding edge nodes of each retrieval keyword in the target knowledge graph, and calculate the semantic distance between the retrieval keywords according to the distance weights, the embedding similarity, and the sum of the embedding vectors.

[0086] In this embodiment, the reference knowledge graph is a graph constructed from data in the main field to which the retrieval object belongs. Taking a document as the retrieval object as an example, the reference knowledge graph can be a knowledge graph constructed based on CNKI data. Obtain the reference knowledge graph corresponding to the target knowledge graph, and determine the distance weights between all the retrieval keywords according to the reference knowledge graph. The distance weight is the distance weight value between the entities corresponding to the retrieval keywords in the reference knowledge graph, and the entity is the same as the retrieval keyword. If there is no entity in the reference knowledge graph that is the same as the retrieval keyword, select the entity with the highest similarity to the retrieval keyword as the entity corresponding to the retrieval keyword in the reference knowledge graph. Obtain the preset reference distance between two entities in the reference knowledge graph, and the distance weight between the retrieval keywords corresponding to the entity can be determined according to the preset reference distance.

[0087] When obtaining the distance weights between all retrieval keywords, calculate the embedding similarity between the retrieval keywords and the sum of the embedding vectors of the corresponding edge nodes of each retrieval keyword in the target knowledge graph; among them, the embedding similarity can be obtained by calculating the cosine similarity of the embedding vectors corresponding to different retrieval keywords; the sum of the embedding vectors is the sum of the embedding vectors of all edge nodes connected by the entity corresponding to the retrieval keyword in the target knowledge graph. When obtaining the embedding similarity and the sum of the embedding vectors, the semantic distance between the retrieval keywords can be calculated according to the distance weight, the embedding similarity, and the sum of the embedding vectors. The calculation formula of the semantic distance is as follows:

[0088]

[0089] Among them, is the distance weight, sim(E i ,E j ) is the embedding similarity, is the sum of the embedding vectors of entity E i , is the sum of the embedding vectors of entity E j .

[0090] In this embodiment, by obtaining the reference knowledge graph corresponding to the target knowledge graph, determining the weight distance between the retrieval keywords according to the reference knowledge graph, and calculating the semantic distance between the retrieval keywords according to the weight distance, the embedding similarity, and the sum of the embedding vectors, the accurate calculation of the semantic distance between the retrieval keywords is realized, so that the documents corresponding to the retrieval keywords can be further accurately retrieved through the semantic distance, and the accuracy of document retrieval is improved.

[0091] In some optional implementation manners of this embodiment, the step of determining the distance weight between the retrieval keywords according to the reference knowledge graph includes:

[0092] Obtain the category attributes and levels of the retrieval keywords in the reference knowledge graph, and determine the distance weight between the retrieval keywords according to the category attributes and the levels.

[0093] In this embodiment, the category attribute is the category of the entity corresponding to each retrieval keyword. For example, for the entities apple and banana, the category attributes of these two entities are both fruit categories; the hierarchy is the position level of the entity corresponding to each retrieval keyword in the reference knowledge graph. When constructing the reference knowledge graph, the position levels of each entity can be divided according to the relationships between different entities. Obtain the category attribute and hierarchy of each retrieval keyword in the reference knowledge graph, and determine the distance weight between the retrieval keywords according to the category attribute and hierarchy. Specifically, the distance weight is the distance measurement weight of the entity corresponding to the retrieval keyword in the reference knowledge graph. Through this distance weight, the semantic distance between the retrieval keywords can be accurately calculated. When obtaining the category attribute and hierarchy, obtain the corresponding distance weight according to the category attribute and hierarchy. The distance weight can be obtained by looking up the weight values corresponding to different category attributes and hierarchies in the weight preset table.

[0094] In this embodiment, by obtaining the category attribute and hierarchy of the retrieval keyword in the reference knowledge graph, and determining the distance weight between the retrieval keywords according to the category attribute and hierarchy, the semantic distance between the retrieval keywords can be accurately calculated through the distance weight, improving the accuracy of the semantic distance.

[0095] In some alternative implementation manners of this embodiment, the step of determining the distance weight between the retrieval keywords according to the category attribute and the hierarchy includes:

[0096] Judge whether the category attributes between each of the retrieval keywords are the same, and whether the hierarchies between each of the retrieval keywords are the same. When the category attributes of the retrieval keywords are the same and the hierarchies are the same, determine that the distance weight between the retrieval keywords is the preset weight;

[0097] When the category attributes are different, or the category attributes are the same and the hierarchies are different, obtain the common superior entity of the retrieval keywords in the reference knowledge graph, calculate the hierarchical distance between the superior entity and the retrieval keywords, and calculate the distance weight between the retrieval keywords according to the hierarchical distance.

[0098] In this embodiment, when obtaining the category attributes and levels of all retrieval keywords in the reference knowledge graph, it is determined whether the category attributes between each pair of retrieval keywords are the same. If the category attributes of two retrieval keywords are the same, it is determined whether their levels are the same, that is, it is determined whether the two retrieval keywords are at the same level in the reference knowledge graph; if the category attributes of two retrieval keywords are the same and their levels are also the same, it is determined that the weight distance between the two retrieval keywords is a preset weight, such as 1. If the category attributes of two retrieval keywords are different, or the category attributes are the same but the levels are different, the common parent entity of the two retrieval keywords in the reference knowledge graph is obtained; for example, for two retrieval keywords at different levels with the same category attribute, balcony and room, the common parent entity of the two retrieval keywords in the reference knowledge graph is building. The level distances between the parent entity and the two retrieval keywords are calculated; based on the level distances, the distance weights between the two retrieval keywords are calculated, and the calculation formula of the distance weights is as follows:

[0099]

[0100] where λ is a preset parameter, es is the common parent entity of the retrieval keywords in the reference knowledge graph, E i is entity i of the retrieval keyword in the reference knowledge graph, E j is entity j of the retrieval keyword in the reference knowledge graph, d(es, E i ) is the level distance from entity i to the parent entity, and d(es, E j ) is the level distance from entity j to the parent entity.

[0101] In this embodiment, by determining the category attributes and levels of the retrieval keywords and calculating the distance weights between the retrieval keywords differently according to the category attributes and levels, the calculation accuracy of the distance weights is further improved, so that the distance weights can more accurately reflect the semantic distances between different retrieval keywords.

[0102] In some optional implementation manners of this embodiment, the step of extracting the subject words of each document in the to-be-retrieved document set includes:

[0103] Obtain the number of words of each document, sort the documents in ascending order according to the number of words to obtain a document queue;

[0104] Obtain the number of subject words corresponding to the document with the lowest order in the document queue, use the number of subject words corresponding to the document with the lowest order as the lowest threshold, and based on the lowest threshold, incrementally increase the number of subject words of other documents in the document queue in the order of arrangement of the document queue until the number of subject words reaches a preset maximum threshold;

[0105] Extract the subject terms of the documents in the document queue in sequence according to the order and quantity from the lowest threshold to the highest threshold.

[0106] In this embodiment, when extracting the subject terms of each document in the document set to be retrieved, obtain the number of words of each document, where the number of words is the total number of words included in the document; sort the documents in the document set to be retrieved from low to high according to the number of words to obtain a document queue. Obtain the number of subject terms corresponding to the document with the lowest order in the document queue, and use the number of subject terms corresponding to the document with the lowest order as the lowest threshold; based on the lowest threshold, increment the number of subject terms of other documents in the document queue by an equal difference quantity in the arrangement order of the documents in the document queue until the number of subject terms reaches a preset maximum threshold. For example, the lowest threshold is TW min , and the number of subject terms of other documents in the document queue is incremented by 1 in sequence until the preset maximum threshold TW is reached. max . Then, extract the subject terms of the documents in the document queue in sequence according to the order and quantity from the lowest threshold to the highest threshold, and the extraction quantity of the subject terms of each document is the number of subject terms of the document in the document queue.

[0107] In this embodiment, the subject terms of the documents in the document set to be retrieved are extracted through the document queue, realizing the sequential extraction of the subject terms of the documents, and improving the extraction efficiency and accuracy of the subject terms of the documents.

[0108] In some alternative implementation manners of this embodiment, the step of extracting feature vectors from the maximum semantic subgraph based on the graph convolutional neural network includes:

[0109] Calculate the adjacency matrix and in-out degree matrix of the maximum semantic subgraph;

[0110] Obtain a preset weight matrix, and calculate the feature vector through the graph convolutional neural network according to the weight matrix, the adjacency matrix, and the in-out degree matrix.

[0111] In this embodiment, when obtaining the maximum semantic subgraph, calculate the adjacency matrix and in-out degree matrix of the maximum semantic subgraph, and then obtain a preset weight matrix. Calculate the feature vector through the graph convolutional neural network according to the weight matrix, the adjacency matrix, and the in-out degree matrix. The calculation formula of the feature vector is as follows:

[0112]

[0113] L (0) =X

[0114] Where Let \(A\) be the adjacency matrix, \(D\) be the in - out degree matrix, \(W_0\) be the preset weight matrix, and \(X\) be the embedding vector of the nodes in the maximum semantic sub - graph.

[0115] In this embodiment, by constructing a graph convolutional neural network, the features of the maximum semantic sub - graph can be accurately and efficiently extracted through this graph convolutional neural network, further improving the accuracy of feature extraction.

[0116] In some alternative implementation manners of this embodiment, before the step of calculating the semantic distance between the retrieval keywords based on the target knowledge graph, the following steps are further included:

[0117] Retrieve the target knowledge graph to determine whether the retrieval keywords exist in the target knowledge graph;

[0118] When the retrieval keywords do not exist in the target knowledge graph, obtain a preset pre - trained language model, input the retrieval keywords and the word segments in the document word segment set into the pre - trained language model respectively, and calculate the first representation vector and the second representation vector;

[0119] According to the first representation vector and the second representation vector, calculate the word similarity between the retrieval keywords and the word segments, determine the word segments with the word similarity greater than or equal to the preset similarity as candidate keywords, and replace the retrieval keywords with the candidate keywords.

[0120] In this embodiment, before calculating the semantic distance between the retrieval keywords based on the target knowledge graph, the target knowledge graph is retrieved to determine whether the retrieval keywords exist in the target knowledge graph. If the retrieval keywords do not exist in the target knowledge graph, a preset pre-trained language representation model (i.e., bert, Bidirectional Encoder Representation from Transformers) is obtained, and the received retrieval keywords are input into the pre-trained language model. After passing through the output layer of the pre-trained language model, a first representation vector is output; the word segments in the document word segment set are input into the pre-trained language model, and after passing through the output layer of the pre-trained language model, a second representation vector is output. The cosine similarity between the first representation vector and the second representation vector is calculated, and this cosine similarity is the word similarity between the retrieval keyword and the word segment; a preset similarity is obtained. When the word similarity between the retrieval keyword and the word segment is greater than or equal to the preset similarity, the word segment with a word similarity greater than or equal to the preset similarity is used as a candidate keyword, and the retrieval keyword is replaced with the candidate keyword; when calculating the semantic distance between the retrieval keywords, the calculation is performed through the candidate keyword. If the word similarity between the retrieval keyword and the word segment is less than the preset similarity, there is no need to replace the retrieval keyword, and when calculating the semantic distance between the retrieval keywords, the original retrieval keyword is still used for the calculation.

[0121] In this embodiment, when the retrieval keywords do not exist in the target knowledge graph, the word similarity between the retrieval keywords and the word segments in the document word segment set is calculated, and the retrieval keywords are replaced according to the word similarity, so that all received retrieval keywords can be accurately searched through the target knowledge graph, improving the retrieval range based on keywords.

[0122] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through computer-readable instructions, and the computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0123] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially according to the indication of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit and can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0124] Further reference is made to Figure 4 , as an implementation of the method shown above Figure 2 , an embodiment of a document retrieval device based on a knowledge graph is provided in this application. This device embodiment corresponds to Figure 2 the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0125] As Figure 4 shown, the document retrieval device 400 based on a knowledge graph described in this embodiment includes: a construction module 401, a first calculation module 402, a selection module 403, a second calculation module 404, an extraction module 405, and a confirmation module 406. Among them:

[0126] The construction module 401 is used to obtain a set of documents to be retrieved, perform word segmentation processing on the documents in the set of documents to be retrieved to obtain a set of document word segments, and construct a target knowledge graph based on the set of document word segments;

[0127] In this embodiment, the set of documents to be retrieved is a set of pre-collected documents. Obtain the set of documents to be retrieved, and perform word segmentation processing on the documents in the set of documents to be retrieved to obtain a set of document word segments. Specifically, a preset word segmentation tool can be used to perform word segmentation on the content of each document to obtain the word segments corresponding to different documents, and the word segments corresponding to all the documents in the set of documents to be retrieved are aggregated to obtain a set of document word segments. When the set of document word segments is obtained, entity and relationship extraction are performed on the word segments in the set of document word segments through named entity recognition to obtain the entities corresponding to each word segment and the relationships between different entities; then, triples are constructed according to the entities and relationships, such as triples in the format of (entity 1, relationship, entity 2), and the triples are input into a graph database (such as neo4j), and a target knowledge graph is output based on the graph database.

[0128] The first calculation module 402 is configured to, when receiving multiple retrieval keywords, calculate the semantic distance between the retrieval keywords based on the target knowledge graph, and determine two retrieval keywords corresponding to the maximum semantic distance as two central keywords;

[0129] In some alternative implementation manners of this embodiment, the first calculation module 402 includes:

[0130] The first acquisition unit is configured to acquire the reference knowledge graph corresponding to the target knowledge graph, and determine the distance weights between all the retrieval keywords according to the reference knowledge graph;

[0131] The calculation unit is configured to calculate the embedding similarity between the retrieval keywords, and the sum of the embedding vectors of the corresponding edge nodes of each retrieval keyword in the target knowledge graph, and calculate the semantic distance between the retrieval keywords according to the distance weights, the embedding similarity, and the sum of the embedding vectors.

[0132] In some alternative implementation manners of this embodiment, the first acquisition unit includes:

[0133] The confirmation subunit is configured to acquire the category attributes and levels of the retrieval keywords in the reference knowledge graph, and determine the distance weights between the retrieval keywords according to the category attributes and the levels.

[0134] In some alternative implementation manners of this embodiment, the confirmation subunit includes:

[0135] The judgment subunit is configured to judge whether the category attributes between each pair of the retrieval keywords are the same, and whether the levels between each pair of the retrieval keywords are the same. When the category attributes of the retrieval keywords are the same and the levels are the same, determine the distance weight between the retrieval keywords as a preset weight;

[0136] The first calculation subunit is configured to, when the category attributes are different, or the category attributes are the same and the levels are different, acquire the common superior entity of the retrieval keywords in the reference knowledge graph, calculate the hierarchical distance between the superior entity and the retrieval keywords, and calculate the distance weight between the retrieval keywords according to the hierarchical distance.

[0137] In this embodiment, the retrieval keyword is the keyword retrieved from the input document. By tokenizing the retrieval statement input by the user, the retrieval keyword can be obtained. When multiple retrieval keywords are received, the semantic distance between different retrieval keywords is calculated based on the target knowledge graph, and the retrieval keywords are filtered based on this semantic distance. Specifically, the semantic distance is the semantic distance between two different retrieval keywords. When the retrieval keyword is obtained, the retrieval keyword is matched with the entity nodes in the target knowledge graph to determine the entity nodes that match the retrieval keyword. Then, the semantic distance between different entity nodes is calculated, and the semantic distance between different entity nodes in the target knowledge graph is used as the semantic distance between the matching retrieval keywords. When determining the entity nodes that match the retrieval keyword, the embedding vector of the entity node is calculated, and the embedding vector can be obtained by performing one-hot encoding on the retrieval keyword through the bag of words (BOW). Then, according to the calculation formula of cosine similarity, the cosine similarity between different embedding vectors is calculated, and this cosine similarity is the semantic distance between the entity nodes. When the semantic distances between all retrieval keywords are obtained, the two retrieval keywords corresponding to the maximum semantic distance are used as the two central keywords.

[0138] The selection module 403 is configured to respectively construct a first sub-graph and a second sub-graph based on the two central keywords, calculate the number of nodes in the first sub-graph and the second sub-graph respectively, and select the maximum semantic sub-graph from the first sub-graph and the second sub-graph according to the number of nodes.

[0139] In this embodiment, when determining the central keywords, a first sub-graph and a second sub-graph are constructed based on the central keywords. Among them, the central keywords are the two optimal retrieval keywords selected through semantic distance. One central keyword constructs one sub-graph, and the first sub-graph and the second sub-graph can be respectively constructed according to the two central keywords. Specifically, any one of the central keywords is arbitrarily selected as the center of the circle, and the maximum semantic distance is used as the radius to circle the first sub-graph from the target knowledge graph; the remaining central keyword is used as the center of the circle, and the maximum semantic distance is used as the radius to circle the second sub-graph from the target knowledge graph. The first sub-graph and the second sub-graph include different numbers of entity nodes. The number of nodes in the first sub-graph and the second sub-graph is obtained, and the sub-graph with the largest number of nodes in the first sub-graph and the second sub-graph is used as the maximum semantic sub-graph. As Figure 3 shown, Figure 3 is a schematic diagram of the maximum semantic sub-graph. Among them, the semantic distance between kw1 and kw3 is the longest, and kw1 and kw3 are the two central keywords. The first sub-graph (i.e., sub Figure 1 ) is obtained based on kw1, and the second sub-graph (i.e., sub Figure 2 ) is obtained based on kw2; the number of nodes in the second sub-graph is greater than that in the first sub-graph, and the second sub-graph is the maximum semantic sub-graph.

[0140] A second computing module 404, configured to obtain a preset graph convolutional neural network, and perform feature extraction on the maximum semantic subgraph based on the graph convolutional neural network to obtain a feature vector;

[0141] In some optional implementation manners of this embodiment, the second computing module 404 includes:

[0142] A second computing subunit, configured to calculate an adjacency matrix and an in-degree / out-degree matrix of the maximum semantic subgraph;

[0143] A third computing subunit, configured to obtain a preset weight matrix, and calculate the feature vector through the graph convolutional neural network according to the weight matrix, the adjacency matrix, and the in-degree / out-degree matrix.

[0144] In this embodiment, when obtaining the maximum semantic subgraph, a preset graph convolutional neural network is obtained, and feature extraction is performed on the maximum semantic subgraph based on the graph convolutional neural network to obtain a feature vector. Specifically, a graph convolutional neural network (GCN, Graph Convolutional Network) is a convolutional neural network used to process graph structures. Through this graph convolutional neural network, feature extraction can be performed on the input maximum semantic subgraph to obtain a feature vector. Specifically, when obtaining the maximum semantic subgraph, calculate the adjacency matrix and the in-degree / out-degree matrix of the maximum semantic subgraph; then, according to the graph convolutional neural network, calculate the adjacency matrix, the in-degree / out-degree matrix, and the embedding vectors of the nodes in the maximum semantic subgraph to obtain a Laplacian matrix feature vector, and this Laplacian matrix feature vector is the feature vector of the maximum semantic subgraph.

[0145] An extraction module 405, configured to extract the topic words of each document in the document set to be retrieved, and calculate the topic embedding vectors of the topic words;

[0146] In some optional implementation manners of this embodiment, the extraction module 405 includes:

[0147] A sorting unit, configured to obtain the number of words of each document, sort the documents in ascending order according to the number of words to obtain a document queue;

[0148] A second obtaining unit, configured to obtain the number of topic words corresponding to the document with the lowest order in the document queue, use the number of topic words corresponding to the document with the lowest order as the lowest threshold, and incrementally increase the number of topic words of other documents in the document queue in the arrangement order of the document queue based on the lowest threshold until the number of topic words reaches a preset maximum threshold;

[0149] An extraction unit is configured to sequentially extract the subject words of the documents in the document queue according to the order and quantity from the lowest threshold to the highest threshold.

[0150] In this embodiment, the subject word is the core word of each document. The subject word can be a word whose occurrence times in the document are greater than or equal to a times threshold, or a word whose semantic weight is greater than or equal to a weight threshold, or a word whose total weight obtained by calculating the occurrence times and semantic weight in the document is greater than or equal to a total weight threshold. Extract the subject words of each document in the document set to be retrieved, and calculate the embedding vector of the subject word of each document. The embedding vector of the subject word is the subject embedding vector. The subject embedding vector can be calculated through a preset pre-trained language model (bert). Specifically, when the subject word of the document is obtained, the subject word is input into the pre-trained language model, and the embedding vector of each subject word is obtained based on the embedding encoding of the pre-trained language model.

[0151] A confirmation module 406 is configured to calculate the vector similarity between each subject embedding vector and the feature vector, determine that the subject embedding vector with the vector similarity greater than or equal to a preset similarity threshold is the target embedding vector, and use the document corresponding to the target embedding vector as the target retrieval document.

[0152] In this embodiment, based on the calculation formula of cosine similarity, calculate the vector similarity between the subject embedding vector and the feature vector, and determine whether the vector similarity is greater than or equal to a preset similarity threshold; if the vector similarity is greater than or equal to the preset similarity threshold, determine that the subject embedding vector corresponding to the vector similarity is the target embedding vector; if the vector similarity is less than the preset similarity threshold, determine that the subject embedding vector corresponding to the vector similarity is not the target embedding vector. When the target embedding vector is obtained, use the document corresponding to the target embedding vector as the target retrieval document.

[0153] In some optional implementation manners of this embodiment, the above-mentioned document retrieval device 400 based on the knowledge graph further includes:

[0154] A retrieval module is configured to retrieve the target knowledge graph to determine whether the retrieval keyword exists in the target knowledge graph;

[0155] A third calculation module is configured to, when the retrieval keyword does not exist in the target knowledge graph, obtain a preset pre-trained language model, input the retrieval keyword and the word segments in the document word segment set into the pre-trained language model respectively, and calculate to obtain a first representation vector and a second representation vector;

[0156] A replacement module, configured to calculate the word similarity between the retrieval keyword and the word segmentation according to the first representation vector and the second representation vector, determine the word segmentations with the word similarity greater than or equal to a preset similarity as candidate keywords, and replace the retrieval keyword with the candidate keyword.

[0157] In this embodiment, before calculating the semantic distance between retrieval keywords based on the target knowledge graph, the target knowledge graph is retrieved to determine whether the retrieval keyword exists in the target knowledge graph. If the retrieval keyword does not exist in the target knowledge graph, a preset pre-trained language representation model (i.e., Bert, Bidirectional Encoder Representation from Transformers) is obtained, and the received retrieval keyword is input into the pre-trained language model. After passing through the output layer of the pre-trained language model, a first representation vector is output; the word segmentations in the document word segmentation set are input into the pre-trained language model, and after passing through the output layer of the pre-trained language model, a second representation vector is output. Calculate the cosine similarity between the first representation vector and the second representation vector, and this cosine similarity is the word similarity between the retrieval keyword and the word segmentation; obtain a preset similarity. When the word similarity between the retrieval keyword and the word segmentation is greater than or equal to the preset similarity, the word segmentations with the word similarity greater than or equal to the preset similarity are used as candidate keywords, and the retrieval keyword is replaced with the candidate keyword; when calculating the semantic distance between retrieval keywords, the calculation is performed through the candidate keyword. If the word similarity between the retrieval keyword and the word segmentation is less than the preset similarity, there is no need to replace the retrieval keyword, and when calculating the semantic distance between retrieval keywords, the original retrieval keyword is still used for the calculation.

[0158] The document retrieval device based on the knowledge graph proposed in this embodiment realizes the efficient screening of the target retrieval document, reduces the interference of irrelevant documents on the retrieval result, saves the machine learning cost, and improves the efficiency and accuracy of the target document retrieval.

[0159] To solve the above technical problems, an embodiment of the present application also provides a computer device. For details, please refer to Figure 5 , Figure 5 which is the basic structural block diagram of the computer device in this embodiment.

[0160] The computer device 6 includes a memory 61, a processor 62, and a network interface 63 that are communicatively connected to each other via a system bus. It should be noted that only the computer device 6 with components 61-63 is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Among them, those skilled in the art of the present technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0161] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can interact with the user through a keyboard, a mouse, a remote control, a touchpad, or a voice control device, etc.

[0162] The memory 61 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 61 can be an internal storage unit of the computer device 6, such as the hard disk or memory of the computer device 6. In other embodiments, the memory 61 can also be an external storage device of the computer device 6, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., equipped on the computer device 6. Of course, the memory 61 can also include both the internal storage unit and the external storage device of the computer device 6. In this embodiment, the memory 61 is generally used to store the operating system and various application software installed on the computer device 6, such as computer-readable instructions of the document retrieval method based on the knowledge graph. In addition, the memory 61 can also be used to temporarily store various data that have been output or will be output.

[0163] In some embodiments, the processor 62 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 62 is generally used to control the overall operation of the computer device 6. In this embodiment, the processor 62 is used to run the computer-readable instructions stored in the memory 61 or process data, such as running the computer-readable instructions of the document retrieval method based on the knowledge graph.

[0164] The network interface 63 may include a wireless network interface or a wired network interface, which is generally used to establish a communication connection between the computer device 6 and other electronic devices.

[0165] The computer device proposed in this embodiment realizes the efficient screening of target retrieval documents, reduces the interference of irrelevant documents on the retrieval results, saves the machine learning cost, and improves the efficiency and accuracy of target document retrieval.

[0166] The present application also provides another implementation manner, that is, to provide a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor, so that the at least one processor executes the steps of the document retrieval method based on the knowledge graph as described above.

[0167] The computer-readable storage medium proposed in this embodiment realizes the efficient screening of target retrieval documents, reduces the interference of irrelevant documents on the retrieval results, saves the machine learning cost, and improves the efficiency and accuracy of target document retrieval.

[0168] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which may be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in the various embodiments of the present application.

[0169] Obviously, the embodiments described above are only a part of the embodiments of this application, rather than all embodiments. The preferred embodiments of this application are given in the drawings, but they do not limit the patent scope of this application. This application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of this application more thorough and comprehensive. Although this application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements on some of the technical features. Any equivalent structure that makes use of the content of the specification and drawings of this application, directly or indirectly applied in other related technical fields, is equally within the scope of patent protection of this application.

Claims

1. A document retrieval method based on a knowledge graph, characterized in that, It includes the following steps: Obtain the collection of documents to be retrieved, perform word segmentation on the documents in the collection of documents to be retrieved to obtain a document word segmentation set, and construct a target knowledge graph based on the document word segmentation set; When receiving multiple retrieval keywords, calculate the semantic distance between the retrieval keywords based on the target knowledge graph, and determine the two retrieval keywords corresponding to the maximum semantic distance as two central keywords; Construct a first subgraph and a second subgraph respectively based on the two central keywords, calculate the number of nodes in the first subgraph and the second subgraph respectively, and select the maximum semantic subgraph from the first subgraph and the second subgraph according to the number of nodes; Obtain a preset graph convolutional neural network, and perform feature extraction on the maximum semantic subgraph based on the graph convolutional neural network to obtain a feature vector; Extract the topic words of each document in the collection of documents to be retrieved, and calculate the topic embedding vectors of the topic words; Calculate the vector similarity between each topic embedding vector and the feature vector, determine the topic embedding vectors whose vector similarity is greater than or equal to a preset similarity threshold as target embedding vectors, and use the documents corresponding to the target embedding vectors as target retrieval documents; Among them, the step of constructing a first subgraph and a second subgraph respectively based on the two central keywords includes: Arbitrarily select one of the central keywords as the center of the circle, use the maximum semantic distance as the radius, and circle the first subgraph from the target knowledge graph; Use the remaining central keyword as the center of the circle, use the maximum semantic distance as the radius, and circle the second subgraph from the target knowledge graph; The step of selecting the maximum semantic subgraph from the first subgraph and the second subgraph according to the number of nodes includes: Use the subgraph with the largest number of nodes in the first subgraph and the second subgraph as the maximum semantic subgraph.

2. The document retrieval method based on a knowledge graph according to claim 1, characterized in that, The step of calculating the semantic distance between the retrieval keywords based on the target knowledge graph includes: Obtain the reference knowledge graph corresponding to the target knowledge graph, and determine the distance weights between all the retrieval keywords according to the reference knowledge graph; Calculate the embedding similarity between the retrieval keywords, and the sum of the embedding vectors of the corresponding edge nodes of each retrieval keyword in the target knowledge graph, and calculate the semantic distance between the retrieval keywords according to the distance weights, the embedding similarity, and the sum of the embedding vectors.

3. The document retrieval method based on a knowledge graph according to claim 2, characterized in that, The step of determining the distance weights between the retrieval keywords according to the reference knowledge graph includes: Obtain the category attributes and levels of the retrieval keywords in the reference knowledge graph, and determine the distance weights between the retrieval keywords according to the category attributes and the levels.

4. The document retrieval method based on a knowledge graph according to claim 3, characterized in that, The step of determining the distance weights between the retrieval keywords according to the category attributes and the levels includes: Judge whether the category attributes between each retrieval keyword are the same, and whether the levels between each retrieval keyword are the same. When the category attributes of the retrieval keywords are the same and the levels are the same, determine the distance weight between the retrieval keywords as the preset weight; When the category attributes are different, or the category attributes are the same and the levels are different, obtain the common superior entities of the retrieval keywords in the reference knowledge graph, calculate the hierarchical distance between the superior entities and the retrieval keywords, and calculate the distance weight between the retrieval keywords according to the hierarchical distance.

5. The document retrieval method based on a knowledge graph according to claim 1, characterized in that, The step of extracting the subject words of each document in the document set to be retrieved includes: Obtain the number of words in each document, sort the documents in ascending order according to the number of words to obtain a document queue; Obtain the number of subject words corresponding to the document with the lowest order in the document queue, use the number of subject words corresponding to the document with the lowest order as the lowest threshold, and incrementally increase the number of subject words of other documents in the document queue in the order of the document queue based on the lowest threshold until the number of subject words reaches a preset maximum threshold; Extract the subject words of the documents in the document queue in sequence according to the order and quantity from the lowest threshold to the maximum threshold.

6. The document retrieval method based on a knowledge graph according to claim 1, characterized in that, The step of extracting feature vectors based on the graph convolutional neural network for the maximum semantic subgraph includes: Calculate the adjacency matrix and in-out degree matrix of the maximum semantic subgraph; Obtain a preset weight matrix, and calculate the feature vectors through the graph convolutional neural network according to the weight matrix, the adjacency matrix, and the in-out degree matrix.

7. The document retrieval method based on a knowledge graph according to claim 1, characterized in that, Before the step of calculating the semantic distance between the retrieval keywords based on the target knowledge graph, it further includes: Retrieve the target knowledge graph to determine whether the retrieval keywords exist in the target knowledge graph; When the retrieval keywords do not exist in the target knowledge graph, obtain a preset pre-trained language model, input the retrieval keywords and the tokens in the document token set into the pre-trained language model respectively, and calculate the first representation vector and the second representation vector; Calculate the word similarity between the retrieval keywords and the tokens according to the first representation vector and the second representation vector, determine the tokens with the word similarity greater than or equal to the preset similarity as candidate keywords, and replace the retrieval keywords with the candidate keywords.

8. A document retrieval device based on a knowledge graph, characterized in that It includes: A construction module for obtaining a document set to be retrieved, performing word segmentation on the documents in the document set to be retrieved to obtain a document token set, and constructing a target knowledge graph based on the document token set; A first calculation module for, when receiving multiple retrieval keywords, calculating the semantic distance between the retrieval keywords based on the target knowledge graph, and determining the two retrieval keywords corresponding to the maximum semantic distance as two central keywords; A selection module for respectively constructing a first subgraph and a second subgraph based on the two central keywords, respectively calculating the number of nodes in the first subgraph and the second subgraph, and selecting the maximum semantic subgraph from the first subgraph and the second subgraph according to the number of nodes; A second calculation module for obtaining a preset graph convolutional neural network, and extracting features of the maximum semantic subgraph based on the graph convolutional neural network to obtain feature vectors; An extraction module for extracting the subject words of each of the to-be-retrieved document sets and calculating the subject embedding vectors of the subject words; A confirmation module for calculating the vector similarity between each of the subject embedding vectors and the feature vector, determining the subject embedding vectors with the vector similarity greater than or equal to a preset similarity threshold as target embedding vectors, and taking the documents corresponding to the target embedding vectors as target retrieval documents; Wherein, the first calculation module is further configured to arbitrarily select one of the central keywords as the center of a circle, and use the maximum semantic distance as the radius to delineate the first subgraph from the target knowledge graph; use the remaining central keywords as the center of a circle, and use the maximum semantic distance as the radius to delineate the second subgraph from the target knowledge graph; The first calculation module is further configured to use the subgraph with the largest number of nodes in the first subgraph and the second subgraph as the maximum semantic subgraph.

9. A computer device, comprising a memory and a processor, wherein computer-readable instructions are stored in the memory, and when the processor executes the computer-readable instructions, the steps of the document retrieval method based on the knowledge graph according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that Computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by a processor, the steps of the knowledge-graph-based document retrieval method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Graph-based document retrieval method and system and related components thereof

    CN112836029A

  • Document-based retrieval method and device

    CN113094519A