Text retrieval method, medium, computer equipment and program product

Through the combination of vector database and graph network, the dual expansion of user query entities is achieved, the contradiction between retrieval efficiency and comprehensiveness in GraphRAG technology is solved, and the real-time and accuracy of retrieval is improved.

CN120429430AInactive Publication Date: 2025-08-05ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510942193.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-08-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing GraphRAG technology is difficult to ensure the comprehensiveness of the search while improving the search efficiency, mainly because the processing of large language models takes a long time.

Method used

The dual entity expansion mechanism of vector database and graph network is adopted to expand the entities in user query through vector database, and then further expand in graph network, and finally retrieve the target text block from the knowledge base.

Benefits of technology

It significantly improves the efficiency and response speed of the search, while ensuring the comprehensiveness of the search. Through two expansions, the information obtained is more comprehensive, reducing the time-consuming process of large models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429430A_ABST
    Figure CN120429430A_ABST
Patent Text Reader

Abstract

A text retrieval method, medium, computer device and program product, the method comprising: receiving a user query, and obtaining an entity in the user query; querying a target vector from a pre-established vector database, wherein the target vector comprises a first entity vector corresponding to an entity in the user query and a second entity vector associated with the first entity vector; determining a target entity, and determining a target sub-graph from a pre-established graph network; the target entity comprises an entity corresponding to the target vector; nodes in the graph network correspond to vectors in the vector database, and the nodes in the graph network and the vectors corresponding to the nodes in the vector database correspond to the same entity; the target sub-graph comprises a first node corresponding to the target entity and a second node associated with the first node; and retrieving the target text block from the knowledge base based on the entity corresponding to each node in the target sub-graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of search enhancement generation technology, and in particular to a text search method, medium, computer device and program product. Background Art

[0002] Retrieval-Augmented Generation (RAG) is a technical framework that combines retrieval and generation. Traditional RAG relies primarily on textual data sources and lacks the ability to effectively utilize complex relationships and global information. Therefore, GraphRAG has been proposed as an innovative solution, combining the advantages of graph structures with the RAG framework to better understand and utilize the associations and structural features in data. By representing knowledge in a graphical format, GraphRAG can consider a wider range of contextual information and connections between concepts, thereby improving the quality and accuracy of retrieval generation. To ensure comprehensive retrieval, existing GraphRAG implementations typically use large models to recursively retrieve entities related to the user query and then retrieve text blocks based on these entities. However, processing large language models is time-consuming, resulting in low retrieval efficiency. GraphRAG implementations in related art struggle to achieve comprehensiveness while improving retrieval efficiency. Summary of the Invention

[0003] In a first aspect, an embodiment of the present application provides a search method, the method comprising: Receive a user query and obtain entities in the user query; Querying a target vector from a pre-established vector database, the target vector including a first entity vector corresponding to the entity in the user query and a second entity vector associated with the first entity vector; Determine a target entity and determine a target subgraph from a pre-established graph network; the target entity includes an entity corresponding to the target vector; the nodes of the graph network are used to represent entities, and the edges in the graph network are used to represent relationships between entities; the nodes in the graph network correspond to vectors in the vector database, and the nodes in the graph network and the vectors corresponding to the nodes in the vector database correspond to the same entity; the target subgraph includes a first node corresponding to the target entity and a second node associated with the first node; The target text block is retrieved from the knowledge base based on the entities corresponding to each node in the target subgraph.

[0004] In a second aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any embodiment of the present application.

[0005] In a third aspect, an embodiment of the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in any embodiment of the present application when executing the computer program.

[0006] In a fourth aspect, an embodiment of the present application provides a computer program product, including a computer program, which implements the method described in any embodiment of the present application when executed by a processor.

[0007] In an embodiment of the present application, the retrieval of the target text block is achieved based on two technologies: a vector database and a graph network. After obtaining the user query, the target vector is retrieved from the vector database. On the one hand, the target vector includes not only the first entity vector corresponding to the entity in the user query, but also the second entity vector that matches the first entity vector. In this way, the entity in the user query is expanded, making the subsequent retrieved information more comprehensive. On the other hand, the time consumption of the vector matching process is much lower than the time consumption of calling a large model for data processing, thereby effectively improving the real-time performance and response speed of the retrieval. Further, based on the target entity, the target subgraph is determined from the graph network. On the one hand, the target subgraph includes not only the first node corresponding to the target entity corresponding to the target vector, but also the second node whose distance to the first node is less than a preset distance threshold. This further expands the entity in the user query, thereby further improving the comprehensiveness of the subsequent retrieved information. On the other hand, the time consumption of the process of determining the target subgraph from the graph network is much lower than the time consumption of calling a large model for data processing, thereby effectively improving the real-time performance and response speed of the retrieval. In summary, this application not only significantly improves the retrieval efficiency but also ensures the comprehensiveness of the retrieval information through the dual entity expansion mechanism of vector and graph network.

[0008] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The drawings herein are incorporated into the specification and constitute a part of this application. These drawings illustrate embodiments consistent with this application and, together with the specification, are used to illustrate the technical solutions of this application.

[0010] Figure 1 It is a flowchart of the retrieval method of an embodiment of the present application.

[0011] Figure 2It is a schematic diagram of the relationship between the vector database and the graph network in an embodiment of the present application.

[0012] Figure 3 It is an overall flow chart of an embodiment of the present application.

[0013] Figure 4 It is a block diagram of the retrieval device of an embodiment of the present application.

[0014] Figure 5 It is a schematic diagram of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION

[0015] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0016] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "said" and "the" used in this application and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items. In addition, the term "at least one" herein represents any combination of at least two of any one or more of a plurality of.

[0017] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0018] In order to enable people in this technical field to better understand the technical solutions in the embodiments of the present application, and to make the above-mentioned purposes, features and advantages of the embodiments of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application are further described in detail below with reference to the accompanying drawings.

[0019] Retrieval-Augmented Generation (RAG) is a technical framework that combines retrieval and generation. Applications of RAG include, but are not limited to, question-answering systems and information retrieval systems based on large language models (LLMs). In question-answering systems, RAG retrieves text blocks relevant to the user's query, generates prompts based on these blocks and the user's query, and then feeds these prompts into the LLM, which generates responses based on the prompts. In information retrieval systems, RAG retrieves text blocks relevant to the target entity and presents these blocks directly to the user.

[0020] GraphRAG (Graph-based Retrieval-Augmented Generation) is an advanced version of RAG. It combines the advantages of graph structure and RAG framework, aiming to better understand and utilize the correlation and structural features in the data. By representing knowledge in a graphical way, GraphRAG can consider a wider range of background information and connections between concepts when generating answers, thereby further improving the quality and accuracy of the content. In related technologies, in order to ensure the comprehensiveness of retrieval, the GraphRAG implementation recursively obtains entities related to the user query through a large model, and retrieves text blocks based on the obtained entities. However, the processing of large language models takes a long time, resulting in low retrieval efficiency. The GraphRAG implementation in related technologies makes it difficult to ensure the comprehensiveness of retrieval while improving retrieval efficiency.

[0021] Based on this, the present application proposes a retrieval method based on a dual entity expansion mechanism of vectors and graph networks. First, the entities in the user query are expanded based on the vector database. Then, the entities after the first expansion are expanded again based on the graph network, and then the target text block is retrieved based on the entities after the second expansion. On the one hand, through the two expansions, the acquired entities are more comprehensive, so that the subsequent retrieved information is more comprehensive; on the other hand, the entity expansion process based on the vector database and the graph network is much less time-consuming than calling a large model for data processing, thereby effectively improving the real-time performance of the retrieval. In summary, the present application can ensure the comprehensiveness of the retrieval while improving the retrieval efficiency. The technical details of the embodiments of the present application are illustrated below with reference to the accompanying drawings.

[0022] like Figure 1 As shown, the retrieval method of the embodiment of the present application includes: Step S12: receiving a user query and obtaining entities in the user query; Step S14: querying a target vector from a pre-established vector database, where the target vector includes a first entity vector corresponding to the entity in the user query and a second entity vector associated with the first entity vector; Step S16: Determine a target entity and determine a target subgraph from a pre-established graph network; the target entity includes an entity corresponding to the target vector; the nodes of the graph network are used to represent entities, and the edges in the graph network are used to represent relationships between entities; the nodes in the graph network correspond to the vectors in the vector database, and the nodes in the graph network and the vectors corresponding to the nodes in the vector database correspond to the same entity; the target subgraph includes a first node corresponding to the target entity and a second node associated with the first node; Step S18: Retrieve the target text block from the knowledge base based on the entities corresponding to each node in the target subgraph.

[0023] In step S12, the user query refers to a keyword, phrase, or sentence entered by the user to express their information needs. The user query is the starting point of the retrieval process, and relevant information can be retrieved from the knowledge base or document collection based on the user query. The user query can be part of the prompt information input into the large language model, a search keyword input into the retrieval system, or search information input into the browser.

[0024] In some embodiments, user queries can be parsed using an entity extraction model to identify entities within the query. An entity extraction model is a natural language processing model designed to identify meaningful entities from text, such as names of people, places, organizations, dates, times, quantities, and so on. By extracting entities from user queries, we can accurately understand user needs and intent, thereby providing more precise search services.

[0025] In step S14 and step S16, a vector database and a graph network may be pre-established. Nodes in the graph network correspond to vectors in the vector database, and a node in the graph network and a vector corresponding to the node in the vector database correspond to the same entity.

[0026] In some embodiments, see Figure 2, it is possible to obtain raw data including several entities, identify several entities from the raw data (denoted as entity 1, entity 2, ..., entity N), perform vectorization processing on each identified entity, obtain a vector corresponding to each entity (denoted as vector 1, vector 2, ..., vector N), and associate the entity identification information (such as the entity name) with the vector corresponding to the entity. A vector database is established based on the associated entity identification information and vectors. In this way, each vector in the vector database can be made to correspond to the entity in the raw data. Furthermore, when the raw data is updated, the vectors in the vector database can also be updated synchronously. For example, after obtaining raw data including a new entity, an association relationship between the vector corresponding to the new entity and the entity identification information can be generated in the above manner, and the associated entity identification information and vector can be added to the vector database.

[0027] In some embodiments, in addition to entities, the original data may also include relationships between entities (hereinafter referred to as entity relationships). Entities and entity relationships can be identified from the original data, and a graph network can be generated using the entities as nodes and the entity relationships as edges. Furthermore, when the original data is updated, the nodes and / or edges in the graph network can be updated synchronously. After obtaining the original data including the new entity, the node corresponding to the new entity can be generated in the above manner, and based on the entity relationship between the new entity and the original entity, an edge can be established between the node corresponding to the new entity and the node corresponding to the original entity.

[0028] Since the vectors in the vector database correspond to entities, and the nodes in the graph network also correspond to entities, the nodes in the vector database also correspond to the nodes in the graph network.

[0029] After establishing the vector database, a target vector is retrieved from the vector database. The target vector may include not only the first entity vector corresponding to the entity in the user query, but also a second entity vector associated with the first entity vector. The second entity vector associated with the first entity vector may be a vector whose similarity to the first entity vector is greater than a preset similarity threshold.

[0030] The entities in the user query can be converted into entity vectors using an embedding model. Based on the converted entity vectors, a nearest neighbor search can be performed in a vector database to determine a number of target vectors. For example, the similarity between each vector in the vector database and the vectors converted from the entities in the user query can be determined. If the similarity between a vector in the vector database and the vector converted from the entity in the user query exceeds a preset similarity threshold, the vector is determined as the target vector. It will be appreciated that the method for determining the target vector is not limited to this, and this is merely an example.

[0031] By using both the first entity vector and the second entity vector as target vectors, the entities in the user query can be expanded, making the representation of the entities in the user query more diverse, thereby improving the comprehensiveness of subsequent searches. For example, the entity in the user query may be "acute pharyngitis." By using both the first entity vector corresponding to the entity "acute pharyngitis" and the second entity vector associated with the first entity vector as target vectors, the entity "acute pharyngitis" can be expanded to other entities with similar meanings, such as "acute pharyngitis." In this way, subsequent searches can not only retrieve text blocks that include the entity "acute pharyngitis," but also retrieve text blocks that include other entities such as "acute pharyngitis." This reduces the situation where relevant text blocks cannot be retrieved due to slight differences between the entity in the user query and the entity in the text block, thereby improving the comprehensiveness of the search.

[0032] In some embodiments, the number of entities in the user query may be greater than or equal to 1. In embodiments where the number of entities in the user query is greater than 1, a target vector may be queried from the vector database for each entity in the user query based on the above approach.

[0033] After determining the target vector, the target vector can be mapped to the target entity based on the correspondence between the vector and the entity. After obtaining the target entity, the target entity can be further expanded based on the graph network. Specifically, a target subgraph can be determined from the graph network. The target subgraph includes a first node corresponding to the target entity and second nodes associated with the first node. The second nodes associated with the first node can include nodes whose distance from the first node is less than a preset distance threshold. Since nodes in a graph network represent entities on their edges, and edges represent entity relationships, the distance between nodes can reflect the closeness of the association between entities. The closer the distance, the closer the association between entities. Designating nodes whose distance from the first node is less than a threshold as second nodes associated with the first node can expand understanding of the target entity. For example, assuming the first node corresponds to the disease "acute pharyngitis," an entity relationship can represent the relationship between the disease and symptoms. The second nodes can include nodes corresponding to entities representing symptoms (such as "sore throat"). Once the target subgraph includes these nodes, subsequent search processes can retrieve not only text blocks covering the disease "acute pharyngitis" but also text blocks covering symptoms related to "acute pharyngitis." In summary, when determining the target subgraph from a graph network, incorporating nodes whose distance from the first node is less than a preset threshold into the target subgraph as the second node can enrich the representation of the target entity, expand the analysis perspective, and make subsequent analysis more comprehensive and in-depth.

[0034] It will be appreciated that determining associated nodes based on the distance between nodes is only one optional method for determining associated nodes. In other examples, associated nodes may also be determined based on other methods, such as factors such as the connectivity between nodes (whether an edge exists between nodes) and / or the type of edge between nodes, which are not listed here.

[0035] Furthermore, in addition to the entity corresponding to the target vector, the target entity may also include synonymous entities of the entity corresponding to the target vector. A synonymous entity of an entity refers to an entity that has the same or similar meaning as the entity. For example, if the entity corresponding to the target vector is apple, its synonymous entity may include "apple". Synonymous entities can be obtained by performing synonym aggregation on the entity corresponding to the target vector.

[0036] In some embodiments, semantic relationships between entities can be recorded using a graph network. Specifically, edges in a graph network can include first-type edges and second-type edges. The first-type edge represents the contextual relationship between the entities corresponding to the node within a text block. For example, assuming a text block contains the following: "Symptoms of a cold include runny nose," the entities within the text block include "cold" and "runny nose," and the contextual relationship between these two entities is "includes," meaning the first-type edge represents this "includes" relationship. The second-type edge represents the semantic relationship between the entities corresponding to the nodes. If two entities are synonymous, a second-type edge exists between the nodes corresponding to these two entities. For example, assuming there are three entities: "apple," "apple," and "banana," a second-type edge exists between the entity corresponding to "apple" and the entity corresponding to "apple," but no second-type edge exists between the entity corresponding to "apple" and the entity corresponding to "banana," nor between the entity corresponding to "apple" and the entity corresponding to "banana." Based on this, the first node can include a first target node and a second target node. The entity corresponding to the first target node is the entity corresponding to the target vector. The second target node is the node connected to the first target node via a second-type edge. For example, assuming the entity corresponding to the target vector is entity A, the first target node includes the node corresponding to entity A. assuming the second target node is node B, node B is the node connected to node A via a second-type edge. In this way, synonymous entities of an entity can also be expanded to the target entity, allowing subsequent searches to retrieve text blocks containing synonymous entities, further improving the comprehensiveness of the search.

[0037] In step S18, target text blocks can be retrieved from the knowledge base based on the entities corresponding to each node in the target subgraph. Since the entities in the user query are expanded twice, the retrieved target text blocks can include not only text blocks related to the entities in the user query, but also text blocks related to the expanded entities.

[0038] In some embodiments, the weights of each node in the target subgraph can be determined, and several candidate text blocks can be retrieved. Each retrieved candidate text block includes an entity corresponding to at least one node in the target subgraph, and based on the weights of each node in the target subgraph, the target text block is filtered out from the several candidate text blocks.

[0039] The weight of a node can represent the correlation between the entity corresponding to the node and the entity in the user query. After two expansions of the vector dimension and the graph network dimension, although the comprehensiveness of the retrieval can be improved, it may also cause the retrieval results to include text blocks with low correlation with the entities in the user query. In order to improve the correlation between the retrieved target text block and the entity in the user query, this embodiment determines a weight for each node in the target subgraph, and determines the target text block from the candidate text blocks based on the weight. Specifically, the probability of a candidate text block being determined as the target text block can be positively correlated with the weight of the node corresponding to the entity included in the candidate text block. That is, the candidate text block containing the entity corresponding to the high-weight node is more likely to become the target text block. In this way, the final retrieval results will be affected by the correlation between the node and the entity in the user query, highlighting important information, filtering irrelevant content, and optimizing the retrieval quality.

[0040] The weights of each node in the target subgraph can be gradually attenuated outward from the first node. That is, in the target subgraph, the first node has the highest weight, followed by its first-degree neighbors, then its second-degree neighbors, and so on. Since the first node directly corresponds to the target entity, the correlation between the first node and the target entity is relatively high. The target entity is the entity in the user query obtained by expanding the vector database. Therefore, the correlation between the first node and the entity in the user query is also relatively high. The second node is an entity further expanded from the first node, and its correlation with the target entity is relatively low, and thus its correlation with the entity in the user query is also relatively low. Furthermore, as the distance between the second node and the first node increases, this correlation further decreases. By gradually attenuating the weights of each node in the target subgraph outward from the first node, the changing trend of the correlation between each node in the target subgraph and the entity in the user query can be better reflected, so that the search results are primarily influenced by nodes with higher correlation to the entity in the user query, thereby improving the correlation between the search results and the entity in the user query.

[0041] In some embodiments, the weights of each node in the target subgraph can be updated multiple times through iterations, and the weights updated in the last iteration are determined as the weights of each node in the target subgraph. In each iteration, each node in the target subgraph passes the weights to other nodes to update the weights of other nodes. The weight of any node is passed to the neighboring nodes and the first node of the node according to a preset ratio. In the above process, the weights of multiple nodes can be obtained in each iteration, and efficient weight acquisition can be achieved, thereby improving retrieval efficiency. Optionally, the above weight acquisition process can be implemented based on the TrustRank algorithm.

[0042] After obtaining the weights of each node in the target subgraph, the target text block can be screened out from the candidate text blocks based on the correlation between the candidate text blocks and the user query. The correlation between any candidate text block and the user query is determined in the following manner: obtaining word frequency information of entities corresponding to each node in the target subgraph in the candidate text block, weighting the weights of the corresponding nodes based on the word frequency information of the entities corresponding to each node in the target subgraph in the candidate text block to obtain a weighted weight of the corresponding node, and obtaining the correlation between the candidate text block and the user query based on the weighted weights of each node in the target subgraph.

[0043] The term frequency information of an entity in a candidate text block is positively correlated with the number of times the entity appears in the candidate text block. The term frequency information of the entity in the candidate text block can be directly determined by the number of times the entity appears in the candidate text block. Alternatively, the density of the entity's appearance in the candidate text block can be determined as the term frequency information of the entity in the candidate text block. The density of the entity's appearance in the candidate text block can be determined based on the ratio of the number of times the entity appears in the candidate text block to the length of the candidate text block. Generally, the more frequently an entity appears in a text, the more important it is in the text. Therefore, the term frequency information of an entity in a candidate text block can reflect the importance of the entity in the candidate text block. The higher the importance of an entity in a candidate text block, the higher the relevance between the candidate text block and the entity. Therefore, by obtaining the term frequency information of an entity in a candidate text block and using this term frequency information to weight the corresponding nodes, it is possible to effectively filter out text blocks that contain some relevant entities but with a shallow connection, or simply mention the entity, thereby improving the accuracy of retrieval results.

[0044] For example, assuming that a candidate text block includes the entity corresponding to node A in the target subgraph (denoted as entity A), the entity corresponding to node B (denoted as entity B), and the entity corresponding to node C (denoted as entity C), the word frequency information F of entity A in the candidate text block can be obtained. A , the word frequency information F of entity B in the candidate text block B , and the word frequency information F of entity C in the candidate text block C Based on F A Weight the weight corresponding to node A , and get the weighted weight , based on F B Weight the weight corresponding to node B , and get the weighted weight , based on F C Weight the weight corresponding to node C , and get the weighted weight Then, based on 、 and Get the relevance between the candidate text block and the user query. For example, 、 and The sum or mean of the two determines the relevance between the candidate text block and the user query.

[0045] In the above manner, the relevance between each candidate text block and the user query can be obtained, and accordingly, the top K candidate text blocks with the highest to lowest relevance to the user query are selected as target text blocks.

[0046] After determining the target text blocks, the target text blocks can be sorted in descending order of relevance to the user query and displayed. Alternatively, prompt information can be generated based on the target text blocks and the user query, and the prompt information can be input into the large language model, so that the large language model generates response information based on the prompt information.

[0047] The following combination Figure 3 The overall process of an embodiment of the present application is described. The overall process includes the following steps: (1) Entity recognition and expansion: Step 1: Use an efficient entity extraction model (or large language model) to parse the user query provided by the user and automatically identify the key entities.

[0048] Step 2: Convert these key entities into vector form through an embedding model (or large language model) to facilitate subsequent similarity matching.

[0049] Step 3: Perform nearest neighbor search operations within a specially constructed vector database to discover more potentially related entities, thereby effectively expanding the entity set.

[0050] (2) Subgraph generation: Based on the target entities obtained in the previous stage, a target subgraph containing the nodes corresponding to these target entities (i.e., first nodes) and their directly or indirectly associated nodes (i.e., second nodes) is retrieved from a pre-built graph network. Compared to traditional relational database query methods, this method significantly reduces the number of data connection operations while ensuring extremely high query efficiency (typically completed in milliseconds).

[0051] (3) Weight calculation: For the target entities and corresponding target subgraphs generated in the above steps, the TrustRank algorithm is applied to evaluate the importance of each entity and, based on this, select the top N most relevant results (N nodes ranked from highest to lowest weight). This algorithm not only considers the links between entities but also comprehensively considers the weight distribution across the entire target subgraph, ensuring highly relevant search results.

[0052] (4) Text block relevance ranking: Based on the results of the TrustRank algorithm, all candidate text blocks containing the specified entity are quickly located in the knowledge base. The weights are then summed based on the word frequency information of each entity to obtain the candidate text block score (which is used to indicate the relevance of the text block to the user query). Finally, the candidate text blocks are ranked according to their scores. The purpose of this is to provide users with a list of documents sorted by importance, helping them find the target text block more quickly.

[0053] (5) Intelligent question and answer generation: The sorted target text blocks, along with the original user query, are submitted to the large model, which generates an answer based on contextual understanding. This approach fully leverages the knowledge accumulated in previous steps, ensuring the quality and accuracy of the response.

[0054] If the final requirement is to return a list of text blocks, then in step 4, directly output the top M text blocks and their rankings and return them as the final result without executing the large model call in step 5.

[0055] This application is the first to build an end-to-end online GraphRAG query process. Relying on graph analysis and data mining systems, it integrates advanced technologies such as graph databases, vector databases, and multiple machine learning models to create a complex and efficient online query path that can directly generate large model output results enhanced by GraphRAG based on user queries. This application has the following advantages: (1) Relationship priority: Multiple steps in this application are based on the graph structure. In particular, in the subgraph extraction and TrustRank calculation stages, subgraph expansion and weight propagation are achieved through the edges of the graph network, effectively overcoming the problem of ignoring the connection between nodes in the traditional RAG method.

[0056] (2) Global Scoring Mechanism: When using the TrustRank algorithm, this application performs multiple iterative calculations based on the topology of the entire graph. The final weight obtained by each node not only reflects the influence of its direct neighbors, but also comprehensively considers information from all other nodes in the entire network, achieving a more comprehensive evaluation.

[0057] (3) Efficient resource utilization: This application only needs to store entities and their connections in the graph database, and only saves entity vector representations in the vector library, without the need for additional data storage. In addition, only slight memory consumption is consumed when executing the TrustRank algorithm. Therefore, overall, this method has low requirements for system resources.

[0058] (4) Limiting large model calls: Throughout the entire process, the large model is only called once, in the final content generation phase; and in the early entity recognition phase, it may require up to two calls. This means that the large model will not be used more than three times throughout the entire process, significantly reducing reliance on large models and effectively improving retrieval efficiency.

[0059] See also Figure 4 , the present application also provides a text retrieval device, the device comprising: Receiving module 102, configured to receive a user query and obtain entities in the user query; A query module 104 is configured to query a target vector from a pre-established vector database, wherein the target vector includes a first entity vector corresponding to the entity in the user query and a second entity vector associated with the first entity vector; Determination module 106 is configured to determine a target entity and determine a target subgraph from a pre-established graph network; the target entity includes the entity corresponding to the target vector; the nodes of the graph network are used to represent entities, and the edges of the graph network are used to represent relationships between entities; the nodes in the graph network correspond to the vectors in the vector database, and the nodes in the graph network and the vectors corresponding to the nodes in the vector database correspond to the same entity; the target subgraph includes a first node corresponding to the target entity and a second node associated with the first node; The retrieval module 108 is configured to retrieve a target text block from a knowledge base based on entities corresponding to each node in the target subgraph.

[0060] The specific implementation details of the above device embodiment are detailed in the above method embodiment and will not be repeated here.

[0061] An embodiment of the present application further provides a computer device, which comprises at least a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method described in any of the aforementioned embodiments is implemented.

[0062] Figure 5 2 shows a more specific hardware structure diagram of a computer device provided in an embodiment of the present application. The device may include: a processor 202, a memory 204, an input / output interface 206, a communication interface 208, and a bus 210. The processor 202, the memory 204, the input / output interface 206, and the communication interface 208 are connected to each other within the device via the bus 210.

[0063] The processor 202 can be implemented using a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application. The processor 202 can also include a graphics card, such as an Nvidia Titan X graphics card or an 1080Ti graphics card.

[0064] The memory 204 can be implemented in the form of a read-only memory (ROM), a random access memory (RAM), a static storage device, a dynamic storage device, etc. The memory 204 can store an operating system and other application programs. When the technical solutions provided in the embodiments of the present application are implemented through software or firmware, the relevant program code is stored in the memory 204 and is called and executed by the processor 202.

[0065] The input / output interface 206 is used to connect to input / output modules to enable information input and output. The input / output modules can be configured as components within the device (not shown) or externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, and various sensors, while output devices may include a display, speaker, vibrator, indicator light, and the like.

[0066] The communication interface 208 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, Wi-Fi, Bluetooth, etc.).

[0067] The bus 210 comprises a pathway for transmitting information between various components of the device, such as the processor 202 , the memory 204 , the input / output interface 206 , and the communication interface 208 .

[0068] It should be noted that although the above device only shows the processor 202, memory 204, input / output interface 206, communication interface 208, and bus 210, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of the present application, and does not necessarily include all the components shown in the figure.

[0069] An embodiment of the present application provides a computer program product, including a computer program, which implements the method described in any embodiment of the present application when executed by a processor.

[0070] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any of the aforementioned embodiments.

[0071] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented using any method or technology to store information. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computer device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0072] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is merely illustrative, wherein the modules described as separate components may or may not be physically separated, and the functions of each module can be implemented in the same one or more software and / or hardware when implementing the embodiment of this application. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the embodiment. Those of ordinary skill in the art can understand and implement it without paying any creative work.

[0073] The above is only a specific implementation of the embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the embodiment of the present application. These improvements and modifications should also be regarded as the scope of protection of the embodiment of the present application.

Claims

1. A text retrieval method, comprising: Receive a user query and obtain entities in the user query; Querying a target vector from a pre-established vector database, the target vector including a first entity vector corresponding to the entity in the user query and a second entity vector associated with the first entity vector; Determine a target entity and determine a target subgraph from a pre-established graph network; the target entity includes an entity corresponding to the target vector; the nodes of the graph network are used to represent entities, and the edges in the graph network are used to represent relationships between entities; the nodes in the graph network correspond to vectors in the vector database, and the nodes in the graph network and the vectors corresponding to the nodes in the vector database correspond to the same entity; the target subgraph includes a first node corresponding to the target entity and a second node associated with the first node; The target text block is retrieved from the knowledge base based on the entities corresponding to each node in the target subgraph.

2. The method according to claim 1, wherein obtaining the entity in the user query comprises: The user query is parsed through an entity extraction model to obtain entities in the user query.

3. The method according to claim 1, wherein querying the target vector from a pre-established vector database comprises: Converting entities in the user query into entity vectors through an embedding model; Based on the converted entity vectors, a nearest neighbor search is performed in the vector database to determine a plurality of target vectors.

4. The method according to claim 1, wherein the target entity further comprises: A synonymous entity of the entity corresponding to the target vector, wherein the synonymous entity is obtained by performing synonym aggregation on the entity corresponding to the target vector.

5. The method according to claim 4, wherein the edges in the graph network include first-type edges and second-type edges, wherein the first-type edges are used to represent the contextual relationship between entities corresponding to nodes in a text block, and the second-type edges are used to represent the semantic relationship between entities corresponding to nodes, and the entities corresponding to multiple nodes connected by the second-type edges are synonymous entities; the first node includes: A first target node, where the entity corresponding to the first target node is the entity corresponding to the target vector; as well as A second target node, wherein the second target node is connected to the first target node via an edge of a second type.

6. The method according to claim 1, wherein the step of retrieving a target text block from a knowledge base based on entities corresponding to each node in the target subgraph comprises: Determining the weight of each node in the target subgraph; The weight of a node is used to represent the correlation between the entity corresponding to the node and the entity in the user query; wherein the weight of each node in the target subgraph gradually decays outward from the first node as the center; Retrieving a plurality of candidate text blocks, wherein each retrieved candidate text block includes an entity corresponding to at least one node in the target subgraph; A target text block is selected from the candidate text blocks based on the weights of the nodes in the target subgraph.

7. The method according to claim 6, wherein determining the weight of each node in the target subgraph comprises: Iteratively updating the weights of the nodes in the target subgraph multiple times, and determining the weights updated in the last iteration as the weights of the nodes in the target subgraph; During each iteration, each node in the target subgraph transfers its weight to other nodes to update the weights of the other nodes; wherein the weight of any node is transferred to the neighboring nodes of the node and the first node respectively according to a preset ratio.

8. The method according to claim 6, wherein the step of selecting a target text block from the plurality of candidate text blocks based on the weights of the nodes in the target subgraph comprises: Based on the relevance between the candidate text blocks and the user query, a target text block is selected from the candidate text blocks; wherein the relevance between any candidate text block and the user query is determined based on the following method: Obtaining word frequency information of entities corresponding to each node in the target subgraph in the candidate text block; the word frequency information of an entity in the candidate text block is positively correlated with the number of times the entity appears in the candidate text block; Weighting the weights of the corresponding nodes based on the word frequency information of the entities corresponding to the nodes in the target subgraph to obtain a weighted weight of the corresponding nodes; The relevance between the candidate text block and the user query is obtained based on the weighted weights of the nodes in the target subgraph.

9. The method according to claim 1, wherein the number of the target text blocks is greater than 1; the method further comprising: sorting the target text blocks according to the order of relevance between the target text blocks and the user query from high to low; Display the sorted target text blocks.

10. The method according to claim 1, further comprising: generating prompt information based on the target text block and the user query; The prompt information is input into a large language model so that the large language model generates response information based on the prompt information.

11. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.

12. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 10 when executing the computer program.

13. A computer program product comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • SVO entity information retrieval system

    CN115210704A

  • Information retrieval method and device

    CN119848201A

  • Question and answer retrieval method

    CN120179773A

  • Information technology service consultation platform based on intelligent software

    CN120179803A

  • Searching for indirect entities using a knowledge graph

    US20250103651A1