Knowledge graph enhanced retrieval application system and method

By using large models to generate extended query and combining knowledge graphs and historical information in the knowledge graph enhancement retrieval application system, the problem of semantic inconsistency in data fusion is solved, the accuracy and retrieval efficiency of knowledge graphs are improved, and the diverse needs of users are met.

CN119938932APending Publication Date: 2025-05-06CHINA TELECOM CORP LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411978590.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing technology may have semantic inconsistency problems in the data fusion process of knowledge graph construction and retrieval, resulting in poor accuracy of the built knowledge graph and cannot meet the diverse needs of users.

Method used

Provide a knowledge graph enhancement retrieval application system and method, including public domain databases, private domain databases, extended query modules, LLM-agent interaction modules, memory Prompt settings modules and integrated language modules, and generate extended queries through large models, combining knowledge graphs and historical information to generate final question-and-answer results to ensure semantic consistency.

Benefits of technology

It improves the accuracy and retrieval efficiency of the knowledge graph, solves the accuracy problem caused by semantic inconsistency, and meets the diverse needs of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938932A_ABST
    Figure CN119938932A_ABST
Patent Text Reader

Abstract

The invention provides a knowledge graph enhanced retrieval application system and method. The expansion query module is arranged to expand potential query of attributes of nodes and relationships to be queried, and then query of the nodes and the relationships is carried out through the knowledge graph; lLM-agent interactive auxiliary query is carried out, after nodes and relation attributes corresponding to queried entities are obtained, interaction with the knowledge graph is carried out in an interactive mode, and graph query is completed; according to the method, knowledge graph query content and historical information are transmitted to the large model, one-time question and answer is completed, and compared with an existing scheme, semantic consistency in the data fusion process of knowledge graph construction and retrieval is guaranteed, so that the accuracy of the constructed knowledge graph is improved, and the user experience is improved. Therefore, the problem that the accuracy of the constructed knowledge graph is poor due to the fact that the reason of semantic inconsistency possibly exists in the data fusion process of knowledge graph construction and retrieval in an existing scheme is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of knowledge graph technology, and more specifically, to a knowledge graph enhanced retrieval application system and method. Background Art

[0002] Knowledge graph is a data integration technology that is different from traditional forms such as XML (Extensible Markup Language) and RDF (Resource Description Framework). It can be understood as representing the relationship between entities through nodes and edges, thereby realizing the construction and visualization of multi-layer association relationships, and then efficiently retrieving association relationships.

[0003] Knowledge graphs are usually constructed by data fusion, knowledge extraction, and manual construction. Data fusion is the process of integrating heterogeneous information from different data sources to form a unified semantic view, thereby constructing a knowledge graph. However, problems such as semantic inconsistency may occur during data fusion, resulting in poor accuracy of the constructed knowledge graph and failure to meet the diverse needs of users.

[0004] Therefore, there is an urgent need for a way to enhance retrieval efficiency and accuracy to solve the problem that existing solutions may have semantic inconsistencies in the process of knowledge graph construction and retrieval data fusion, resulting in poor accuracy of the constructed knowledge graph. Summary of the invention

[0005] The main purpose of this application is to provide a knowledge graph enhanced retrieval application system and method, so as to at least solve the problem that in the existing solutions, there may be semantic inconsistencies in the process of knowledge graph construction and retrieval data fusion, resulting in poor accuracy of the constructed knowledge graph.

[0006] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a knowledge graph enhanced retrieval application system is provided, the system comprising:

[0007] Public domain databases that store pre-built knowledge graphs;

[0008] Private domain database, used to update knowledge in specific fields in real time;

[0009] An extended query module, used to generate an extended query system for the node and relationship attribute values ​​of the knowledge graph based on the big model;

[0010] An LLM-agent interaction module, wherein the LLM-agent interaction module cooperates with the extended query module to obtain node attribute names using a pre-trained large model to enhance interaction with the knowledge graph;

[0011] A memory prompt setting module is used to set a specific prompt template, simplify the historical question and answer information and convert it into a vector form, and a vector database is connected to the memory prompt setting module to store the vectorized historical question and answer information;

[0012] The integrated language module is used to combine the knowledge graph query results and historical information to generate the final question and answer results.

[0013] Optionally, the private domain database ensures the real-time and accuracy of knowledge in a specific domain through a regular or event-driven update strategy; the extended query module generates an extended query based on the large model that can identify and locate entities and relationships in the knowledge graph, accurate to the node and relationship attribute values.

[0014] Optionally, the LLM-agent interaction module parses natural language queries and executes Cypher query statements to enhance interaction with the knowledge graph.

[0015] Optionally, the Prompt template set by the memory Prompt setting module is used to simplify historical question and answer information and realize conversion into vector form.

[0016] According to another aspect of the present application, a knowledge graph enhanced retrieval application method applied to any of the above systems is provided, the method comprising:

[0017] Receive user query requests and generate extended queries based on the large model;

[0018] Retrieve matching entities and relationships from public and private databases;

[0019] Setting a target Prompt template to convert historical question and answer information into a vector form, obtaining vectorized information, and storing the vectorized information in a vector database;

[0020] Integrate retrieval results and historical information to generate the final question-answering results.

[0021] Optionally, receiving a user query request and generating an extended query based on the large model includes:

[0022] According to the query target of the user query request, a query statement is constructed to search for relevant nodes and relationships in the knowledge graph, wherein the query statement includes query keywords, condition restrictions, and logical operators;

[0023] Input the constructed query statement into the knowledge graph for query, wherein the knowledge graph matches relevant nodes and relationships according to the query statement and returns the query result;

[0024] According to the query results returned by the knowledge graph, the nodes and relationships with the highest relevance to the query target are filtered out.

[0025] Optionally, in the process of generating an extended query based on the large model, the method further includes:

[0026] The large model is trained using deep learning technology, and the large model can automatically learn the relationship between the attribute name and value of the node. In the query process, according to the attribute value of a node input, the large model encodes the query value according to the knowledge obtained from the training, and converts the code into a representation form that the large model can recognize;

[0027] The large model performs correlation matching based on the encoded query value and the pre-learned node attribute name;

[0028] Obtain the node attribute name of the model returned by the large model that has the highest correlation with the query value.

[0029] Optionally, in the process of retrieving matching entities and relationships from the public domain database and the private domain database, the method further includes:

[0030] Screening and matching data in the public domain database and the private domain database according to preset conditions to find the relationship between the data;

[0031] When performing a relational query, first query the associated nodes to find the source or target of the data;

[0032] Connecting the associated nodes through an association operation to establish a relationship between the associated nodes;

[0033] The associated query result is obtained, wherein the associated query result includes data that meets the preset conditions, the attribute name and attribute value of each associated node.

[0034] Optionally, in the process of setting the target Prompt template, the method further includes:

[0035] Asking questions to the language model-based dialogue system in a natural language manner, wherein the language model-based dialogue system uses a pre-trained model and the knowledge graph to perform interactive queries according to the questions;

[0036] Receive an interactive query result returned by the language model-based dialogue system.

[0037] Optionally, the historical question and answer information is converted into a vector form to obtain vectorized information, including:

[0038] Simplify the historical question and answer information to obtain simplified question and answer information;

[0039] The simplified question and answer is converted into a vector form to obtain the vectorized information, and the vectorized information is stored in the Qdrant vector library to optimize storage and query efficiency.

[0040] By applying the technical solution of the present application, an extended query module is set up to expand the potential query of the attributes of the nodes and relationships to be queried, and then the nodes and relationships are queried through the knowledge graph; the LLM-agent interactive auxiliary query obtains the node and relationship attributes corresponding to the queried entity, and then uses an interactive method to interact with the knowledge graph to complete the graph query; the knowledge graph query content and historical information are transmitted to the big model to complete a question and answer. Compared with the existing solutions, the semantic consistency in the process of data fusion of knowledge graph construction and retrieval is guaranteed, thereby improving the accuracy of the constructed knowledge graph, and further solving the problem that the existing solutions may have semantic inconsistencies in the process of data fusion of knowledge graph construction and retrieval, resulting in poor accuracy of the constructed knowledge graph. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The drawings constituting part of the present application are used to provide a further understanding of the present application. The exemplary embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0042] Figure 1 A schematic diagram showing the principle of a knowledge graph enhanced retrieval application system provided in an embodiment of the present application is shown;

[0043] Figure 2 A flow chart of a knowledge graph enhanced retrieval application method provided according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0044] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0045] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0046] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0047] As introduced in the background technology, the construction methods of knowledge graphs usually include data fusion, knowledge extraction, and manual construction, among which data fusion is to integrate heterogeneous information from different data sources to form a unified semantic view, thereby constructing a knowledge graph. However, there may be problems such as semantic inconsistency in the process of data fusion, resulting in poor accuracy of the constructed knowledge graph and failure to meet the diverse needs of users. In order to solve the problem that the existing solutions may have semantic inconsistencies in the process of data fusion in the construction and retrieval of knowledge graphs, resulting in poor accuracy of the constructed knowledge graph, the embodiments of the present application provide a knowledge graph enhanced retrieval application system and method.

[0048] About the existing solutions: In the field of information retrieval, the question-answering system is an advanced form. It uses the intent recognition in natural language understanding technology to understand the customer's question-answering intention and knowledge extraction to complete the specific questions of the customer's questions and answers. Question-answering corpus, knowledge graph, public and private domain knowledge base, network knowledge, etc. can all be used as the knowledge base of the question-answering system. Through natural language technology, the customer's question-answering needs are understood, and concise and effective answers are returned in combination with the knowledge base. Compared with search engines, question-answering systems have more flexible question-answering knowledge base forms, more accurate answers, more convenient deployment forms and better question-answering experience. They are also technologies that have received attention and developed rapidly in recent years. The mobile network is divided into two parts: the wireless network and the core network. The wireless knowledge question-answering assistant aims to reduce the pressure of telecom operators, network operation and maintenance teams or technical support teams when solving wireless-related problems, improve solution efficiency, reduce operating costs, reduce service interruption time, and improve customer satisfaction. The wireless network accumulates a large amount of work order data, operation and maintenance knowledge and experience. These wireless data contain rich knowledge and information, but also contain a lot of redundant information. How to extract useful knowledge and remove redundant information is a prerequisite for improving the accuracy and user experience of the wireless question-answering assistant. For the existing wireless knowledge graph, how to efficiently use and accurately mine the hidden information between nodes is one of the keys to improving the accuracy of question answering. The emergence of large models and prompt words provides effective assistance in solving the utilization of wireless knowledge graphs.

[0049] Since training a large model suitable for this field requires a lot of computing resources and high-quality question-answer corpus pairs, knowledge retrieval enhancement can help improve the accuracy of question-answering based on the open source large model and the existing knowledge base, and meet the question-answering needs with fewer resources. Currently, knowledge question-answering faces the following challenges:

[0050] Requires additional training of multiple models: When processing queries, traditional methods need to identify the query intent, queried entities, and entity relationships, and require training of corresponding models, which can easily lead to inaccurate recognition problems.

[0051] Inconsistency between entities and relationships: After entities and relationships are extracted, they need to be normalized. The most appropriate entities and relationships are matched by similarity, which can easily cause inconsistency between query entities, relationships and knowledge graph entities and relationships.

[0052] Inaccurate subgraph query: The traditional similar subgraph query method is to perform community detection and clustering on the extracted query entities and relationships to determine the most similar subgraph. However, due to the performance of the model and the scale of the knowledge graph, it is easy to cause inaccurate queries.

[0053] Limitations of knowledge: Knowledge graph node entities and relationships are extracted from professional technical documents or work orders, which misses a lot of knowledge and is difficult to update and maintain.

[0054] The problem of knowledge forgetting in multi-round question-answering: It is difficult to implement multi-round question-answering based on knowledge graphs, and when multiple rounds of question-answering are generated, the current question-answering is difficult to associate with the previous rounds of question-answering.

[0055] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0056] The present application provides a knowledge graph enhanced retrieval application system, which includes:

[0057] Public domain databases that store pre-built knowledge graphs;

[0058] Private domain database, used to update knowledge in specific fields in real time;

[0059] The public domain database uses the existing knowledge graph library, and the private domain database uses the work orders, technical specifications, and case documents that are updated and published locally in real time;

[0060] An extended query module, used to generate an extended query system for the node and relationship attribute values ​​of the above knowledge graph based on the big model;

[0061] Large models usually refer to models with a large number of parameters and high complexity in the field of machine learning or artificial intelligence. These models usually require a lot of data and computing resources to train and deploy, but can achieve relatively good performance on some tasks. Large models can be used for a variety of different tasks, including natural language processing, computer vision, speech recognition and other fields. Some well-known large models include BERT, GPT-3, ResNet, etc. However, since large models require a lot of resources to train and deploy, their use also faces some challenges, such as computing resource consumption, data privacy and other issues.

[0062] LLM-agent interaction module, the LLM-agent interaction module cooperates with the extended query module to obtain node attribute names using a pre-trained large model to enhance interaction with the knowledge graph;

[0063] The LLM-Agent interaction module is a module for implementing natural language processing and dialogue management, and is usually used to build a dialogue system. This module implements interaction with users by converting natural language input into a form that can be processed by a computer and generating corresponding responses based on the dialogue strategy designed by the system. The LLM-Agent interaction module can include sub-modules such as speech recognition, natural language understanding, dialogue management, and natural language generation. Through the collaborative work of these sub-modules, the functions of the intelligent dialogue system are realized.

[0064] A memory prompt setting module is used to set a specific prompt template, simplify the historical question and answer information and convert it into a vector form. The vector database is connected to the memory prompt setting module and is used to store the vectorized historical question and answer information.

[0065] The Memory Prompt setting module is a tool for users to create memory prompts for specific topics or situations. Users can choose different setting options according to their needs, such as selecting the topic, keywords, time range, etc. of the memory prompt. Through this module, users can more easily remember specific information or events and improve memory efficiency.

[0066] The integrated language module is used to combine the above knowledge graph query results and historical information to generate the final question and answer results.

[0067] In the above system, an extended query module is set up to expand the potential query of the attributes of the nodes and relationships to be queried, and then the nodes and relationships are queried through the knowledge graph; the LLM-agent interactive auxiliary query obtains the node and relationship attributes corresponding to the queried entity, and then uses an interactive method to interact with the knowledge graph to complete the graph query; the knowledge graph query content and historical information are transmitted to the big model to complete a question and answer. Compared with the existing scheme, the semantic consistency in the process of data fusion of knowledge graph construction and retrieval is guaranteed, thereby improving the accuracy of the constructed knowledge graph, and further solving the problem that the existing scheme may have semantic inconsistencies in the process of data fusion of knowledge graph construction and retrieval, resulting in poor accuracy of the constructed knowledge graph.

[0068] The principle of this system is as follows Figure 1 As shown in the figure, LLM stands for "Language Model". In natural language processing, a language model is a model used to evaluate the probability of a sentence or text, usually used to predict the next word or sentence. The LLM in the LLM-Agent interaction module refers to the interaction module between the agent and the user based on the language model. This module can analyze and understand the text input by the user and generate corresponding responses.

[0069] In one embodiment of the present application, the private domain database ensures the real-time and accuracy of knowledge in a specific field through a periodic or event-driven update strategy; the extended query module generates an extended query based on the large model that can identify and locate entities and relationships in the knowledge graph, accurate to the node and relationship attribute values.

[0070] Specifically, for extended queries on attribute values, in order to better query the knowledge graph and accurately identify the attribute values ​​of nodes and relationships, separate queries are generated for node and relationship attribute values.

[0071] In one embodiment of the present application, the above-mentioned LLM-agent interaction module parses natural language queries and executes Cypher query statements to enhance the interaction with the knowledge graph.

[0072] Specifically, Cypher is a language for querying graph databases that is similar to SQL but is specifically designed for graph databases, thereby enhancing interaction with knowledge graphs.

[0073] In one embodiment of the present application, the Prompt template set by the above-mentioned memory Prompt setting module is used to simplify historical question and answer information, realize vector form conversion, enhance the relevance of multiple rounds of questions and answers, and ensure the accuracy of the question and answer assistant's answers.

[0074] In this embodiment, a knowledge graph enhanced retrieval application method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0075] Figure 2 is a flow chart of a knowledge graph enhanced retrieval application method provided according to an embodiment of the present application. Figure 2 As shown, the method comprises the following steps:

[0076] Step S101, receiving a user query request and generating an extended query based on a large model;

[0077] The specific implementation method of receiving user query requests and generating extended queries based on the large model is as follows:

[0078] According to the query target of the user query request, a query statement is constructed to search for relevant nodes and relationships in the knowledge graph, wherein the query statement includes query keywords, condition restrictions, and logical operators;

[0079] Input the constructed query statement into the knowledge graph for query, wherein the above knowledge graph matches relevant nodes and relationships according to the query statement and returns the query result;

[0080] According to the query results returned by the knowledge graph, the nodes and relationships with the highest relevance to the query target are filtered out.

[0081] In order to better perform knowledge graph queries, the query nodes and relationships must be accurately located. The method of generating queries separately for knowledge graph nodes and relationships is adopted. For questions such as "What are the reasons for high RRC reconstruction?", the large language model will be prompted to perform potential extended queries on Description, Reason and Solution nodes and Description-Reason, Description-Solution relationships.

[0082] The Description node describes the specific details of the problem or situation, the Reason node explains why the problem or situation occurred, and the Solution node proposes a solution or suggestion to solve the problem or improve the situation. The Description-Reason relationship is the association between the description of the problem or situation and the reason for its occurrence, describing the specific details and causes of the problem. The Description-Solution relationship is the association between the description of the problem or situation and the solution, describing the specific details that need to be solved and the solution.

[0083] In one embodiment of the present application, in the process of generating an extended query based on a large model, the method further includes:

[0084] The large model is trained by using deep learning technology. The large model can automatically learn the relationship between the attribute name and value of the node. In the query process, according to the attribute value of a node input, the large model encodes the query value according to the knowledge obtained by training, and converts the code into a representation form that can be recognized by the large model.

[0085] The above large model performs relevance matching based on the encoded query value and the pre-learned node attribute name;

[0086] Get the node attribute name that is most relevant to the query value in the model returned by the large model.

[0087] Specifically, after the extended query is generated, the node query obtains the attribute name of the node through the pre-trained encoder large model. The large model encodes the query value and performs correlation matching through the encoded node attribute name. The queried node includes the node type and attribute name. The prompt for the node extended query is shown in Table 1.

[0088] Table 1

[0089]

[0090] Step S102, retrieving matching entities and relationships from the public domain database and the private domain database;

[0091] The establishment of a private domain database is a sign that distinguishes this field from the open source field. The data source of the public domain database is a collection of excellent cases. Experts select excellent cases and manually mark them through professional operation and maintenance personnel, and save the marked entities to the wireless knowledge graph library. The data sources of the private domain database are in various forms: text, images, audio data, work orders handled by operation and maintenance personnel, group release documents, and important files are the data sources of the private domain database. In order to ensure the timeliness of the data, the private domain database is updated monthly. At the same time, in order to ensure the quality of the private domain database, when updating the database, some data with high similarity to the database (for example, higher than the similarity threshold) is deleted.

[0092] In one embodiment of the present application, in the process of retrieving matching entities and relationships from the public domain database and the private domain database, the method further includes:

[0093] Screening and matching data in the public domain database and the private domain database according to preset conditions to find the relationship between the data;

[0094] When performing a relational query, first query the associated nodes to find the source or target of the data;

[0095] Connecting the above-mentioned associated nodes through an association operation to establish a relationship between the above-mentioned associated nodes;

[0096] The associated query result is obtained, wherein the associated query result includes data that meets the preset conditions, the attribute name and attribute value of each associated node.

[0097] Specifically, by querying the associated nodes first, unnecessary data scanning and screening can be reduced, thereby improving query efficiency; ensuring data accuracy: by establishing the relationship between associated nodes, the data source in the query results can be ensured to be accurate, avoiding data duplication or errors; the associated query results include the attribute names and attribute values ​​of each associated node, which can more intuitively show the relationship between the data and improve the readability and comprehensibility of the data; through the associated query, data that meets the requirements can be obtained according to the preset conditions, meeting business needs, and facilitating further analysis and processing; after the node expansion query, a relationship query is performed, and the queried relationships and associated nodes are used to answer questions. The queried relationships include node attribute names, and the number is indefinite. The prompt for the relationship expansion query is shown in Table 2.

[0098] Table 2

[0099]

[0100]

[0101] RRC high reconstruction refers to the process of high-level reconstruction and recovery of affected areas after a major disaster or accident. This reconstruction usually includes repairing damaged infrastructure, rebuilding houses and buildings, and rebuilding communities and economies. RRC high reconstruction aims to help disaster-stricken areas return to normal life as soon as possible and improve their ability to resist future disasters.

[0102] Step S103, setting a target Prompt template to convert the historical question and answer information into a vector form, obtain vectorized information, and store the vectorized information in a vector database;

[0103] In one embodiment of the present application, in the process of setting the target Prompt template, the method further includes:

[0104] Asking questions to the language model-based dialogue system in a natural language manner, wherein the language model-based dialogue system uses the pre-trained model and the knowledge graph to perform interactive queries based on the questions;

[0105] Receive the interactive query result returned by the language model-based dialogue system.

[0106] Specifically, the dialogue system based on the language model can quickly understand the questions raised by the user, and query through the pre-trained model and knowledge graph, so as to quickly give accurate answers, saving the user's time; the dialogue system based on the language model can perform interactive queries based on the user's questions and the information in the knowledge graph to ensure that the answers provided are accurate and enhance the credibility of the information; by interactively querying the dialogue system based on the language model, users can more conveniently obtain the required information, improving the user experience and satisfaction; the dialogue system based on the language model combines the pre-trained model and the knowledge graph to provide broader and deeper knowledge, helping users to understand information in more fields; the dialogue system based on the language model can perform personalized interactive queries based on the user's questions and needs, and provide users with customized services and solutions.

[0107] After obtaining the nodes and relationships related to the question, LLM-agent is used to perform interactive queries with the knowledge graph library, which facilitates viewing the specific query process and locating the problem when it occurs. The prompts used for LLM-agent interactive auxiliary queries are shown in Table 3.

[0108] Table 3

[0109]

[0110]

[0111] python_cypher is a cryptographic algorithm or tool implemented in the Python programming language.

[0112] Markdown is a lightweight markup language that simplifies text layout and formatting. The format of Markdown is simple and easy to understand. You can use some simple symbols and tags to achieve basic layout effects, such as titles, lists, links, pictures, etc.

[0113] In one embodiment of the present application, the historical question and answer information is converted into a vector form to obtain vectorized information, including:

[0114] Simplify the historical question and answer information to obtain simplified question and answer information;

[0115] The simplified questions and answers are converted into vector form to obtain the vectorized information, and the vectorized information is stored in the Qdrant vector library to optimize storage and query efficiency.

[0116] Specifically, the Qdrant vector database is used to store historical questions and answers, and targeted prompt templates are set to enhance the relevance of multiple rounds of questions and answers and ensure the accuracy of the question and answer assistant's answers.

[0117] Step S104, integrating the search results and historical information to generate the final question-and-answer results.

[0118] In the above steps, an extended query is generated based on the big model for the nodes and relationships to be queried, the potential query of their attributes is expanded, and then the nodes and relationships are queried through the knowledge graph; matching entities and relationships are retrieved from the public domain database and the private domain database, and after obtaining the node and relationship attributes corresponding to the queried entity, the graph query is completed by interacting with the knowledge graph in an interactive manner; the knowledge graph query content and historical information are delivered to the big model to complete a question and answer, and the target Prompt template is set to convert the historical question and answer information into a vector form to obtain vectorized information, and the above vectorized information is stored in the vector database to enhance the relevance of multiple rounds of questions and answers and ensure the accuracy of the question and answer assistant's answers. Compared with the existing scheme, the semantic consistency in the process of data fusion of knowledge graph construction and retrieval is guaranteed, thereby improving the accuracy of the constructed knowledge graph, and further solving the problem that the existing scheme may have semantic inconsistencies in the process of data fusion of knowledge graph construction and retrieval, resulting in poor accuracy of the constructed knowledge graph.

[0119] In order to better query the knowledge graph and accurately identify the attribute values ​​of nodes and relationships, this application expands a single query question into node query and relationship query; uses a large model and prompt (prompt is usually defined as a prompt or instruction to stimulate thinking or action) to generate extended queries for nodes and relationships to enhance the accuracy of knowledge graph queries; LLM-agent interactive auxiliary query obtains the node and relationship attributes corresponding to the queried entity, and then uses an interactive tool calling method to interact with the knowledge graph to complete the graph query; in order to enhance the accuracy of multiple rounds of question and answer, a special memory prompt is designed to streamline the information of the previous round, vectorize it and store it in the Qdrant vector library, and an update prompt is designed to update the memory vector and play the role of deleting memory.

[0120] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0121] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute the knowledge graph enhanced retrieval application method.

[0122] An embodiment of the present invention provides a processor, which is used to run a program, wherein the knowledge graph enhanced retrieval application method is executed when the program is running.

[0123] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.

[0124] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0125] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0126] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0127] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0128] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0129] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0130] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0131] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0132] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:

[0133] 1) The knowledge graph enhanced retrieval application system of the present application, by setting up an extended query module for the nodes and relationships to be queried, expands the potential query of their attributes, and then queries the nodes and relationships through the knowledge graph; the LLM-agent interactive auxiliary query obtains the node and relationship attributes corresponding to the queried entity, and then uses an interactive method to interact with the knowledge graph to complete the graph query; the knowledge graph query content and historical information are transmitted to the big model to complete a question and answer. Compared with the existing solutions, the semantic consistency in the process of knowledge graph construction and retrieval data fusion is guaranteed, thereby improving the accuracy of the constructed knowledge graph, and further solving the problem that the existing solutions may have semantic inconsistencies in the process of knowledge graph construction and retrieval data fusion, resulting in poor accuracy of the constructed knowledge graph.

[0134] 2) The knowledge graph enhanced retrieval application method of the present application generates an extended query based on the big model for the nodes and relationships to be queried, expands the potential query of their attributes, and then queries the nodes and relationships through the knowledge graph; retrieves matching entities and relationships from the public domain database and the private domain database, obtains the node and relationship attributes corresponding to the queried entity, and then uses an interactive method to interact with the knowledge graph to complete the graph query; delivers the knowledge graph query content and historical information to the big model, completes a question and answer, sets the target Prompt template to convert the historical question and answer information into a vector form, obtains vectorized information, and stores the above vectorized information in the vector database, enhances the relevance of multiple rounds of questions and answers, and ensures the accuracy of the question and answer assistant's answers. Compared with the existing scheme, it ensures the semantic consistency in the process of knowledge graph construction and retrieval data fusion, thereby improving the accuracy of the constructed knowledge graph, and further solves the problem that the existing scheme may have semantic inconsistencies in the process of knowledge graph construction and retrieval data fusion, resulting in poor accuracy of the constructed knowledge graph.

[0135] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A knowledge graph enhanced retrieval application system, characterized in that: include: Public domain databases that store pre-built knowledge graphs; Private domain database, used to update knowledge in specific fields in real time; An extended query module, used to generate an extended query system for the node and relationship attribute values ​​of the knowledge graph based on the big model; An LLM-agent interaction module, wherein the LLM-agent interaction module cooperates with the extended query module to obtain node attribute names using a pre-trained large model to enhance interaction with the knowledge graph; A memory prompt setting module is used to set a specific prompt template, simplify the historical question and answer information and convert it into a vector form, and a vector database is connected to the memory prompt setting module to store the vectorized historical question and answer information; The integrated language module is used to combine the knowledge graph query results and historical information to generate the final question and answer results.

2. The system according to claim 1, characterized in that The private domain database ensures the real-time and accuracy of knowledge in a specific field through a regular or event-driven update strategy; the extended query module generates extended queries based on the large model, which can identify and locate entities and relationships in the knowledge graph, accurate to the node and relationship attribute values.

3. The system according to claim 1, characterized in that The LLM-agent interaction module parses natural language queries and executes Cypher query statements to enhance the interaction with the knowledge graph.

4. The system according to claim 1, characterized in that The Prompt template set by the memory Prompt setting module is used to simplify the historical question and answer information and realize the conversion in vector form.

5. A knowledge graph enhanced retrieval application method applied to the system of any one of claims 1 to 4, characterized in that: include: Receive user query requests and generate extended queries based on the large model; Retrieve matching entities and relationships from public and private databases; Setting a target Prompt template to convert historical question and answer information into a vector form, obtaining vectorized information, and storing the vectorized information in a vector database; Integrate retrieval results and historical information to generate the final question-answering results.

6. The method according to claim 5, characterized in that Receive user query requests and generate extended queries based on the large model, including: According to the query target of the user query request, a query statement is constructed to search for relevant nodes and relationships in the knowledge graph, wherein the query statement includes query keywords, condition restrictions, and logical operators; Input the constructed query statement into the knowledge graph for query, wherein the knowledge graph matches relevant nodes and relationships according to the query statement and returns the query result; According to the query results returned by the knowledge graph, the nodes and relationships with the highest relevance to the query target are filtered out.

7. The method according to claim 5, characterized in that In the process of generating an extended query based on the large model, the method further includes: The large model is trained using deep learning technology, and the large model can automatically learn the relationship between the attribute name and value of the node. In the query process, according to the attribute value of a node input, the large model encodes the query value according to the knowledge obtained from the training, and converts the code into a representation form that the large model can recognize; The large model performs correlation matching based on the encoded query value and the pre-learned node attribute name; Obtain the node attribute name of the model returned by the large model that has the highest correlation with the query value.

8. The method according to claim 5, characterized in that In the process of retrieving matching entities and relationships from the public domain database and the private domain database, the method further includes: Screening and matching data in the public domain database and the private domain database according to preset conditions to find the relationship between the data; When performing a relational query, first query the associated nodes to find the source or target of the data; Connecting the associated nodes through an association operation to establish a relationship between the associated nodes; The associated query result is obtained, wherein the associated query result includes data that meets the preset conditions, the attribute name and attribute value of each associated node.

9. The method according to claim 5, characterized in that In the process of setting the target Prompt template, the method further includes: Asking questions to the language model-based dialogue system in a natural language manner, wherein the language model-based dialogue system uses a pre-trained model and the knowledge graph to perform interactive queries according to the questions; Receive an interactive query result returned by the language model-based dialogue system.

10. The method according to claim 5, characterized in that Convert historical question-and-answer information into vector form to obtain vectorized information, including: Simplify the historical question and answer information to obtain simplified question and answer information; The simplified question and answer is converted into a vector form to obtain the vectorized information, and the vectorized information is stored in the Qdrant vector library to optimize storage and query efficiency.

Citation Information

Cited By

  • Decision-making auxiliary method and device based on large model

    CN120654795A

  • Building physical fuzzy query method based on large language model and medium

    CN121579558A

  • Multipath mixed knowledge retrieval enhancement generation method and system applied to vertical field

    CN122152969A