Retrieval enhancement generation method and device based on knowledge graph, equipment and medium

By combining knowledge graphs and large language models, the problem of insufficient accuracy of vectorized RAG libraries when dealing with complex knowledge queries is solved, achieving deeper knowledge understanding and more accurate answer generation.

CN120277206AActive Publication Date: 2025-07-08ZHONGDIAN DATA IND CO LTD

Patent Information

Application Number
CN202510766359.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The existing vectorization-based RAG library is insufficiently accurate when processing inter-entity calculations and internal logical relationship queries, and cannot meet complex knowledge needs.

Method used

The search-enhanced generation method based on knowledge graph is adopted. By obtaining knowledge graph patterns and querying problem texts, pre-processing and pattern matching is used for large language models, query statements are generated and querying in the knowledge graph database to obtain query results.

Benefits of technology

It improves the accuracy of the search results, can handle inter-entity calculations and internal logical relationship queries, and broadens the scope of application of RAG technology in complex knowledge processing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277206A_ABST
    Figure CN120277206A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a retrieval enhancement generation method and device based on a knowledge graph, equipment and a medium, and relates to the technical field of artificial intelligence. Processing the knowledge graph mode to obtain a target knowledge graph mode, and preprocessing the query question text to obtain a target question text; inputting the target question text and the target knowledge graph mode into a preset large language model for processing, and obtaining mode information matched with the target question text from the target knowledge graph mode; and generating a query statement based on the mode information and the target problem text, and performing query in a knowledge graph database corresponding to the knowledge graph mode based on the query statement to obtain a query result. By adopting the technical scheme, the semantics and the background of the problem can be deeply understood through the combination of the knowledge graph mode and the large language model, so that a more accurate query statement is generated, complex knowledge query is supported, and the retrieval accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a retrieval-enhanced generation method, apparatus, device, and medium based on a knowledge graph. Background Art

[0002] RAG (Retrieval-Augmented Generation) is a technology that combines external knowledge retrieval and text generation to improve the quality and accuracy of text generation. Currently, the method for constructing RAG is mainly based on a vectorized RAG library of texts. By converting text data into vector form, it is stored in a vector database; during retrieval, by calculating the similarity between the query vector and the vectors in the vector database, the most relevant text fragments are returned, and then text generation is performed by combining these fragments; for example, in a question-and-answer system, when a user inputs a question, the question-and-answer system vectorizes the question and retrieves relevant answer fragments in the RAG library, and then generates a complete answer.

[0003] The existing vector-based RAG library has limitations. It can only complete content retrieval at the literal and semantic levels and cannot further query or calculate results based on the existing description content and data; for example, when it is necessary to query the calculation relationship between entities (such as the merger time of two companies, the calculation of kinship between people, etc.), or when it involves internal logical relationship queries (such as causal relationships, conditional relationships, etc.), relying solely on the vector RAG library for retrieval often cannot answer or answer accurately. Summary of the Invention

[0004] To solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a retrieval-enhanced generation method, apparatus, device, and medium based on a knowledge graph.

[0005] An embodiment of the present disclosure provides a retrieval-enhanced generation method based on a knowledge graph, including: obtaining a knowledge graph schema and a query question text; processing the knowledge graph schema according to a preset format to obtain a target knowledge graph schema, and preprocessing the query question text to obtain a target question text; inputting the target question text and the target knowledge graph schema into a preset large language model for processing, and obtaining pattern information matching the target question text from the target knowledge graph schema; generating a query statement based on the pattern information and the target question text, and querying in a knowledge graph database corresponding to the knowledge graph schema based on the query statement to obtain a query result.

[0006] Optionally, processing the knowledge graph schema according to a preset format to obtain a target knowledge graph schema includes: obtaining predefined node identifiers and node relationship attribute identifiers; identifying nodes and relationships between nodes in the knowledge graph schema based on the node identifiers and the node relationship attribute identifiers to obtain the target knowledge graph schema.

[0007] Optionally, processing the knowledge graph schema according to a preset format to obtain a target knowledge graph schema includes: generating metadata based on each node, the relationship type between nodes, constraint conditions, and hierarchical conditions in the knowledge graph schema to obtain the target knowledge graph schema.

[0008] Optionally, preprocessing the query problem text to obtain a target problem text includes: deleting stop words in the query problem text to obtain a problem text to be processed; performing part-of-speech tagging on the problem text to be processed to obtain the target problem text.

[0009] Optionally, inputting the target problem text and the target knowledge graph schema into a preset large language model for processing, and obtaining pattern information matching the target problem text from the target knowledge graph schema includes: the large language model traversing the target knowledge graph schema based on the target problem text, calculating the similarity between the target problem text and each node in the target knowledge graph schema based on a preset semantic matching model, and using nodes with a similarity greater than or equal to a preset similarity threshold as target nodes; taking all the target nodes and the relationships between the target nodes as the pattern information.

[0010] Optionally, the method further includes: obtaining the problem category of the target problem text; adjusting the model parameters of the semantic matching model and the similarity threshold based on the problem category.

[0011] Optionally, generating a query statement based on the pattern information and the target problem text includes: performing named entity recognition on the target problem text to obtain at least one entity; mapping the at least one entity to the pattern information and generating the query statement based on a preset query logic.

[0012] Optionally, querying in a knowledge graph database corresponding to the knowledge graph schema based on the query statement to obtain a query result includes: sending the query statement to the knowledge graph database to query and obtain query entities, query entity relationships, and query entity attributes; integrating the query entities, the query entity relationships, and the query entity attributes according to a preset processing rule to obtain the query result.

[0013] An embodiment of the present disclosure also provides a retrieval enhanced generation device based on a knowledge graph. The device includes: an acquisition module for acquiring a knowledge graph schema and a query problem text; a first processing module for processing the knowledge graph schema in a preset format to obtain a target knowledge graph schema; a second processing module for preprocessing the query problem text to obtain a target problem text; an input module for inputting the target problem text and the target knowledge graph schema into a preset large language model for processing, and obtaining pattern information matching the target problem text from the target knowledge graph schema; a generation module for generating a query statement based on the pattern information and the target problem text; and a query module for querying in a knowledge graph database corresponding to the knowledge graph schema based on the query statement to obtain a query result.

[0014] An embodiment of the present disclosure also provides an electronic device. The electronic device includes: a processor; a memory for storing executable instructions executable by the processor; and the processor for reading the executable instructions from the memory and executing the instructions to implement the retrieval enhanced generation method based on a knowledge graph provided by an embodiment of the present disclosure.

[0015] An embodiment of the present disclosure also provides a computer-readable storage medium. The storage medium stores a computer program, and the computer program is used to execute the retrieval enhanced generation method provided by an embodiment of the present disclosure.

[0016] An embodiment of the present disclosure also provides a computer program product, including a computer program, wherein the computer program, when executed by a processor, implements the retrieval enhanced generation method provided by an embodiment of the present application.

[0017] The technical solutions provided by the embodiments of the present disclosure have the following advantages compared with the prior art: The retrieval-enhanced generation solution based on a knowledge graph provided by the embodiments of the present disclosure obtains a knowledge graph schema and query problem text; processes the knowledge graph schema in a preset format to obtain a target knowledge graph schema, and preprocesses the query problem text to obtain a target problem text; inputs the target problem text and the target knowledge graph schema into a preset large language model for processing, and obtains schema information matching the target problem text from the target knowledge graph schema; generates a query statement based on the schema information and the target problem text, and queries in the knowledge graph database corresponding to the knowledge graph schema based on the query statement to obtain a query result. By combining the knowledge graph schema with the large language model, it is possible to more deeply understand the semantics and background of the problem, thereby generating more accurate query statements and improving the quality of retrieval results; it can handle complex knowledge requirements such as entity calculations and internal logical relationship queries, making up for the deficiencies of traditional vector RAG libraries; it makes full use of the rich knowledge structure and relationship information in the knowledge graph, can dig out deeper knowledge, and provide more comprehensive and accurate answers. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and that the original elements and elements are not necessarily drawn to scale.

[0019] Figure 1 It is a schematic flowchart of a method for retrieval-enhanced generation based on a knowledge graph provided by an embodiment of the present disclosure; Figure 2 It is a schematic flowchart of another method for retrieval-enhanced generation based on a knowledge graph provided by an embodiment of the present disclosure; Figure 3 It is a schematic structural diagram of a device for retrieval-enhanced generation based on a knowledge graph provided by an embodiment of the present disclosure; Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0021] It should be understood that the various steps described in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0022] As used herein, the term "comprising" and its variations are open-ended, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0023] It should be noted that the concepts such as "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0024] It should be noted that the modifications of "one" and "a plurality" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0025] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0026] Specifically, the existing technical solution, the existing vector-based RAG library has limitations; it can only complete content retrieval at the literal and semantic levels, and cannot further query or calculate results based on the existing description content and data; that is to say, relying solely on the vector RAG library retrieval, it is often impossible to answer or accurately answer.

[0027] To address the above problems, how to improve the accuracy of RAG retrieval, and how to implement problems such as entity calculation and internal logical relationship query that cannot be answered by the vector RAG library, so as to broaden the application scope of the RAG technology in complex knowledge processing scenarios; the retrieval enhancement generation method based on the knowledge graph proposed in the embodiments of the present disclosure can improve the accuracy of RAG retrieval, and at the same time add functions such as entity calculation and internal logical relationship query that cannot be achieved by the vector RAG library, and can handle more complex knowledge query requirements.

[0028] Figure 1The following is a schematic flowchart of a retrieval enhanced generation method based on a knowledge graph provided by an embodiment of the present disclosure. This method can be executed by a retrieval enhanced generation device based on a knowledge graph, where the device can be implemented by software and / or hardware and is generally integrated in an electronic device. As Figure 1 shown, the method includes: Step 101, obtain a knowledge graph schema and a query problem text.

[0029] In an embodiment of the present disclosure, a knowledge graph schema refers to in a knowledge graph, where the structure, type, and relationship of knowledge are defined through the knowledge graph schema, and it is used to describe the organization method of data. The knowledge graph schema is usually preset based on a knowledge system in a specific field. In an embodiment of the present disclosure, the knowledge graph schema can be input according to actual query needs.

[0030] In an embodiment of the present disclosure, the query problem text can be input manually or by voice according to actual query needs.

[0031] Step 102, process the knowledge graph schema according to a preset format to obtain a target knowledge graph schema, and preprocess the query problem text to obtain a target problem text.

[0032] In an embodiment of the present disclosure, after obtaining the knowledge graph schema, there are many ways to process the knowledge graph schema according to a preset format to obtain a target knowledge graph schema. In one embodiment, obtain predefined node identifiers and node relationship attribute identifiers; based on the node identifiers and node relationship attribute identifiers, identify the nodes and the relationships between the nodes in the knowledge graph schema to obtain a target knowledge graph schema; in another embodiment, generate metadata based on each node, the relationship type, constraint conditions, and hierarchical conditions between nodes in the knowledge graph schema to obtain a target knowledge graph schema.

[0033] Furthermore, there are also many ways to preprocess the query problem text to obtain a target problem text. In some embodiments, delete the stop words in the query problem text to obtain a problem text to be processed; perform part-of-speech tagging on the problem text to be processed to obtain a target problem text; in other embodiments, segment the query problem text to obtain multiple keywords, screen the multiple keywords to obtain target keywords, then perform part-of-speech tagging on the target keywords, and form the tagged target keywords into a target problem text.

[0034] Step 103, input the target problem text and the target knowledge graph schema into a preset large language model for processing, and obtain pattern information matching the target problem text from the target knowledge graph schema.

[0035] In the embodiments of the present disclosure, the preset large language model refers to a large-scale language model based on deep learning that can understand and generate natural language.

[0036] In the embodiments of the present disclosure, inputting the target problem text and the target knowledge graph schema into the preset large language model for processing, and obtaining the schema information matching the target problem text from the target knowledge graph schema can be understood as determining one or more target nodes and the relationships between the target nodes from multiple nodes in the target knowledge graph schema for the target problem text through the large language model, so as to determine the schema information matching the target problem text.

[0037] Step 104: Generate a query statement based on the schema information and the target problem text, and perform a query in the knowledge graph database corresponding to the knowledge graph schema based on the query statement to obtain a query result.

[0038] In the embodiments of the present disclosure, after obtaining the schema information matching the target problem text, a query statement can be generated based on the schema information and the target problem text. Specifically, named entity recognition can be performed on the target problem text to obtain at least one entity, map the at least one entity to the schema information, and generate a query statement based on the preset query logic.

[0039] Furthermore, perform a query in the knowledge graph database corresponding to the knowledge graph schema based on the query statement to obtain a query result. Specifically, query entities, query entity relationships, query entity attributes, etc. can be obtained as the query result based on the query statement in the knowledge graph database.

[0040] The retrieval enhanced generation solution based on the knowledge graph provided by the embodiments of the present disclosure includes: obtaining a knowledge graph schema and a query problem text; processing the knowledge graph schema in a preset format to obtain a target knowledge graph schema, and preprocessing the query problem text to obtain a target problem text; inputting the target problem text and the target knowledge graph schema into the preset large language model for processing to obtain schema information matching the target problem text from the target knowledge graph schema; generating a query statement based on the schema information and the target problem text, and performing a query in the knowledge graph database corresponding to the knowledge graph schema based on the query statement to obtain a query result. By combining the knowledge graph schema with the large language model, it is possible to more deeply understand the semantics and background of the problem, thereby generating more accurate query statements and improving the quality of retrieval results; it can handle complex knowledge requirements such as entity calculations and internal logical relationship queries, making up for the deficiencies of traditional vector RAG libraries; making full use of the rich knowledge structure and relationship information in the knowledge graph, it is possible to dig out deeper knowledge and provide more comprehensive and accurate answers.

[0041] Figure 2The flowchart of another retrieval - enhanced generation method based on a knowledge graph provided by an embodiment of the present disclosure. On the basis of the above - mentioned embodiment, the retrieval - enhanced generation method based on the knowledge graph is further optimized. As Figure 2 shown, the method further includes: Step 201, obtain the knowledge graph schema and the query problem text.

[0042] Step 202, obtain the predefined node identifiers and node - relationship attribute identifiers, and based on the node identifiers and node - relationship attribute identifiers, identify the nodes and the relationships between the nodes in the knowledge graph schema to obtain the target knowledge graph schema.

[0043] Step 203, generate metadata based on each node, the relationship type between nodes, constraint conditions, and hierarchical conditions in the knowledge graph schema to obtain the target knowledge graph schema.

[0044] After step 201, step 202 or step 203 can be executed, Figure 2 This is only an example, and the present disclosure makes no specific restrictions.

[0045] Specifically, when obtaining the knowledge graph schema and the query problem text, before inputting the knowledge graph schema into the large - language model, it is necessary to normalize the knowledge graph schema to ensure that the knowledge graph schema conforms to the format that the large - language model can accept. For example, the nodes and relationships in the knowledge graph schema are annotated with a specific markup language so that the large - language model can clearly identify them.

[0046] As an example, design a specific JSON - LD extension format; while retaining the generality of JSON - LD, this format adds specific annotations for the knowledge graph schema; for example, through the custom "@schemaNode" and "@schemaRelation" attributes, the nodes and relationships in the knowledge graph schema are clearly identified; as another example, develop a metadata description mechanism for describing the structure and semantics of the knowledge graph schema; this metadata not only includes the type information of the nodes and relationships in the knowledge graph schema, but also includes their hierarchical relationships and constraint conditions, such as the constraint conditions between nodes, for example, for the same drug for the same disease, it cannot be both an indication drug and a contraindicated drug at the same time; thus, the large - language model can utilize this metadata to quickly understand the overall structure and meaning of the knowledge graph schema, and thus perform subsequent processing more effectively.

[0047] Step 204, delete the stop words in the query problem text to obtain the to - be - processed problem text, and perform part - of - speech tagging on the to - be - processed problem text to obtain the target problem text.

[0048] Specifically, the input query problem text also needs to be preprocessed, including operations such as removing stop words and performing part-of-speech tagging, to improve the accuracy of the large language model's understanding of the problem. Then, the preprocessed target problem text and the normalized target knowledge graph pattern are passed to the large language model according to a specific input format (such as the format of general prompts, such as role, requirements, limitations, tasks, inputs, outputs, etc.).

[0049] Step 205: The large language model traverses the target knowledge graph pattern based on the target problem text, calculates the similarity between the target problem text and each node in the target knowledge graph pattern based on a preset semantic matching model, and takes the nodes with similarity greater than or equal to the preset similarity threshold as target nodes, and takes all target nodes and the relationships between the target nodes as pattern information.

[0050] Specifically, the large language model filters the knowledge graph pattern information. After receiving the input target problem text and target knowledge graph pattern, the large language model will first perform semantic understanding and analysis on the target problem text; it will use its pre-trained language understanding ability to judge the key concepts and relationships involved in the problem; among them, a concept is a generalization of instances, for example, the concepts of type 2 diabetes and coronary heart disease are "diseases". By giving the target knowledge graph pattern to the large language model, the large language model can understand the content from the conceptual level to the instance level.

[0051] Specifically, based on the understanding of the target problem text, the large language model will traverse the target knowledge graph pattern, and determine the nodes and relationships in the target knowledge graph pattern related to the target problem text through a series of semantic matching algorithms and attention mechanisms; for example, for a target problem text about "the product release time of XX company", the large language model will screen out the nodes related to the entity "XX company" from the target knowledge graph pattern, and the relationships related to "product release time", such as the attribute relationship of "release time". Among them, corresponding semantic matching algorithms and attention mechanisms are designed according to different business scenarios and application scenarios.

[0052] Specifically, in the screening process, in order to improve the accuracy and efficiency of screening, some optimization strategies will also be adopted; for example, a similarity threshold is set, and when a certain part of the target knowledge graph pattern (such as the word "data component", which is part of the target problem text "whether the data component is part of a national data infrastructure construction plan") has a semantic similarity with the problem exceeding the similarity threshold, it will be included in the screening results.

[0053] In some embodiments, the problem category of the target problem text is obtained, and the model parameters and similarity threshold of the semantic matching model are adjusted based on the problem category.

[0054] Specifically, the problem category can be understood as the complexity of the problem. By setting the model parameters and similarity thresholds of different semantic matching models for different problem categories, the processing accuracy is further ensured while improving the processing efficiency.

[0055] Specifically, according to the type and complexity of the target problem text, the screening parameters and rules are dynamically adjusted; for simple problems, a fast shallow matching strategy is adopted; for complex problems, a more in-depth reasoning and matching mechanism is enabled to ensure that relevant pattern information can be accurately screened in different scenarios; that is to say, if only the instance-level screening parameter rules are required for simple problems, and the concept level needs to be used for screening complex problems.

[0056] Among them, based on the semantic matching algorithm of deep learning, a neural network using the Transformer architecture is used to construct a semantic matching model; the semantic matching model takes the target problem text and the pattern fragments of the target knowledge graph pattern as inputs, and calculates the semantic similarity between the problem and the pattern fragments through the multi-head attention mechanism, so as to screen out the most relevant pattern information.

[0057] Therefore, the ontology reasoning mechanism (knowledge graph pattern) of the knowledge graph is introduced. During the screening process, not only direct semantic matching is considered, but also the ontology reasoning rules (knowledge graph pattern) of the knowledge graph are used. For example, if the "XX Company" is defined as a subclass of the "Technology Company" in the knowledge graph pattern, when the target problem text involves the "Technology Company", relevant pattern information related to the "XX Company" can also be included in the screening scope through ontology reasoning, improving the comprehensiveness of the screening.

[0058] Step 206: Perform named entity recognition on the target problem text to obtain at least one entity, map the at least one entity to the pattern information, and generate a query statement based on the preset query logic.

[0059] Specifically, according to the screened pattern information and the target problem text, the large language model will generate corresponding query statements; the generation process follows the syntax rules of the knowledge graph query language. Taking SPARQL as an example, it needs to accurately construct triple patterns.

[0060] More specifically, map the entities and relationships in the target problem text to the pattern information, that is, the nodes and relationships of the screened target knowledge graph; then, according to the structure of the target knowledge graph and the query requirements, determine the query logic, such as whether multi-hop queries are required, whether conditional filtering is required, etc.

[0061] For example, for the question of "the time when XX Company released mobile phone A", a SPARQL query statement similar to "SELECT?time WHERE {<XX Company><released product><mobile phone A>.<mobile phone A><release time>?time}" will be generated.

[0062] Step 207: Send the query statement to the knowledge graph database to query and obtain query entities, query entity relationships, and query entity attributes, and integrate the query entities, query entity relationships, and query entity attributes according to preset processing rules to obtain a query result.

[0063] Specifically, send the generated query statement to the knowledge graph database for execution; after receiving the query request, the knowledge graph database performs data retrieval in the knowledge graph database according to the requirements of the query statement.

[0064] Thus, after the retrieval is completed, the knowledge graph database will return a query result. These query results may be a series of entities, relationships, or attribute values. The large language model will process and integrate the returned results and convert them into answers in natural language form; among them, the integration is to give the query results to the large language model, and the context long window of the large language model can accept the text length.

[0065] When generating an answer, the large language model will also consider the integrity, accuracy, and readability of the answer. For example, if the query result contains multiple relevant pieces of information, the large language model will reasonably organize and sort this information to generate a smooth and well-organized answer to feedback to the user.

[0066] In summary, compared with the vector-based RAG library, the present disclosure has the following advantages: higher retrieval accuracy, by combining the knowledge graph schema with the large language model, it can more deeply understand the semantics and background of the question, thereby generating more accurate query statements and improving the quality of retrieval results; supporting complex knowledge queries, being able to handle complex knowledge requirements such as entity calculations and internal logical relationship queries, making up for the deficiencies of traditional vector RAG libraries and broadening the application fields of RAG technology; better knowledge utilization, making full use of the rich knowledge structure and relationship information in the knowledge graph, enabling the RAG system to dig out deeper knowledge and provide more comprehensive and accurate answers.

[0067] Figure 3 It is a schematic structural diagram of a retrieval enhanced generation device based on a knowledge graph provided by an embodiment of the present disclosure. This device can be implemented by software and / or hardware and is generally integrated in an electronic device. As Figure 3 shown, this device includes: An acquisition module 301, configured to acquire a knowledge graph schema and query problem text.

[0068] The first processing module 302 is configured to process the knowledge graph schema according to a preset format to obtain a target knowledge graph schema.

[0069] The second processing module 303 is configured to preprocess the query problem text to obtain a target problem text.

[0070] The input module 304 is configured to input the target problem text and the target knowledge graph schema into a preset large language model for processing, and obtain pattern information matching the target problem text from the target knowledge graph schema.

[0071] The generation module 305 is configured to generate a query statement based on the pattern information and the target problem text.

[0072] The query module 306 is configured to query in the knowledge graph database corresponding to the knowledge graph schema based on the query statement to obtain a query result.

[0073] Optionally, the first processing module 302 is specifically configured to: obtain predefined node identifiers and node relationship attribute identifiers; identify the nodes in the knowledge graph schema and the relationships between the nodes based on the node identifiers and the node relationship attribute identifiers to obtain the target knowledge graph schema.

[0074] Optionally, the first processing module 302 is specifically configured to: generate metadata based on each node, the relationship type between nodes, constraint conditions, and hierarchical conditions in the knowledge graph schema to obtain the target knowledge graph schema.

[0075] Optionally, the second processing module 303 is specifically configured to: delete stop words in the query problem text to obtain a problem text to be processed; perform part-of-speech tagging on the problem text to be processed to obtain the target problem text.

[0076] Optionally, the input module 304 is specifically configured to: the large language model traverses the target knowledge graph schema based on the target problem text, calculates the similarity between the target problem text and each node in the target knowledge graph schema based on a preset semantic matching model, and uses the nodes with similarity greater than or equal to a preset similarity threshold as target nodes; uses all the target nodes and the relationships between the target nodes as the pattern information.

[0077] Optionally, the apparatus further includes: an acquisition and adjustment module, configured to acquire the problem category of the target problem text, and adjust the model parameters of the semantic matching model and the similarity threshold based on the problem category.

[0078] Optionally, the generating module 305 is specifically configured to: perform named entity recognition on the target problem text to obtain at least one entity; map the at least one entity to the pattern information, and generate the query statement based on a preset query logic.

[0079] Optionally, the query module 306 is specifically configured to: send the query statement to the knowledge graph database to query and obtain query entities, query entity relationships, and query entity attributes; integrate the query entities, the query entity relationships, and the query entity attributes according to a preset processing rule to obtain the query result.

[0080] The retrieval enhanced generation device based on a knowledge graph provided by an embodiment of the present disclosure can execute the retrieval enhanced generation method based on a knowledge graph provided by any embodiment of the present disclosure, and has functional modules and beneficial effects corresponding to the execution of the method.

[0081] An embodiment of the present disclosure also provides a computer program product, including a computer program / instructions, which when executed by a processor, implement the retrieval enhanced generation method based on a knowledge graph provided by any embodiment of the present disclosure.

[0082] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Specifically refer to Figure 4 , which shows a schematic structural diagram of an electronic device 400 suitable for implementing the present disclosure. The electronic device 400 in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The electronic device shown is only an example, and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0083] As Figure 4 shown, the electronic device 400 may include a processing device 401 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in the ROM 402 (ROM is a read-only memory) or a program loaded from a storage device 408 into the RAM 403 (RAM is a random access memory). In the RAM 403, various programs and data required for the operation of the electronic device 400 are also stored. The processing device 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. The I / O interface 405 (I / O is input / output) is also connected to the bus 404.

[0084] Typically, the following devices can be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 can allow the electronic device 400 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 4 the electronic device 400 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.

[0085] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the knowledge graph-based retrieval enhancement generation method of the embodiment of the present disclosure are executed.

[0086] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0087] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0088] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; it can also exist separately without being assembled into the electronic device.

[0089] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The foregoing programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, execute as a stand-alone software package, execute partially on the user's computer and partially on a remote computer, or execute entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0090] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0091] The units described in the embodiments of the present disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation on the unit itself in some cases.

[0092] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.

[0093] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0094] According to one or more embodiments of the present disclosure, the present disclosure provides an electronic device, comprising: a processor; a memory for storing executable instructions of the processor; the processor, configured to read the executable instructions from the memory and execute the instructions to implement any one of the knowledge graph-based retrieval enhanced generation methods provided by the present disclosure.

[0095] According to one or more embodiments of the present disclosure, the present disclosure provides a computer-readable storage medium storing a computer program for executing any one of the knowledge graph-based retrieval enhanced generation methods provided by the present disclosure.

[0096] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, technical solutions formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in the present disclosure.

[0097] Moreover, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0098] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A retrieval augmented generation method based on a knowledge graph, characterized in that, Including: Obtain a knowledge graph schema and a query problem text; Process the knowledge graph schema according to a preset format to obtain a target knowledge graph schema, and preprocess the query problem text to obtain a target problem text; Input the target problem text and the target knowledge graph schema into a preset large language model for processing, and obtain pattern information matching the target problem text from the target knowledge graph schema; Generate a query statement based on the pattern information and the target problem text, and query in the knowledge graph database corresponding to the knowledge graph schema based on the query statement to obtain a query result.

2. The method according to claim 1, wherein The processing the knowledge graph schema according to a preset format to obtain a target knowledge graph schema includes: Obtain predefined node identifiers and node relationship attribute identifiers; Identify the nodes and the relationships between nodes in the knowledge graph schema based on the node identifiers and the node relationship attribute identifiers to obtain the target knowledge graph schema.

3. The method according to claim 1, wherein The processing the knowledge graph schema according to a preset format to obtain a target knowledge graph schema includes: Generate metadata based on each node, the relationship type between nodes, constraint conditions, and hierarchical conditions in the knowledge graph schema to obtain the target knowledge graph schema.

4. The method according to claim 1, wherein The preprocessing the query problem text to obtain a target problem text includes: Delete stop words in the query problem text to obtain a problem text to be processed; Perform part-of-speech tagging on the problem text to be processed to obtain the target problem text.

5. The method according to claim 1, characterized in that The inputting the target problem text and the target knowledge graph schema into a preset large language model for processing, and obtaining pattern information matching the target problem text from the target knowledge graph schema includes: The large language model traverses the target knowledge graph schema based on the target problem text, calculates the similarity between the target problem text and each node in the target knowledge graph schema based on a preset semantic matching model, and uses the nodes with similarity greater than or equal to a preset similarity threshold as target nodes; Use all the target nodes and the relationships between the target nodes as the pattern information.

6. The method according to claim 5, characterized in that, The method further includes: Obtain the problem category of the target problem text; Adjust the model parameters of the semantic matching model and the similarity threshold based on the problem category.

7. The method according to claim 1, wherein The generating a query statement based on the pattern information and the target problem text includes: Perform named entity recognition on the target problem text to obtain at least one entity; Map the at least one entity to the pattern information, and generate the query statement based on a preset query logic.

8. The method according to claim 1, characterized in that, The querying in the knowledge graph database corresponding to the knowledge graph schema based on the query statement to obtain a query result includes: Send the query statement to the knowledge graph database to query and obtain query entities, query entity relationships, and query entity attributes; Integrate the query entities, the query entity relationships, and the query entity attributes according to preset processing rules to obtain the query result.

9. An electronic device, characterized in that, The electronic device includes: Processor; A memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the knowledge graph-based retrieval enhanced generation method according to any one of claims 1-8 above.

10. A computer-readable storage medium, characterized in that, A storage medium stores a computer program, and the computer program is used to execute the knowledge graph-based retrieval enhanced generation method according to any one of claims 1-8 above.

Citation Information

Patent Citations

  • Retrieval enhancement method and device based on knowledge graph, equipment and medium

    CN117668157A

  • Retrieval enhancement generation system and method based on knowledge graph

    CN117973540A

  • Database query method and device, electronic equipment and nonvolatile storage medium

    CN119226315A

  • Large model and knowledge graph fusion method, application method and system

    CN119226529A

  • Cypher-stack type alignment generation method and device based on large language model

    CN120104110A

Cited By

  • Intelligent agent navigation effect enhancement method and equipment based on scene map matching and medium

    CN120740609A

  • Automatic construction method and system for dynamic mode knowledge graph

    CN120745784A

  • Retrieval enhancement generation method based on document knowledge base and knowledge graph

    CN120780849A

  • Search enhancement generation method based on document knowledge base and knowledge graph

    CN120780849B

  • Text generation graph query statement-based model training method and device, computer equipment, readable storage medium and program product

    CN120910309A