Text generation method and device, electronic equipment and readable medium
By constructing a lightweight knowledge graph and performing two-layer search, the problem of insufficient retrieval quality dependence and reasoning capabilities in search enhancement generation technology is solved, and high-quality text generation and multi-hop reasoning capabilities are achieved to adapt to user needs in different fields.
Patent Information
- Application Number
- CN202510454838.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-08-19
AI Technical Summary
The performance of search-enhanced generation technology is highly dependent on the search quality of the external knowledge base, resulting in inaccurate or uncorrelated text quality, and the basic RAG's inference ability is limited, making it impossible to handle complex multi-hop reasoning tasks.
By extracting content keywords of user text requirements information, a lightweight knowledge graph is constructed, a double-layer search based on content keywords and topic keywords is performed, and a similarity evaluation is combined to generate target requirements text.
It improves the quality of retrieval text, can handle complex multi-hop reasoning tasks, adapts to user problems in different fields, reduces the construction and maintenance costs of knowledge graphs, and improves the real-time and accuracy of text generation.
Smart Images

Figure CN120509478A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of language technology, and in particular to a text generation method, a text generation device, an electronic device, and a computer-readable medium. Background Art
[0002] Retrieval-Augmented Generation (RAG) is an AI technology that combines information retrieval and language generation. It retrieves relevant information about the user's desired text from an external knowledge base and uses this information as context input into a language model. The language model can then generate the desired text based on this contextual information, thereby enhancing the accuracy, relevance, and timeliness of the text generated by the language model.
[0003] However, the performance of retrieval-augmented generation techniques is highly dependent on the quality of retrieval of external knowledge bases. If the retrieved information is inaccurate or irrelevant, the quality of the text generated by the language model will be affected. Summary of the Invention
[0004] The embodiments of the present application provide a text generation method, device, electronic device, and computer-readable storage medium to address the problem that the performance of retrieval-enhanced generation technology is highly dependent on the retrieval quality of an external knowledge base. If the retrieved information is inaccurate or irrelevant, the quality of the text generated by the language model will be affected.
[0005] The present application discloses a method for generating text, including:
[0006] Extracting at least one content keyword from the user's text demand information;
[0007] Acquire a text to be retrieved associated with the text requirement information, and divide the text to be retrieved into at least one text block;
[0008] Constructing a knowledge graph of the at least one text block, and querying the knowledge graph based on the content keyword to obtain query information associated with the content keyword in the knowledge graph;
[0009] taking a text block corresponding to the query information in the at least one text block as a first target text block;
[0010] evaluating a similarity between the text requirement information and the at least one text block, and determining a second target text block in the at least one text block based on the similarity;
[0011] Based on at least one of the query information, the first target text block, the second target text block, and the text requirement information, a preset language model is used to generate the user's target requirement text.
[0012] Optionally, the content keywords include content entity keywords and / or subject keywords.
[0013] Optionally, constructing the knowledge graph of the at least one text block includes:
[0014] Extracting text entity keywords from the at least one text block, and determining association relationships between the text entity keywords based on the at least one text block;
[0015] Determining relationship keywords for describing the association relationship between the text entity keywords, and determining description text for describing the text entity keywords and the association relationship between the text entity keywords;
[0016] The knowledge graph is constructed based on the text entity keywords, the association relationship between the text entity keywords, the relationship keywords and at least one of the description texts.
[0017] Optionally, querying the knowledge graph based on the content keyword to obtain query information associated with the content keyword in the knowledge graph includes:
[0018] querying the knowledge graph based on the content entity keyword to obtain a target text entity keyword associated with the content entity keyword, and determining at least one of an association relationship, a relationship keyword, and a description text associated with the target text entity keyword based on the knowledge graph; and / or,
[0019] The knowledge graph is queried based on the subject keyword to obtain a target relationship keyword associated with the subject keyword, and at least one of a text entity keyword, an association relationship, and a description text associated with the target relationship keyword is determined based on the knowledge graph.
[0020] Optionally, generating the user's target requirement text using a preset language model based on at least one of the query information, the first target text block, the second target text block, and the text requirement information includes:
[0021] Merging the first target text block and the second target text block, and deleting duplicate text in the first target text block and the second target text block to obtain a third target text block;
[0022] The target requirement text is generated using the language model based on at least one of the query information, the third target text block, and the text requirement information.
[0023] Optionally, constructing the knowledge graph of the at least one text block includes:
[0024] The knowledge graph is constructed using a preset lightweight retrieval enhancement generation technology.
[0025] Optionally, the evaluating the similarity between the text requirement information and the at least one text block includes:
[0026] Mapping the text requirement information and the at least one text block respectively using a preset vector model to obtain a requirement information vector corresponding to the text requirement information and a text block vector corresponding to the text block;
[0027] A similarity calculation is performed on the demand information vector and the text block vector to obtain a similarity between the text demand information and the at least one text block.
[0028] The present application also discloses a text generation device, including:
[0029] A content keyword extraction module, configured to extract at least one content keyword from the user's textual demand information;
[0030] a segmentation module, configured to obtain a text to be retrieved associated with the text requirement information, and segment the text to be retrieved into at least one text block;
[0031] a knowledge graph construction module, configured to construct a knowledge graph for the at least one text block, and query the knowledge graph based on the content keywords to obtain query information associated with the content keywords in the knowledge graph;
[0032] The first target text block is used as a module, and is used to take the text block corresponding to the query information in the at least one text block as the first target text block;
[0033] a second target text block determining module, configured to evaluate the similarity between the text requirement information and the at least one text block, and determine a second target text block in the at least one text block based on the similarity;
[0034] The text generation module is used to generate the user's target requirement text using a preset language model based on the query information, the first target text block, the second target text block and at least one of the text requirement information.
[0035] Optionally, the content keywords include content entity keywords and / or subject keywords.
[0036] Optionally, the knowledge graph construction module includes:
[0037] a text entity keyword extraction submodule, configured to extract text entity keywords from the at least one text block and determine association relationships between the text entity keywords based on the at least one text block;
[0038] a relationship keyword determination submodule, configured to determine relationship keywords for describing the association relationship between the text entity keywords, and to determine description text for describing the text entity keywords and the association relationship between the text entity keywords;
[0039] The first knowledge graph construction submodule is used to construct the knowledge graph based on the text entity keywords, the association relationship between the text entity keywords, the relationship keywords and at least one of the description text.
[0040] Optionally, the knowledge graph construction module includes:
[0041] a query submodule, configured to query the knowledge graph based on the content entity keywords, obtain target text entity keywords associated with the content entity keywords, and determine at least one of an association relationship, a relationship keyword, and a description text associated with the target text entity keywords based on the knowledge graph; and / or,
[0042] The knowledge graph is queried based on the subject keyword to obtain a target relationship keyword associated with the subject keyword, and at least one of a text entity keyword, an association relationship, and a description text associated with the target relationship keyword is determined based on the knowledge graph.
[0043] Optionally, the text generation module includes:
[0044] a merging submodule, configured to merge the first target text block and the second target text block, and delete duplicate text in the first target text block and the second target text block, to obtain a third target text block;
[0045] The target requirement text generation submodule is configured to generate the target requirement text using the language model based on at least one of the query information, the third target text block, and the text requirement information.
[0046] Optionally, the knowledge graph construction module includes:
[0047] The second knowledge graph construction submodule is used to construct the knowledge graph using a preset lightweight retrieval enhancement generation technology.
[0048] Optionally, the second target text block determination module includes:
[0049] a mapping submodule, configured to map the text requirement information and the at least one text block respectively using a preset vector model to obtain a requirement information vector corresponding to the text requirement information and a text block vector corresponding to the text block;
[0050] The similarity calculation submodule is configured to perform similarity calculation on the demand information vector and the text block vector to obtain the similarity between the text demand information and the at least one text block.
[0051] The embodiment of the present application further discloses an electronic device, comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;
[0052] The memory is used to store computer programs;
[0053] The processor is used to implement the method described in the embodiment of the present application when executing the program stored in the memory.
[0054] The embodiments of the present application also disclose one or more computer-readable media having instructions stored thereon, which, when executed by one or more processors, enable the processors to perform the method described in the embodiments of the present application.
[0055] The embodiments of the present application include the following advantages:
[0056] In an embodiment of the present application, at least one content keyword is extracted from the user's text demand information; the text to be retrieved associated with the text demand information is obtained, and the text to be retrieved is divided into at least one text block; a knowledge graph of at least one text block is constructed, and the knowledge graph is queried based on the content keywords to obtain query information associated with the content keywords in the knowledge graph; the text block corresponding to the query information in at least one text block is used as the first target text block; the similarity between the text demand information and the at least one text block is evaluated, and based on the similarity, the second target text block in the at least one text block is determined; based on the query information, the first target text block, the second target text block and at least one of the text demand information, the user's target demand text is generated using a preset language model. Retrieving the text to be retrieved based on the knowledge graph improves the retrieval quality of the text to be retrieved, thereby improving the quality of the text generated by the language model. In addition, by searching the text to be retrieved based on the knowledge graph, query information related to the content keywords in the user's text demand information is obtained, and then the first target text block corresponding to the query information in at least one text block of the text to be retrieved is determined, and relevant information is extracted from multiple information sources. The information is comprehensively analyzed to draw the final conclusion, which solves the problem that the reasoning ability of basic RAG is limited and cannot handle complex multi-hop reasoning tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 This is a flowchart of the steps of a text generation method provided in an embodiment of the present application;
[0058] Figure 2 This is a structural block diagram of a text generation device provided in an embodiment of the present application;
[0059] Figure 3 is a block diagram of an electronic device provided in an embodiment of the present application;
[0060] Figure 4 It is a schematic diagram of a computer-readable medium provided in an embodiment of the present application. DETAILED DESCRIPTION
[0061] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0062] To facilitate understanding of the technical solutions and technical effects of the embodiments of the present application, the relevant technologies of the present application are briefly described below.
[0063] Retrieval-Augmented Generation (RAG) is an AI technology that combines information retrieval and language generation. RAG is used in question-answering systems, content generation, knowledge management, and other fields. It can generate the text users need, and is particularly suitable for scenarios requiring real-time knowledge updates or processing domain-specific data.
[0064] In related technologies, retrieval enhancement generation technology includes basic RAG and GraphRAG (graph retrieval enhancement generation).
[0065] Basic RAG is the most basic retrieval-enhanced generation framework that can be applied to question-answering systems. Basic RAG retrieves document fragments (chunks) from external knowledge bases to obtain relevant information about user questions, merges this information with the user question, and uses it as input to the language model to generate answers that better match the user question. However, basic RAG has at least the following problems:
[0066] First, the performance of basic RAG is highly dependent on the retrieval quality of the external knowledge base. If the retrieved information is inaccurate or irrelevant, the quality of the text generated by the language model will be affected.
[0067] Second, basic RAGs have limited reasoning capabilities and are unable to handle complex multi-hop reasoning tasks. These tasks often require multiple steps or reasoning stages, extracting relevant information from multiple sources and synthesizing and analyzing this information to reach a final conclusion. In other words, basic RAGs lack an understanding of global context and struggle to handle problems that require synthesizing information from multiple sources.
[0068] Third, basic RAG is difficult to dynamically adjust the retrieval strategy according to the complexity of user questions, and cannot flexibly respond to the relevant information retrieval of user questions in different fields.
[0069] Before retrieving information from document fragments in an external knowledge base, GraphRAG technology scans the entire knowledge base, constructs a knowledge graph, and builds a graph index within the knowledge graph. A graph index is a data structure used to quickly retrieve information from a knowledge graph and locate and access document fragments related to the external knowledge base. A knowledge graph includes at least one node and the relationships between nodes. GraphRAG can use community detection algorithms to identify multiple closely related nodes in the knowledge graph. GraphRAG technology can improve the retrieval quality of external knowledge bases and the quality of text generated by language models through graph indexes and community detection algorithms. However, GraphRAG technology also has at least the following problems:
[0070] First, high resource requirements: Building and processing knowledge graphs requires a large amount of storage and computing resources, which makes GraphRAG technology difficult to apply to resource-constrained devices.
[0071] Second, high implementation complexity: The construction and maintenance costs of knowledge graphs are high, and development and optimization are difficult.
[0072] Third, graph updates are difficult: Updating knowledge graphs in real time is complex. GraphRAG technology, which uses knowledge graphs for information retrieval and text generation, lags behind the latest data. This makes GraphRAG difficult to apply to tasks requiring high real-time performance.
[0073] Reference Figure 1 , shows a flowchart of the steps of a text generation method provided in an embodiment of the present application, which may specifically include the following steps:
[0074] Step 101: extract at least one content keyword from the user's text demand information;
[0075] In an embodiment of the present application, a user's textual request information can be obtained. The user's textual request information can be in the form of questions, for example: I want to know more about the technical specifications and camera features of brand A mobile phone? What is the motto of university B? Based on material C, please tell me the meaning of the protagonist's behavior D in material C, etc.
[0076] In the embodiment of the present application, a large language model can be used to extract content keywords from the user's text demand information. Content keywords refer to words that can represent the main content and core meaning of the text demand information.
[0077] In some embodiments of the present application, the content keywords include content entity keywords and / or subject keywords.
[0078] In embodiments of the present application, a large language model can be used to extract content keywords from a user's textual demand information. Content keywords include low-level keywords that focus on specific entities or details in the textual demand information and / or high-level keywords that focus on the overall concept or theme of the textual demand information. Low-level keywords are also called content entity keywords, and high-level keywords are also called theme keywords.
[0079] In a specific example, a large language model can be used to extract content entity keywords and / or topic keywords from the text request "I want to learn more about the technical specifications and camera features of Brand A mobile phone?" The large language model extracts content entity keywords such as "Brand A mobile phone," "Technical specifications," and "Camera features," while the extracted topic keywords include "smartphones" and "consumer electronics."
[0080] Step 102: obtaining a text to be retrieved associated with the text requirement information, and dividing the text to be retrieved into at least one text block;
[0081] In the embodiment of the present application, the text to be retrieved that is associated with the text requirement information can be obtained. For example, for the text requirement information "What is the motto of University B?", the text to be retrieved can be relevant materials of University B.
[0082] In an embodiment of the present application, a large language model can be used to extract text entity keywords in the text to be retrieved, determine the association relationship between the text entity keywords, and use the text entity keywords and the association relationship to build a knowledge graph. Since the large language model has an input character length limit, the length of the text to be retrieved may be very long, making it difficult to process the entire content of the text to be retrieved in a single input, so the text to be retrieved needs to be segmented into at least one text block of fixed length. It should be noted that the text to be retrieved is segmented according to the length, content logic and semantics of the text to be retrieved, thereby improving the quality of the segmented text blocks.
[0083] Step 103: construct a knowledge graph of the at least one text block, and query the knowledge graph based on the content keyword to obtain query information associated with the content keyword in the knowledge graph;
[0084] In an embodiment of the present application, retrieval enhancement generation technology can be used to construct a knowledge graph of at least one text block, and the knowledge graph can be queried based on content keywords in the user's text demand information to obtain query information associated with the content keywords in the knowledge graph.
[0085] In some embodiments of the present application, constructing the knowledge graph of the at least one text block includes:
[0086] The knowledge graph is constructed using a preset lightweight retrieval enhancement generation technology.
[0087] In an embodiment of the present application, the retrieval enhancement generation technology for constructing a knowledge graph and querying the knowledge graph can be: LightRAG (lightweight retrieval enhancement generation).
[0088] In the embodiment of the present application, the knowledge graph is constructed using LightRAG technology, which makes the construction and update costs of the knowledge graph lower, and solves the problem that the construction and processing of the knowledge graph using GraphRAG technology requires a large amount of storage and computing resources, the construction and maintenance costs of the knowledge graph are high, and the real-time updating of the knowledge graph is relatively complex, resulting in the results of information retrieval and text generation based on the knowledge graph by GraphRAG technology lagging behind the latest data.
[0089] In some embodiments of the present application, constructing the knowledge graph of the at least one text block includes:
[0090] Extracting text entity keywords from the at least one text block, and determining association relationships between the text entity keywords based on the at least one text block;
[0091] Determining relationship keywords for describing the association relationship between the text entity keywords, and determining description text for describing the text entity keywords and the association relationship between the text entity keywords;
[0092] The knowledge graph is constructed based on the text entity keywords, the association relationship between the text entity keywords, the relationship keywords and at least one of the description texts.
[0093] In an embodiment of the present application, after the text to be retrieved is divided into at least one text block of fixed length, at least one text block can be input into the large language model in sequence, and the preset first prompt template can be used to extract the text entity keywords in the at least one text block, and based on the at least one text block, the association relationship between the text entity keywords is determined, and a description text for describing the text entity keywords and the association relationship between the text entity keywords is generated.
[0094] The first prompt template may be in the form of a question and answer, including a question and an answer. The first prompt template may be:
[0095] Question: Please extract the text entity keywords from the text block, determine the association relationship between the text entity keywords, and generate descriptive text to describe the text entity keywords and the association relationship between the text entity keywords.
[0096] answer: Output of text entity keywords in the text block, the associations between text entity keywords, and descriptive text used to describe the text entity keywords and the associations between text entity keywords.
[0097] In the embodiment of the present application, the large language model can also output relationship keywords for describing the association relationship between text entity keywords, wherein the association keywords can be abstract keywords such as "meaning".
[0098] In an embodiment of the present application, a knowledge graph of at least one text block can be constructed based on text entity keywords, the association relationship between text entity keywords, relationship keywords, and at least one of the descriptive texts. A knowledge graph is essentially a semantic network that organizes information into nodes (entities) and edges (relationships), where nodes represent text entity keywords and edges represent the association relationship between these text entity keywords. The knowledge graph includes text entity keywords, the association relationship between text entity keywords, relationship keywords, and descriptive text.
[0099] In some embodiments of the present application, querying the knowledge graph based on the content keyword to obtain query information associated with the content keyword in the knowledge graph includes:
[0100] querying the knowledge graph based on the content entity keyword to obtain a target text entity keyword associated with the content entity keyword, and determining at least one of an association relationship, a relationship keyword, and a description text associated with the target text entity keyword based on the knowledge graph; and / or,
[0101] The knowledge graph is queried based on the subject keyword to obtain a target relationship keyword associated with the subject keyword, and at least one of a text entity keyword, an association relationship, and a description text associated with the target relationship keyword is determined based on the knowledge graph.
[0102] In an embodiment of the present application, the user's text demand information includes content entity keywords and / or subject keywords.
[0103] In an embodiment of the present application, a two-level search can be performed on the knowledge graph, including low-level search and high-level search.
[0104] Low-level retrieval (Local Query) refers to querying the knowledge graph based on content entity keywords to obtain the target text entity keywords associated with the content entity keywords in the knowledge graph. Then, the neighborhood relationship (such as one-hop neighbor node) of the target text entity keywords is searched in the knowledge graph to determine the neighborhood relationship of the target text entity keywords, the association relationship keywords associated with the target text entity keywords, and at least one of the description texts used to describe the target text entity keywords and the association relationship associated with the target text entity keywords. Among them, the one-hop neighbor node refers to the text entity keyword that is directly connected to the target text entity keyword through an edge. At least one of the target text entity keywords, the association relationship associated with the target text entity keywords, the relationship keywords, and the description text is the query information obtained based on the content entity keyword retrieval.
[0105] Advanced search (Global Query) refers to querying the knowledge graph based on subject keywords, obtaining target relationship keywords associated with the subject keywords in the knowledge graph, and then determining at least one of the text entity keywords associated with the target relationship keywords, the corresponding association relationships, and the descriptive text describing the association relationships corresponding to the target relationship keywords and the text entity keywords associated with the target relationship keywords. At least one of the target relationship keywords, the text entity keywords associated with the target relationship keywords, the association relationships, and the descriptive text constitutes the query information obtained based on the subject keyword search.
[0106] It should be noted that the content entity keywords and topic keywords in the user's textual demand information, as well as the description text, relationship keywords, and text entity keywords in the knowledge graph, can all be mapped into high-dimensional vectors. When searching the knowledge graph based on content entity keywords and / or topic keywords, this can be considered as a determination of whether the vectors are related.
[0107] In addition, the query information retrieved from the knowledge graph is extracted by the large language model from at least one text block of the text to be retrieved.
[0108] Step 104: using the text block corresponding to the query information in the at least one text block as a first target text block;
[0109] In the embodiment of the present application, since the query information retrieved from the knowledge graph is extracted by the large language model from at least one text block of the text to be retrieved, the query information and the text block have a corresponding relationship.
[0110] In an embodiment of the present application, a text block corresponding to query information obtained based on content entity keyword retrieval and / or query information obtained based on subject keyword retrieval may be used as the first target text block.
[0111] Step 105: evaluating the similarity between the text requirement information and the at least one text block, and determining a second target text block in the at least one text block based on the similarity;
[0112] In an embodiment of the present application, in addition to performing a knowledge graph-based search on the text to be retrieved, a search based on text block similarity matching can also be performed. Specifically, the similarity between the text demand information and at least one text block can be evaluated. The higher the similarity between the text demand information and the text block, the stronger the correlation between the text block and the text demand information. Based on the similarity, the text block with the highest similarity among the at least one text block is selected as the second target text block.
[0113] In some embodiments of the present application, the evaluating the similarity between the text requirement information and the at least one text block includes:
[0114] Mapping the text requirement information and the at least one text block respectively using a preset vector model to obtain a requirement information vector corresponding to the text requirement information and a text block vector corresponding to the text block;
[0115] A similarity calculation is performed on the demand information vector and the text block vector to obtain a similarity between the text demand information and the at least one text block.
[0116] In an embodiment of the present application, a preset vector model can be used to map the text requirement information and at least one text block into a high-dimensional vector, thereby obtaining a requirement information vector corresponding to the text requirement information and a text block vector corresponding to the text block. A similarity calculation is then performed on the requirement information vector and the text block vector to obtain the similarity between the text requirement information and the at least one text block.
[0117] Step 106 : Based on the query information, the first target text block, the second target text block, and at least one of the text requirement information, generate the user's target requirement text using a preset language model.
[0118] In an embodiment of the present application, at least one of the query information, the first target text block, the second target text block, and the text requirement information is input into a preset language model to generate the user's target requirement text.
[0119] In some embodiments of the present application, generating the user's target requirement text using a preset language model based on at least one of the query information, the first target text block, the second target text block, and the text requirement information includes:
[0120] Merging the first target text block and the second target text block, and deleting duplicate text in the first target text block and the second target text block to obtain a third target text block;
[0121] The target requirement text is generated using the language model based on at least one of the query information, the third target text block, and the text requirement information.
[0122] In an embodiment of the present application, after performing a knowledge graph search on the text to be retrieved to obtain a first target text block and query information obtained by content entity keyword search and / or query information obtained by subject keyword search, and performing a text block similarity matching search on the text to be retrieved to obtain a second target text block, the query information obtained by content entity keyword search and the query information obtained by subject keyword search can be merged and duplicates can be removed to obtain target query information. The first target text block and the second target text block can also be merged and duplicate text in the first target text block and the second target text block can be deleted to obtain a third target text block.
[0123] In an embodiment of the present application, at least one of the target query information, the third target text block, and the text requirement information can be input into the large language model, and the large language model calls an appropriate second prompt template to output the user's target requirement text.
[0124] The second prompt template may be in the form of a question and answer, including a question and an answer. The second prompt template may be:
[0125] question: Please output the user's target demand text based on the target query information retrieved based on content entity keywords and / or subject keywords, the third target text block, and at least one of the user's text demand information.
[0126] answer: The user's target requirement text output.
[0127] In an embodiment of the present application, at least one content keyword is extracted from the user's text demand information; the text to be retrieved associated with the text demand information is obtained, and the text to be retrieved is divided into at least one text block; a knowledge graph of at least one text block is constructed, and the knowledge graph is queried based on the content keywords to obtain query information associated with the content keywords in the knowledge graph; the text block corresponding to the query information in at least one text block is used as the first target text block; the similarity between the text demand information and the at least one text block is evaluated, and based on the similarity, the second target text block in the at least one text block is determined; based on the query information, the first target text block, the second target text block and at least one of the text demand information, the user's target demand text is generated using a preset language model. Retrieving the text to be retrieved based on the knowledge graph improves the retrieval quality of the text to be retrieved, thereby improving the quality of the text generated by the language model. In addition, by searching the text to be retrieved based on the knowledge graph, query information related to the content keywords in the user's text demand information is obtained, and then the first target text block corresponding to the query information in at least one text block of the text to be retrieved is determined, and relevant information is extracted from multiple information sources. The information is comprehensively analyzed to draw the final conclusion, which solves the problem that the reasoning ability of basic RAG is limited and cannot handle complex multi-hop reasoning tasks.
[0128] In a specific example, the user's textual demand information is "Who are the previous presidents of University B?", but in the text to be retrieved, there is no text that can directly answer this question. In the text to be retrieved, the first paragraph may mention the first president of University B, the second paragraph mentions the second president of University B, the third paragraph mentions the third president of University B, and so on, and the last paragraph mentions the current president of University B. If the text to be retrieved is divided into multiple text blocks, and each paragraph is in a different text block, the basic RAG cannot handle the problem of extracting relevant information from multiple text blocks and comprehensively analyzing this information to draw a final conclusion. This application is based on the knowledge graph for retrieval. In the process of constructing the knowledge graph in the large language model, the previous presidents of University B and University B will be proposed as text entity keywords, and the association relationship edge between the previous presidents of University B and University B will be established. Then, in the knowledge graph, by querying the previous presidents of University B, you can easily get the correct answer to the question, and better answer multi-hop questions or questions with clues throughout the text.
[0129] In the embodiment of the present application, the method of retrieving the text to be retrieved based on the knowledge graph can also solve the problem that the basic RAG has poor effect when processing the user's demand text information involving "abstract class" problems.
[0130] In a specific example, the user's text demand information is "Based on material C, please tell me what the meaning of the protagonist's behavior D in material C is." When constructing the knowledge graph, the large language model generates corresponding descriptions for the extracted text entity keywords and the association relationships between the text entity keywords, and additionally generates relationship keywords for the association relationships. The relationship keywords can be abstract words, such as meaning. When searching the knowledge graph based on the subject keywords in the text demand information, the target relationship keywords can be retrieved, and then the text entity keywords, association relationships, and at least one of the description texts used to describe the association relationships corresponding to the target relationship keywords and the text entity keywords associated with the target relationship keywords are determined in the knowledge graph. Then, based on this query information, the language model can achieve high-quality output of user demand text for "abstract" questions.
[0131] In an embodiment of the present application, the knowledge graph is searched based on at least one content keyword in the user's text demand information, so that relevant information retrieval can be used to deal with user problems in different fields.
[0132] In the embodiment of the present application, the knowledge graph is constructed using LightRAG technology, which makes the construction and update costs of the knowledge graph lower, and solves the problems that the construction and processing of the knowledge graph using GraphRAG technology requires a large amount of storage and computing resources, the construction and maintenance costs of the knowledge graph are high, and the real-time updating of the knowledge graph is relatively complex, and the results of information retrieval and text generation based on the knowledge graph by GraphRAG technology lag behind the latest data.
[0133] In an embodiment of the present application, when using LightRAG technology to construct a knowledge graph, the text to be retrieved is segmented according to its length, content logic and semantics, thereby improving the quality of the segmented text blocks, and avoiding the problem of poor quality of text blocks caused by LightRAG technology simply segmenting the text to be retrieved based on its length to generate text blocks.
[0134] In an embodiment of the present application, the knowledge graph constructed using LightRAG technology may not contain all text entity keywords in at least one text block, and two entity words with the same meaning but different descriptions (such as tomato and tomatoes) may be regarded as two text entity keywords, resulting in low retrieval accuracy for retrieved texts based on the knowledge graph and poor text quality generated by the language model. This application addresses the incompleteness that LightRAG technology may exhibit when constructing a knowledge graph, and utilizes a retrieval method based on text block similarity matching to alleviate the problem of low retrieval accuracy for retrieved texts based on the knowledge graph.
[0135] In a specific example, the text to be retrieved is an introductory manual about University B, which contains a sentence "The motto of University B is 'Self-improvement and Virtue'." The user's text demand information is the question "What is the motto of University B?" There may be omissions in the process of constructing the knowledge graph using LightRAG technology. If "motto" is not extracted as a text entity keyword, relevant materials cannot be retrieved when searching based on the knowledge graph. At this time, the search is combined with the search strategy based on text block similarity matching, which can effectively detect the text block containing "The motto of University B is 'Self-improvement and Virtue'". Therefore, compared with the LightRAG technology that answers questions based only on knowledge graph retrieval, the search strategy of this application has significantly improved its ability to answer precise matching questions.
[0136] It should be noted that for the method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present application are not limited by the order of the actions described, because according to the embodiments of the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present application.
[0137] Reference Figure 2 , shows a structural block diagram of a text generation device provided in an embodiment of the present application, which may specifically include the following modules:
[0138] Content keyword extraction module 201, used to extract at least one content keyword from the user's text demand information;
[0139] A segmentation module 202 is configured to obtain a text to be retrieved associated with the text requirement information and segment the text to be retrieved into at least one text block;
[0140] A knowledge graph construction module 203 is configured to construct a knowledge graph for the at least one text block, and query the knowledge graph based on the content keywords to obtain query information associated with the content keywords in the knowledge graph;
[0141] The first target text block serves as a module 204, configured to use a text block corresponding to the query information in the at least one text block as a first target text block;
[0142] a second target text block determining module 205 for evaluating the similarity between the text requirement information and the at least one text block, and determining a second target text block in the at least one text block based on the similarity;
[0143] The text generation module 206 is configured to generate the user's target requirement text using a preset language model based on the query information, the first target text block, the second target text block, and at least one of the text requirement information.
[0144] In an optional embodiment of the present application, the content keywords include content entity keywords and / or subject keywords.
[0145] In an optional embodiment of the present application, the knowledge graph construction module includes:
[0146] a text entity keyword extraction submodule, configured to extract text entity keywords from the at least one text block and determine association relationships between the text entity keywords based on the at least one text block;
[0147] a relationship keyword determination submodule, configured to determine relationship keywords for describing the association relationship between the text entity keywords, and to determine description text for describing the text entity keywords and the association relationship between the text entity keywords;
[0148] The first knowledge graph construction submodule is used to construct the knowledge graph based on the text entity keywords, the association relationship between the text entity keywords, the relationship keywords and at least one of the description text.
[0149] In an optional embodiment of the present application, the knowledge graph construction module includes:
[0150] a query submodule, configured to query the knowledge graph based on the content entity keywords, obtain target text entity keywords associated with the content entity keywords, and determine at least one of an association relationship, a relationship keyword, and a description text associated with the target text entity keywords based on the knowledge graph; and / or,
[0151] The knowledge graph is queried based on the subject keyword to obtain a target relationship keyword associated with the subject keyword, and at least one of a text entity keyword, an association relationship, and a description text associated with the target relationship keyword is determined based on the knowledge graph.
[0152] In an optional embodiment of the present application, the text generation module includes:
[0153] a merging submodule, configured to merge the first target text block and the second target text block, and delete duplicate text in the first target text block and the second target text block, to obtain a third target text block;
[0154] The target requirement text generation submodule is configured to generate the target requirement text using the language model based on at least one of the query information, the third target text block, and the text requirement information.
[0155] In an optional embodiment of the present application, the knowledge graph construction module includes:
[0156] The second knowledge graph construction submodule is used to construct the knowledge graph using a preset lightweight retrieval enhancement generation technology.
[0157] In an optional embodiment of the present application, the second target text block determination module includes:
[0158] a mapping submodule, configured to map the text requirement information and the at least one text block respectively using a preset vector model to obtain a requirement information vector corresponding to the text requirement information and a text block vector corresponding to the text block;
[0159] The similarity calculation submodule is configured to perform similarity calculation on the demand information vector and the text block vector to obtain the similarity between the text demand information and the at least one text block.
[0160] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0161] In addition, the present invention also provides an electronic device, such as Figure 3 As shown, it includes a processor 301, a communication interface 302, a memory 303 and a communication bus 304, wherein the processor 301, the communication interface 302, and the memory 303 communicate with each other through the communication bus 304.
[0162] Memory 303, used for storing computer programs;
[0163] The processor 301 is configured to execute the program stored in the memory 303, and implement the following steps:
[0164] Extracting at least one content keyword from the user's text demand information;
[0165] Acquire a text to be retrieved associated with the text requirement information, and divide the text to be retrieved into at least one text block;
[0166] Constructing a knowledge graph of the at least one text block, and querying the knowledge graph based on the content keyword to obtain query information associated with the content keyword in the knowledge graph;
[0167] taking a text block corresponding to the query information in the at least one text block as a first target text block;
[0168] evaluating a similarity between the text requirement information and the at least one text block, and determining a second target text block in the at least one text block based on the similarity;
[0169] Based on at least one of the query information, the first target text block, the second target text block, and the text requirement information, a preset language model is used to generate the user's target requirement text.
[0170] In an optional embodiment of the present application, the content keywords include content entity keywords and / or subject keywords.
[0171] In an optional embodiment of the present application, constructing the knowledge graph of the at least one text block includes:
[0172] Extracting text entity keywords from the at least one text block, and determining association relationships between the text entity keywords based on the at least one text block;
[0173] Determining relationship keywords for describing the association relationship between the text entity keywords, and determining description text for describing the text entity keywords and the association relationship between the text entity keywords;
[0174] The knowledge graph is constructed based on the text entity keywords, the association relationship between the text entity keywords, the relationship keywords and at least one of the description texts.
[0175] In an optional embodiment of the present application, querying the knowledge graph based on the content keyword to obtain query information associated with the content keyword in the knowledge graph includes:
[0176] querying the knowledge graph based on the content entity keyword to obtain a target text entity keyword associated with the content entity keyword, and determining at least one of an association relationship, a relationship keyword, and a description text associated with the target text entity keyword based on the knowledge graph; and / or,
[0177] The knowledge graph is queried based on the subject keyword to obtain a target relationship keyword associated with the subject keyword, and at least one of a text entity keyword, an association relationship, and a description text associated with the target relationship keyword is determined based on the knowledge graph.
[0178] In an optional embodiment of the present application, generating the user's target requirement text using a preset language model based on at least one of the query information, the first target text block, the second target text block, and the text requirement information includes:
[0179] Merging the first target text block and the second target text block, and deleting duplicate text in the first target text block and the second target text block to obtain a third target text block;
[0180] The target requirement text is generated using the language model based on at least one of the query information, the third target text block, and the text requirement information.
[0181] In an optional embodiment of the present application, constructing the knowledge graph of the at least one text block includes:
[0182] The knowledge graph is constructed using a preset lightweight retrieval enhancement generation technology.
[0183] In an optional embodiment of the present application, the evaluating the similarity between the text requirement information and the at least one text block includes:
[0184] Mapping the text requirement information and the at least one text block respectively using a preset vector model to obtain a requirement information vector corresponding to the text requirement information and a text block vector corresponding to the text block;
[0185] A similarity calculation is performed on the demand information vector and the text block vector to obtain a similarity between the text demand information and the at least one text block.
[0186] The communication bus mentioned in the terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.
[0187] The communication interface is used for communication between the above terminal and other devices.
[0188] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.
[0189] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.
[0190] like Figure 4 As shown, in another embodiment provided in the present application, a computer-readable storage medium 401 is also provided, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer executes a text generation method described in the above embodiment.
[0191] In another embodiment provided by the present application, a computer program product including instructions is also provided, which, when executed on a computer, enables the computer to execute a text generation method described in the above embodiment.
[0192] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0193] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0194] Each embodiment in this specification is described in a related manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiment is generally similar to the method embodiment, so the description is relatively simple. For related parts, refer to the description of the method embodiment.
[0195] The above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are included in the scope of protection of the present application.
Claims
1. A text generation method, characterized in that: include: Extracting at least one content keyword from the user's text demand information; Acquire a text to be retrieved associated with the text requirement information, and divide the text to be retrieved into at least one text block; Constructing a knowledge graph of the at least one text block, and querying the knowledge graph based on the content keyword to obtain query information associated with the content keyword in the knowledge graph; taking a text block corresponding to the query information in the at least one text block as a first target text block; evaluating a similarity between the text requirement information and the at least one text block, and determining a second target text block in the at least one text block based on the similarity; Based on at least one of the query information, the first target text block, the second target text block, and the text requirement information, a preset language model is used to generate the user's target requirement text.
2. The method according to claim 1, characterized in that The content keywords include content entity keywords and / or subject keywords.
3. The method according to claim 2, characterized in that The constructing of the knowledge graph of the at least one text block includes: Extracting text entity keywords from the at least one text block, and determining association relationships between the text entity keywords based on the at least one text block; Determining relationship keywords for describing the association relationship between the text entity keywords, and determining description text for describing the text entity keywords and the association relationship between the text entity keywords; The knowledge graph is constructed based on the text entity keywords, the association relationship between the text entity keywords, the relationship keywords and at least one of the description texts.
4. The method according to claim 3, characterized in that The querying the knowledge graph based on the content keyword to obtain query information associated with the content keyword in the knowledge graph includes: querying the knowledge graph based on the content entity keyword to obtain a target text entity keyword associated with the content entity keyword, and determining at least one of an association relationship, a relationship keyword, and a description text associated with the target text entity keyword based on the knowledge graph; and / or, The knowledge graph is queried based on the subject keyword to obtain a target relationship keyword associated with the subject keyword, and at least one of a text entity keyword, an association relationship, and a description text associated with the target relationship keyword is determined based on the knowledge graph.
5. The method according to claim 4, characterized in that The generating the user's target requirement text by using a preset language model based on at least one of the query information, the first target text block, the second target text block, and the text requirement information includes: Merging the first target text block and the second target text block, and deleting duplicate text in the first target text block and the second target text block to obtain a third target text block; The target requirement text is generated using the language model based on at least one of the query information, the third target text block, and the text requirement information.
6. The method according to claim 1, characterized in that The constructing of the knowledge graph of the at least one text block includes: The knowledge graph is constructed using a preset lightweight retrieval enhancement generation technology.
7. The method according to claim 1, characterized in that The evaluating the similarity between the text requirement information and the at least one text block includes: Mapping the text requirement information and the at least one text block respectively using a preset vector model to obtain a requirement information vector corresponding to the text requirement information and a text block vector corresponding to the text block; A similarity calculation is performed on the demand information vector and the text block vector to obtain a similarity between the text demand information and the at least one text block.
8. A text generation device, characterized in that: include: A content keyword extraction module, configured to extract at least one content keyword from the user's textual demand information; a segmentation module, configured to obtain a text to be retrieved associated with the text requirement information, and segment the text to be retrieved into at least one text block; a knowledge graph construction module, configured to construct a knowledge graph for the at least one text block, and query the knowledge graph based on the content keywords to obtain query information associated with the content keywords in the knowledge graph; The first target text block is used as a module, and is used to take the text block corresponding to the query information in the at least one text block as the first target text block; a second target text block determining module, configured to evaluate the similarity between the text requirement information and the at least one text block, and determine a second target text block in the at least one text block based on the similarity; The text generation module is used to generate the user's target requirement text using a preset language model based on the query information, the first target text block, the second target text block and at least one of the text requirement information.
9. An electronic device, characterized in that: comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; The memory is used to store computer programs; The processor is configured to implement the method according to any one of claims 1 to 7 when executing a program stored in the memory.
10. One or more computer-readable media having instructions stored thereon, which, when executed by one or more processors, cause the processors to perform the method of any one of claims 1-7.
Citation Information
Patent Citations
Multi-document intelligent question and answer method and system based on large language model
CN118394897A
Knowledge retrieval enhancement-based characteristic agricultural product standardized file content generation method
CN118779407A
Retrieval enhancement generation-based retrieval method, product, equipment and medium
CN119003795A
Intelligent question and answer method, device and equipment based on knowledge graph and medium
CN119129738A
Method and system for generating enhanced knowledge questions and answers for mixed retrieval of heterogeneous database
CN119311831A