Text generation method and device, electronic equipment, storage medium and product
By performing multi-dimensional sorting of candidate text fragments and training the target generation model using reinforcement learning, the problems of low text matching and information omission in existing technologies are solved, and high-quality text generation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZHONGKE JINDEZHU INTELLIGENT TECH CO LTD
- Filing Date
- 2025-12-29
- Publication Date
- 2026-05-05
AI Technical Summary
Existing text generation methods, when directly responding to user queries based on generation models, suffer from problems such as low matching degree between generated text and actual user needs, and easy occurrence of factual deviations or omission of key information. In particular, the generated answers are inaccurate when faced with noisy search results.
By acquiring multiple candidate text fragments corresponding to the user's query request, performing multi-dimensional score weighted sorting, filtering out the target text fragments, and inputting them into the target generation model trained by retrieval enhancement fine-tuning and reinforcement learning, the output text is generated.
It significantly improves the matching degree and accuracy between the output text and the user's query request, avoids factual bias and omission of key information, and generates high-quality response text.
Smart Images

Figure CN121981084A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a text generation method, apparatus, electronic device, storage medium, and product. Background Technology
[0002] In existing text generation solutions, those that directly respond to user queries using a generative model typically input all candidate text fragments returned by the retrieval into the model. This approach suffers from low matching accuracy between the generated text and the user's actual needs, and is prone to factual bias or omission of key information. Although some solutions attempt to combine retrieval and generation models, noise in the retrieval results still leads to inaccurate output text, making it difficult to generate high-quality output text. Summary of the Invention
[0003] This disclosure provides a text generation method, apparatus, electronic device, storage medium, and product to address problems in the related art.
[0004] A first aspect of this disclosure provides a text generation method, the method comprising: Retrieve multiple candidate text fragments corresponding to the user's query request; Sort multiple candidate text fragments to obtain a candidate list; Select at least one target text fragment from the candidate list; The target text fragment is input into the target generation model, which generates the output text corresponding to the user's query request. The target generation model is trained through retrieval enhancement fine-tuning and reinforcement learning.
[0005] In one embodiment, obtaining multiple candidate text fragments corresponding to a user query request includes: Get the user's query request; The target knowledge base is retrieved based on the user's query request, resulting in multiple candidate text fragments corresponding to the user's query request.
[0006] In one embodiment, a search of the target knowledge base is performed based on a user query request to obtain multiple candidate text fragments corresponding to the user query request, including: Based on the user's query request, keyword matching and semantic similarity retrieval are performed on the target knowledge base to obtain the first candidate text and the second candidate text corresponding to the user's query request. The first and second candidate texts are merged to obtain multiple candidate text fragments corresponding to the user's query request.
[0007] In one embodiment, multiple candidate text fragments are sorted to obtain a candidate list, including: Determine the basic relevance score between each candidate text fragment and the user query request, the matching degree score between the metadata information associated with each candidate text fragment and the user query request, and the edit distance relevance score between the user query request and the file identifier; The target relevance score is determined based on the basic relevance score, the matching degree score, and the edit distance relevance score. Multiple candidate text fragments are sorted according to their target relevance scores to obtain a candidate list.
[0008] In one embodiment, the method provided in this disclosure, prior to inputting the target text fragment into the target generation model and generating the output text corresponding to the user's query request, includes: Obtain the training dataset, which includes user query requests, positive samples associated with user query requests, negative samples associated with user query requests, and standard output text corresponding to user query requests. Supervised fine-tuning of the initial generated model using the training dataset; The target generative model is obtained by iteratively optimizing the supervised fine-tuned initial generative model based on a reinforcement learning strategy.
[0009] In one embodiment, the reinforcement learning strategy employs a multi-dimensional reward mechanism, which includes at least one of the following: a reward based on the reasonableness of the length of the training output text generated during model training, a reward based on the matching degree between the training output text generated during model training and the standard output text, a credibility reward based on the citation priority of the target text fragment, and a reward based on the format standardization of the training output text generated during model training.
[0010] A second aspect of this disclosure provides a text generation apparatus, the apparatus comprising: The acquisition unit is used to acquire multiple candidate text fragments corresponding to the user's query request; The sorting unit is used to sort multiple candidate text fragments to obtain a candidate list; A selection unit is used to select at least one target text fragment from a candidate list; The generation unit is used to input the target text fragment into the target generation model and generate the output text corresponding to the user's query request. The target generation model is obtained through retrieval enhancement fine-tuning and reinforcement learning training.
[0011] A third aspect of this disclosure provides an electronic device comprising: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the methods described in the first aspect of this disclosure.
[0012] A fourth aspect of this disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform the methods described in the first aspect of this disclosure.
[0013] A fifth aspect of this disclosure provides a computer program product including a computer program that, when executed by a processor, implements the methods described in the first aspect of this disclosure.
[0014] In summary, this disclosure proposes a text generation method, which includes: obtaining multiple candidate text fragments corresponding to a user query request; sorting the multiple candidate text fragments to obtain a candidate list; selecting at least one target text fragment from the candidate list; inputting the target text fragment into a target generation model to generate output text corresponding to the user query request, wherein the target generation model is obtained through retrieval enhancement fine-tuning and reinforcement learning training.
[0015] According to the solution provided in this disclosure, target text fragments are obtained by acquiring candidate text fragments and sorting and filtering them, providing accurate information support for model generation. At the same time, the target generation model is optimized through retrieval enhancement fine-tuning and reinforcement learning training, which significantly improves the ability to utilize retrieval information, ensures the matching degree and accuracy of the output text with the user's query request, avoids factual bias and omission of key information, and thus ensures the matching degree and accuracy of the output text with the user's query request, resulting in high-quality output text.
[0016] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0018] Figure 1 A flowchart illustrating a text generation method provided in an embodiment of this disclosure; Figure 2 A flowchart illustrating a method for determining a candidate list provided in an embodiment of this disclosure; Figure 3 A flowchart illustrating yet another text generation method provided in this disclosure embodiment; Figure 4 This is a schematic diagram of the structure of a text generation device provided in an embodiment of the present disclosure; Figure 5 This is a schematic diagram of the hardware composition structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0019] Embodiments of this disclosure are described in detail below. Examples of these embodiments are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting this disclosure.
[0020] As enterprise customer service scales continuously expand, call center or online customer service agents need to handle massive and diverse user inquiries. To ensure standardized service quality and improve response efficiency, enterprises generally establish structured and unstructured knowledge bases. When answering customer calls or responding to online inquiries, agents need to quickly retrieve relevant information from the knowledge base and organize it into accurate answers to provide to the customer.
[0021] The following is a brief introduction to several methods for text generation in related technologies: Currently, mainstream agent assistance systems typically employ the following workflow: a. Retrieval Phase: The system receives customer questions (Query) input by agents and retrieves several potentially relevant information fragments from the knowledge base using technologies such as keyword matching and vector semantic retrieval, which serve as a set of candidate answers.
[0022] b. Presentation stage: The system presents the retrieved list of candidate answers to the operators.
[0023] c. Manual Judgment and Organization Stage: The agent needs to quickly read and understand multiple candidate answers, judge their relevance, accuracy and timeliness, and then manually extract key information from them to organize them into a coherent, accurate final answer that fits the current context.
[0024] To further improve efficiency, the industry has begun to introduce generative artificial intelligence technology. After retrieving a set of candidate answers, the system inputs these candidate answers as context into a large language model, which then directly generates a complete and easy-to-read answer, thereby reducing the workload of agents in organizing their language.
[0025] Despite the introduction of generative models, existing customer service knowledge base assistants based on retrieval enhancement still have significant shortcomings in the core aspect of accurately generating answers, mainly in the following aspects: a. Noise in the search results leads to inaccurate generated answers. The candidate answer set returned by the knowledge base retrieval step usually contains multiple pieces of information, which may have weak relevance, be redundant, or even contradict each other. If directly input into the generation model, the model may use noisy information to generate "illusory" answers that omit key information or contain factual errors.
[0026] b. The system lacks a module for fine-grained ranking and confidence assessment before generation. It cannot determine the semantic relevance of multiple candidate answers to the current query, or their priority within the knowledge base. This ambiguity in priority leads to a lack of clear and reliable guiding principles for the generation model when integrating information. It can only process all candidate information homogeneously, resulting in unstable and inaccurate generation.
[0027] The above solution has the following drawbacks: Existing customer service knowledge base assistants are prone to confusion when faced with redundant and contradictory candidate answers, and the generation model can only process all candidate information in a homogeneous manner due to the lack of a dominant reference, resulting in unstable output.
[0028] To address the shortcomings of related technologies, this disclosure obtains target text fragments by acquiring candidate text fragments and sorting and filtering them, providing accurate information support for model generation. At the same time, the target generation model is optimized through retrieval enhancement fine-tuning and reinforcement learning training, significantly improving its ability to utilize retrieval information, ensuring the matching degree and accuracy of the output text with the user's query request, avoiding factual bias and omission of key information, and thus ensuring the matching degree and accuracy of the output text with the user's query request, resulting in high-quality output text.
[0029] The present disclosure will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0030] The text generation method provided in this disclosure can be applied to scenarios that require generating accurate response text based on preset knowledge, such as intelligent customer service, knowledge base question and answer, and information consultation response. For example, it can be applied to scenarios such as automatic reply to user inquiries by enterprise online customer service systems and accurate response to user queries by intelligent question answering robots. The execution subject of this method can be a server with data processing and model running capabilities, or a terminal device that integrates relevant algorithms and storage modules.
[0031] like Figure 1 As shown, Figure 1 This is a flowchart illustrating a text generation method provided in an embodiment of this disclosure. The text generation method provided in this embodiment includes the following steps: Step 101: Obtain multiple candidate text fragments corresponding to the user's query request; In one embodiment, a user query request refers to a text-based request initiated by a user to obtain specific information or solve a specific problem. It is an input signal that triggers the entire text generation process, such as a text message in an intelligent customer service scenario where a user asks "how to report a lost account".
[0032] In one embodiment, the candidate text fragment is a text unit that is potentially related to the user's query request, and may be a document paragraph, a summary of key information, etc.
[0033] In one embodiment, a hybrid retrieval mode can be used to obtain multiple candidate text fragments corresponding to the user's query request. Specifically, keyword matching retrieval and semantic similarity retrieval can be performed simultaneously. Keyword matching retrieval extracts the core keywords in the user's query request and traverses the knowledge base to filter text fragments containing the corresponding keywords. Semantic similarity retrieval converts the query request and knowledge base text into semantic vectors through a semantic encoding model, calculates vector similarity to filter relevant text fragments, and finally merges the two types of retrieval results to obtain candidate text fragments.
[0034] In one embodiment, multiple candidate text fragments corresponding to the user's query request can also be obtained based on a rule-matching retrieval method. Specifically, matching rules between query requests and knowledge base texts under different business scenarios can be preset. By parsing the scenario type of the user's query request, text fragments can be retrieved from the knowledge base of the corresponding scenario as candidate text fragments.
[0035] In one embodiment, a retrieval model can also be used to obtain multiple candidate text fragments corresponding to the user's query request. Specifically, a pre-trained retrieval model, such as a dual encoder model, can be used to directly rank the relevance between the user's query request and the knowledge base text, and select the text fragments ranked higher as candidate text fragments.
[0036] Step 102: Sort the multiple candidate text fragments to obtain a candidate list; In one embodiment, the candidate list is a list formed by arranging multiple candidate text fragments in order of their relevance to the user's query request, and is used to determine the priority of the candidate text fragments.
[0037] In one embodiment, a multi-dimensional score weighted sorting method can be used to sort multiple candidate text fragments to obtain a candidate list. Specifically, the basic relevance score between the candidate text fragments and the user's query request, the matching degree score between the candidate text fragments' associated metadata (such as title and summary) and the query request, and the edit distance similarity score between the query request and the file identifier corresponding to the text fragments can be calculated respectively. Preset weights are assigned to each score and summed to obtain a comprehensive score. The candidate list is formed by sorting the comprehensive scores in descending order.
[0038] In one embodiment, a ranking model-based ranking method can also be used to rank multiple candidate text segments to obtain a candidate list. Specifically, the combination information of the user query request and the candidate text segments can be input into a pre-trained ranking model, such as a cross-encoder model. The model outputs the relevance score of each candidate text segment, and the candidate list is obtained by arranging the scores in descending order.
[0039] In one embodiment, multiple candidate text fragments can be sorted using a rule-based priority ranking method to obtain a candidate list. Specifically, a relevance evaluation rule can be preset, such as text fragments containing core keywords having a higher priority than fragments that are only semantically related, and the most recently updated text fragments having a higher priority than older text fragments. The candidate text fragments are then sorted according to the rule priority to form a candidate list.
[0040] Step 103: Select at least one target text fragment from the candidate list; In one embodiment, the target text fragment is a text fragment selected from the candidate list that has core value in responding to a user query request.
[0041] In one embodiment, at least one target text fragment can be selected from the candidate list in a fixed number. Specifically, the selection number can be preset, such as the first 3 or the first 5, and the top preset number of candidate text fragments in the candidate list can be selected as the target text fragments.
[0042] In one embodiment, a threshold can be used to select at least one target text segment from the candidate list. Specifically, a relevance score threshold can be preset, and candidate text segments with scores higher than the threshold can be selected as target text segments. If all segments have scores lower than the threshold, the segment with the highest score can be selected.
[0043] Step 104: Input the target text fragment into the target generation model to generate the output text corresponding to the user's query request. The target generation model is obtained through retrieval enhancement fine-tuning and reinforcement learning training.
[0044] In one embodiment, the target generation model is an artificial intelligence model with text generation capabilities. After retrieval enhancement fine-tuning and reinforcement learning training, it can generate output text that accurately matches the user's query request based on the input target text fragment.
[0045] In one embodiment, retrieval enhancement fine-tuning refers to optimizing the parameters of the initial generation model by introducing knowledge base retrieval information, thereby enhancing the model's ability to utilize retrieval information and improving the accuracy and relevance of the generated text; reinforcement learning training is a training method that guides the model to iteratively optimize by setting a reward mechanism, so that the model gradually learns the generation strategy that meets the preset requirements during the text generation process, thereby improving the quality of the output text.
[0046] In one embodiment, the output text is the final text result generated by the target generation model based on the target text fragment, used to respond to a user query request.
[0047] In one embodiment, the selected target text fragments and the user query request are input together into the target generation model, which has been trained by retrieval enhancement fine-tuning and reinforcement learning. The model combines the core information of the target text fragments with the user's needs to generate logically coherent and accurate output text.
[0048] In one embodiment, output format constraints can be passed to the model at the same time as the input target text fragment and user query request, such as the output being in point-by-point format and the output word count being controlled within 50 characters, to generate output text that meets the requirements.
[0049] This application aims to improve the quality and accuracy of text generation for user queries by combining retrieval enhancement and model optimization. First, multiple candidate text fragments associated with the user query are obtained. These fragments are then ordered to form a candidate list for easy filtering. Next, at least one key target text fragment is selected from this list and used as the basis for generation. This target text fragment is then fed into a target generation model trained through retrieval enhancement fine-tuning and reinforcement learning optimization. The model then combines the effective information from the target text fragment to generate output text that matches the user query.
[0050] By retrieving and filtering relevant candidate text fragments, the model can be generated. Combined with targeted model training and optimization, the matching degree and accuracy of the output text with the user's query request are improved. At the same time, the information sources for text generation are broadened, and the risk of missing key information or factual deviations during the model generation process is reduced.
[0051] In one embodiment, obtaining multiple candidate text fragments corresponding to a user query request includes: Get the user's query request; In one embodiment, user query requests can be obtained through interface calls. Specifically, text requests sent by users, such as web-based customer service, mobile apps, and mini-programs, can be received via HyperText Transfer Protocol (HTTP) / HTTPS interfaces. The request data is then parsed and cleaned, such as removing special symbols and filtering invalid characters, to obtain standardized user query requests. Alternatively, user query requests can be obtained through message queues. Specifically, message middleware such as RabbitMQ or Kafka can be used to receive user query requests, enabling asynchronous reception and buffering of requests to avoid system overload. After receiving, the requests are then standardized. User query requests can also be obtained through local input. Specifically, user query requests can be received on local terminal devices, such as dedicated query terminals or maintenance equipment, via keyboard input, file import, etc., and converted into standardized text after format validation.
[0052] The target knowledge base is retrieved based on the user's query request, resulting in multiple candidate text fragments corresponding to the user's query request.
[0053] In one embodiment, the target knowledge base refers to pre-built structured data that stores text data related to a specific domain or business. The data within the target knowledge base undergoes standardization processing, such as paragraph splitting and information annotation, to facilitate efficient retrieval and matching.
[0054] In one embodiment, the target knowledge base can be a knowledge base such as an enterprise customer service knowledge base, a product manual knowledge base, or an industry policy knowledge base.
[0055] In one embodiment, a user query request can be split into multiple search terms, all text data in the target knowledge base can be traversed, and text fragments containing the search terms can be matched and used as candidate text fragments.
[0056] In one embodiment, the text data in the target knowledge base is stored according to a preset structure, such as partitioning by business type, question type, and document type. The category to which the user's query request belongs can be determined first through semantic parsing, and then the search can be performed only in the data partition of the corresponding category to filter relevant text fragments as candidate text fragments.
[0057] In one embodiment, a semantic coding model, such as Bidirectional Encoder Representations from Transformers (BERT), can be used to convert user query requests into query vectors. At the same time, all text fragments in the target knowledge base can be converted into semantic vectors and a vector index can be built. During retrieval, the similarity between the query vector and each text fragment vector is calculated, and text fragments with similarity higher than a preset threshold are selected as candidate text fragments.
[0058] When retrieving multiple candidate text fragments corresponding to a user's query request, the system first receives the user's query request through a pre-defined interactive interface. This query request is captured by the system in text form and converted into a standardized data format suitable for retrieval. The target knowledge base is a pre-built structured data storage system where the original documents are split into paragraphs. Each paragraph is encapsulated as an independent retrieval slice (fragment). In addition to retaining the original text content, each retrieval slice is associated with extended information such as related query statements, core keywords, and paragraph titles extracted from the original content, in order to eliminate semantic differences between the query request and the document content. Based on the standardized user query request, the system calls the knowledge base's retrieval interface to perform a retrieval operation. During the retrieval process, the system simultaneously matches and calculates the query request with the original content and extended information of the retrieval slices, filtering out all retrieval slices that are related to the query request. These retrieval slices are the multiple candidate text fragments corresponding to the user's query request.
[0059] By performing structured segmentation of knowledge base documents and supplementing them with extended information, the comprehensiveness of retrieval matching is improved, avoiding the omission of relevant information due to semantic differences.
[0060] In one embodiment, a search of the target knowledge base is performed based on a user query request to obtain multiple candidate text fragments corresponding to the user query request, including: Based on the user's query request, keyword matching and semantic similarity retrieval are performed on the target knowledge base to obtain the first candidate text and the second candidate text corresponding to the user's query request. In one embodiment, keyword matching retrieval involves extracting core keywords from the user's query request and traversing and filtering text data containing these keywords in the target knowledge base; semantic similarity retrieval involves converting the user's query request and the knowledge base text into semantic vectors of a unified dimension through a semantic encoding model, calculating the similarity between the vectors, such as cosine similarity, and filtering texts that meet the similarity criteria.
[0061] In one embodiment, the first candidate text is a set of candidate texts selected from the target knowledge base by keyword matching retrieval that contain the core keywords of the user's query request, and each text unit has a literal association with the query request; the second candidate text is a set of candidate texts selected from the target knowledge base by semantic similarity retrieval that have a semantic association with the user's query request, and the text unit may not contain the query keywords, but its core meaning matches the query requirements.
[0062] In one embodiment, keyword matching retrieval of the target knowledge base is performed based on user query requests. This can be precise keyword matching, where the user query request is split into core keywords using a word segmentation tool. For example, the process of withdrawing housing provident funds is split into "housing provident fund," "withdrawal," and "process." The text in the target knowledge base is traversed, and text containing all or part of the core keywords is selected to form the first candidate text. Alternatively, keyword matching retrieval of the target knowledge base based on user query requests can be fuzzy keyword matching. Based on precise matching, synonyms and near-synonyms are allowed for keyword substitution. For example, "open account" is replaced with "open a new account." The search terms are expanded using a thesaurus, and then matching and filtering are performed to broaden the coverage of the first candidate text.
[0063] In one embodiment, the user query request and the target knowledge base text can be converted into multi-dimensional semantic vectors using pre-trained models such as Sentence-BERT and ERNIE, respectively. The cosine similarity between the vectors is calculated, and a similarity threshold such as 0.65 is set to filter texts with similarity higher than the threshold as second candidate texts. Alternatively, based on knowledge graph retrieval, a domain knowledge graph can be constructed, mapping the user query request to entities and relations in the knowledge graph. Texts containing the corresponding entities and relations can be retrieved from the knowledge base to form second candidate texts.
[0064] The first and second candidate texts are merged to obtain multiple candidate text fragments corresponding to the user's query request.
[0065] In one embodiment, the fusion processing of the first candidate text and the second candidate text refers to the process of integrating and optimizing the first candidate text and the second candidate text, with the aim of merging the two types of search results, removing redundant information, and obtaining comprehensive and high-quality candidate text fragments.
[0066] In one embodiment, duplicate texts between the first and second candidate texts can be removed first by text hashing or content similarity comparison, such as texts that are completely identical or have a similarity higher than 0.95. Then, the remaining texts are directly merged to obtain candidate text fragments.
[0067] In one embodiment, a search source weight can be assigned to each text in the first and second candidate texts after deduplication, such as a keyword matching search weight of 0.4 and a semantic similarity search weight of 0.6. The comprehensive score is calculated by combining the ranking score of the text in their respective search results, and the top N texts are selected as candidate text fragments after being sorted in descending order of the comprehensive score.
[0068] For example, when retrieving a target knowledge base based on a user query request, keyword matching retrieval is first performed. The user query request is broken down into multiple core keywords. The original content and related extended information of all retrieval slices in the target knowledge base are traversed to accurately filter out retrieval slices containing character combinations that are completely consistent with these core keywords, thus obtaining the first candidate text. Simultaneously, semantic similarity retrieval is performed. A pre-trained semantic encoding model is used to convert the user query request and each retrieval slice into semantic vectors of a unified dimension. The cosine similarity between each pair is calculated, and retrieval slices with similarity values higher than a preset threshold are selected to obtain the second candidate text. When fusing the two sets of candidate texts, duplicate retrieval slices are first removed by string comparison. Then, a ranking inverse fusion algorithm is used to calculate the inverse ranking of each remaining retrieval slice in the first and second candidate texts and sum them to obtain a fusion score. Finally, the retrieval slices are sorted from high to low according to the fusion score, and the top preset number of retrieval slices are selected as the final multiple candidate text fragments.
[0069] Keyword matching ensures the accuracy of candidate texts, while semantic similarity retrieval broadens the coverage of relevant information. The combination of the two effectively balances the accuracy and comprehensiveness of the retrieval. By integrating deduplication and ranking, the quality of candidate text fragments is further optimized, avoiding information redundancy.
[0070] In one embodiment, such as Figure 2 As shown, multiple candidate text fragments are sorted to obtain a candidate list, including: Step 201: Determine the basic relevance score between each candidate text fragment and the user query request, the matching degree score between the metadata information associated with each candidate text fragment and the user query request, and the edit distance relevance score between the user query request and the file identifier. In one embodiment, the basic relevance score is a quantitative indicator of the direct correlation between the candidate text fragment and the user's query request, reflecting the matching degree between the two at the semantic and content levels; metadata information is auxiliary descriptive information associated with the candidate text fragment, which can help determine its relevance to the query request. The metadata information can be the file title, summary, publication time, category tag, etc., corresponding to the text; the matching degree score is a quantitative indicator of the fit between the metadata information associated with the candidate text fragment and the user's query request. By supplementing the evaluation of the relevance of the text fragment through matching at the metadata level, the comprehensiveness of the evaluation can be improved.
[0071] In one embodiment, the file identifier is information used to uniquely identify the original file to which the candidate text fragment belongs, such as file name, file number, file path, etc.; the edit distance relevance score is a quantitative indicator calculated based on edit distance (the minimum number of single-character insertion, deletion, or replacement operations required to convert one string into another string), used to evaluate the similarity between the user query request and the file identifier, and indirectly reflect the association potential between the text fragment and the query request.
[0072] In one embodiment, the candidate text fragment can be concatenated with the user's query request and input into a pre-trained cross-encoder model. The model directly outputs the probability value in the 0-1 interval as the basic relevance score. The higher the probability value, the stronger the semantic matching degree. Alternatively, it can be calculated based on keyword overlap. The number of overlaps between the candidate text fragment and the core keywords of the user's query request is counted, and the result of normalizing by the number of overlapping keywords / total number of keywords is used to obtain the score in the 0-1 interval as the basic relevance score.
[0073] In one embodiment, keywords can be extracted from the metadata (such as the title) of candidate text fragments, and the overlap with the keywords in the user's query request can be calculated. After normalization, a matching score can be obtained. Alternatively, the metadata text and the user's query request can be converted into semantic vectors through the Sentence-BERT model, and the cosine similarity can be calculated. The similarity value can be used as the matching score. If the metadata contains business category tags, such as social security processing and housing provident fund withdrawal, it can also be determined whether the tags are consistent with the business category to which the user's query request belongs. If they are consistent, the score is 1. If they are inconsistent, a score in the range of 0-1 is assigned according to the degree of association.
[0074] In one embodiment, the edit distance between the user query request and the file identifier (such as the file name) can be calculated first, and then the score in the 0-1 range can be obtained by normalizing the result of 1 - edit distance / maximum character length of both. The higher the score, the stronger the similarity. Alternatively, different weights can be assigned to different character operations (insertion, deletion, replacement) (such as replacement having a higher weight than insertion / deletion), and the weighted edit distance can be calculated and then normalized to obtain the edit distance relevance score. Alternatively, the edit distance relevance score can be obtained by combining the length of the longest common substring and the edit distance, i.e., the length of the longest common substring / maximum character length of both × 0.5 + (1 - edit distance / maximum character length of both) × 0.5".
[0075] Step 202: Determine the target relevance score based on the basic relevance score, matching degree score, and edit distance relevance score; In one embodiment, the target relevance score is a final relevance quantification index obtained by combining the basic relevance score, the matching degree score, and the edit distance relevance score, which can comprehensively reflect the association priority between the text fragment and the user's query request.
[0076] In one embodiment, fixed weights are preset for the basic relevance score, matching degree score, and edit distance relevance score according to business needs. For example, the weight of the basic relevance score is 0.6, the weight of the matching degree score is 0.3, and the weight of the edit distance relevance score is 0.1. The target relevance score is obtained by calculating the target relevance score as: target relevance score × 0.6 + matching score × 0.3 + edit distance score × 0.1.
[0077] In one embodiment, the weights can also be dynamically adjusted based on the type of user query request. For example, short query requests focus on keyword matching, increasing the weight of the matching degree score; long query requests focus on semantic matching, increasing the weight of the basic relevance score, and then the weighted sum is performed to obtain the target relevance score.
[0078] Step 203: Sort the multiple candidate text fragments according to the target relevance score to obtain a candidate list.
[0079] In one embodiment, all candidate text fragments can be arranged in descending order of target relevance score, and text fragments with the same score can be sorted according to the retrieval time, i.e., the most recently retrieved one is first, forming a candidate list. Alternatively, the order can be adjusted by adding business rules on the basis of sorting in descending order of target relevance score. For example, in the government affairs scenario, "official release text" has higher priority than "third-party interpretation text", and in the enterprise customer service scenario, "latest business process text" has higher priority than "historical process text", ultimately forming a candidate list.
[0080] For example, when sorting multiple candidate text fragments, a basic relevance score between each candidate text fragment and the user's query request is first calculated using a relevance assessment model. This model receives the concatenated content of the user's query request and the candidate text fragments, and outputs a probability value between 0 and 1 as the basic relevance score. The higher the probability value, the stronger the direct relevance between the two. Next, the matching degree score between the metadata information associated with the candidate text fragments and the user's query request is calculated. The metadata information includes file subheadings, summaries, etc. The matching degree score between 0 and 1 is obtained by statistically analyzing and normalizing the number of overlapping 2-gram character combinations between the two. Simultaneously, the edit distance relevance score between the user's query request and the file identifier corresponding to the candidate text fragments is calculated. The minimum number of operations required to insert, delete, or replace a single character is used as the edit distance, and this score is normalized using the formula "1 - edit distance / maximum length of the two strings". Subsequently, the basic relevance score, matching degree score, and edit distance relevance score are weighted and summed using preset weights. The weight allocation is pre-configured according to the business scenario requirements. The summation result is the target relevance score for each candidate text segment. Finally, all candidate text segments are sorted from high to low according to the target relevance score to form a candidate list.
[0081] By comprehensively calculating the target relevance score through multi-dimensional scoring, the ranking results more accurately and reliably reflect the degree of association between candidate text fragments and query requests compared to single-dimensional evaluation. The weighted summation calculation method takes into account the importance of different dimensions, making the ranking results more in line with actual application needs.
[0082] In one embodiment, before inputting the target text fragment into the target generation model and generating the output text corresponding to the user's query request, the text generation method includes: Obtain the training dataset, which includes user query requests, positive samples associated with user query requests, negative samples associated with user query requests, and standard output text corresponding to user query requests. In one embodiment, the training dataset is a structured collection of data used to train and optimize the generative model, containing the input data required for model training and the corresponding reference output.
[0083] In one embodiment, positive samples are text fragments that are directly related to the user's query request and can provide valid information for generating a response. For example, when a user queries "social security transfer process", the text fragments in the knowledge base about the social security transfer steps are positive samples. Negative samples are text fragments that are unrelated to the user's query request, have conflicting information, or cannot support the generation of a valid response. They are used to help the model distinguish between valid and invalid information. For example, the text fragments of the medical insurance reimbursement process corresponding to the above query are negative samples. The standard output text is an ideal response text that is pre-set for the user's query request and meets the requirements of accurate information, logical coherence, and standardized format.
[0084] In one embodiment, the initial generation model is a basic text generation model that has not been specifically trained. It has basic text generation capabilities but is not adapted to specific query response scenarios, such as general GPT series models, BART models, etc.
[0085] In one embodiment, real user query requests can be extracted from historical user query logs, corresponding positive and negative samples can be manually selected from the knowledge base, and standard output text can be written by domain experts. After data cleaning, such as deduplication and filtering of invalid data, a training dataset is formed. Alternatively, based on existing knowledge base text, simulated user query requests and corresponding positive and negative samples can be generated in batches through a text generation model, then manually verified and corrected, and finally matched with standard output text to form a training dataset. Alternatively, publicly available question-and-answer datasets, such as MS MARCO and DuReader, can be introduced, and the format can be adapted according to the structure of user query, positive and negative samples, and standard output. After supplementing with specific domain data (such as government affairs and finance), a training dataset is formed.
[0086] Supervised fine-tuning of the initial generated model using the training dataset; In one embodiment, supervised fine-tuning refers to the training process of optimizing the initial generation model parameters through backpropagation on a labeled training dataset (standard output text), so that the model gradually learns the mapping relationship between user queries, sample information and standard output; reinforcement learning strategy refers to a training strategy that guides the model to iteratively optimize by setting a reward mechanism, without relying on fixed labels, but adjusting parameters based on the quality feedback of the model's generation results, so that the model gradually optimizes its generation behavior.
[0087] In one embodiment, the training dataset can be divided into a training set and a validation set according to a preset ratio (e.g., 9:1). The training set is used to fine-tune all parameters of the initial generated model, and the validation set is used to monitor the model performance to avoid overfitting. Alternatively, the initial generated model can be fine-tuned using general domain training data first, and then fine-tuned a second time using training data from a specific domain (e.g., enterprise customer service, government services) to gradually improve the model's adaptability to specific scenarios. Alternatively, efficient fine-tuning methods such as Low-Rank Adaptation (LoRA) and Adapter can be used to adjust only some parameters (rather than all parameters) of the initial generated model, reducing training resource consumption while ensuring the model's ability to learn text generation for specific scenarios.
[0088] The target generative model is obtained by iteratively optimizing the supervised fine-tuned initial generative model based on a reinforcement learning strategy.
[0089] In one embodiment, the quality of the model-generated results can be scored manually, such as accuracy, fluency, and relevance. The scores are then converted into reward signals to guide the model to adjust its parameters. After multiple iterations, the target generation model is obtained. Alternatively, an automatic reward mechanism can be used to reinforce learning. The reward value is calculated using preset automatic evaluation indicators, such as semantic similarity to the standard output, reasonableness of text length, and format standardization. The model parameters are then iteratively optimized based on the reward value to obtain the target generation model.
[0090] In acquiring the target generation model, a training dataset is first constructed. This dataset contains user query requests covering various common consultation questions in business scenarios. Positive samples associated with the query requests are relevant text fragments selected from the target knowledge base that accurately respond to the corresponding queries. Negative samples are distracting text fragments unrelated to the query requests or containing conflicting information. The standard output text is a logically rigorous reference text generated based on the core information of the positive samples, conforming to business expression standards. Subsequently, the initial generation model is fine-tuned under supervision using this training dataset. User query requests, corresponding positive and negative samples are used as model inputs, and the standard output text is used as the supervision target. The network parameters of the model are iteratively adjusted using the backpropagation algorithm, enabling the model to initially learn the ability to filter effective information from the input text and generate corresponding response text. After fine-tuning, the model is iteratively optimized using a reinforcement learning strategy. The optimization objective is initially set as the accuracy, reliability, and compliance of the output text. The model's output results are collected through multiple generation iterations, and the output results are evaluated and fed back to the model using a pre-set reward mechanism. The model parameters are continuously adjusted to optimize the generation strategy. After a pre-set number of iterations, the target generation model is finally obtained.
[0091] Supervised fine-tuning using a training dataset containing both positive and negative samples lays the foundation for the model to distinguish between valid and distracting information. Combined with iterative optimization of reinforcement learning strategies, the accuracy and adaptability of the model's generated text are further improved, ensuring that the target generation model can stably output high-quality text that meets business requirements.
[0092] In one embodiment, the reinforcement learning strategy employs a multi-dimensional reward mechanism, which includes at least one of the following: a reward based on the reasonableness of the length of the training output text generated during model training, a reward based on the matching degree between the training output text generated during model training and the standard output text, a credibility reward based on the citation priority of the target text fragment, and a reward based on the format standardization of the training output text generated during model training.
[0093] In one embodiment, the multi-dimensional reward mechanism refers to evaluating the quality of the training output text generated by the model from multiple independent dimensions and assigning reward values accordingly. By using comprehensive reward signals to guide model optimization, the one-sidedness of single-dimensional evaluation is avoided.
[0094] In one embodiment, the reward for the reasonable length of the training output text is a reward dimension based on whether the character length of the training output text conforms to a preset reasonable range. This is used to guide the model to generate text of moderate length and compact information, avoiding excessively long and redundant text or excessively short and incomplete text.
[0095] In one embodiment, the reward for the matching degree between the training output text and the standard output text is a reward dimension that measures the semantic and core information-level fit between the training output text and the preset standard output text.
[0096] In one embodiment, the credibility reward for the reference priority of the target text fragment is based on the reward dimension of the priority of the target text fragment referenced by the model in the candidate list, which guides the model to prioritize referencing target text fragments with higher relevance and improves the information reliability of the generated text.
[0097] In one embodiment, the reward for the standardization of the training output text format is a reward dimension for evaluating whether the training output text conforms to the preset structured format requirements, which is used to ensure that the model generates text with a uniform format that is easy to read and understand.
[0098] In one embodiment, the training output text refers to the intermediate text result generated by the initial generative model based on the input data during the reinforcement learning training process.
[0099] In one embodiment, a reasonable length range (e.g., 20-60 characters) can be preset. When the length of the training output text is within the range, the reward value is set to 1; when the length is less than the lower limit of the range (e.g., <20 characters), the reward value is calculated as "length / 20" (0-1 range); when the length is greater than the upper limit of the range (e.g., >60 characters), the reward value is calculated as "60 / length" (0-1 range); the length range can also be subdivided into multiple gradients (e.g., 20-30 characters, 31-45 characters, 46-60 characters), with corresponding reward values of 0.8, 1, and 0.8 respectively.
[0100] In one embodiment, the training output text and the standard output text can be converted into semantic vectors using the Sentence-BERT model, and the cosine similarity can be calculated. The similarity value (0-1 range) can be directly used as the reward value, with higher similarity resulting in higher rewards. Alternatively, the core information points (such as key steps and essential conditions) of the standard output text can be extracted, and the number of core information points contained in the training output text can be counted. The reward value can then be calculated based on the number of core information points contained in the training output text divided by the total number of core information points.
[0101] In one embodiment, the correspondence between candidate list ranking and reward value can be preset, such as 0.6 for ranking 1st, 0.3 for ranking 2nd, 0.1 for ranking 3rd, and 0 for ranking ≥4th. If the model references multiple fragments, the reward value corresponding to the highest ranking is taken. Alternatively, based on the ranking gradient reward, if the model references multiple fragments ranked in the top 3 at the same time, the reward value is calculated according to the superposition rule, which encourages the model to integrate information from multiple highly relevant fragments.
[0102] In one embodiment, format validation rules can be preset, such as requiring the inclusion of core conclusions, specific explanations, and blank lines between paragraphs. Validation is performed using regular expressions or format validation tools. Full compliance earns 1 point, while non-compliance earns 0 points. Alternatively, format specifications can be subdivided into basic specifications, such as no garbled characters and correct punctuation, and advanced specifications, such as complete tags and neat layout. Meeting basic specifications earns 0.5 points, meeting advanced specifications earns an additional 0.5 points, and partial compliance deducts points proportionally.
[0103] In one embodiment, fixed weights can be preset for each dimension (such as matching degree 0.4, length 0.2, reference priority 0.2, format 0.2), and the final reward value = reward value of each dimension × sum of corresponding weights; the weights can also be adjusted according to the application scenario (such as increasing the format specification weight to 0.3 in government scenarios; increasing the matching degree weight to 0.5 in customer service scenarios), and the matching weight configuration can be adjusted according to the scenario category.
[0104] The multi-dimensional reward mechanism guides model training from multiple aspects, including text length, information accuracy, citation credibility, and format standardization. This avoids the model generating redundant, erroneous, or non-standard text, thereby improving the quality, stability, and practicality of the target generation model's output text.
[0105] like Figure 3 As shown, Figure 3 This is a flowchart illustrating another text generation method provided in this disclosure, comprising three stages: retrieval and recall, candidate list sorting, and answer generation. First, in the retrieval and recall stage, a user query is received, and a hybrid retrieval is performed based on a customer service knowledge base to obtain relevant text fragments. Next, in the candidate list sorting stage, the candidate list is first sorted using Qwen3 and Reranker, and then combined with a similarity score based on a code strategy to complete the sorting based on similarity scores. Finally, in the answer generation stage, the top 3 candidate lists are selected and input into the generation model to generate the answer and return it. The specific process is as follows: 1. Retrieval and recall.
[0106] User Query: The question entered by the customer service representative.
[0107] Customer service knowledge base: structured data of underlying knowledge documents. Meanwhile, to bridge the semantic gap between queries and document paragraph content, for each sample, in addition to calculating similarity between the original paragraph and the query for retrieval, a large model extracts the query, keywords, and title from the original paragraph to match the input query. Each document paragraph is structured as a sample, representing a slice in the customer service knowledge base.
[0108] Hybrid retrieval: Retrieves relevant slice samples from the customer's knowledge base based on the user's query. To ensure recall, hybrid retrieval is used.
[0109] 2. Sort the candidate list.
[0110] Ranking based on Qwen3-Reranker: The Qwen3-Reranker-4B is used to rank the recall results. When ranking, the model input is the Query and concatenated content containing more information about the sliced samples (such as file names, file subheadings, original paragraphs, etc.). A Prompt is set to increase the priority of the preferred data in the ranking.
[0111] Similarity score based on code strategy: Based on the similarity score output during Qwen3-Reranker sorting, a post-strategy operation based on code is added to calculate the edit distance between the query and the file name and the 2-gram matching degree between the query and the original paragraph, and a score is added to the similarity value.
[0112] Similarity score sorting in descending order: Before the input to the generative model to generate the answer, the similarity values of the slice samples are sorted in descending order. Finally, the top 3 candidate lists are input into the generative model as a reference to generate the final answer.
[0113] 3. Answer generation.
[0114] TOP3 Candidate List: Take the top 3 samples from the sorted list and add them to the Prompt of the generative model as the basis for generating the answer.
[0115] Generative model: A generative expert model trained on Qwen3-8B using Retrieval-Augmented Fine-Tuning (RAFT) and reinforcement learning methods.
[0116] Generate Answer Return: Returns the generated final answer.
[0117] Specifically, 1. The retrieval and recall process introduces a large model to perform structured processing of text paragraphs from a customer service perspective.
[0118] a. When constructing the customer service knowledge base, to mitigate the semantic differences between queries and document paragraphs and improve retrieval recall, a large-scale model is introduced for each paragraph to extract potential queries, keywords, and paragraph titles from the customer service perspective. During the retrieval phase, in addition to similarity matching between the original paragraph content and the query, these extracted queries, keywords, and titles also participate in the retrieval matching to enhance recall. Each document paragraph is structured as an independent sample and serves as a retrieval slice in the knowledge base.
[0119] b. Hybrid Search Hybrid retrieval is a combination of semantic retrieval and full-text retrieval. The proportion of text slices from each can be controlled by setting a ratio. Full-text retrieval only finds keywords that exactly match the user's query, resulting in greater precision. Semantic retrieval, on the other hand, not only finds keywords that exactly match the user's question but also keywords with semantic similarities. For example, when a user searches for "food," semantic search can also find related keywords such as "delicious" and "tasty," greatly enriching the diversity of search results.
[0120] In practical use, the maximum number of search results is set to 100, the proportion of full-text search is 0.4 (40 slices), and the proportion of semantic search is 0.6 (60 slices).
[0121] The Reciprocal Rank Fusion (RRF) algorithm is used to rank the combined results of semantic and full-text searches. RRF is a simple and effective ranking fusion method often used to combine multiple search result lists into a final ranking. The RRF algorithm calculates the reciprocal rank of each candidate item in multiple ranking lists, then adds these reciprocals to obtain a fusion score, and finally re-ranks the candidates based on this score. It can provide a comprehensive ranking in a relatively short time, ensuring a certain level of search effectiveness while having relatively low system performance overhead.
[0122] 2. Sort the candidate list.
[0123] After the retrieval and recall, the recall results are finely sorted, with the aim of ranking the most preferred data into the top 3.
[0124] a.Qwen3-Reranker sorting: The core of the standard workflow of Qwen3-Reranker-4B is to determine the relevance between a "document" and a "query." This requires constructing a specific input format, and then having the model determine the probability of outputting "yes" or "no." The input template typically looks like this: The input template for Qwen3-Reranker is: text=[ {"role": "system", "content": "Judge whether the Document meets therequirements based on the Query and the Instruct provided. Note that the answer can only be \"yes\" or \"no\"."}; {"role": "user", "content": f" <instruct>: {instruction}\n\n <query>:{query}\n\n <document>: {doc}"}; ] The Instruct in the input template is a prompt word that can be manually written, such as adding role information, task descriptions, and precautions, to help the model determine whether the Document can satisfy the query, thus affecting the probability of outputting "yes" or "no".
[0125] During the ranking phase, the Qwen3-Reranker-4B model can be used to re-rank the recall results. The model's input consists of the user query and concatenated content containing rich contextual information from sliced samples, such as filenames, file headings, and original paragraphs. By designing appropriate prompts, the model can be guided to prioritize the ranking of preferred data, thereby optimizing the overall retrieval performance.
[0126] b. Similarity score based on code strategy and sorting similarity scores in descending order.
[0127] After obtaining the probability of "yes" or "no" from the Qwen3-Reranker-4B's judgment output, the mathematical expression for determining the similarity score is as follows:
[0128] Where true_score is the probability of judging "yes" and false_score is the probability of judging "no".
[0129] Based on the similarity score mentioned above, add the score calculated using the following method, including: 2-gram normalized matching degree calculation formula: Q is a multiset of all 2-grams in the query (with possible repetitions), and D is a multiset of all 2-grams in the document. For any 2-gram g, let its frequency in Q be denoted as . In D, it is .
[0130] but:
[0131] The higher the score, the more completely the document contains the character combination characteristics of the query.
[0132] Similarity calculation based on edit distance normalization: Edit distance represents the minimum number of single-character editing operations (insertion, deletion, replacement) required to transform one string into another. q is the query string, len(q) is its length, f is the filename string, len(f) is its length, and lev(q,f) is the edit distance between the two. The mathematical expression is as follows: but:
[0133] The higher the score, the stronger the relevance of the filename to the slice to the query.
[0134] c. After the above process, sort the similarity scores of the sliced samples in descending order, and put the content that is more relevant to the query at the top.
[0135] 3. Answer generation.
[0136] Even though the top 3 candidate slices have been obtained after ranking, they may still contain content that is weakly related to the query. To reduce the interference of this noise on the generative model, a combination of RAFT and reinforcement learning can be used to train the generative model, thereby improving its ability to generate correct answers based on the retrieval results.
[0137] a. Generative model training, including: The core idea of i.RAFT is as follows: 1. Training data: Question (Q): Clearly define the specific problem to be solved; Document set (D): Includes documents directly related to the problem (positive samples) and irrelevant distracting documents (negative samples); Answer (A): A chain-of-thought style answer generated from the positive sample documents, emphasizing the reasoning process and cited sources.
[0138] 2. Training process: The model is trained using supervised fine-tuning, enabling it to generate answers from provided documents and questions.
[0139] The model is trained to identify and ignore irrelevant documents, referencing only those relevant to the question.
[0140] Emphasis is placed on generative reasoning processes, such as chain-of-thought (CoT), to improve the logic and credibility of the model.
[0141] RAFT improves the model's accuracy and robustness in question-answering tasks by introducing positive and negative sample documents and chain-like reasoning. ii. Drawing on the adversarial training concept of RAFT, using the training data constructed therein and setting corresponding reward functions, the generative model is optimized through reinforcement learning to ensure that its output is consistent with the business objectives.
[0142] 1. Use reinforcement learning algorithms: Group Relative Policy Optimization (GRPO) is a reinforcement learning strategy. Its core idea is to sample multiple outputs from the old policy for each problem and then optimize the new policy based on the rewards of these outputs.
[0143] The objective function of GRPO is as follows:
[0144] In the formula, the first term is the policy gradient term, which encourages the model to generate action sequences with higher rewards; the second term is the pruning term, which limits the magnitude of policy updates and prevents policy collapse; the third term is the Kullback-Leibler Divergence (KL) penalty term, which prevents the new policy from deviating too far from the reference normal policy; G represents the total number of training samples; i represents the i-th sample; |o i | represents the length of the i-th output sequence, t is the time step, and o t Let q be the token generated in step t, and q be the input query. Let represent the advantage function estimate of the i-th sample at step t, and β be the penalty coefficient.
[0145] 2. Reward functions, including length reward function, result reward, process credibility reward, and format reward.
[0146] Length reward function: Without sacrificing the correctness of the answer, it encourages the model to generate "concise and complete" text, suppresses redundancy and over-generation, and avoids information loss due to excessive pursuit of brevity.
[0147]
[0148] A constant maximum reward is given within the entire interval of 10-50 characters, and the reward decreases sharply outside the interval; is the interval boundary value (10 or 50) closest to L.
[0149] Result reward: Each sample has a real label. The model's output is checked. If the output matches the real label, the reward is 1; otherwise, it is 0.
[0150] Credible process reward: This reward guides the model to prioritize higher-ranked search results and establish the correct inference chain.
[0151]
[0152] The reward is discrete. It only focuses on the ranking of the fragment cited by the model for the first time, and does not add rewards for subsequent citations, thereby forcing the model to develop the habit of "prioritizing the use of the best information".
[0153] Formatting Reward: To ensure the model output conforms to the predefined JavaScript Object Notation (JSON) format, we introduce a format checking mechanism during training: the json_repair library is used to validate the model output. If the output format is correct, a reward value of 1 is given; otherwise, 0 is given. In the early stages of training (after approximately 2-3 epochs), if the model can consistently output correctly formatted JSON (i.e., the formatting reward remains at 1), it is considered that the model has mastered the target output structure. Thereafter, the formatting reward will be removed in subsequent training sessions to focus on optimizing semantic relevance and alignment with business objectives.
[0154] Finally, by combining RAFT training data with a customized reward function, the reinforcement learning training of the generative model was completed. This training significantly improved the accuracy and stability of the model's output, ultimately creating a knowledge base assistant capable of accurately, reliably, and automatically generating high-quality answers. The application of this assistant will effectively reduce the workload of customer service representatives and improve overall customer service efficiency and quality.
[0155] In summary, the solution provided in this public disclosure is as follows: First, by acquiring candidate text fragments and sorting and filtering them, target text fragments are obtained, providing accurate information support for model generation. At the same time, the target generation model is optimized through retrieval enhancement fine-tuning and reinforcement learning training, which significantly improves the ability to utilize retrieval information, effectively ensuring the matching degree and accuracy of the output text with the user's query request, avoiding factual bias and omission of key information, and thus ensuring the matching degree and accuracy of the output text with the user's query request, resulting in high-quality output text.
[0156] To implement the text generation method provided in this disclosure, this disclosure also provides a text generation apparatus, such as... Figure 4 As shown. Figure 4 This is a schematic diagram of the structure of a text generation device provided in an embodiment of the present disclosure. The text generation device 400 includes: The acquisition unit 401 is used to acquire multiple candidate text fragments corresponding to the user's query request; The sorting unit 402 is used to sort multiple candidate text fragments to obtain a candidate list; Selection unit 403 is used to select at least one target text fragment from the candidate list; The generation unit 404 is used to input the target text fragment into the target generation model and generate the output text corresponding to the user's query request. The target generation model is obtained through retrieval enhancement fine-tuning and reinforcement learning training.
[0157] In one embodiment, the acquisition unit 401 is specifically used for: Get the user's query request; The target knowledge base is retrieved based on the user's query request, resulting in multiple candidate text fragments corresponding to the user's query request.
[0158] In one embodiment, the acquisition unit 401 is specifically used for: Based on the user's query request, keyword matching and semantic similarity retrieval are performed on the target knowledge base to obtain the first candidate text and the second candidate text corresponding to the user's query request. The first and second candidate texts are merged to obtain multiple candidate text fragments corresponding to the user's query request.
[0159] In one embodiment, the sorting unit 402 is specifically used for: Determine the basic relevance score between each candidate text fragment and the user query request, the matching degree score between the metadata information associated with each candidate text fragment and the user query request, and the edit distance relevance score between the user query request and the file identifier; The target relevance score is determined based on the basic relevance score, the matching degree score, and the edit distance relevance score. Multiple candidate text fragments are sorted according to their target relevance scores to obtain a candidate list.
[0160] In one embodiment, the text generation apparatus 400 further includes an optimization unit, which is used to: Obtain the training dataset, which includes user query requests, positive samples associated with user query requests, negative samples associated with user query requests, and standard output text corresponding to user query requests. Supervised fine-tuning of the initial generated model using the training dataset; The target generative model is obtained by iteratively optimizing the supervised fine-tuned initial generative model based on a reinforcement learning strategy.
[0161] In one embodiment, the reinforcement learning strategy employs a multi-dimensional reward mechanism, which includes at least one of the following: a reward based on the reasonableness of the length of the training output text generated during model training, a reward based on the matching degree between the training output text generated during model training and the standard output text, a credibility reward based on the citation priority of the target text fragment, and a reward based on the format standardization of the training output text generated during model training.
[0162] It should be noted that the text generation device provided in the above embodiments is only illustrated by the division of the above program modules. In practical applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the text generation device can be divided into different program modules to complete all or part of the processing described above. In addition, the text generation device provided in the above embodiments and the text generation method embodiments provided in this disclosure belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0163] Figure 5 This is a schematic diagram of the hardware composition structure of the electronic device provided in the embodiments of this disclosure, such as... Figure 5 As shown, the electronic device 500 includes at least one processor 502; and a memory 501 communicatively connected to the at least one processor 502; wherein the memory 501 stores instructions executable by the at least one processor 502, the instructions being executed by the at least one processor 502 to implement the steps of the text generation method of the present disclosure embodiments.
[0164] Optionally, the electronic device may specifically be a text generation device according to the embodiments of this application, and the electronic device may implement the corresponding processes implemented by the text generation device in the various methods of the embodiments of this application. For the sake of brevity, it will not be described in detail here.
[0165] It is understood that the electronic device also includes a communication interface 503. Various components in the electronic device are coupled together via a bus system 504. It is understood that the bus system 504 is used to implement communication between these components. In addition to a data bus, the bus system 504 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 5 The general designated all buses as Bus System 504.
[0166] It is understood that memory 501 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memory 501 described in this embodiment of the invention is intended to include, but is not limited to, these and any other suitable types of memory.
[0167] The methods disclosed in the above embodiments can be applied to or implemented by processor 502. Processor 502 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above methods can be completed by integrated logic circuits in the hardware of processor 502 or by instructions in software form. Processor 502 may be a general-purpose processor, DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 502 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, specifically memory 501. Processor 502 reads information from memory 501 and, in conjunction with its hardware, completes the steps of the aforementioned methods.
[0168] In an exemplary embodiment, the electronic device may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), FPGAs, general-purpose processors, controllers, MCUs, microprocessors, or other electronic components to perform the aforementioned method.
[0169] This disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, which are used to cause a computer to execute the steps of the text generation method of the present invention.
[0170] Optionally, the computer-readable storage medium can be applied to the text generation apparatus in the embodiments of this application, and the computer instructions cause the computer to execute the corresponding processes implemented by the text generation apparatus in the various methods of the embodiments of this application. For the sake of brevity, these will not be described in detail here.
[0171] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the text generation method provided in this embodiment of the invention.
[0172] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.
[0173] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.
[0174] In addition, in the various embodiments of the present invention, each functional unit can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0175] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.
[0176] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.
[0177] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.< / document> < / query> < / instruct>
Claims
1. A text generation method, characterized in that, include: Retrieve multiple candidate text fragments corresponding to the user's query request; The candidate text fragments are sorted to obtain a candidate list; Select at least one target text fragment from the candidate list; The target text fragment is input into the target generation model to generate output text corresponding to the user's query request. The target generation model is obtained through retrieval enhancement fine-tuning and reinforcement learning training.
2. The method according to claim 1, characterized in that, The step of obtaining multiple candidate text fragments corresponding to the user's query request includes: Get the user's query request; Based on the user's query request, the target knowledge base is retrieved to obtain multiple candidate text fragments corresponding to the user's query request.
3. The method according to claim 2, characterized in that, The step of retrieving the target knowledge base based on the user query request to obtain multiple candidate text fragments corresponding to the user query request includes: Based on the user query request, keyword matching retrieval and semantic similarity retrieval are performed on the target knowledge base to obtain a first candidate text and a second candidate text corresponding to the user query request. The first candidate text and the second candidate text are fused to obtain multiple candidate text fragments corresponding to the user's query request.
4. The method according to claim 1, characterized in that, The process of sorting the plurality of candidate text fragments to obtain a candidate list includes: Determine the basic relevance score between each candidate text fragment and the user query request, the matching degree score between the metadata information associated with each candidate text fragment and the user query request, and the edit distance relevance score between the user query request and the file identifier; The target relevance score is determined based on the basic relevance score, the matching degree score, and the edit distance relevance score. The candidate text fragments are sorted according to the target relevance score to obtain the candidate list.
5. The method according to claim 1, characterized in that, Before inputting the target text fragment into the target generation model to generate the output text corresponding to the user query request, the method includes: Obtain a training dataset, which includes user query requests, positive samples associated with user query requests, negative samples associated with user query requests, and standard output text corresponding to user query requests. The initial generated model was then fine-tuned under supervision using the training dataset. The target generative model is obtained by iteratively optimizing the supervised fine-tuned initial generative model based on a reinforcement learning strategy.
6. The method according to claim 5, characterized in that, The reinforcement learning strategy employs a multi-dimensional reward mechanism, which includes at least one of the following: a reward based on the reasonableness of the length of the training output text generated during model training, a reward based on the matching degree between the training output text generated during model training and the standard output text, a credibility reward based on the citation priority of the target text fragment, and a reward based on the format standardization of the training output text generated during model training.
7. A text generation device, characterized in that, include: The acquisition unit is used to acquire multiple candidate text fragments corresponding to the user's query request; A sorting unit is used to sort the plurality of candidate text fragments to obtain a candidate list; A selection unit is used to select at least one target text fragment from the candidate list; The generation unit is used to input the target text fragment into the target generation model and generate output text corresponding to the user query request. The target generation model is obtained through retrieval enhancement fine-tuning and reinforcement learning training.
8. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 6.
10. A computer program product comprising a computer program that, when executed by a processor, implements the method of any one of claims 1 to 6.