Question and answer method, training method of question and answer large model, related equipment and program product
By training a large model using reinforcement learning and utilizing a configured knowledge base and a referential consistency reward function, the problem of knowledge illusion in the content generated by large language models is solved, thereby improving the accuracy and reliability of the generated answers.
Patent Information
- Application Number
- CN202511184894.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-11-21
AI Technical Summary
Large language models suffer from knowledge illusion when generating content, especially in open-domain question answering and long text generation scenarios. Existing retrieval-enhanced generation techniques struggle to effectively address the issue of insufficient matching between fabricated knowledge citation sources and answer content.
A large model is trained using reinforcement learning. Knowledge fragments related to the question are retrieved through a configured knowledge base. A reward function is designed based on reference consistency and fact consistency to encourage the knowledge references generated by the model to be consistent with the truth labels and the answer content to be consistent with the semantics of the knowledge fragments.
It improves the accuracy and reliability of content generated by large models, effectively alleviates the knowledge illusion problem, and enhances information retrieval capabilities and the consistency of answer results.
Smart Images

Figure CN120994796A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and more particularly, to a question answering method, a training method of a question answering large model, related equipment and a program product. BACKGROUND
[0002] In recent years, large language models have shown strong generation capabilities in the field of natural language processing, but the knowledge illusion problem existing in the generated content has become a key bottleneck restricting the landing of technology. Knowledge illusion refers to the generation of seemingly reasonable but actually incorrect content by the model in the absence of factual basis, and this phenomenon is particularly significant in open-domain question answering and long text generation scenarios.
[0003] Existing research shows that the generation of knowledge illusion is not only due to the timeliness limitation of the pre-training data of the model, but also related to the lack of real-time knowledge verification of the self-recurrence generation mechanism. Retrieval-augmented generation technology combines external knowledge base retrieval with text generation, providing a new technical path to alleviate knowledge illusion. While generating answers, it can also generate knowledge reference sources. However, there are still problems such as fabricating knowledge reference sources and insufficient matching between generated answer content and reference knowledge. SUMMARY
[0004] In view of the above problems, the present application is proposed to provide a question answering method, a training method of a question answering large model, related equipment and a program product to alleviate the large model illusion problem and improve the reliability of the generated content. The specific scheme is as follows:
[0005] In a first aspect, a question answering method is provided, comprising:
[0006] obtaining question data;
[0007] retrieving a knowledge fragment related to the question data in a configured knowledge base;
[0008] generating feedback information by a large model trained through reinforcement learning, with reference to the question data and the knowledge fragment, the feedback information including answer content and knowledge reference sources;
[0009] wherein the reward function of the large model reinforcement learning process is determined based on reference consistency and / or factual consistency, the reference consistency being used to encourage the model to generate knowledge reference sources consistent with the knowledge reference sources in the true value label, and the factual consistency being used to encourage the model to generate answer content consistent with the semantics of the knowledge fragment corresponding to the generated knowledge reference sources.
[0010] In a possible design, in a third implementation manner of the first aspect of the embodiment of the present application, the knowledge base includes knowledge segments obtained by splitting the original knowledge corpus at two or more splitting granularities.
[0011] In a possible design, in a third implementation manner of the first aspect of the embodiment of the present application, the process of generating feedback information by the large model trained through reinforcement learning, with reference to the question data and the knowledge segments, includes:
[0012] indicating the large model to filter, from the knowledge segments related to the question data, a knowledge segment to be referenced for answering the question data, to obtain a filtered knowledge segment;
[0013] indicating the large model to generate feedback information with reference to the question data and the filtered knowledge segment.
[0014] In a possible design, in a third implementation manner of the first aspect of the embodiment of the present application, before indicating the large model to filter, from the knowledge segments, a knowledge segment to be referenced for answering the question data, the method further includes:
[0015] performing fine-grained splitting on a knowledge segment of a set coarse granularity in the knowledge segments, to obtain a split knowledge segment;
[0016] adding a reference source to each knowledge segment.
[0017] In a possible design, in a third implementation manner of the first aspect of the embodiment of the present application, the process of indicating the large model to filter, from the knowledge segments, a knowledge segment to be referenced for answering the question data, to obtain a filtered knowledge segment, includes:
[0018] obtaining a first prompt instruction prompt template, the prompt template including a task instruction, a knowledge slot, and a question slot, the task instruction being used to instruct the large model to filter, from the knowledge segments in the knowledge slot, a knowledge segment to be referenced for answering question data in the question slot, and to generate answering content based on the filtered knowledge segment;
[0019] filling the knowledge segments related to the question data into the knowledge slot, filling the question data into the question slot, to obtain a first prompt instruction prompt, and sending the first prompt instruction prompt into the large model to obtain a filtered knowledge segment output by the large model.
[0020] In a possible design, in a third implementation manner of the first aspect of the embodiment of the present application, before indicating the large model to generate feedback information with reference to the question data and the filtered knowledge segment, the method further includes:
[0021] De-duplicate the knowledge fragments after the screening.
[0022] In a possible design, in another implementation manner of the first aspect of the embodiment of the present application, the method further includes:
[0023] In response to the trigger of the knowledge reference source, the knowledge fragment corresponding to the knowledge reference source is obtained and displayed.
[0024] In a second aspect, a training method of a question and answer large model is provided, including:
[0025] Obtaining question and answer training data, the question and answer training data including an input sample and a true value label, the input sample including a sample question and a related knowledge fragment, and the true value label including an answer content marked with a knowledge reference source;
[0026] Sending the input sample into the question and answer large model to obtain an output of the question and answer large model, and training the question and answer large model in a manner of reinforcement learning, a reward function of the reinforcement learning being determined based on reference consistency and / or fact consistency, the reference consistency being used to encourage the knowledge reference source generated by the model to be consistent with the knowledge reference source in the true value label, and the fact consistency being used to encourage the answer content generated by the model to be consistent with the semantics of the knowledge fragment corresponding to the knowledge reference source.
[0027] In a possible design, in another implementation manner of the second aspect of the embodiment of the present application, the calculation process of the reference consistency includes:
[0028] Based on the knowledge reference source generated by the question and answer large model and the knowledge reference source in the true value label, a first evaluation index value is calculated, the first evaluation index being used to measure the consistency of the knowledge reference source generated by the question and answer large model and the knowledge reference source in the true value label.
[0029] In a possible design, in another implementation manner of the second aspect of the embodiment of the present application, the calculation process of the fact consistency includes:
[0030] Respective semantic vectors of the answer content and the knowledge reference source generated by the question and answer large model are obtained, and a vector similarity of the two semantic vectors is calculated, the vector similarity representing the fact consistency.
[0031] In a possible design, in another implementation manner of the second aspect of the embodiment of the present application, the reward function is obtained by weighted summation of the reference consistency and the fact consistency.
[0032] In a possible design, in another implementation manner of the second aspect of the embodiment of the present application, the process of obtaining the question and answer training data includes:
[0033] obtaining a sample question, and retrieving a knowledge fragment related to the sample question in a configured knowledge base, to form an input sample composed of the sample question and the related knowledge fragment;
[0034] indicating the large model to filter a knowledge fragment to be referenced for answering the sample question from the knowledge fragment related to the sample question, to obtain a filtered knowledge fragment;
[0035] indicating the large model to generate a final result by referencing the sample question and the filtered knowledge fragment, the final result including answer content marked with a knowledge reference source, and using the final result as a true value label corresponding to the input sample.
[0036] In a possible design, in another implementation manner of the second aspect of the embodiments of the present application, the process of obtaining a sample question includes:
[0037] indicating the large model to ask questions on the provided corpus document, to obtain a question generated by the large model as the sample question.
[0038] In a third aspect, an electronic device is provided, including a memory and a processor.
[0039] The memory is configured to store a program.
[0040] The processor is configured to execute the program, to implement each step of the question-answering method described in any one of the preceding first aspects of the present application, or to implement each step of the training method of the question-answering large model described in any one of the preceding second aspects of the present application.
[0041] In a fourth aspect, a readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement each step of the question-answering method described in any one of the preceding first aspects of the present application, or to implement each step of the training method of the question-answering large model described in any one of the preceding second aspects of the present application.
[0042] In a fifth aspect, a computer program product is provided, which includes a computer program, and the computer program is executed by a processor to implement each step of the question-answering method described in any one of the preceding first aspects of the present application, or to implement each step of the training method of the question-answering large model described in any one of the preceding second aspects of the present application.
[0043] By the above technical solution, the application adopts a retrieval enhancement generation scheme based on a large model, which can generate answer content and knowledge reference sources simultaneously. The large model used by the application is pre-trained through reinforcement learning, and in order to improve the correctness and factuality of the knowledge reference sources generated by the large model, the reward function of the reinforcement learning process is determined based on reference consistency and / or fact consistency. The reference consistency is used to encourage the knowledge reference sources generated by the model to be consistent with the knowledge reference sources in the true value label, and the fact consistency is used to encourage the answer content generated by the model to be consistent with the semantics of the knowledge segment corresponding to the generated knowledge reference sources. By reinforcing the learning training of the large model according to the above reward function, the accuracy and reliability of the content generated by the large model can be improved, and the large model hallucination problem can be effectively alleviated. BRIEF DESCRIPTION OF DRAWINGS
[0044] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a description of preferred embodiments, and are not meant to limit the present application. Moreover, the same reference numerals in the attached drawings refer to the same or like components throughout the several drawings. In the drawings:
[0045] Figure 1 An embodiment system architecture schematic diagram of a question and answer method and a question and answer large model training method provided by the application;
[0046] Figure 2 A question and answer large model training process schematic diagram provided by the application;
[0047] Figure 3 A question and answer training data generation flow schematic diagram provided by the application;
[0048] Figure 4 A question and answer method flow schematic diagram provided by the application;
[0049] Figure 5 A structure schematic diagram of an electronic device provided by the application. DETAILED DESCRIPTION
[0050] Before introducing the scheme of the application, first, the related terms involved in the scheme of the application are introduced.
[0051] Reinforcement Learning: In the reinforcement learning framework, there is usually an "agent" and an "environment". The agent decides an action a based on its policy π(s) at each step, and the environment gives a new state and a reward r based on the action. The agent collects the reward and continues to the next step. This cycle constitutes a time series process until the termination condition is reached (such as achieving the goal or timeout, etc.).
[0052] In language models (especially large language models, LLM), a "question" (such as a text prompt) can also be regarded as a state given by the environment, and the model (agent) outputs the next token (action), and then repeats until a complete answer is generated. A quality score for the entire answer is given by a human or an additional reward model, or a local reward is given at each token (or step) moment. Although large language models seem to have some differences from traditional Markov Decision Processes (MDP) in reinforcement learning, they can also be abstracted as a state-action-reward-state-action mechanism in essence.
[0053] State s: For language models, the generated token sequence (and the current question) can be regarded as a compressed state; in traditional reinforcement learning RL, it is some vector or feature observed by the environment.
[0054] Action a: In the language model generation scenario, the action can be "selecting the next token from the vocabulary"; in the robot or game environment, it is "moving, rotating, jumping", etc.
[0055] Reward r: an indicator of good or bad. In language model alignment, it is common to train a reward model to score; or directly use rules to judge whether the answer is correct, etc.
[0056] Policy π: the probability distribution function π(a|s) of how the agent selects action a in state s. In language models, this is the conditional distribution of generating each token.
[0057] Value function and advantage function: In typical policy gradient methods such as PPO, a value function is usually introduced, which roughly represents how much reward can be expected in the future in the current state; or further, the "advantage function" can be introduced after each action, which measures "how much better this action is than the average level". If only the direct guidance of the reward is used in training, each sample may have a large variance, and the convergence is slow. The introduction of the value function can reduce the training variance and improve the training efficiency.
[0058] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0059] It is understood that before using the technical solutions disclosed in the various embodiments of this application, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this application in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0060] It is understood that the data (including but not limited to the data itself, the acquisition or use of the data) involved in the technical solutions disclosed in the various embodiments of this application shall comply with the requirements of relevant laws, regulations and related provisions.
[0061] This application provides a question-answering method based on a large model, and a training method for the question-answering large model, which can be applied to, for example... Figure 1 The system architecture shown may include a terminal 100 and a server 200. The server 200 may include one or more servers (…). Figure 1 (This example uses a server as an illustration).
[0062] Either terminal 100 or server 200 can be used independently to execute the large-model-based question-answering method or the question-answering large-model training method provided in the embodiments of this application. Alternatively, terminal 100 and server 200 can also be used collaboratively to execute the large-model-based question-answering method or the question-answering large-model training method provided in the embodiments of this application.
[0063] The following description Figure 1 The product form of the mid-terminal 100;
[0064] The terminal 100 in this application embodiment can be a mobile phone, tablet computer, learning machine, teaching large screen, wearable device, vehicle-mounted device, conference terminal, augmented reality (AR) / virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), etc., and this application embodiment does not impose any restrictions on it.
[0065] First, a training method for a large question-answering model provided in the embodiments of this application will be introduced. Taking the application of this method to a computer device as an example, the computer device can specifically be... Figure 1 The system consists of terminal 100 or a combination of terminal 100 and server 200. (Refer to...) Figure 2 The training method for this large question-answering model specifically includes the following steps:
[0066] First, obtain the question-answering training data, which includes input samples and truth labels. The input samples include sample questions and related knowledge fragments, and the truth labels include answer content marked with knowledge citation sources.
[0067] The knowledge fragments related to the sample question can be retrieved from the configured knowledge base. These knowledge fragments serve as retrieval enhancement information to assist the large model in generating the answer to the sample question.
[0068] The truth labels corresponding to the input samples are the reference answers to the sample questions. These can be manually labeled or obtained through other means. The truth labels include the answer content corresponding to the sample question and the knowledge reference sources. The knowledge reference sources represent the sources of knowledge referenced by the large model when generating the answer content. These can be tags (such as numbers) of knowledge fragments or index addresses of knowledge fragments.
[0069] The input samples from the aforementioned question-answering training data are fed into the large-scale question-answering model to obtain its output, which may include the answer content and the source of the knowledge references. The large-scale question-answering model is then trained using reinforcement learning.
[0070] The reward function for reinforcement learning can be determined based on one or more of the following:
[0071] Consistency in citations and consistency in facts are established.
[0072] Reference consistency is used to encourage the knowledge reference sources generated by the model to be consistent with the knowledge reference sources in the truth labels.
[0073] Fact consistency is used to encourage the semantics of the answers generated by the model to remain consistent with the semantics of the knowledge fragments corresponding to the knowledge references generated.
[0074] The question-answering large model training method provided in this embodiment, by performing reinforcement learning training on the large model according to the above reward function, can enable the output of the question-answering large model to have accurate information retrieval capabilities, and alleviate the problem of factual errors generated by the question-answering large model, ensuring the consistency between the answer results and the cited knowledge.
[0075] In some possible implementations, the reward function can further include answer content consistency, which is used to encourage the model to generate answer content consistent with the answer content in the ground truth label. In this way, the accuracy of the answer content of the trained large-scale question answering model can be further improved.
[0076] The embodiments of the present application train the large-scale question answering model in a reinforcement learning manner. The reinforcement learning algorithm used can be various, such as the proximal policy optimization (PPO) algorithm, the group relative policy optimization (GRPO) algorithm, and the like.
[0077] Taking the GRPO reinforcement learning algorithm as an example:
[0078] When the GRPO reinforcement learning algorithm is used to train the large-scale question answering model, a group of outputs {o1, o2, ……, o G} is sampled for the same input sample q. i The reward function score r i is calculated for each output o i . After being relativized, it is regarded as the advantage function for each token, and finally a target with a ratio is maximized. The objective function of the GRPO algorithm is as follows:
[0079]
[0080]
[0081]
[0082] Among them, represents the group relative reward, that is, all tokens of the output o i share the same score, and their good or bad is measured based on the average level within the group. The reward function r cite is calculated by reference consistency and / or fact consistency.
[0083] The embodiments introduce an optional calculation method of reference consistency, specifically:
[0084] Based on the knowledge reference source generated by the large-scale question answering model and the knowledge reference source in the ground truth label, a first evaluation index value is calculated, which is used to measure the consistency of the knowledge reference source generated by the large-scale question answering model and the knowledge reference source in the ground truth label.
[0085] Among them, the first evaluation index can adopt various types of indexes, including but not limited to: precision, recall, F value, and the like.
[0086] Taking the F value as an example in the embodiments, the reference consistency s citeThe F value of the knowledge reference source generated by the large-scale question-answering model and the knowledge reference source in the true value label can be calculated.
[0087] The F value can measure the overall balance of the knowledge reference source generated by the large-scale question-answering model in precision and recall, and better reflect the reference consistency.
[0088] The embodiment further introduces an optional calculation method of fact consistency, specifically:
[0089] The semantic vectors of the answer content and the knowledge reference source generated by the large-scale question-answering model are obtained respectively, and the vector similarity of the two semantic vectors is calculated, which represents the fact consistency. Specifically, the fact consistency s fact The cosine similarity of the semantic vectors of the answer content and the knowledge reference source generated by the large-scale question-answering model can be calculated.
[0090] In some possible implementations, the reward function can comprehensively consider the reference consistency s cite and the fact consistency s fact , for example:
[0091] The reference consistency s cite and the fact consistency s fact are weighted and summed to obtain the reward function r i :
[0092] r i =αs cite +βs fact
[0093] In some embodiments of the present application, the acquisition process of the question-answering training data in the large-scale question-answering model training process is described.
[0094] The question-answering training data of the present application is composed of input samples and corresponding true value labels. The input samples include sample questions and related knowledge fragments, and the true value labels include answer content marked with knowledge reference sources.
[0095] In some examples, the above question-answering training data can be obtained through public data sets, the Internet and the like.
[0096] In addition, considering that the above question-answering training data is difficult to obtain in some fields, the present embodiment further provides a generation scheme of question-answering training data. In combination with Figure 3 As shown, it can specifically include the following steps:
[0097] S1, acquire a sample question q, and retrieve a knowledge fragment R related to the sample question q in a configured knowledge base, to form an input sample composed of the sample question q and the related knowledge fragment R.
[0098] wherein the sample question can be provided by a user, or obtained from a public dataset, the Internet, etc. In addition, the embodiment also provides a way to obtain a sample question, that is:
[0099] Generating a sample question by a large model. Specifically, instructing the large model to ask questions about the provided corpus document, obtaining the question generated by the large model as the sample question.
[0100] wherein the corpus document can be a document content related to the field to which the question and answer large model is to be applied. By asking questions about the corpus document by the large model, a plurality of sample questions can be generated.
[0101] Table 1 below provides an example of a prompt instruction prompt for instructing the large model to ask questions about the corpus document:
[0102] Table 1
[0103]
[0104] The present application is pre-configured with a knowledge base, and the knowledge base stores a plurality of knowledge segments. The knowledge segments R related to the sample question q can be retrieved in the configured knowledge base.
[0105] In a possible implementation, the knowledge segments can be vectorized using a semantic vector representation model and stored in the knowledge base. Similarly, the sample question is vectorized, and the knowledge segments related to the sample question are retrieved from the knowledge base according to the vector similarity. The knowledge segments that meet the set similarity requirement are taken as the knowledge segments R related to the sample question, and the retrieval enhancement information is obtained. The set similarity requirement can be any one or a combination of multiple conditions such as similarity greater than a threshold, top N knowledge segments in similarity ranking, etc.
[0106] S2, instructing the large model to filter the knowledge segments to be referred to for answering the sample question from the knowledge segments R related to the sample question, and obtaining the filtered knowledge segments r.
[0107] In order to make the knowledge segments referred to by the large model when generating an answer more accurate, improve the factuality of the model's answer, and reduce the model's hallucination problem, the knowledge segments R retrieved in the foregoing step are further called by the large model for filtering, and the more accurate knowledge segments r to be referred to for answering the sample question are filtered by the large model, removing part of the redundant information in the knowledge segments R obtained in the previous step.
[0108] In some possible implementations, the task instruction in the prompt instruction prompt for instructing the large model to filter the knowledge segments can be used to instruct the large model to filter the knowledge segments to be referred to for answering the sample question from the knowledge segments R related to the sample question.
[0109] In other possible implementations, to improve the accuracy of the knowledge fragments screened by the large model and avoid the problem of large model fabrication, the task instruction in the prompt instruction for instructing the large model to screen the knowledge fragments can be specifically used to instruct the large model to screen the knowledge fragments to be referred to for answering the sample question from the input knowledge fragments, and generate the answer content based on the screened knowledge fragments. That is, while instructing the large model to perform the task of screening the knowledge fragments to be referred to for answering the sample question from the input knowledge fragments, the large model is further required to generate the answer content based on the screened knowledge fragments.
[0110] This processing method (letting the large model complete the two tasks of "screening knowledge fragments" and "generating answers based on screening results" at the same time) is of great help to the process of screening knowledge fragments, mainly in the following aspects:
[0111] Strengthen the "purpose" orientation of screening:
[0112] When the model knows that a coherent and accurate answer will be generated based on these screened knowledge fragments, it will have a deeper understanding that the ultimate purpose of screening is not simply to find knowledge fragments, but to find those that can actually be used to build answers.
[0113] Promote more comprehensive and "adequate" screening:
[0114] Generation-driven coverage: To generate answer content, the model will realize the need to cover all aspects of the question. This will encourage it to actively search for all necessary knowledge fragments that can answer the question, provide background information, or support core arguments in screening, rather than just a few most relevant ones.
[0115] Prevent information omission: If only the screening task is done, the model may be satisfied with finding a few key fragments and stopping. But the existence of the generation task will force the model to assess whether the screened knowledge fragments are sufficient to support a decent answer, and if not, it will tend to find more relevant knowledge fragments.
[0116] Provide immediate feedback and dynamic adjustment:
[0117] Generation process as verification of screening: When trying to integrate the screened knowledge fragments into an answer, the model will verify in real time whether these knowledge fragments are really useful.
[0118] If it finds that the knowledge fragments are insufficient to answer some part of the question, the model may re-examine the unselected knowledge fragments and try to find supplementary information.
[0119] If contradictions are found between the pieces, the model will sense the difficulty in generating, which may prompt it to more carefully assess the consistency of knowledge pieces or choose more reliable sources in the filtering stage.
[0120] Dynamic optimization: Generating tasks provide a closed-loop feedback for filtering tasks. During the process of "trying to answer", the model can continuously fine-tune and optimize its judgment of which knowledge pieces are most valuable.
[0121] In summary, the method of this embodiment places knowledge filtering in a specific application scenario (generating answers). Generating tasks provide goal-oriented, coverage requirements, and immediate verification for filtering tasks. This prompts the model to pay more attention to the practicality, completeness, integrability, and value in actual answer construction when filtering, thereby significantly improving the quality and effectiveness of knowledge filtering, making it more consistent with the needs of the final answer. Essentially, the generation task becomes a powerful constraint and feedback mechanism for optimizing the filtering process.
[0122] Table 2 below provides an example of a prompt instruction prompt indicating that the large model filters knowledge pieces:
[0123] Table 2
[0124]
[0125] After the above steps, the filtered knowledge pieces can be obtained, which are used to guide the large model to generate the final answer in the next step.
[0126] In some possible implementations, after obtaining the filtered knowledge pieces through the large model, the filtered knowledge pieces can be further processed to remove redundant knowledge pieces, thereby reducing the redundant knowledge pieces sent to the large model.
[0127] S3, instructing the large model to generate the final result by referring to the sample question and the filtered knowledge pieces, the final result including answer content marked with knowledge reference sources, and the final result as the true value label corresponding to the input sample.
[0128] Specifically, after the above steps, more accurate filtered knowledge pieces related to the sample question can be obtained, so the large model can be called to generate the final result by referring to the sample question and the filtered knowledge pieces, that is, to generate answer content marked with knowledge reference sources. The final result is used as the true value label corresponding to the input sample, and the input sample and the true value label form the question and answer training data.
[0129] Table 3 below provides an example of a prompt instruction prompt indicating that the large model generates the final result:
[0130] Table 3
[0131]
[0132] The embodiment provides an example of answer content generated by using a large model and containing a knowledge reference source: "Experiments show that this method only needs a small amount of fine-tuning, and can maintain stable performance in tasks such as password retrieval and long document summary, and support model expansion from 7B to 65B [^3^]". Wherein, "[^3^]" is a knowledge reference source mark.
[0133] By using the generation method of the question and answer training data provided in the embodiment, high-quality question and answer training data can be batch-generated. In the embodiment, after the preliminary relevant knowledge fragments are retrieved from the knowledge base based on the sample question, the knowledge fragments are not directly used to guide the large model to generate the final answer. Instead, the knowledge fragments are further screened by the large model to screen the knowledge fragments required for answering the sample question, so as to realize the simplification of the knowledge fragments, and then the high-quality knowledge fragments can be used to assist the large model to generate high-quality answer content and knowledge reference sources, and finally high-quality question and answer training data is obtained.
[0134] In some embodiments of the present application, the knowledge base pre-configured by the present application is introduced.
[0135] The present application obtains one or more types of original knowledge corpus and parses the original knowledge corpus content.
[0136] For different types of original knowledge corpus, the original knowledge corpus content can be extracted by matching the parsing method. The types of original knowledge corpus include but are not limited to different file types such as docx, PDF, txt, xlm, etc.
[0137] After the original knowledge corpus content is parsed, the original knowledge corpus content can be split into knowledge fragments and added to the knowledge base.
[0138] In one possible implementation, the original knowledge corpus content can be divided into fixed-length knowledge fragments according to the fixed length and added to the knowledge base.
[0139] Although this fixed-length segmentation method is simple and easy to implement, it may lead to incomplete semantics and information loss. Since this method ignores the sentence and paragraph boundaries in natural language, it may split related information, resulting in the inability to capture key information in the knowledge corpus completely when generating answers. In addition, important context information may be cut into different knowledge fragments and cannot be fully utilized in the generation stage, ultimately affecting the accuracy and coherence of the generated content. These problems are particularly prominent in processing complex texts or tasks requiring cross-paragraph understanding, limiting the application effect of the existing retrieval augmentation RAG method in deep information acquisition and complex problem solving.
[0140] To this end, the embodiment provides a multi-granularity segmentation strategy, specifically:
[0141] In the embodiment, the original corpus content can be segmented by two or more different segmentation granularities to obtain knowledge segments of different segmentation granularities, which are added to the knowledge base.
[0142] The different segmentation granularities include, but are not limited to, chapter-level segmentation granularity, section-level segmentation granularity, paragraph-level segmentation granularity, fixed-length-level segmentation granularity, etc.
[0143] Considering that the retrieved knowledge segments will be further screened in subsequent steps, a relatively coarse segmentation granularity can be selected in the embodiment. In the embodiment, the original knowledge corpus content is segmented according to chapter-level segmentation granularity and fixed-length-level segmentation granularity respectively to obtain chapter-level knowledge segments and fixed-length-level knowledge segments.
[0144] The multi-granularity segmentation method proposed in the embodiment can flexibly divide the original knowledge corpus content according to chapter-level and fixed-length-level according to the structure of the original knowledge corpus content. This segmentation method not only considers the natural hierarchical structure of the document, but also preserves the coherence and integrity of the information at each granularity level, avoiding the problems of incomplete semantics and missing information caused by the single fixed-length segmentation method.
[0145] In some examples, there may be no explicit chapter, structure information annotation in the obtained original knowledge corpus, especially in academic papers, books or reports, chapter titles and content may be presented in different formats or styles. Therefore, the original knowledge corpus can be accurately divided into chapters by using chapter recognition technology in the embodiment. In some possible implementations, the embodiment can use a sequence labeling model for chapter recognition, thereby effectively capturing the text structure and chapter information in the original knowledge corpus, providing a basis for subsequent chapter-level segmentation.
[0146] By using chapter-level segmentation granularity, it is helpful to process relevant knowledge content according to the logical structure of the document, facilitating the identification and extraction of chapter titles and their corresponding content. While the fixed-length-level segmentation granularity provides a general segmentation method, which is suitable for processing documents without explicit structure identification, ensuring that the size of the knowledge segments is relatively uniform in all cases, thereby optimizing the processing effect of subsequent knowledge retrieval.
[0147] The knowledge segments of different granularities obtained after the above multi-granularity segmentation can be stored in the same knowledge base, or can be stored in different granularity knowledge bases respectively. For example, the knowledge segments segmented by each segmentation granularity are stored in the knowledge base corresponding to the granularity.
[0148] For the case where knowledge segments of different granularities are stored in the same knowledge base:
[0149] In the question and answer training data generation method introduced in the foregoing embodiments, the process of retrieving the knowledge segment R related to the sample question q in the knowledge base in step S1 can be based on vector retrieval technology, and knowledge segments R that meet a set similarity condition in similarity to the sample question q are retrieved in the same knowledge base, such as the top N1 knowledge segments R that exceed a threshold in similarity to the sample question q and have the highest similarity, as the retrieval enhanced information.
[0150] Taking segmentation according to chapter level and fixed length level as examples, the chapter level knowledge segments and the fixed length level knowledge segments obtained after segmentation can be stored in the same knowledge base.
[0151] For the case of storing knowledge segments of different granularities in knowledge bases of different granularities:
[0152] In the question and answer training data generation method introduced in the foregoing embodiments, the process of retrieving the knowledge segment R related to the sample question q in the knowledge base in step S1 can be based on vector retrieval technology, and knowledge segments R that meet a set similarity condition in similarity to the sample question q are retrieved in the same knowledge base, such as the top N1 knowledge segments R that exceed a threshold in similarity to the sample question q and have the highest similarity, as the retrieval enhanced information.
[0153] In the knowledge base of each granularity, knowledge segments that meet a set similarity condition in similarity to the sample question q are retrieved, such as the top N2 knowledge segments that exceed a threshold in similarity to the sample question q and have the highest similarity. Further, the knowledge segments retrieved from the knowledge bases of different granularities can be combined to form a final retrieval result, that is, the knowledge segment R related to the question data, as the retrieval enhanced information. Alternatively, the knowledge segments retrieved from the knowledge bases of different granularities can be sorted according to the similarity to the sample question q, and the top M knowledge segments with the highest similarity are selected to form the final retrieval result, that is, the knowledge segment R related to the question data, as the retrieval enhanced information.
[0154] The similarity conditions set for the knowledge bases of different granularities can be the same or different.
[0155] Taking segmentation according to chapter level and fixed length level as examples, the chapter level knowledge segments and the fixed length level knowledge segments obtained after segmentation are stored in different knowledge bases. In the subsequent knowledge segment retrieval stage, an optional retrieval strategy is as follows: in the chapter level knowledge base, the top N1 knowledge segments that exceed a threshold in similarity to the sample question q and have the highest similarity are retrieved, and in the fixed length level knowledge base, the top N2 knowledge segments that exceed a threshold in similarity to the sample question q and have the highest similarity are retrieved, where N1 and N2 can be the same or different. Further, the N1 knowledge segments and the N2 knowledge segments can be combined to form the retrieval result R, that is, the retrieval enhanced information. Alternatively, the N1 knowledge segments and the N2 knowledge segments can be sorted according to the similarity to the sample question q, and the top M knowledge segments with the highest similarity are selected to form the final retrieval result R, where M < N1+N2.
[0156] As can be known from the foregoing scheme introduction, the knowledge segments in the knowledge base can be one or more knowledge segments of different granularities. That is, in the question and answer training data generation method introduced in the foregoing embodiments, the knowledge segment R related to the sample question q retrieved in step S1 in the knowledge base can include one or more knowledge segments of different granularities. When the knowledge segment R contains a knowledge segment of a set coarse granularity (such as a fixed-length-level granularity), the processing of splitting the knowledge segment into knowledge segments of finer granularity can be added in this embodiment. Taking the example of the knowledge base introduced in the foregoing in which the chapter-level knowledge segments and the fixed-length-level knowledge segments are included, the fixed-length-level knowledge segments in the retrieved knowledge segment R can be split, and the splitting can be performed in paragraph or other fine-granularity units. The chapter-level knowledge segments can not be split.
[0157] For the knowledge segments processed in the foregoing manner, a reference source can be further added to each knowledge segment. For example, a sequence number in the form of [^i^] can be added to each knowledge segment as a knowledge reference source.
[0158] By splitting the retrieved coarse-granularity knowledge segments, the subsequent step S2 of screening, by the large model, the knowledge segments to be referred to for answering the sample question from the knowledge segments can be performed in a finer granularity, the accuracy of the knowledge segment screening can be ensured, the redundant knowledge information can be removed, and thus the large model can be guided to generate more accurate and reliable answers.
[0159] In some possible implementations, in the case where the knowledge segments R related to the question data retrieved from the knowledge base in the foregoing step S1 include knowledge segments of different granularities, an optional implementation manner of step S2 of instructing the large model to screen, from the knowledge segments R related to the sample question, the knowledge segments to be referred to for answering the sample question to obtain the screened knowledge segments r is introduced as follows:
[0160] Specifically, the knowledge segments of different granularities have differentiated values. For example, the paragraph-level knowledge segments support fine-granularity fact checking, the chapter-level knowledge segments help to grasp the knowledge context, and the document-level knowledge segments can provide a basis for macro credibility assessment. In order to enable the large model to fully utilize the knowledge segments of different granularities, the knowledge segments of different granularities related to the sample question retrieved can be respectively marked with granularity information and sent to the large model, to instruct the large model to screen, from the knowledge segments R of different granularities, the knowledge segments to be referred to for answering the sample question, to obtain the screened knowledge segments r.
[0161] Another example of the prompt instruction prompt for instructing the large model to screen the knowledge segments is provided in Table 4 as follows:
[0162] Table 4
[0163]
[0164] The prompt shown in Table 4 only takes the knowledge pieces of two granularities of paragraph level and chapter level as examples.
[0165] As can be seen from the comparison of the prompt examples shown in Table 4 and Table 2, the prompt instruction prompt for instructing the large model to screen the knowledge pieces in the embodiment example clearly labels the knowledge pieces of different granularities retrieved, so that the large model can autonomously select the knowledge pieces of appropriate granularity according to specific circumstances.
[0166] In some embodiments of the present application, a large model-based question answering method is further provided, which can be based on the large model trained by reinforcement learning introduced in the foregoing embodiments to perform a knowledge question answering task.
[0167] Referring to Figure 4 A large model-based question answering method is provided, specifically including the following steps:
[0168] Step S100, acquiring question data.
[0169] Step S110, retrieving knowledge pieces related to the question data in a configured knowledge base.
[0170] The knowledge base can include knowledge pieces of one segmentation granularity, or knowledge pieces of two or more segmentation granularities. The structure and construction method of the knowledge base can refer to the introduction of the related embodiments in the foregoing description, and will not be repeated here.
[0171] Step S120, generating feedback information including answer content and knowledge reference source by the large model trained by reinforcement learning, with reference to the question data and the knowledge pieces, and determining the reward function of the reinforcement learning process based on reference consistency and / or fact consistency.
[0172] The reference consistency is used to encourage the model to generate a knowledge reference source consistent with the knowledge reference source in the true value label, and the fact consistency is used to encourage the model to generate answer content consistent with the semantics of the knowledge piece corresponding to the generated knowledge reference source.
[0173] The features related to the reinforcement learning training of the large model can refer to the introduction of the related embodiments of the question answering model training method in the foregoing description, and will not be repeated here.
[0174] Since the large model is trained by reinforcement, the factual problems generated by the large model can be alleviated, and when applied to a knowledge question answering task, the accuracy and reliability of the content generated by the large model can be improved, effectively alleviating the hallucination problem of the large model.
[0175] Similar to the generation manner of the true value label adopted in the question and answer training data generation process introduced in the foregoing embodiments (as shown in the flow) Figure 3 The embodiment introduces step S120, an optional implementation manner of generating feedback information including answer content and knowledge reference source by the large model trained through reinforcement learning, with reference to the question data and the knowledge fragments.
[0176] In the embodiment, a two-level calling manner is adopted:
[0177] First, the large model is called and instructed to filter the knowledge fragments to be referred to for answering the question data from the retrieved knowledge fragments related to the question data, to obtain the filtered knowledge fragments.
[0178] In the embodiment, in order to make the knowledge fragments referred to by the large model when generating the feedback information more accurate, improve the factuality of the model answer, and reduce the hallucination problem of the model, the large model is further called to filter the knowledge fragments retrieved in step S110, and the large model filters more accurate knowledge fragments to be referred to for answering the question data, and removes part of the redundant information.
[0179] In some possible implementations, the task instruction in the prompt instruction prompt for instructing the large model to filter the knowledge fragments can be used to instruct the large model to filter the knowledge fragments to be referred to for answering the question data from the knowledge fragments related to the question data.
[0180] In another possible implementation, in order to improve the accuracy of the knowledge fragments filtered by the large model and avoid the problem of fabrication by the large model, the task instruction in the prompt instruction prompt for instructing the large model to filter the knowledge fragments can be specifically used to instruct the large model to filter the knowledge fragments to be referred to for answering the question data from the input knowledge fragments, and generate answer content based on the filtered knowledge fragments. That is, while instructing the large model to perform the task of filtering the knowledge fragments to be referred to for answering the question data from the input knowledge fragments, the large model is further required to generate answer content based on the filtered knowledge fragments.
[0181] In some embodiments of the present application, an optional implementation manner is provided for the process of instructing the large model to filter the knowledge fragments to be referred to for answering the question data from the knowledge fragments in the above steps, and specifically includes:
[0182] A first prompt instruction prompt template is obtained, the prompt template includes a task instruction, a knowledge slot and a question slot, and the task instruction is used to instruct the large model to filter the knowledge fragments to be referred to for answering the question data in the question slot from the knowledge fragments in the knowledge slot, and generate answer content based on the filtered knowledge fragments.
[0183] Fill the knowledge fragment into the knowledge slot, fill the question data into the question slot, obtain a first prompt instruction prompt, and input the first prompt instruction prompt into the large model to obtain a screened knowledge fragment output by the large model.
[0184] The task instruction of the above example places the knowledge screening task in a specific and concrete application scenario (generating an answer). The generation task provides goal orientation, coverage requirements, and immediate verification for the screening task. This encourages the model to pay more attention to the practicality, completeness, integrability, and value in actual answer construction when screening, thereby significantly improving the quality and effectiveness of knowledge screening and making it more consistent with the needs of the final answer. Essentially, the generation task becomes a powerful constraint and feedback mechanism for optimizing the screening process. The related prompt instruction prompt examples can refer to Table 2 shown in the foregoing text, which will not be described here again.
[0185] In some possible implementations, after obtaining the screened knowledge fragment through the large model, the screened knowledge fragment can be further processed for deduplication, thereby reducing redundant knowledge fragments input into the large model.
[0186] Further, the large model is called again and instructed to generate feedback information with reference to the question data and the screened knowledge fragment.
[0187] Specifically, the more accurate screened knowledge fragment related to the question data can be obtained through the foregoing processing, and therefore the large model can be called to generate the final feedback information, that is, the answer content marked with the knowledge reference source.
[0188] The prompt instruction prompt example for calling the large model to generate the feedback information can refer to Table 3 shown in the foregoing text, which will not be described here again.
[0189] By using the question and answer method provided in this embodiment, after the preliminary relevant knowledge fragments are retrieved from the knowledge base based on the question data, the knowledge fragments are not directly used to guide the large model to generate the final feedback information. Instead, the large model is further used to screen the knowledge fragments, and the knowledge fragments required for reference by the question data are screened, so that the knowledge fragments are simplified, and then the high-quality knowledge fragments can be used to assist the large model to generate high-quality answer content and knowledge reference sources, and finally high-quality feedback information is obtained.
[0190] Since the knowledge pieces in the knowledge base can be one or more knowledge pieces of different granularities. That is, the knowledge pieces retrieved in step S110 in the knowledge base can include one or more knowledge pieces of different granularities. When the knowledge pieces include knowledge pieces of a set coarse granularity (such as a fixed-length-level granularity), the embodiment can further include a process of splitting the knowledge pieces into knowledge pieces of a finer granularity. Taking the knowledge base introduced above as an example, which includes chapter-level knowledge pieces and fixed-length-level knowledge pieces, the fixed-length-level knowledge pieces retrieved can be split into knowledge pieces of a finer granularity, such as a paragraph or other fine-grained unit. The chapter-level knowledge pieces can not be split.
[0191] For the knowledge pieces processed above, a reference source can be further added to each knowledge piece. For example, a serial number in the form of [^i^] is added to each knowledge piece as a knowledge reference source.
[0192] The embodiment can split the retrieved coarse-grained knowledge pieces, so that in subsequent step S120, the large model can select knowledge pieces to be referenced by the knowledge pieces of the question data when filtering the knowledge pieces. The embodiment can ensure the accuracy of the knowledge piece filtering, remove redundant knowledge information, and guide the large model to generate more accurate and reliable feedback information.
[0193] Based on the introduction of the knowledge base construction process above, the knowledge corpus can be divided into knowledge pieces of different granularities. The knowledge pieces of different granularities obtained after the multi-granularity division can be stored in the same knowledge base or different knowledge bases of different granularities. For example, each knowledge piece obtained after the division of different granularities can be stored in a knowledge base of a corresponding granularity.
[0194] For the case where knowledge pieces of different granularities are stored in the same knowledge base:
[0195] The foregoing step S110, which retrieves knowledge pieces related to the question data in the configured knowledge base, can be based on a vector retrieval technology to retrieve knowledge pieces in the same knowledge base that satisfy a set similarity condition with the question data, such as the top N1 knowledge pieces with a similarity to the question data exceeding a threshold value, as enhanced retrieval information.
[0196] Taking the division into chapter-level and fixed-length-level granularities as an example, the chapter-level knowledge pieces and the fixed-length-level knowledge pieces obtained after the division can be stored in the same knowledge base.
[0197] For the case where knowledge pieces of different granularities are stored in different knowledge bases:
[0198] The foregoing step S110, retrieving the knowledge pieces related to the question data in the configured knowledge base, an optional implementation manner is as follows:
[0199] Retrieving the knowledge pieces related to the question data in the knowledge base of each granularity respectively, such as the top N2 knowledge pieces with similarity to the question data exceeding a threshold. Further, the knowledge pieces retrieved from the knowledge bases of different granularities can be combined to form the final retrieval result, i.e., the knowledge pieces related to the question data, as the retrieval enhancement information. Alternatively, the knowledge pieces retrieved from the knowledge bases of different granularities can be sorted according to the similarity to the question data, and the top M knowledge pieces with the highest similarity are selected to form the final retrieval result, i.e., the knowledge pieces related to the question data, as the retrieval enhancement information.
[0200] The similarity conditions set for the knowledge bases of different granularities can be the same or different.
[0201] Taking the division according to the chapter level and the fixed length level as examples, the chapter-level knowledge pieces and the fixed-length-level knowledge pieces obtained after the division are stored in different knowledge bases. In the subsequent knowledge piece retrieval stage, an optional retrieval strategy is as follows: retrieving the top N1 knowledge pieces with similarity to the question data exceeding a threshold in the chapter-level knowledge base, and retrieving the top N2 knowledge pieces with similarity to the question data exceeding a threshold in the fixed-length-level knowledge base. N1 and N2 can be the same or different. Further, the N1 knowledge pieces and the N2 knowledge pieces can be combined to form the retrieval result, i.e., the retrieval enhancement information. Alternatively, the N1 knowledge pieces and the N2 knowledge pieces can be sorted according to the similarity to the question data, and the top M knowledge pieces with the highest similarity are selected to form the final retrieval result, where M < N1+N2.
[0202] In some possible implementations, in the case that the knowledge pieces related to the question data R retrieved from the knowledge base in the foregoing step S110 include knowledge pieces of different granularities, an optional implementation manner of calling and instructing the large model to filter the knowledge pieces to be referred to for answering the question data from the retrieved knowledge pieces related to the question data in the foregoing step is as follows:
[0203] Specifically, knowledge segments of different granularities have differentiated values, such as paragraph-level knowledge segments supporting fine-grained fact checking, chapter-level knowledge segments helping to grasp the knowledge context, and document-level knowledge segments providing the basis for macro credibility assessment. In order to enable the large model to fully utilize knowledge segments of different granularities, the retrieved knowledge segments of different granularities related to the question data can be respectively labeled with granularity information and sent to the large model, to instruct the large model to filter the knowledge segments to be referred by the question data from the knowledge segments of different granularities, to obtain the filtered knowledge segments. The prompt examples for instructing the large model to filter the knowledge segments to be referred by the question data from the retrieved knowledge segments related to the question data can refer to Tables 2 and 4 shown in the foregoing, and will not be described here again.
[0204] In some question and answer scenarios, the feedback information generated by the large model can be output for display. Further, in response to a trigger of a knowledge reference source in the feedback information, the knowledge segment corresponding to the knowledge reference source can be obtained and displayed for the user to refer. At the same time, the credibility of the answer content is also improved.
[0205] The embodiments of the present application also provide an electronic device. Referring to Figure 5 Fig. 1 shows a structural schematic diagram suitable for implementing the electronic device in the embodiments of the present application. The electronic device in the embodiments of the present application can include, but is not limited to, fixed terminals such as mobile phones, tablet computers, learning machines, teaching large screens, wearable devices, and the like. Figure 5 The electronic device shown is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present application.
[0206] As shown in Figure 4 The electronic device can include a processing device (such as a central processing unit, a graphics processing unit, etc.) 1, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 2 or loaded from a storage device 8 to a random access memory (RAM) 3, to implement the question and answer method or the training method of the question and answer large model of the foregoing embodiments of the present application. In the powered-on state of the electronic device, the RAM 3 also stores various programs and data required for the operation of the electronic device. The processing device 1, the ROM 2, and the RAM 3 are connected to each other through a bus 4. An input / output (I / O) interface 5 is also connected to the bus 4.
[0207] Generally, the following devices can be connected to the I / O interface 5: input devices 6 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 7 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 8 including, for example, a memory card, a hard disk, and the like; and communication devices 9. The communication devices 9 can allow the electronic device to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 The electronic device is shown with various devices, but it is understood that all of the shown devices are not required to be implemented or present. More or fewer devices can alternatively be implemented or present.
[0208] The embodiments of the present application further provide a computer program product comprising computer readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the question-answering methods or the training method of the question-answering large model provided by the embodiments of the present application.
[0209] The embodiments of the present application further provide a computer readable storage medium carrying one or more computer programs, which, when executed by an electronic device, can cause the electronic device to implement any of the question-answering methods or the training method of the question-answering large model provided by the embodiments of the present application.
[0210] In addition, it should be noted that the device embodiments described above are only schematic and that one part or all of the components illustrated as separate components can be combined, and that as a unit illustrated or described as a separate component can or can not be physically separate, that is, can be located in one place or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments of the present application. In addition, the device embodiments provided by the present application in the drawings, the connection relationship between the modules indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.
[0211] Those skilled in the art can clearly understand, through the description of the foregoing embodiments, that the present application can be implemented by means of software and the necessary universal hardware, and of course can also be implemented by means of special hardware including special integrated circuits, special CPUs, special memories, special components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure for implementing the same function can be various, such as analog circuits, digital circuits, or special circuits. However, for the present application, software program implementation is a better embodiment. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., including a plurality of instructions for causing a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in various embodiments of the present application.
[0212] In the above embodiments, all or part can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part can be implemented in the form of a computer program product.
[0213] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another, for example, the computer instructions can be transferred from one website, computer, training device or data center to another through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. integrated with one or more available media. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
[0214] The embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. Each embodiment can be combined as needed, and the same or similar parts refer to each other.
Claims
1. A question and answer method, characterized by, The method comprises: obtaining question data; retrieving knowledge segments related to the question data in a configured knowledge base; generating feedback information by a large model trained through reinforcement learning, with reference to the question data and the knowledge segments, the feedback information including answer content and knowledge reference sources; wherein a reward function of the large model reinforcement learning process is determined based on reference consistency and / or fact consistency, the reference consistency being used to encourage the model to generate knowledge reference sources consistent with knowledge reference sources in a true value label, and the fact consistency being used to encourage the model to generate answer content consistent with the semantics of the knowledge segments corresponding to the generated knowledge reference sources.
2. The method of claim 1, wherein, The knowledge base includes knowledge segments obtained by splitting original knowledge corpus at two or more splitting granularities.
3. The method according to claim 1 or 2, characterized in that, The process of generating feedback information by the large model trained through reinforcement learning, with reference to the question data and the knowledge segments, comprises: instructing the large model to filter knowledge segments to be referenced for answering the question data from the knowledge segments related to the question data, to obtain filtered knowledge segments; instructing the large model to generate feedback information with reference to the question data and the filtered knowledge segments.
4. The method of claim 3, wherein, Before instructing the large model to filter knowledge segments to be referenced for answering the question data from the knowledge segments, the method further comprises: performing fine-grained splitting on knowledge segments of a set coarse-grained granularity in the knowledge segments, to obtain split knowledge segments; adding reference sources to each knowledge segment.
5. The method of claim 3, wherein, The process of instructing the large model to filter knowledge segments to be referenced for answering the question data from the knowledge segments, to obtain filtered knowledge segments, comprises: obtaining a first prompt instruction template, the prompt instruction template including a task instruction, a knowledge slot and a question slot, the task instruction being used to instruct the large model to filter knowledge segments to be referenced for answering question data in the question slot from knowledge segments in the knowledge slot, and generate answer content based on the filtered knowledge segments; filling the knowledge segments related to the question data into the knowledge slot, filling the question data into the question slot, to obtain a first prompt instruction prompt, and inputting the first prompt instruction prompt into the large model to obtain filtered knowledge segments output by the large model.
6. The method of claim 3, wherein, Before instructing the large model to generate feedback information with reference to the question data and the filtered knowledge segments, the method further comprises: performing deduplication processing on the filtered knowledge segments.
7. The method of claim 1, wherein, The method further comprises: in response to triggering of the knowledge reference sources, obtaining and displaying knowledge segments corresponding to the knowledge reference sources.
8. A method for training a large question and answer model, characterized in that, The method comprises: obtaining question and answer training data, the question and answer training data including input samples and true value labels, the input samples including sample questions and related knowledge segments, and the true value labels including answer content marked with knowledge reference sources; The input sample is input into the large question-answering model to obtain an output of the large question-answering model, and the large question-answering model is trained in a reinforcement learning manner, and a reward function of the reinforcement learning is determined based on reference consistency and / or fact consistency, the reference consistency is used to encourage the knowledge reference source generated by the model to be consistent with the knowledge reference source in the true value label, and the fact consistency is used to encourage the answer content generated by the model to be consistent with the semantics of the knowledge segment corresponding to the generated knowledge reference source.
9. The method of claim 8, wherein, The calculation process of the reference consistency comprises: a first evaluation index value is calculated based on the knowledge reference source generated by the large question-answering model and the knowledge reference source in the true value label, and the first evaluation index is used to measure the consistency of the knowledge reference source generated by the large question-answering model and the knowledge reference source in the true value label.
10. The method of claim 8, wherein, The calculation process of the fact consistency comprises: respectively, the semantic vectors of the answer content and the knowledge reference source generated by the large question-answering model are obtained, and a vector similarity of the two semantic vectors is calculated, and the vector similarity represents the fact consistency.
11. The method of claim 8, wherein, The reward function is obtained by weighted summation of the reference consistency and the fact consistency.
12. The method according to any one of claims 8-11, characterized in that, The process of obtaining the question-answering training data comprises: a sample question is obtained, and a knowledge segment related to the sample question is retrieved in a configured knowledge base, and an input sample is formed by the sample question and the related knowledge segment; the large model is instructed to screen a knowledge segment to be referred to for answering the sample question from the knowledge segment related to the sample question, to obtain a screened knowledge segment; the large model is instructed to generate a final result by referring to the sample question and the screened knowledge segment, and the final result comprises answer content marked with a knowledge reference source, and the final result is used as a true value label corresponding to the input sample.
13. The method of claim 12, wherein, The process of obtaining the sample question comprises: the large model is instructed to ask questions about a provided corpus document, and a question generated by the large model is obtained as the sample question.
14. An electronic device, comprising: comprise: a memory and a processor; the memory is used to store a program; the processor is used to execute the program, and implement each step of the question-answering method according to any one of claims 1-7, or implement each step of the training method of the large question-answering model according to any one of claims 8-13.
15. A readable storage medium, having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement each step of the question-answering method according to any one of claims 1-7, or implement each step of the training method of the large question-answering model according to any one of claims 8-13.
16. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement each step of the question-answering method according to any one of claims 1-7, or implement each step of the training method of the large question-answering model according to any one of claims 8-13.
Citation Information
Cited By
A process optimization method based on reinforcement learning for decoupling information extraction and generation
CN122388131A