Question and answer method and device, equipment, storage medium and product
By performing feedback data-driven performance evaluation and reference question and answer example updates to large language models, the problem of insufficient adaptability of the model in dynamically changing environments is solved, rapid response and performance improvement is achieved, and maintenance costs and delays are reduced.
Patent Information
- Application Number
- CN202510773644.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-11
AI Technical Summary
When facing the dynamic changes in user problem distribution and demand, large language models are difficult to adapt quickly, resulting in a degradation of answer performance. The existing technology takes a long time by re-labeling data and training models, and it is difficult to quickly respond to the changing needs of actual scenarios.
Generate answers to questions through large language models, collect feedback data for performance evaluation, and update reference question and answer examples when they are below the performance threshold to guide the model to generate answers, build a closed-loop optimization process, and ensure that the reference question and answer examples of the model are in line with the current problem distribution and user needs.
The model is quickly responded and improved in dynamic changing environments, reduced maintenance costs and response delays, improved the flexibility and stability of the model, and solved the problem of insufficient adaptability.
Smart Images

Figure CN120336489A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of natural language processing technology, and in particular to a question-answering method, device, equipment, storage medium and product. Background Art
[0002] With its powerful natural language understanding and generation capabilities, the large language model has been widely used in question-answering tasks, such as intelligent customer service, knowledge retrieval, etc. It can generate corresponding answers based on input questions to meet users' diverse information needs.
[0003] However, in actual applications, the distribution of input questions and the answers that users expect are dynamically changing. For example, the language style and subject areas of users' questions will change over time, and users' requirements for the accuracy and expression of answers are also constantly increasing. However, the model will have difficulty adapting to new changes, resulting in a gradual decline in answering performance.
[0004] At present, related technologies usually adopt the method of re-labeling data and retraining models to adapt the models to new input and output requirements. However, since labeling data and training models takes a long time, it is difficult to quickly respond to the dynamic changes in actual scenarios.
[0005] The above contents are only used to assist in understanding the technical solution of the present application and do not constitute an admission that the above contents are prior art. Summary of the invention
[0006] The main purpose of this application is to provide a question-answering method, device, equipment, storage medium and product that can quickly respond to the dynamically changing needs of actual scenarios and improve the stability of the model in practical applications.
[0007] To achieve the above purpose, the present application proposes a question-answering method, which comprises: Generate, by means of a large language model, an answer to at least one received question based on a reference question-and-answer example, and collect feedback data corresponding to the generated at least one answer; Based on the feedback data, a performance evaluation is performed on the large language model to obtain performance parameters; When the performance parameter is lower than a performance threshold, the reference question and answer example is updated based on the feedback data, and the updated reference question and answer example is used to guide the large language model to subsequently generate answers to questions.
[0008] Optionally, updating the reference question and answer example based on the feedback data includes: Based on the feedback data, obtaining a standard answer to the at least one question; The questions and the standard answers to the questions are used to form candidate question and answer examples; Updating the reference Q&A examples based on multiple constructed candidate Q&A examples so that the answering performance of the large language model based on the updated reference Q&A examples meets the performance conditions.
[0009] Optionally, the large language model is a student large model, and the performance of the student large model is inferior to that of the teacher large model; obtaining the standard answers to the at least one question based on the feedback data includes: Regenerating optimized answers to the at least one question based on the feedback data through the teacher large model; Taking the optimized answers generated by the teacher large model for the at least one question as the standard answers to the at least one question.
[0010] Optionally, obtaining the standard answers to the at least one question based on the feedback data includes: Extracting the corrected answers corresponding to each answer from the feedback data corresponding to the at least one answer; Taking the corrected answers corresponding to each answer as the standard answers to each question.
[0011] Optionally, updating the reference Q&A examples based on multiple constructed candidate Q&A examples so that the answering performance of the large language model based on the updated reference Q&A examples meets the performance conditions includes: Generating multiple groups of test examples based on the multiple candidate Q&A examples, and each group of test examples includes at least one candidate Q&A example; Obtaining the performance parameters of the large language model on the test data set based on each group of test examples respectively; Taking the group of test examples with the optimal performance parameters among the multiple groups of test examples as the updated reference Q&A examples.
[0012] Optionally, evaluating the performance of the large language model based on the feedback data to obtain performance parameters includes: Selecting test questions from the multiple questions received by the large language model in the current time period; Constituting a test data set with the test questions and the feedback data corresponding to the answers to the test questions; Determining the performance parameters based on the answers corresponding to the test questions in the test data set and the feedback data corresponding to the answers to the test questions.
[0013] Optionally, the selecting test questions from the multiple questions received by the large language model in the current time period includes: Classifying the multiple questions received by the large language model in the current time period to obtain multiple question sets, and the types of questions included in each question set are different; Extract at least one test question from each set of questions respectively.
[0014] Optionally, selecting test questions from the multiple questions received by the large language model during the current time period includes: Randomly select a target number of test questions from the multiple questions received by the large language model during the current time period.
[0015] Optionally, after updating the reference Q&A examples based on the feedback data when the performance parameter is lower than the performance threshold, the method further includes: Obtain the performance parameters of the large language model on the test data set based on the reference Q&A examples before and after the update respectively. When the performance parameter corresponding to the reference Q&A example after the update is better than the performance parameter corresponding to the reference Q&A example before the update, deploy the reference Q&A example after the update as the guiding example for the large language model to generate answers subsequently.
[0016] Optionally, the test data set includes multiple test subsets of different types. When the performance parameter corresponding to the reference Q&A example after the update is better than the performance parameter corresponding to the reference Q&A example before the update, deploying the reference Q&A example after the update as the guiding example for the large language model to generate answers subsequently includes: When, for any test subset, the performance parameter corresponding to the reference Q&A example after the update is better than the performance parameter corresponding to the reference Q&A example before the update, deploy the reference Q&A example after the update as the guiding example for the large language model to generate answers subsequently.
[0017] In addition, to achieve the above object, the present application also proposes a Q&A device, and the device includes: A feedback collection module, configured to generate answers to at least one received question based on a reference Q&A example through a large language model, and collect feedback data corresponding to the at least one generated answer; A performance evaluation module, configured to perform performance evaluation on the large language model based on the feedback data to obtain performance parameters; An example update module, configured to update the reference Q&A example based on the feedback data when the performance parameter is lower than the performance threshold, and the reference Q&A example after the update is used to guide the large language model to generate answers to questions subsequently.
[0018] Optionally, the example update module includes: A standard answer acquisition unit, configured to obtain the standard answers to the at least one question based on the feedback data; Candidate example construction unit for constructing candidate Q&A examples from each question and the standard answer to each question; Example update unit for updating the reference Q&A examples based on multiple constructed candidate Q&A examples, so that the answering performance of the large language model based on the updated reference Q&A examples meets the performance conditions.
[0019] Optionally, the large language model is a student large model, and the performance of the student large model is inferior to that of the teacher large model; The standard answer acquisition unit is used to regenerate the optimized answer to the at least one question through the teacher large model based on the feedback data; and use the optimized answer generated by the teacher large model for the at least one question as the standard answer to the at least one question.
[0020] Optionally, the standard answer acquisition unit is used to extract the corrected answer corresponding to each answer from the feedback data corresponding to the at least one answer; and use the corrected answer corresponding to each answer as the standard answer to each question.
[0021] Optionally, the example update unit is used to generate multiple sets of test examples based on the multiple candidate Q&A examples, and each set of test examples includes at least one candidate Q&A example; obtain the performance parameters of the large language model on the test data set based on each set of test examples respectively; and use the set of test examples with the optimal performance parameters among the multiple sets of test examples as the updated reference Q&A examples.
[0022] Optionally, the performance evaluation module includes: Test question selection unit for selecting test questions from multiple questions received by the large language model during the current time period; Test set construction unit for constructing a test data set from the test questions and the feedback data corresponding to the answers to the test questions; Performance parameter determination unit for determining the performance parameters based on the answers corresponding to the test questions in the test data set and the feedback data corresponding to the answers to the test questions.
[0023] Optionally, the test question selection unit is used to classify multiple questions received by the large language model during the current time period to obtain multiple question sets, and the types of questions included in each question set are different; and extract at least one test question from each question set respectively.
[0024] Optionally, the test question selection unit is used to randomly select a target number of test questions from multiple questions received by the large language model during the current time period.
[0025] Optionally, the device further includes: An example deployment module, configured to obtain performance parameters of the large language model on a test data set based on the reference question-answer examples before update and the reference question-answer examples after update respectively; and deploy the reference question-answer examples after update as guiding examples for the large language model to generate answers subsequently when the performance parameters corresponding to the reference question-answer examples after update are better than those corresponding to the reference question-answer examples before update.
[0026] Optionally, the test data set includes multiple test subsets of different types, The example deployment module is configured to deploy the reference question-answer examples after update as guiding examples for the large language model to generate answers subsequently when the performance parameters corresponding to the reference question-answer examples after update are better than those corresponding to the reference question-answer examples before update for any test subset.
[0027] In addition, to achieve the above object, the present application further provides a question-answer device, where the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the question-answer method as described above.
[0028] In addition, to achieve the above object, the present application further provides a storage medium, where the storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and the computer program, when executed by a processor, implements the steps of the question-answer method as described above.
[0029] In addition, to achieve the above object, the present application further provides a computer program product, where the computer program product includes a computer program, and the computer program, when executed by a processor, implements the steps of the question-answer method as described above.
[0030] One or more technical solutions proposed by the present application have at least the following technical effects: The Q&A solution provided by this application uses a large language model to generate answers to at least one received question based on reference Q&A examples, and collects feedback data corresponding to at least one answer generated by the large language model. Then, based on this feedback data, the performance of the large language model is evaluated to obtain performance parameters. This solution can timely perceive the performance of the model when facing new inputs and achieve dynamic monitoring of the model's running state. When the obtained performance parameters are lower than the performance threshold, the reference Q&A examples are automatically updated based on the feedback data, and the updated reference Q&A examples are used to guide the large language model to generate answers to questions subsequently. Since the reference Q&A examples are continuously updated based on the feedback data, it is ensured that the reference Q&A examples used by the model always adapt to the current question distribution or user needs, thereby improving the answer performance. This solution constructs a closed-loop optimization process from execution, feedback, evaluation to example update, enabling the model to continuously improve the answer performance without modifying its internal parameters. Compared with the traditional method that relies on re-labeling data and model retraining, this solution significantly reduces the maintenance cost and response delay, effectively solves the problem of insufficient adaptability of the large language model when facing dynamic changes in user question distributions and requirements, and improves the flexibility and stability of the model in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0032] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0033] Figure 1 Schematic diagram of an implementation environment of the Q&A method of this application; Figure 2 Schematic flowchart provided by the first embodiment of the Q&A method of this application; Figure 3 Schematic diagram of the detailed steps of step S30 in the second embodiment of the Q&A method of this application; Figure 4 Schematic diagram of the detailed steps of step S20 in the third embodiment of the Q&A method of this application; Figure 5 Schematic diagram of the new steps in the fourth embodiment of the Q&A method of this application; Figure 6 Schematic diagram of a Q&A system provided by this application; Figure 7It is a schematic diagram of the module structure of the Q&A device according to an embodiment of the present application; Figure 8 It is a schematic diagram of the device structure of the hardware operating environment involved in the Q&A method according to an embodiment of the present application.
[0034] The implementation, functional features and advantages of the present application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. Detailed implementation manners
[0035] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0036] In order to better understand the technical solutions of the present application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific implementation manners.
[0037] Figure 1 It is a schematic diagram of an implementation environment provided by an embodiment of the present application. Refer to Figure 1 , this implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 are connected through a wireless or wired network. Exemplarily, a target application provided by the server 102 is installed on the terminal 101, and the terminal 101 can implement functions such as data transmission and message interaction through this target application.
[0038] Exemplarily, the terminal 101 is a computer, a mobile phone, a tablet computer or other terminal. Exemplarily, the target application is a target application in the operating system of the terminal 101 or a target application provided by a third party. For example, the target application is a search application, a shopping application, an online learning application, etc. This target application has a Q&A function and can give corresponding answers based on the questions raised by the user. Exemplarily, the server 102 is the background server corresponding to this target application. Correspondingly, the server 102 is a search application server, a shopping application server, an online learning application server, etc.
[0039] In this application, the terminal 101 is used to receive the questions input by the user and forward the questions input by the user to the server. The server 102 is used to generate answers to at least one received question based on the reference Q&A examples, and collect the feedback data corresponding to the at least one generated answer. Then, based on the feedback data, the performance of the large language model is evaluated to obtain performance parameters. When the performance parameters are lower than the performance threshold, the current reference Q&A examples are updated based on the feedback data, so that the updated reference Q&A examples better meet the current answer requirements. The updated reference Q&A examples are used to guide the large language model to generate answers to questions subsequently. That is, subsequently, when receiving the user questions sent by the terminal 101, the server 102 uses the large language model to generate answers to the questions based on the updated reference answer examples. Thus, the problem of insufficient adaptability of the large language model in the face of dynamically changing user question distributions and requirements is solved.
[0040] Alternatively, the above solution can also be completed by the terminal 101 alone. Alternatively, the terminal 101 completes it through the installed target application. The embodiments of this application do not limit this.
[0041] The Q&A method provided in this application is applicable to various scenarios. For example, an intelligent customer service system: using a large language model as an intelligent customer service, and evaluating the quality of the answers of the intelligent customer service by collecting the interaction data between the user and the intelligent customer service, that is, the feedback data. When the answer quality does not meet the requirements, the reference Q&A examples used to guide the model to perform the Q&A task are continuously optimized based on this feedback, so as to change the answer strategy of the intelligent customer service and improve the ability and efficiency to solve user problems.
[0042] For example, an education tutoring platform: using a large language model as an intelligent teacher, and adjusting and improving the reference Q&A examples used to guide the model to perform the Q&A task based on the feedback from students on the answers provided by the intelligent teacher, so as to change the teaching content and methods to better meet the needs of learners and improve the learning effect.
[0043] For example, enterprise internal knowledge management: the Q&A system within the enterprise uses a large language model to help internal employees quickly find the information they need or solve problems, and uses the feedback after employees use it to optimize the Q&A examples referred to by the large language model when performing the Q&A task in the Q&A system, so as to improve the accuracy of the Q&A system and the work efficiency of employees.
[0044] Figure 2 It is a schematic flowchart of the first embodiment of the Q&A method of this application. Refer to Figure 2 , taking the server as the execution entity as an example, the Q&A method includes the following steps S10~S30: Step S10: Based on the large language model, generate answers to at least one received question based on reference Q&A examples, and collect feedback data corresponding to the at least one generated answer.
[0045] The large language model is a deep learning model trained based on a vast amount of text, such as GPT (Generative Pre-trained Transformer), etc., which has the ability to understand and generate natural language and is used to perform Q&A tasks.
[0046] The reference Q&A examples are a set of demonstration data used to guide the large language model to answer, including "question - answer" pairs. For example: Question: "How to set up a router?" Answer: "First, connect the power supply, then access the management page, and then..." The reference Q&A examples can be artificially constructed or selected from the historical answer record pairs of the large language model. The embodiments of this application do not limit this.
[0047] The feedback data is the opinion or evaluation of the user or the system on the answer of the large language model, which can include annotation information on whether the answer is correct, the reason for the wrong answer, the score of the answer, the content suggested for modification, the content suggested for supplementation, the corrected answer, etc.
[0048] Generating an answer to at least one received question based on the reference Q&A examples by the large language model means: integrating the reference Q&A examples into the model prompt, and then inputting the question and the model prompt into the large language model to obtain the answer generated by the large language model.
[0049] Exemplarily, collecting the feedback data corresponding to at least one answer generated by the large language model includes collecting explicit feedback from users. For example, through the front - end interface, that is, the answer display interface of the terminal, prompt the user to give feedback on the answer content. For example, prompt the user to mark whether the answer is correct, point out the wrong content in the answer, rate the answer, give the corrected answer, etc.
[0050] Exemplarily, collecting the feedback data corresponding to at least one answer generated by the large language model can also include collecting implicit feedback from users. For example, determine whether the user repeats asking the same question. In the case where the user repeats asking the same question, integrate the answers generated by the model multiple times to obtain the corrected answer to this question; determine whether the user clicks to view more answers, or modifies the question and asks again, or the page stay time is short, and take these as feedback data indicating dissatisfaction with the current answer.
[0051] Exemplarily, to ensure the timeliness and effectiveness of the feedback data, only the feedback data corresponding to the answers generated by the model within the current time period can be collected. For example, the feedback data corresponding to the answers generated in the past quarter or the feedback data corresponding to the answers generated in the past month. Focusing on recent data can more accurately locate the real-time problems of the model and improve the pertinence of subsequent optimization.
[0052] Step S20: Based on the feedback data, perform a performance evaluation on the large language model to obtain performance parameters.
[0053] The performance parameter is a quantitative indicator for measuring the question-and-answer ability of the model. For example, the performance parameter can include answer accuracy rate, F1 value, relevance, user satisfaction score, etc. Among them, the F1 value is the harmonic mean of the precision rate and the recall rate.
[0054] Exemplarily, based on the feedback data, perform a performance evaluation on the large language model to obtain performance parameters, including: extracting the answer scores in the feedback data corresponding to at least one answer, averaging all the answer scores to obtain the user satisfaction score, and using this satisfaction score as the performance parameter of the model.
[0055] Exemplarily, based on the feedback data, perform a performance evaluation on the large language model to obtain performance parameters, including: extracting the corrected answers in the feedback data corresponding to at least one answer, and determining the answer accuracy rate of the model based on the similarity between at least one answer generated by the model and the corresponding corrected answer. Among them, the higher the similarity between the answer generated by the model and the corresponding corrected answer, the higher the answer accuracy rate of the model.
[0056] Exemplarily, the large language model can also be evaluated for performance from multiple perspectives to obtain performance parameters from multiple perspectives. For example, obtain performance parameters from multiple perspectives such as answer accuracy rate, F1 value, user satisfaction score, etc., so as to be able to comprehensively evaluate the answer performance of the model by combining the performance parameters from multiple perspectives.
[0057] Exemplarily, the large language model can be evaluated for performance regularly or as needed to monitor the running state of the large language model.
[0058] Step S30: In the case where the performance parameter is lower than the performance threshold, update the reference Q&A examples based on the feedback data. The updated reference Q&A examples are used to guide the large language model to generate answers to questions subsequently.
[0059] The performance threshold is a preset performance standard used to measure whether the performance of the large language model has reached the expected goal. If the performance parameter of the model is lower than this threshold, it is considered that improvement is needed. For example, if the performance threshold is that the accuracy rate reaches 90%, then when the accuracy rate of the model is lower than 90%, it is determined that the reference Q&A example is no longer the optimal choice under the current data distribution and needs to be optimized.
[0060] The updated reference Q&A examples are optimized demonstration data updated based on feedback data, containing "question-answer" pairs that better meet current requirements, and are used to improve the quality of model answers.
[0061] Exemplarily, updating the reference Q&A examples based on feedback data includes at least the following implementation methods: Error correction: Identify and correct the incorrect content in the current reference Q&A examples according to the feedback data. For example, if a user feedbacks that there are omissions in the answer about the contraindicated population of a certain drug, relevant information will be supplemented in the reference Q&A examples to ensure the accuracy of the answer. Missing supplement: Analyze the feedback data to find out the question types not covered in the current reference Q&A examples. For new domain questions raised by users, create corresponding "question-answer" pairs in combination with the feedback of model-generated answers and add them to the reference Q&A examples to expand the Q&A coverage. Obsolete removal: Judge the content that is no longer applicable in the reference Q&A examples according to the feedback data and delete it in a timely manner. For example, as the promotion activity ends, obsolete "question-answer" pairs such as "2023 promotion rules" are removed from the examples to avoid interference of stale information on model output.
[0062] After updating the reference Q&A examples based on the feedback data, the subsequent large language model generates answers to questions based on the updated reference Q&A examples. Exemplarily, replace the reference Q&A examples in the model prompt with the updated reference Q&A examples to obtain the updated model prompt. For subsequent questions, input the questions and the updated model prompt into the large language model to generate answers.
[0063] The Q&A solution provided by this application uses a large language model to generate answers to at least one received question based on reference Q&A examples, and collects feedback data corresponding to at least one answer generated by the large language model. Then, based on this feedback data, the performance of the large language model is evaluated to obtain performance parameters. This solution can promptly perceive the performance of the model when facing new inputs and achieve dynamic monitoring of the model's running state. When the obtained performance parameters are lower than the performance threshold, the reference Q&A examples are automatically updated based on the feedback data, and the updated reference Q&A examples are used to guide the large language model to generate answers to subsequent questions. Since the reference Q&A examples are continuously updated based on the feedback data, it is ensured that the reference Q&A examples used by the model are always adapted to the current question distribution or user needs, thereby improving the answer performance. This solution constructs a closed-loop optimization process from execution, feedback, evaluation to example update, enabling the model to continuously improve the answer performance without modifying its internal parameters. Compared with the traditional method that relies on relabeling data and retraining the model, this solution significantly reduces the maintenance cost and response latency, effectively solves the problem of insufficient adaptability of the large language model when facing dynamic changes in user question distributions and requirements, and improves the flexibility and stability of the model in practical applications.
[0064] Based on the above first embodiment, the second embodiment of this application is proposed. For the same or similar content as the first embodiment, reference can be made to the above introduction and will not be repeated hereinafter. Referring to Figure 3 , in the second embodiment, the above step S30 includes steps S301 to S303: Step S301, when the performance parameters are lower than the performance threshold, based on the feedback data, obtain the standard answers to at least one question.
[0065] The standard answer is the accurate, comprehensive and demand-compliant answer content determined after being clearly pointed out by the user for the user's question, and is the core element for constructing high-quality Q&A examples.
[0066] Optionally, the large language model is a student large model, and the performance of the student large model is inferior to that of the teacher large model; obtaining the standard answers to at least one question based on the feedback data includes: using the teacher large model to regenerate optimized answers to at least one question based on the feedback data; using the optimized answers generated by the teacher large model for at least one question as the standard answers to at least one question.
[0067] Among them, the student large model is a large language model with a smaller scale and weaker performance compared to the teacher large model, and it has higher inference efficiency. The teacher large model is a large language model with a larger scale and stronger performance, which is used to generate high-quality answers or as an authoritative knowledge source.
[0068] Exemplarily, input each question, the answer generated by the student large model for this question, and the feedback information corresponding to this answer into the teacher large model to obtain the optimized answer generated by the teacher large model.
[0069] In the embodiments of the present application, since the teacher large model has stronger language understanding and generation capabilities, the optimized answers generated by it are often more accurate, more logical, and more natural in expression. Therefore, by automatically generating the standard answers to each question through the teacher large model, and subsequently using the examples composed of the questions and the standard answers as a reference for the student large model to generate answers, on the one hand, the student large model can be guided in this way, significantly improving its performance in actual applications. On the other hand, there is no need to rely on manual annotation, making the update of the reference Q&A examples more automated and sustainable, and suitable for dynamically changing application scenarios.
[0070] Optionally, based on the feedback data, obtaining the standard answers to at least one question includes: extracting the corrected answers corresponding to each answer from the feedback data corresponding to at least one answer; using the corrected answers corresponding to each answer as the standard answers to each question. It can be understood that the feedback data corresponding to some answers contains clear modification suggestions or correct answer texts, then these information can be directly extracted as the corrected answers. By extracting the corrected answers from the feedback data as the standard answers, it can be ensured that the answers learned by the model are accurate and most in line with the actual needs of users.
[0071] Optionally, the multi-model voting method can also be used to obtain the standard answers to at least one question. Specifically, use multiple large language models to answer the same question, and select the answer with the majority agreement from the multiple generated answers as the standard answer. Compared with the answers of a single model, the multi-model voting method can effectively reduce the risk brought by the wrong reasoning of individual models, thereby improving the accuracy of the finally selected standard answer.
[0072] Exemplarily, questions with poor feedback data can be selectively targeted, and the teacher large model can be used to generate optimized answers as the standard answers. For example, questions for which the feedback data does not carry modification suggestions or corrected answers. For feedback data with better quality, the corrected answers can be directly extracted from it as the standard answers.
[0073] Step S302, form candidate Q&A examples from each question and the standard answers to each question.
[0074] Candidate Q&A examples are "question-answer" pairs composed of user questions and their corresponding standard answers, which are alternative data for updating the reference Q&A examples, and are obtained through screening and sorting of feedback data.
[0075] Step S303: Update the reference Q&A examples based on the constructed multiple candidate Q&A examples, so that the answering performance of the large language model based on the updated reference Q&A examples meets the performance conditions.
[0076] Answering performance is an indicator for quantitatively evaluating the answering quality of a large language model, such as accuracy rate, response speed, answer relevance, user satisfaction, etc., and is used to measure the performance of the model in actual applications.
[0077] The performance condition is a pre-set qualified standard for measuring the answering performance of the model. For example, the accuracy rate needs to reach more than 90%, the average response time does not exceed 2 seconds, and it is the best in multiple control groups, etc., and is used as the basis for judging whether the update of the reference Q&A examples is effective.
[0078] Optionally, updating the reference Q&A examples based on the constructed multiple candidate Q&A examples so that the answering performance of the large language model based on the updated reference Q&A examples meets the performance conditions includes: generating multiple groups of test examples based on the multiple candidate Q&A examples, each group of test examples includes at least one candidate Q&A example, and the candidate Q&A examples included in each group of test examples are different or the number of candidate Q&A examples included is different; obtaining the performance parameters of the large language model based on each group of test examples on the test data set; using the group of test examples with the best performance parameters among the multiple groups of test examples as the updated reference Q&A examples. This means that the standard for the updated reference Q&A examples to meet the performance conditions is that among the multiple groups of test examples, the corresponding performance parameters reach the optimal level.
[0079] Generating multiple groups of test examples based on multiple candidate Q&A examples means: forming multiple groups of test examples by combining the candidate Q&A examples in different ways. For example: if there are candidate Q&A examples A, B, and C, the possible generated test example groups include {A}, {B}, {C}, {A + B}, {A + C}, {B + C}, etc. Then, compare the performance parameters of each group of test examples and find the group with the best performance parameters, such as the group with the highest accuracy rate as the updated reference Q&A examples. It should be noted that the number of candidate Q&A examples included in the test example groups in the embodiments of the present application is not limited and can be set according to actual needs. For example, set the number of candidate Q&A examples included in each test example group to be not less than 2 and not more than 5.
[0080] Among them, the test data set contains multiple test questions and the corresponding standard answers for each test question. Exemplarily, the test data set can be constructed based on the feedback data. Specifically, test questions are selected from the multiple questions received by the large language model during the current time period, and the standard answers are extracted from the feedback data corresponding to the answers to each test question. Then, each test question and the standard answer to each test question are used to form the test data set. Next, the large language model is used to generate answers to each test question in the test data set based on each group of test examples respectively, and the performance parameters of the model, such as the answer accuracy rate, are calculated based on the answers to each test question generated by the model and the standard answers to each test question.
[0081] In the embodiments of the present application, multiple groups of test examples are generated based on multiple candidate Q&A examples, and multiple different examples are used for testing, which can comprehensively cover different combined scenarios of the candidate Q&A examples and fully verify the influence of each example and its combination on the model performance. By selecting the group that can make the model perform best from multiple groups of test examples, it can ensure that the quality and guiding ability of the updated reference Q&A examples reach the optimal level, thereby significantly improving the answer accuracy and overall performance of the model.
[0082] In the embodiments of the present application, when the performance parameter is lower than the performance threshold, based on the feedback data, it can be identified which questions have inaccurate or unsatisfactory answers, so that the standard answers to these questions can be obtained. Then, each question and its standard answer are used to construct candidate Q&A examples, and the reference Q&A examples are updated with them. This method ensures that the model can learn the latest knowledge and user needs based on the reference Q&A examples and is exposed to more diverse questions and scenarios. Thus, the adaptability and flexibility of the model can be enhanced, and the answer accuracy of the model can be significantly improved.
[0083] Based on the first embodiment of the present application above, the third embodiment of the present application is proposed. The same or similar content as the first embodiment can be referred to the above introduction and will not be repeated hereinafter. Referring to Figure 4 , in the third embodiment, step S20 includes steps S201 to S203.
[0084] Step S201, select test questions from the multiple questions received by the large language model during the current time period.
[0085] Optionally, selecting test questions from the multiple questions received by the large language model during the current time period includes: classifying the multiple questions received by the large language model during the current time period to obtain multiple question sets, where the types of questions included in each question set are different; extracting at least one test question from each question set respectively.
[0086] Exemplarily, a plurality of questions received within the current time period can be divided into multiple types such as technology, health, education, sports, food, work, beauty, fashion, travel, etc. The specific classification method can be determined according to actual needs, and the embodiments of the present application do not limit this.
[0087] Exemplarily, test questions can be evenly selected from each question set, that is, the same number of test questions are randomly selected from each question set. This ensures that the test questions can cover a wide range of question types, thereby comprehensively evaluating the comprehensive performance of the model on various types of questions and avoiding the neglect of certain domain questions due to sample bias. Or, test questions are selected from each question set according to the proportion of the number of questions in each question set, such that the number of questions in each question set is positively correlated with the number of test questions selected from each question set. The larger the number of questions in a certain question set, it indicates that the questions included in this question set belong to a category of questions that users are most concerned about or most common. Correspondingly, the more test questions are extracted from this question set, the more accurately the answering performance of the model for these high-frequency questions can be evaluated, thereby ensuring the answering quality of these high-frequency questions.
[0088] Optionally, selecting test questions from a plurality of questions received by the large language model within the current time period includes: randomly extracting a target number of test questions from the plurality of questions received by the large language model within the current time period. This way, there is no need for complex classification, scoring, or sorting logic, and only a random algorithm is needed to complete the selection of test questions, thereby being able to save costs and improve the evaluation efficiency.
[0089] Optionally, selecting test questions from a plurality of questions received by the large language model within the current time period includes: extracting the answering scores of each question from the feedback data corresponding to the answers to the received plurality of questions; based on the answering scores of the plurality of questions, dividing the plurality of questions into different question sets; respectively extracting at least one test question from each question set, and the more test questions are extracted from the question set with a lower answering score. This can effectively identify and improve the areas where the model performs poorly, thereby targetedly improving the overall performance and user experience of the model.
[0090] Step S202, forming a test data set with the test questions and the feedback data corresponding to the answers to the test questions.
[0091] Step S203, determining performance parameters based on the answers corresponding to the test questions in the test data set and the feedback data corresponding to the answers to the test questions.
[0092] Exemplarily, the standard answers to the test questions are extracted from the feedback data corresponding to the answers to the test questions. Correspondingly, based on the similarity between the answers to the test questions in the test dataset and the standard answers, the performance parameters of the model are determined, such as the answer accuracy rate. Alternatively, the answer scores can also be extracted from the feedback data, and based on the answer scores of all the test questions in the test dataset, the performance parameters of the model are comprehensively determined, such as the satisfaction score of the answers.
[0093] In the embodiments of the present application, using the real user questions and feedback to form the test dataset, rather than relying on the predefined dataset, can more accurately reflect the performance of the model in the actual application scenario, and help identify the advantages and disadvantages of the model when dealing with real-world problems. Thus, it can ensure that after the subsequent optimization of the reference Q&A examples, the answers of the model are more in line with the real needs of users. Moreover, the embodiments of the present application limit that the selected test questions and feedback data should belong to the current time period, that is, the embodiments of the present application select new test questions and feedback data according to different time periods, so as to timely capture the performance fluctuations of the model over time and environmental changes, and ensure that the evaluation results always reflect the latest state of the model.
[0094] Based on the first embodiment of the present application above, the fourth embodiment of the present application is proposed. This embodiment mainly describes the process of validating the effectiveness of the updated reference Q&A examples. The same or similar content as the first embodiment can be referred to the above introduction and will not be repeated hereinafter. Refer to Figure 5 , in the fourth embodiment, after step S30, steps S401 to S402 are further included.
[0095] Step S401, obtain the performance parameters of the large language model on the test dataset based on the reference Q&A examples before and after the update, respectively.
[0096] Among them, the reference Q&A examples before the update are the reference Q&A data used by the large language model to guide answer generation before this update and optimization. The acquisition method of the test dataset has been introduced in the above embodiments and will not be repeated here.
[0097] Step S402, in the case that the performance parameters corresponding to the reference Q&A examples after the update are better than the performance parameters corresponding to the reference Q&A examples before the update, deploy the reference Q&A examples after the update as the guiding examples for the large language model to generate answers subsequently.
[0098] Deploying the reference Q&A examples after the update as the guiding examples for the large language model to generate answers subsequently means: officially integrating the updated reference Q&A examples that have passed the effectiveness verification into the production environment, that is, the online inference process, so that it becomes the actual guiding basis for the large language model to generate answers subsequently.
[0099] Exemplarily, in order to further ensure the optimization effect, when the performance parameter corresponding to the updated reference question and answer example is better than the sum of the performance parameter corresponding to the reference question and answer example before the update and the preset performance improvement parameter, the updated reference question and answer example can be deployed as a guiding example for the subsequent generation of answers by the large language model. This is to avoid invalid updates due to slight advantages and ensure efficient use of optimization resources. Among them, the preset performance improvement parameter refers to the minimum performance improvement threshold value that is pre-set in the large language model optimization task and is used to measure whether the updated reference question and answer example is "significantly better" than the old version.
[0100] Optionally, the test data set includes multiple test subsets of different types. When the performance parameters corresponding to the updated reference question and answer examples are better than the performance parameters corresponding to the reference question and answer examples before the update, the updated reference question and answer examples are deployed as guiding examples for the subsequent generation of answers by the large language model, including: when, for any test subset, the performance parameters corresponding to the updated reference question and answer examples are better than the performance parameters corresponding to the reference question and answer examples before the update, the updated reference question and answer examples are deployed as guiding examples for the subsequent generation of answers by the large language model.
[0101] The test subset is a subdivided component of the test data set, divided according to specific dimensions. For example, dimensions such as data type, business scenario, and difficulty level. For example, if the test data set is "customer service question and answer data", the test subset can be divided into "product consultation", "complaint handling", "order inquiry", etc. If divided by difficulty, it can be divided into "simple question subset", "complex logic subset", "multi-round dialogue subset", etc.
[0102] In an embodiment of the present application, when the performance parameters corresponding to the updated reference question and answer examples are better than the performance parameters corresponding to the reference question and answer examples before the update for any test subset, the updated reference question and answer examples are deployed as guide examples for the subsequent generation of answers by the large language model. That is, the updated reference question and answer examples are required to be stable and reliable in all segmented scenarios to ensure the comprehensiveness of model optimization and prevent the model from failing in specific scenarios due to data deviations. Through this comprehensive verification mechanism, the solution provides more stringent quality control standards for the optimization of large language models, ensuring that each deployed guide example can comprehensively upgrade the model performance.
[0103] In an embodiment of the present application, by comparing the performance parameters of the reference question and answer examples before and after the update on the test data set, the updated reference question and answer examples are deployed as guiding examples for the subsequent generation of answers by the large language model only when the performance parameters are improved. This ensures that each update can specifically address the shortcomings of the model, thereby ensuring the effectiveness of the optimization.
[0104] Figure 6It is a schematic diagram of a question-answering system provided by an embodiment of the present application. Refer to Figure 6 , first, a user or an external system asks a question. The question and the model prompt containing the latest reference question-answer examples are input into a large language model, and the large language model generates an answer under the guidance of the model prompt. Then, the user or the external system provides feedback data based on the answer. The feedback collection module collects a recent feedback data set, which includes the questions input into the model recently, the answers generated by the model, and the feedback data corresponding to the answers. Next, the feedback collection module constructs a test data set using the collected feedback data and gives the test data set to the performance monitoring module. The performance monitoring module determines the performance parameters of the model based on the test data set and then gives the performance parameters to the optimization trigger module. The optimization trigger module compares the performance parameters with the performance threshold. When the performance parameters are lower than the performance threshold, it sends an optimization instruction to the example optimization module. When the performance parameters are not lower than the performance threshold, the performance monitoring module does not send an optimization instruction, so that the model continues to perform the question-answering task based on the current model prompt. The example optimization module, when receiving the optimization instruction, constructs candidate question-answer examples using the feedback data set obtained from the feedback collection module and generates multiple groups of test examples based on the candidate question-answer examples. Then, it selects a group of test examples with the best performance on the test data set as the updated reference question-answer examples. Then, the updated reference question-answer examples are input into the validity verification and deployment module. The validity verification and deployment module verifies the validity of the updated reference question-answer examples by comparing the performance parameters of the currently used reference question-answer examples and the updated reference question-answer examples. When the validity verification passes, the reference question-answer examples in the model prompt are replaced with the updated reference question-answer examples to obtain the updated model prompt. Then, the updated model prompt is stored so that the stored model prompt always contains the latest reference question-answer examples. Then, the model prompt is provided to the large language model to guide the large language model to generate an answer to the newly input question.
[0105] The question-answering solution provided by the present application has at least achieved the following beneficial effects: First, automation and low cost. It realizes the automatic monitoring and dynamic optimization of the question-answer examples in the model prompt, significantly reducing the need for and cost of manual intervention. The optimization process does not involve modifying the parameters of the model itself, avoiding the high computing overhead and time cost of large-scale model retraining.
[0106] Second, strong adaptability and continuous improvement. The model can actively adapt to changes in the data distribution, continuously maintain a high performance level, and effectively alleviate the negative impact brought by data drift. Through the closed-loop feedback mechanism, the question-answering system can continuously learn from new interaction data and realize the continuous evolution of the prompting strategy.
[0107] Third, generality. This solution does not depend on a specific model structure and has good generality, and can be applied to a variety of prompt-based model application scenarios.
[0108] Another point to note is that the above examples are only for understanding this application and do not constitute a limitation on the Q&A method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.
[0109] This application also provides a Q&A device. Please refer to Figure 7 , the Q&A device includes: A feedback collection module 10, configured to generate answers to at least one received question based on reference Q&A examples through a large language model, and collect feedback data corresponding to the at least one generated answer; A performance evaluation module 20, configured to perform performance evaluation on the large language model based on the feedback data to obtain performance parameters; An example update module 30, configured to update the reference Q&A examples based on the feedback data when the performance parameters are lower than the performance threshold, and the updated reference Q&A examples are used to guide the large language model to generate answers to questions subsequently.
[0110] Optionally, the example update module 30 includes: A standard answer acquisition unit, configured to obtain standard answers to at least one question based on the feedback data; A candidate example construction unit, configured to form candidate Q&A examples with each question and the standard answer to each question; An example update unit, configured to update the reference Q&A examples based on the constructed multiple candidate Q&A examples, so that the answer performance of the large language model based on the updated reference Q&A examples meets the performance conditions.
[0111] Optionally, the large language model is a student large model, and the performance of the student large model is inferior to that of the teacher large model; A standard answer acquisition unit, configured to regenerate optimized answers to at least one question through the teacher large model based on the feedback data; and use the optimized answers generated by the teacher large model for at least one question as the standard answers to at least one question.
[0112] Optionally, the standard answer acquisition unit is configured to extract the corrected answer corresponding to each answer from the feedback data corresponding to at least one answer; and use the corrected answer corresponding to each answer as the standard answer to each question.
[0113] Optionally, an example update unit is configured to generate multiple sets of test examples based on multiple candidate Q&A examples, where each set of test examples includes at least one candidate Q&A example; obtain performance parameters of the large language model on the test data set respectively based on each set of test examples; and use the set of test examples with the optimal performance parameter among the multiple sets of test examples as the updated reference Q&A example.
[0114] Optionally, the performance evaluation module 20 includes: A test question selection unit is configured to select test questions from multiple questions received by the large language model in the current time period; A test set construction unit is configured to form a test data set with the test questions and the feedback data corresponding to the answers to the test questions; A performance parameter determination unit is configured to determine performance parameters based on the answers corresponding to the test questions in the test data set and the feedback data corresponding to the answers to the test questions.
[0115] Optionally, the test question selection unit is configured to classify multiple questions received by the large language model in the current time period to obtain multiple question sets, where the types of questions included in each question set are different; and extract at least one test question from each question set.
[0116] Optionally, the test question selection unit is configured to randomly select a target number of test questions from multiple questions received by the large language model in the current time period.
[0117] Optionally, the apparatus further includes: An example deployment module is configured to obtain performance parameters of the large language model on the test data set respectively based on the reference Q&A example before update and the reference Q&A example after update; and in the case where the performance parameter corresponding to the reference Q&A example after update is better than the performance parameter corresponding to the reference Q&A example before update, deploy the reference Q&A example after update as the guiding example for the large language model to generate answers subsequently.
[0118] Optionally, the test data set includes multiple test subsets of different types, The example deployment module is configured to, in the case where the performance parameter corresponding to the reference Q&A example after update is better than the performance parameter corresponding to the reference Q&A example before update for any test subset, deploy the reference Q&A example after update as the guiding example for the large language model to generate answers subsequently.
[0119] The Q&A device provided in this application adopts the Q&A method in the above embodiment, which can solve the technical problem that in the related art, the model is adapted to new input and output requirements by re-labeling data and re-training the model, resulting in the model being difficult to quickly respond to the dynamic change requirements of the actual scenario. Compared with the prior art, the beneficial effects of the Q&A device provided in this application are the same as those of the Q&A method provided in the above embodiment, and other technical features in the Q&A device are the same as those disclosed in the method of the above embodiment, which will not be elaborated here.
[0120] This application provides a Q&A device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the Q&A method in the above embodiment.
[0121] Next, refer to Figure 8 , which shows a schematic structural diagram of a Q&A device suitable for implementing the embodiments of this application. The Q&A device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The Q&A device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of this application.
[0122] As Figure 8As shown in the figure, the question-and-answer device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM: Read Only Memory) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the question-and-answer device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the question-and-answer device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a question-and-answer device having various systems, it should be understood that it is not required to implement or have all the shown systems. Instead, more or fewer systems can be implemented or had.
[0123] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiments disclosed in the present application are executed.
[0124] The question-and-answer device provided by the present application adopts the question-and-answer method in the above-mentioned embodiment, and can solve the technical problem that in the related art, the model is made to adapt to new input and output requirements by re-labeling data and re-training the model, resulting in the model being difficult to quickly respond to the dynamic change requirements of the actual scenario. Compared with the prior art, the beneficial effects of the question-and-answer device provided by the present application are the same as those of the question-and-answer method provided by the above-mentioned embodiment, and the other technical features in this question-and-answer device are the same as the features disclosed in the method of the previous embodiment, and will not be elaborated here.
[0125] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0126] As described above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0127] This application provides a computer-readable storage medium with computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the Q&A method in the above embodiments.
[0128] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM: Random Access Memory), read-only memory (ROM: Read Only Memory), erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0129] The above computer-readable storage medium can be included in the Q&A device; or it can exist separately without being assembled into the Q&A device.
[0130] The above computer-readable storage medium carries one or more programs, which, when executed by a question-and-answer device, cause the question-and-answer device to: generate answers to at least one received question based on reference question-and-answer examples through a large language model, and collect feedback data corresponding to the at least one generated answer; perform a performance evaluation on the large language model based on the feedback data to obtain performance parameters; and in the case where the performance parameters are lower than a performance threshold, update the reference question-and-answer examples based on the feedback data, and the updated reference question-and-answer examples are used to guide the large language model to generate answers to questions subsequently.
[0131] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., by connecting through an Internet service provider via the Internet).
[0132] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in an order different from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0133] The modules described in the embodiments of the present application may be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.
[0134] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned question-and-answer method, which can solve the technical problem that in the related art, the model is made to adapt to new input and output requirements by re-labeling data and re-training the model, resulting in the model being difficult to quickly respond to the dynamic change requirements of the actual scenario. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the question-and-answer method provided in the above embodiment, and will not be elaborated here.
[0135] This application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it realizes the steps of the question-and-answer method as described above.
[0136] The computer program product provided by this application can solve the technical problem that in the related art, the model is made to adapt to new input and output requirements by re-labeling data and re-training the model, resulting in the model being difficult to quickly respond to the dynamic change requirements of the actual scenario. Compared with the prior art, the beneficial effects of the computer program product provided by this application are the same as those of the question-and-answer method provided in the above embodiment, and will not be elaborated here.
[0137] The above are only some embodiments of this application, and thus do not limit the patent scope of this application. Any equivalent structural transformation made under the technical concept of this application by using the content of the specification and drawings of this application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of this application.
Claims
1. A question-answering method, characterized in that, The method includes: Based on a large language model, generating answers to at least one received question based on reference Q&A examples, and collecting feedback data corresponding to the at least one generated answer; Based on the feedback data, performing a performance evaluation on the large language model to obtain performance parameters; In the case where the performance parameters are lower than a performance threshold, updating the reference Q&A examples based on the feedback data, and the updated reference Q&A examples are used to guide the large language model to generate answers to questions subsequently.
2. The method according to claim 1, characterized in that, The updating the reference Q&A examples based on the feedback data includes: Based on the feedback data, obtaining the standard answers to the at least one question; Forming candidate Q&A examples with each question and the standard answer to each question; Updating the reference Q&A examples based on the constructed multiple candidate Q&A examples, so that the answer performance of the large language model based on the updated reference Q&A examples meets the performance conditions.
3. The method according to claim 2, wherein The large language model is a student large model, and the performance of the student large model is inferior to that of a teacher large model; the obtaining the standard answers to the at least one question based on the feedback data includes: Through the teacher large model, regenerating optimized answers to the at least one question based on the feedback data; Taking the optimized answers generated by the teacher large model for the at least one question as the standard answers to the at least one question.
4. The method according to claim 2, characterized in that, The obtaining the standard answers to the at least one question based on the feedback data includes: Extracting the corrected answer corresponding to each answer from the feedback data corresponding to the at least one answer; Taking the corrected answer corresponding to each answer as the standard answer to each question.
5. The method according to claim 2, wherein The updating the reference Q&A examples based on the constructed multiple candidate Q&A examples, so that the answer performance of the large language model based on the updated reference Q&A examples meets the performance conditions, includes: Generating multiple groups of test examples based on the multiple candidate Q&A examples, and each group of test examples includes at least one candidate Q&A example; Obtaining the performance parameters of the large language model on a test data set based on each group of test examples respectively; Taking the group of test examples with the optimal performance parameters among the multiple groups of test examples as the updated reference Q&A examples.
6. The method according to claim 1, characterized in that, The performing a performance evaluation on the large language model based on the feedback data to obtain performance parameters includes: Selecting test questions from multiple questions received by the large language model in the current time period; Forming a test data set with the test questions and the feedback data corresponding to the answers to the test questions; Determining the performance parameters based on the answers corresponding to the test questions in the test data set and the feedback data corresponding to the answers to the test questions.
7. A question-and-answer device, characterized in that, The device includes: A feedback collection module, configured to generate answers to at least one received question based on a large language model based on reference Q&A examples, and collect feedback data corresponding to the at least one generated answer; A performance evaluation module, configured to perform a performance evaluation on the large language model based on the feedback data to obtain performance parameters; An example update module, configured to update the reference Q&A example based on the feedback data when the performance parameter is lower than the performance threshold, and the updated reference Q&A example is used to guide the large language model to generate answers to questions subsequently.
8. A question-and-answer device, characterized in that, The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program being configured to implement the steps of the Q&A method according to any one of claims 1 to 6.
9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the Q&A method according to any one of claims 1 to 6 are implemented.
10. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the Q&A method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Self-correcting intelligent teaching assistance method based on instruction-guided large language model
CN116860922A
Question and answer model editing method and device, electronic equipment and storage medium
CN116882450A
Assessment method and device of large language model system and related equipment
CN119179631A
Prompt word verbal skill hot update processing method for intelligent customer service question and answer scene
CN119961273A
Method, apparatus and computer-readable medium for operating chatbot
KR102047385B1