Model training data generation method and electronic equipment

Through self-evaluation and rewriting input content, high-quality training data is generated, which solves the misunderstanding problem of the LLMs model when dealing with fuzzy or unclear user input, and realizes more stable model training and response content generation more in line with human preferences.

CN119988963APending Publication Date: 2025-05-13HANGZHOU ALIBABA INT INTERNET IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411813712.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the intelligent customer service scenario, when the LLMs model is difficult to deal with user input, it may lead to misunderstanding of the model and generate response content that does not meet the user's expectations.

Method used

By obtaining the original input content, the first AI generation model is called to generate the response content and perform self-evaluation. If the evaluation score does not meet the criteria, the original input is rewritten and the response content and self-evaluation are regenerated until the criteria are met. Then, based on the input content and evaluation scores during the rewriting process, positive and negative samples are determined, and triple data is constructed for training the second AI generation model.

Benefits of technology

This method can save resources and costs required to construct training data, the model training results are more stable, and can generate response content more in line with human preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988963A_ABST
    Figure CN119988963A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a model training data generation method and electronic equipment. The method comprises the steps that multiple pieces of original input content are acquired; generating response content corresponding to the original input content by a first AI generation model; performing self-evaluation on the response content by the first AI generation model; if the evaluation score does not meet the condition, calling the first AI generation model to rewrite the original input content, and according to the rewritten input content, re-executing the generation and self-evaluation of the response content until the evaluation score of the response content corresponding to the rewritten input content meets the condition or the cycle index reaches a threshold value; and determining the input contents belonging to the positive samples and the negative samples according to the input contents of multiple versions obtained in the input content rewriting process and the evaluation scores of the corresponding response contents so as to serve as model training data of the second AI generation model. Through the embodiment of the invention, resources and cost required for constructing the training data can be saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of AI generation model technology, and in particular to a method for generating model training data and an electronic device. Background Art

[0002] In the intelligent customer service scenario, LLMs (Large Language Model) has a wide range of application scenarios. For example, in the process of serving foreign users, it can provide AI polishing (that is, polishing the customer service response content through the LLMs model) or assisted reception (providing customer service with suggestions on the response content after the user enters the question) and other services.

[0003] The LLMs model has a strong ability to understand semantics. In real business communication and basic product information questions and answers, after fine-tuning with domain data, it performs well in following instructions and imitating customer service responses. However, in practical applications, the questions (queries) entered by users may be unclear or even misunderstood by the model, so that even if the model itself is powerful, it may be difficult to give a satisfactory answer to the user. Therefore, how to align LLMs with complex human preferences (for example, understanding the unexpressed meaning from the query entered by the user, etc.) has always been a major challenge facing academia and industry. Summary of the invention

[0004] The present application provides a method for generating model training data and an electronic device, which can save the resources and costs required to construct training data, and the model training results are more stable.

[0005] This application provides the following solutions:

[0006] A method for generating model training data, comprising:

[0007] Get multiple raw input contents;

[0008] Calling a first artificial intelligence AI generation model according to the original input content to generate response content corresponding to the original input content;

[0009] Calling the first AI generation model to perform self-evaluation on the response content according to the original input content, the response content, and a preset first evaluation rule;

[0010] If the evaluation score does not meet the conditions, the first AI generation model is called to rewrite the original input content, and the response content generation and self-evaluation are re-executed according to the rewritten input content, until the evaluation score of the response content corresponding to the rewritten input content meets the conditions or the number of cycles reaches the threshold;

[0011] Based on the multiple versions of input content obtained during the input content rewriting process and the evaluation scores of the corresponding response content, the input content belonging to the positive sample and the negative sample is determined, and triple data is constructed based on the original input content, the positive sample input content, and the negative sample input content to be used as model training data for the second AI generation model.

[0012] The first AI generation model and the second AI generation model are models of the same series and / or have the same / similar content generation preferences.

[0013] Among them, the performance of the first AI generated model is better than that of the second AI generated model.

[0014] Among them, the triplet data includes the original input content, the positive sample input content and its corresponding response content, the negative sample input content and its corresponding response content, so that during the training process, the second AI generation model rewrites the input content and generates corresponding response content, and makes the response content corresponding to the rewritten input content tend to the response content corresponding to the positive sample input content, and away from the response content corresponding to the negative sample input content.

[0015] Among them, it also includes:

[0016] When the response content is evaluated by the first AI generation model and the evaluation score does not meet the requirements, the first AI generation model is used to analyze the defects in the response content and provide improvement suggestions for the response content, so that the first AI generation model can rewrite the input content according to the improvement suggestions.

[0017] Among them, it also includes:

[0018] Based on the original input content and the rewritten input content, the first AI generation model is called to evaluate the changes in the rewritten input content relative to the original input content. If the evaluation fails, the rewriting of the input content is triggered.

[0019] A method for generating dialogue reply content, comprising:

[0020] Determine the current question asked by the user in the target session;

[0021] The current question content is rewritten by a second AI generation model to generate a rewritten question content; wherein the second AI generation model is trained in the following manner:

[0022] Get multiple raw input contents;

[0023] Calling a first artificial intelligence AI generation model according to the original input content to generate response content corresponding to the original input content;

[0024] Calling the first AI generation model to perform self-evaluation on the response content according to the original input content, the response content, and a preset first evaluation rule;

[0025] If the evaluation score does not meet the conditions, the first AI generation model is called to rewrite the original input content, and the response content generation and self-evaluation are re-executed according to the rewritten input content, until the evaluation score of the response content corresponding to the rewritten input content meets the conditions or the number of cycles reaches the threshold;

[0026] Determine the input content belonging to the positive sample and the negative sample according to the multiple versions of the input content obtained in the input content rewriting process and the evaluation scores of the corresponding response content, and construct triple data according to the original input content, the positive sample input content, and the negative sample input content for training the second AI generation model;

[0027] Prompt information is generated based on the question content rewritten by the second AI generation model and input into the reasoning model to generate reply content or generate reply content for recommendation.

[0028] A method for optimizing conversation reply content, comprising:

[0029] Determine the answer content to the user's question in the target session;

[0030] The reply content is rewritten by a second AI generation model to generate a rewritten reply content; wherein the second AI generation model is trained in the following manner:

[0031] Get multiple raw input contents;

[0032] Calling a first artificial intelligence AI generation model according to the original input content to generate response content corresponding to the original input content;

[0033] Calling the first AI generation model to perform self-evaluation on the response content according to the original input content, the response content, and a preset first evaluation rule;

[0034] If the evaluation score does not meet the conditions, the first AI generation model is called to rewrite the original input content, and the response content generation and self-evaluation are re-executed according to the rewritten input content, until the evaluation score of the response content corresponding to the rewritten input content meets the conditions or the number of cycles reaches the threshold;

[0035] Determine the input content belonging to the positive sample and the negative sample according to the multiple versions of the input content obtained in the input content rewriting process and the evaluation scores of the corresponding response content, and construct triple data according to the original input content, the positive sample input content, and the negative sample input content for training the second AI generation model;

[0036] Prompt information is generated based on the reply content rewritten by the second AI generation model and input into the reasoning model for optimizing the reply content.

[0037] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of any of the methods described above.

[0038] An electronic device, comprising:

[0039] one or more processors; and

[0040] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of any of the methods described above.

[0041] A computer program product comprises a computer program / computer executable instructions, wherein the computer program / computer executable instructions implement the steps of any of the aforementioned methods when executed by a processor in an electronic device.

[0042] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0043] Through the embodiments of the present application, when constructing training data, the first AI generation model can be used to generate corresponding response content for the original input content, and then the first AI generation model can be used to self-evaluate the response content generated by itself. If the evaluation result does not meet the preset conditions, the original input content can be rewritten by the first AI generation model, and then the response content generation and self-evaluation are re-executed based on the rewritten input content, and this cycle is repeated until the evaluation score of the response content corresponding to the rewritten input content meets the conditions or the number of cycles reaches the threshold. Afterwards, based on the multiple versions of input content obtained during the input content rewriting process and the corresponding evaluation scores of the response content, the input content belonging to the positive sample and the negative sample is determined, and triple data is constructed for training the second AI generation model so that the second AI generation model acquires the ability to rewrite the input content. Through this solution, in the process of obtaining training data, only the same AI generation model is needed to complete the generation of response content, self-evaluation and rewriting of input content to construct training data triplets. There is no need to call a third-party AI generation model or manual labeling. Therefore, it is beneficial to save resources and costs. The construction of positive and negative samples will not be affected by the different standards of labelers, making the model training results more stable and improving model performance.

[0044] In other words, the embodiment of the present application utilizes the knowledge and judgment of the first AI generation model itself to generate high-quality training data through self-evaluation, avoiding the high cost and bias problems of manual labeling and calling third-party AI generation models. Moreover, through the self-evaluation and feedback mechanism, the model can continuously optimize prompts, and only a small number of iterations are required to generate rich, high-quality preference data, which is conducive to improving the efficiency and quality of training data collection. In addition, with sophisticated process design, the solution has high computational efficiency in both training and reasoning stages.

[0045] In a preferred manner, the first AI generation model and the second AI generation model can be from the same series and / or have the same / similar generation preferences. In this way, through the self-evaluation and feedback mechanism of the first AI generation model, high-quality data consistent with the internal logic of the second AI generation model can be generated, thereby adapting to the needs of different second AI generation models and ensuring the effectiveness and pertinence of the optimization prompts.

[0046] In another preferred manner, the response contents corresponding to the positive and negative samples can also be added to the training data. In this way, during the training process, the second generation model needs to consider not only what the modified query looks like, but also what kind of response the modified query will generate. In addition, the training goal can be "making the response content corresponding to the rewritten input content tend to the response content corresponding to the positive sample input content, and away from the response content corresponding to the negative sample input content", so as to obtain a better training effect. In this way, in the mapping process from the original query to the rewritten query, the response that may be triggered by the rewritten query can be fully considered, and a more comprehensive and targeted optimization can be performed.

[0047] In other words, in the process of training the second AI generation model using training data, implicit reasoning can be incorporated based on the training method of reinforcement learning, so that the model can fully consider the actual effect when rewriting the input content, thereby improving the effectiveness and pertinence of the input content rewriting. In addition, thanks to the training method based on reinforcement learning and the paradigm of injecting implicit reasoning, the prompt optimizer trained in the solution provided in the embodiment of the present application shows good generalization ability. On data sets in multiple different fields, the prompt optimizer can generate high-quality optimization prompts that adapt to the characteristics of the field, helping the model to produce more accurate and professional responses.

[0048] In addition, the trained second AI generation model can be used as a query rewriter (or prompt optimizer, etc.), which can be applied to any original input content to automatically generate optimized input content to make it more in line with human preferences. Moreover, the query rewriter can quickly adapt to new tasks, generate high-quality, expected responses, and show good adaptability and flexibility. It provides a general and efficient framework for handling diverse tasks, while improving the ability to follow instructions in the field while taking into account general performance. This pluggable prompt optimizer only requires one additional reasoning during reasoning. Specifically, the original input content is input into the prompt optimizer to generate an optimized prompt, which is then sent to the reasoning model to produce the final response. Compared with the traditional prompt word engineering method, this method significantly reduces the need for multiple calls to large models, thereby reducing the cost of reasoning.

[0049] Of course, any product implementing the present application does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0051] Figure 1 It is a schematic diagram of the system architecture provided by the embodiment of the present application;

[0052] Figure 2 is a flow chart of the first method provided in an embodiment of the present application;

[0053] Figure 3 is a flow chart of the second method provided in an embodiment of the present application;

[0054] Figure 4 is a flowchart of the third method provided in an embodiment of the present application;

[0055] Figure 5 It is a schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0056] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.

[0057] In order to facilitate understanding of the solution provided by the embodiments of the present application, it is first necessary to explain that in the process of using the LLMs model for AI polishing, assisted reception and other services, in order to improve the quality of response content generation, there are usually two ways: "aligned output" and "aligned input". For example, usually supervised fine-tuning, reinforcement learning and other methods can be used to achieve aligned output, and prompt word engineering can be used to achieve aligned input, etc. However, existing implementation solutions all have problems such as data bias and insufficient generalization.

[0058] Supervised fine-tuning improves instruction-following capabilities by maximizing the alignment of generated responses with reference outputs, but the strong constraints it uses (such as the cross-entropy loss function) can undermine the model's native capabilities. In addition, the training data for supervised fine-tuning is usually in the format of query + answer, which places extremely high demands on the quality of data annotation. It often requires manual annotation or the use of high-performance large models to annotate the corresponding answers to specific questions, which is very costly.

[0059] Reinforcement learning provides a new output alignment method after supervised fine-tuning, which only requires acceptance data, namely query (question) + good answer (positive answer sample) + bad answer (negative answer sample), and only requires the relative quality of the answers, thereby reducing the labeling requirements. However, this method also faces some problems, such as high training cost, unstable training, and the preferences reflected by positive and negative sample data in the data set may conflict (such as different standards of labelers), which may affect the performance of the final model.

[0060] In addition, if the input prompt is vague, it is difficult to obtain a satisfactory answer only by aligning the output. Therefore, there are some methods of aligned input in the prior art, namely, prompt word engineering. The so-called prompt word engineering means that the input prompt information is first rewritten, and then the response content is generated. Among them, the rewriting of the prompt information is mainly to rewrite the specific input content. Regarding "input content", it can also have different meanings in different scenarios. For example, in scenarios such as auxiliary reception, it is necessary to provide recommended reply content for customer service. The questions entered by specific users during the conversation belong to the above-mentioned "input content"; or, in the AI ​​polishing scenario, since it is necessary to polish the reply content entered by the customer service, the reply content entered by the customer service can be used as "input content", and so on. In practical applications, the various "input contents" mentioned above can be collectively referred to as "queries". Usually, a prompt template will be provided according to the specific tasks that the large model needs to complete. The prompt template will include a "query" field to be filled in, and may also include content that describes the tasks that the large model needs to complete (for example, telling the model that the user's expression is unclear, please try to say the user's hidden information, etc.). After determining the specific "query", the specific content of the "query" is assembled into the prompt template to generate a specific prompt for input to the large model to guide the output of the large model. Since the query part is input by the user (consumer user or customer service, etc.), there may be unclear or incomplete expressions. Therefore, the query part has optimization requirements. Accordingly, the so-called prompt word engineering is mainly to rewrite the query part to make the query expression more accurate, clear, complete, etc., and then assemble the rewritten query into the prompt template so that the prompt is also optimized.

[0061] There are two main types of traditional prompt word engineering. One is a method that does not rely on model training. In this method, after receiving the user's query, the query is first rewritten through the big model. Afterwards, because the big model has not been specially trained, it is necessary to call the big model for "reflection" to determine whether there is a big gap between the rewritten query and the original query in terms of the meaning expressed. If there is, it needs to be rewritten and re-reflected until the rewritten query meets the requirements, and then the big model is called to reason based on the new query. During this period, at least three calls to the big model are required, which is costly and may cause response delays.

[0062] Another way is to pre-train a large model (which can be called a "rewriter") for rewriting queries. In this way, after receiving the query input by the user, the "rewriter" is directly used to rewrite the input query, and then the reasoner (which can be another large model) is called for reasoning and the response content is generated. In this way, from rewriting to reasoning, only two calls to the large model are needed, so the cost and response delay are relatively low. However, in the process of implementing this application, the inventors of the present application found that the above-mentioned scheme of pre-training the rewriter has at least the following problems:

[0063] First, there are problems in the collection or generation of training data. In order to train the rewriter, some training data needs to be collected in advance. In the existing method, the required training data set is a plurality of data pairs, each of which is composed of (original query, rewritten query). In order to obtain the above training data, some queries can be collected from open source data sets or from the historical conversation records of the business party, or some queries can be constructed by handwriting, etc. After that, the high-performance LLMs model and the ordinary performance LLMs model can be used for reasoning to generate the corresponding response content (answer), and then which response content is better can be marked by manual annotation. Then, the original query and the knowledge that response content 1 is better than response content 2 are told to the high-performance LLMs model to prompt the high-performance LLMs model how to rewrite the query, so that the reasoning model (specifically which model is used for reasoning is unknown to the high-performance LLMs model used for query rewriting) is more likely to answer content similar to response content 1. Then, the high-performance LLMs will rewrite the query according to the above guidance information to obtain the rewritten query. After that, a training data pair can be constructed: (original query, rewritten query). After constructing multiple training data pairs in a similar way, these training data pairs can be used to train the large model used to provide query rewriting function.

[0064] The disadvantage of the above method is that in the process of constructing training data, a high-performance LLMs model needs to be called. This high-performance LLMs model usually belongs to a third-party model, and the call of this model may incur a high cost. In addition, this method also needs to rely on manual annotation, and the process of manual annotation will also face high labor costs and time costs. Furthermore, this method does not consider the differences between different reasoning models in the process of generating training data pairs. Although in theory, the rewriter trained by it can be used with any reasoning model, different reasoning models may have some differences in answer preferences, etc. For example, some reasoning models like to answer questions directly, while some reasoning models like to say some polite words first and then answer questions, and so on. If query rewriting is performed without considering such differences, it is difficult to truly improve the quality of response content generation of the reasoning model, and it may even lose the inherent data distribution of the reasoning model.

[0065] Secondly, there are problems with the model training method. After collecting specific training data, it involves training the model. In the prior art, supervised fine-tuning is used for training. Supervised fine-tuning is a common machine learning method. The idea is to train the model by annotating data so that it can better match the data. In other words, it is equivalent to giving a query and an answer, and letting the model memorize the "correct answer" during the training process. If the answer generated by the model deviates slightly from the given answer, the penalty will be very strong. Therefore, there is a problem that the supervision method is too strict.

[0066] In view of the above situation, the embodiment of the present application provides a corresponding solution. In this solution, firstly, the method of obtaining training data is optimized. Specifically, the query can be rewritten by using the model self-evaluation method and the training data triples can be generated; secondly, in terms of training method, the reinforcement learning training method can be used to achieve implicit reasoning instead of the supervised fine-tuning method.

[0067] Specifically, when collecting training data, you can first obtain multiple original input contents. As mentioned above, this input content can be the aforementioned content that can be used as a query. In different scenarios, it can have different meanings. Taking the auxiliary reception scenario as an example, it can refer to the questions entered by the user during the communication with the customer service. Among them, since the training data is collected, the specific query is not the user input query received during the actual online customer service process, but can be collected and obtained through other means. For example, the original query can be collected from some open source databases, or the original query can be collected from the historical conversation records of the online customer service system (in the auxiliary reception scenario, only the question content of the consumer user in the historical conversation record needs to be used as the query), or it can also be some manually constructed queries, and so on.

[0068] After the above original query is collected, the training data can be constructed by the AI ​​generation model (including the aforementioned LLMs model). That is to say, in the embodiment of the present application, the purpose of collecting training data is to train the AI ​​generation model used for query rewriting. However, in the process of collecting training data, the AI ​​generation model will also be used. In order to distinguish the AI ​​generation model used in the process of collecting training data from the AI ​​generation model for query rewriting that needs to be trained, the former is called the first AI generation model, and the latter is called the second AI generation model. In a preferred implementation, the first AI generation model and the second AI generation model can be models of the same series. For example, it can be an AI generation model provided by the same developer but with different performance, and the performance of the first AI generation model can be better than the second AI generation model. Among them, AI generation models with different performances can usually be determined according to the different parameter scales used. In a series of AI generation models provided by a developer, there may be multiple different versions such as 7B, 14B, and 32B, among which 7B corresponds to 7 billion parameter scales, 14B corresponds to 14 billion parameter scales, and 32B version corresponds to 32 billion parameter scales. Generally speaking, the larger the parameter scale, the better the model performance, and the stronger the natural language understanding, reasoning, generation and other capabilities. In the process of collecting training data, since the links involved are relatively long, including a series of capabilities such as response content generation, self-assessment, and query rewriting, the first AI generation model can be a higher-performance AI generation model, for example, it can be the aforementioned 14B AI generation model, otherwise it may not be able to cope with the processing on the above-mentioned long link. The second AI generation model can be the aforementioned 7B AI generation model, and so on. Alternatively, if the first AI generation model and the second AI generation model do not belong to the same series, but are the same or similar in generation preference, better training results can also be achieved. Of course, if the first generation model and the second generation model are not models of the same series, better results can be obtained compared to the solutions in the prior art.

[0069] Specifically, when the training data is generated by the first AI generation model, the prompt can be constructed according to the original input content, and the first AI generation model can be called so that the response content corresponding to the original input content is generated by the first AI generation model. Then, the response content can be directly evaluated by the first AI generation model, that is, the same model performs self-evaluation on the response content generated by itself without the help of a third-party AI generation model or manual. If the evaluation score does not meet the preset conditions (it can be lower than a certain threshold, for example, when the total score is 10 points, the evaluation score is lower than 5 points, etc.), the original input content can be rewritten by the first AI generation model, and the generation and self-evaluation of the response content can be re-executed according to the rewritten input content. If it still does not meet the conditions, it continues to be rewritten until the evaluation score of the response content corresponding to the rewritten input content meets the conditions. Alternatively, in actual applications, in order to balance the generation quality and computational efficiency, an upper limit on the number of iterations can also be set. For example, after 3 rounds of loops, if the evaluation score still does not meet the conditions, the iteration process can also be forced to end. At this time, if the response content corresponding to the latest version of the rewritten input content is close to the aforementioned threshold, it can be reluctantly added to the training data; otherwise, the input content can be discarded.

[0070] In the above-mentioned process of generating, evaluating, and rewriting the response content, the same AI generation model is used to complete the process. It does not involve calling other third-party models or manual labeling. Therefore, as long as the program is written according to the above-mentioned processing logic, the first AI generation model call and other processes in each step can be automatically completed.

[0071] After completing the above-mentioned response content generation, evaluation, input content rewriting and other processes, the second AI generation model needs to be trained by reinforcement learning in the future, that is, it is necessary to construct a triple of training data, which may include the original input content, and may also include input content belonging to positive samples, and input content belonging to negative samples, so that the second AI generation model can acquire the ability to judge who is better than who during the training process, and only learn preferences without memorizing data. The training penalty is relatively light, which is conducive to retaining the model's own capabilities.

[0072] In an embodiment of the present application, in order to obtain a better training effect, the specifically constructed triples may include: the original input content (query), the positive sample input content and its corresponding response content, and the negative sample input content and its corresponding response content. That is to say, in an embodiment of the present application, the response content corresponding to the positive and negative samples, respectively, can also be used as training data. During subsequent training, the task specified to the second AI generation model through Prompt can be: rewrite the original query, and generate the response that the rewritten query may bring at the same time, that is, not only what the modified query looks like, but also what kind of response the modified query will generate, and the training goal can be "making the response content corresponding to the rewritten input content tend to the response content corresponding to the positive sample input content, away from the response content corresponding to the negative sample input content", so as to obtain a better training effect. In this way, in the mapping process from the original query to the rewritten query, the response that may be triggered by the rewritten query can be fully considered, and a more comprehensive and targeted optimization can be carried out.

[0073] In order to construct the above triples, the problem of how to determine positive samples and negative samples is also involved. Specifically, since the aforementioned link involves the evaluation process of the response content by the first AI generation model, and the evaluation result of the response content can reversely prove the quality of the input content, the positive and negative samples of the input content can be determined according to the evaluation of the response content. Among them, if the evaluation score of the response content obtained for the rewritten input content can meet the preset conditions (for example, if the full score is 5 points, then the score is 4 to 5 points when it meets the preset conditions), it proves that the rewritten input content is more likely to obtain better response content. Therefore, this rewritten input content is a better sample and can be called a positive sample; if some intermediate versions are generated during the rewriting process, for example, the version after the first rewrite, the evaluation score of the response content may still not reach the threshold (for example, if the full score is 5 points, then the score is 1 to 3 points when it does not meet the preset conditions), or even lower than the evaluation score of the response content corresponding to the original version, then the rewritten input content of this intermediate version can be used as a negative sample. If the response content obtained after a certain original input content is rewritten once has reached the threshold, then there may be no intermediate version of the rewriting result. At this time, the original version of the input content can also be used as a negative sample, and so on. In short, through the above method, a triple of the form (original query, positive sample query and its corresponding response content, negative sample query and its corresponding response content) can be obtained.

[0074] After obtaining multiple sets of the above training data triplets, the second AI generation model can be trained by reinforcement learning so that the second AI generation model can acquire the ability to rewrite the input content. Among them, in a preferred manner, implicit reasoning optimization can also be incorporated into the training process, that is, given the original query, the model generates the rewritten query while predicting and generating the response that the rewritten query may trigger. By incorporating implicit reasoning, the model can be prompted to fully consider the actual effect of the rewriting when rewriting the query, that is, what kind of response will be guided by the model. In this way, the second AI generation model can be optimized by preference data, eliminating the explicit reward modeling step, and its goal can be to maximize the probability of preference output (that is, the probability of generating the response content corresponding to the positive sample) while minimizing the probability of non-preference output (that is, the probability of generating the response content corresponding to the negative sample). It is equivalent to letting the model learn the ability to judge which is better (that is, learning preference), and then output more expected response content through this ability.

[0075] It should be noted that, unlike traditional supervised fine-tuning, the above training method is not limited to matching the reference response token by token (information input unit of AI generation model), but evaluates and optimizes the overall optimization effect based on the scoring information in the preference data. This optimization paradigm retains greater expression flexibility for the model, avoids overfitting the expression of the reference response, and helps to improve the generalization ability of the optimizer without destroying the knowledge and data distribution learned in pre-training like the supervised fine-tuning scheme. Therefore, this mechanism helps to improve the effectiveness and pertinence of optimization.

[0076] This method of collecting data by self-evaluating the first AI generation model of the same series or the same / similar production preferences can automatically generate high-quality data that is consistent with the internal logic of the second AI generation model, which is both simple and efficient. Moreover, the training data is generated through the model's own feedback mechanism, avoiding the deviation problem of manual labeling and ensuring the consistency and high quality of the data.

[0077] After training, the second AI generation model can be trained into a "pluggable" query optimizer. Since the query will be assembled into the Prompt template to generate prompt information and then input into the inference model, it can also be called a "prompt optimizer." Given any original query, the optimizer can automatically generate the corresponding optimized query to make it more in line with human preferences (including the ability to complete the meaning that the user failed to express, etc.), and can guide the subsequent inference model to produce high-quality, expected response content.

[0078] From the perspective of system architecture, see Figure 1 The embodiment of the present application mainly involves constructing training data using a first AI generation model to train the query rewriting capability of a second AI generation model. The trained second generation model can be used to rewrite the query actually input by the user, and then assembled into a prompt template to generate a prompt, and a third generation model (i.e., a model used for reasoning) is used to generate corresponding response content based on the prompt assembled from the rewritten query.

[0079] Among them, in the training data construction stage, firstly, multiple original queries can be obtained, and then the response content can be generated by the first AI generation model, and then the first AI generation model can be used to perform self-evaluation on the response content generated by itself (at this time, a prompt can be constructed based on the original query and the response content). In an optional manner, if the evaluation score is relatively low, the first AI generation model can also generate optimization suggestions for the response content; afterwards, the first AI generation model can also perform query rewriting (prompt can be constructed based on the original query, response content, self-evaluation score, optimization suggestions, etc.), and then the response content generation and evaluation steps are performed again through the first AI generation model until the evaluation score is higher than a certain threshold, or the number of iterations exceeds a certain threshold, and so on. Through the above process, the rewritten query with an evaluation score higher than the threshold can be obtained as a positive sample query, the rewritten query with an evaluation score for the threshold or the original query can be obtained as a negative sample query, and then (original query, positive sample query, negative sample query) can be formed into a training data triple. Alternatively, in a preferred implementation, the response content corresponding to the positive and negative sample queries can also be used as training data, that is, the composed triples can be specifically (original query, positive sample query + response content, negative sample query + response content). After obtaining the above training data triples, it can be used to train the second AI generation model, wherein a reinforcement learning training method can be adopted.

[0080] The specific implementation scheme provided in the embodiments of the present application is described in detail below.

[0081] Embodiment 1

[0082] First, this embodiment provides a method for generating model training data. Figure 2 , the method may include:

[0083] S201: Acquire multiple original input contents.

[0084] As mentioned above, the "input content" described in the embodiments of the present application is also called "query", which can be the question content input by the consumer user in the scenario of assisted reception of the intelligent customer service system, or the reply content input by the customer service staff to the user in the scenario of AI polishing, etc. In specific implementation, the original input content can be collected and the subsequent training data can be generated for different scenarios, so as to complete the training of the model for different scenarios, etc.

[0085] There are many different ways to collect original queries. For example, the original queries can be collected from some open source databases, or from historical conversation records of online customer service systems (in assisted reception scenarios, only the questions asked by consumers in historical conversation records need to be used as queries, for example, "What should I do if the merchant does not ship the goods?", etc.), or some manually constructed queries, etc.

[0086] It should be noted here that since the number of original queries required for the model training process may be relatively large, if the number of collected original queries is not enough, the collected original queries can be used as seed queries, and then the number of queries can be expanded through extended generation. That is, the AI ​​generation model can be used to write similar queries based on the seed query, and the AI ​​generation model can be required to write similar questions in the process of writing, imitating the theme of the original query, but not too similar in terms of wording, tone, etc.

[0087] Among them, the specific tasks, requirements and original queries of the AI ​​generation model when performing query expansion can be input to the AI ​​generation model in the form of prompts. For example, in a specific implementation, in order to make the generated query present a richer diversity, the query expansion generation can be performed in multiple times, and a certain number (for example, 5) of original queries are randomly selected from the original query each time, so that the AI ​​generation model can generate similar queries by referring to these original queries. For example, the prompt template input to the AI ​​generation model at this stage can be:

[0088] "instruction:

[0089] You are a professional query generator. Your job is to generate two similar queries based on a given seed query, covering different topics, formats, and complexity. When generating new queries, you can:

[0090] Modify the context, subject, or frame of a seed query

[0091] Explore creative, thought-provoking or unconventional perspectives

[0092] Please generate a total of two novel and interesting queries based on the following five seed queries.

[0093] Seed query: {}

[0094] Seed query: {}

[0095] Seed query: {}

[0096] Seed query: {}

[0097] Seed query: {}

[0098] Please remember:

[0099] The generated query1 should not exceed 15 words / characters.

[0100] The resulting query2 can be slightly longer, but should still be concise.

[0101] Output:

[0102] Generate query1: ''

[0103] Generate query2: ''"

[0104] By assembling a specific seed query into the above prompt template, a specific prompt can be generated, and then input into the AI ​​generation model for query expansion generation. In short, this query written by the AI ​​generation model can be used as a supplement to the original query set. Of course, in practical applications, the original query can also be deduplicated to remove semantically repeated queries, and finally obtain an original query set.

[0105] S202: Calling a first artificial intelligence AI generation model according to the original input content to generate response content corresponding to the original input content.

[0106] After obtaining multiple original queries, positive and negative sample queries can be constructed based on such original queries. Specifically, the response content corresponding to the above original query can be generated by the first AI generation model. For example, in an assisted reception scenario, the query can be the user's question content in an intelligent dialogue scenario, and correspondingly, the specific generated response content can be the reply content used to reply to the user's question content, and so on.

[0107] Specifically, when the response content corresponding to the original query is generated by the first AI generation model, the original query can be assembled into a prompt template to generate a specific prompt, and then input into the first AI generation model to generate the response content. In the prompt template, the specific instruction can be to let the first AI generation model generate the corresponding response content according to the input query, and some precautions and requirements in the generation process can also be described.

[0108] S203: Based on the original input content, the response content and a preset first evaluation rule, the first AI generation model is called to perform a self-evaluation on the response content.

[0109] After generating the corresponding response content according to the original input content, the first AI generation model can also self-evaluate the response content generated by itself. That is to say, in an embodiment of the present application, after the response content for the original query is generated by the first AI generation model, when evaluating it, there is no need for manual intervention or the use of other models. Instead, the first AI generation model can evaluate the response content generated by itself and give an evaluation score. Among them, in order to enable the first AI generation model to have the above-mentioned evaluation capabilities, specific evaluation rules can also be reflected in the prompt template. For example, scoring criteria can be given from multiple angles, etc. For example, specific evaluation rules may include:

[0110] "Score the generated response (1-5 points) on the following dimensions:

[0111] - Accuracy: Does the response accurately answer the question in the prompt?

[0112] - Completeness: Is the response complete and detailed, including all necessary information?

[0113] - Relevance: Is the response content closely related to the prompt topic and does not deviate from the main topic?

[0114] - Fluency: whether the response is grammatically correct, fluent, and logically clear.

[0115] - Safety: Is the response ethical and does not contain inappropriate or harmful content?

[0116] In an optional implementation, after the evaluation score is given, if it is found that the evaluation score is relatively low, for example, below a certain threshold, the first AI generation model can also analyze the defects in the response content and give improvement suggestions. For example, whether the score in one of the above dimensions is relatively low, if so, there may be defects in this dimension, and corresponding improvement suggestions are given. For example, if a response content scores relatively low in the completeness dimension, the improvement suggestion given may be "more complete", etc.

[0117] For example, in one example, the specific prompt template can be:

[0118] "instruction:

[0119] Based on the given query and the relevant response content, evaluate the quality of the response content in terms of accuracy, completeness, relevance, fluency, and safety of the response. Assign scores from 1 to 5, where 1 marks content with significant errors or harmful content and 5 indicates excellent response content that not only meets but also exceeds expectations. Use the following scale for evaluation:

[0120] Score 1: The response is inadequate and contains significant errors or harmful content.

[0121] 2 points: The response is marginally adequate, contains clearly inaccurate or irrelevant information, or contains potentially mildly harmful content.

[0122] 3 points: The response is acceptable, follows instructions, and contains no offensive content, but may lack depth, detail, or insight.

[0123] 4 points: The response is commendable, but has minor issues that prevent it from being outstanding.

[0124] 5 points: The response is exemplary and goes beyond the basic requirements by providing comprehensive, accurate insights and added value beyond the explicit requirements.

[0125] Please evaluate critically and give 5 points for responses that are exceptional in each criterion, with no detectable deficiencies and no need to elaborate on strengths.

[0126] query: {}

[0127] Response content: {}

[0128] Output using the following format:

[0129] Score: [Assign scores based on the above criteria in the format: 'Score: ×', where × is the score.

[0130] Comments: [In a clear and concise sentence, identify at least one dimension that needs improvement, unless a score of 5 was earned]]".

[0131] By assembling the specific query and response content into the above prompt template, a specific prompt can be obtained. By inputting the prompt into the first AI generation model, the first AI generation model can be guided to complete self-assessment. In an optional manner, improvement suggestions on the response content can also be given.

[0132] S204: If the evaluation score does not meet the conditions, the first AI generation model is called to rewrite the original input content, and the response content generation and self-evaluation are re-executed according to the rewritten input content, until the evaluation score of the response content corresponding to the rewritten input content meets the conditions or the number of cycles reaches a threshold.

[0133] After evaluating the response content, if it is found that the evaluation score does not meet the conditions, for example, the evaluation score is lower than a certain threshold (for example, lower than 5 points), the first AI generation model can also be called to rewrite the original input content. That is to say, in the embodiment of the present application, the generation, evaluation and rewriting of the response content can all be completed by the same AI generation model.

[0134] When rewriting a query, the original query, response content, evaluation score, etc. can be assembled into a prompt template to generate a specific prompt. If there are improvement suggestions, the improvement suggestions can also be reflected in the prompt. For example, the specific prompt template can be:

[0135] "enter:

[0136] Original query: {}

[0137] Evaluation score:{}

[0138] Improvement suggestions:{}

[0139] The output format starts with 'Rewritten query:', as follows:

[0140] Rewritten query: [Using the evaluation score as a guide, carefully refine the original query to better match the original intent. Even if it differs from the improvement suggestion, make appropriate adjustments to ensure clarity and conciseness. The revised query should not be 50% or more longer than the original query]".

[0141] After the above rewriting, the generation of response content and self-evaluation can be re-executed according to the rewritten query. If the evaluation score still does not meet the conditions, the rewriting can be continued, and the response content and self-evaluation can be regenerated, and this cycle is repeated until the evaluation score of the response content corresponding to the rewritten input content meets the conditions. Of course, in practical applications, in order to balance the generation quality and computational efficiency, the maximum number of loops can also be set. If after reaching the maximum number of loops (for example, 3 times), even if the evaluation score still does not meet the conditions, the loop can be ended and the corresponding sample can be discarded. Alternatively, if after the last loop, the evaluation score of the corresponding response content barely meets the standard (for example, the threshold is 5 points, the actual evaluation score is 4 points, etc.), it can also form a specific training sample.

[0142] In addition, in an optional manner, after each query is rewritten, the first AI generated model can also evaluate the rewritten query. For example, the following dimensions can be used to evaluate whether the rewritten query meets the requirements:

[0143] - Targetedness: Is the rewritten query clearer and more specific, and can it resolve the ambiguity or omissions in the original prompt?

[0144] - Comprehensiveness: Whether the rewritten query retains all the key information of the original prompt without missing any important content.

[0145] If the optimized prompts cannot meet the above requirements, the first AI generation model can rewrite the query until a satisfactory result is obtained, and then proceed to the next round of response content generation, self-evaluation and other steps.

[0146] This extra layer of evaluation helps to achieve better generation quality.

[0147] S205: Determine the input content belonging to positive samples and negative samples based on multiple versions of input content obtained during the input content rewriting process and the evaluation scores of the corresponding response content, and construct triple data to be used as model training data for the second AI generation model.

[0148] In the processes of response content generation, self-evaluation, query rewriting, etc. in the aforementioned steps, a large number of rewritten queries will be generated, including queries that are ultimately evaluated as meeting the conditions, and some intermediate versions of queries that may be evaluated as not meeting the conditions (for example, an original query has been rewritten twice, the query after the first rewrite belongs to the intermediate version, and the corresponding response content evaluation score is not higher than the threshold, and the response content evaluation score corresponding to the query after the second rewrite is higher than the threshold). Therefore, positive and negative samples can be determined from these queries. For example, for the aforementioned rewritten query, if the corresponding generated response content is evaluated as meeting the conditions, then the query can be used as a positive sample query, and some intermediate versions of the query, because their corresponding generated response content is evaluated as not meeting the conditions, can be used as negative sample queries, and so on. In this way, a triple consisting of the original query, the positive sample query, and the negative sample query can be constructed, and this triple can be used as training data for the second AI generation model.

[0149] For example, suppose in a triple, the negative sample query is:

[0150] “Can you write a poem about a sad love story?”.

[0151] A positive sample query can be:

[0152] “Can you write a poem about a sad love story that has depth and originality, that experiments with different forms and styles, that delves deeply into the characters’ emotions and experiences, while also incorporating a clear story structure and specific details that make the story more engaging?”

[0153] Or, in another example, the negative query is:

[0154] "Analyzing the popularity of the term 'artificial intelligence' over the past five years".

[0155] A positive sample query can be:

[0156] “How has the popularity of the term ‘artificial intelligence’ changed over the past five years, and what factors have contributed to this trend? Can you provide specific examples of the growing popularity of AI technology and discuss the potential positive impacts it may have?”

[0157] The training data triplet obtained above can be used to train the second AI generation model. The specific training method can be a reinforcement learning method. In this way, the model can acquire the ability to judge "who is better than who", only learn preferences, do not need to memorize answers, and the training penalty is relatively small, which is conducive to retaining the ability of the second AI generation model itself.

[0158] In a preferred implementation, the response contents corresponding to the positive and negative samples can also be added to the training data, that is, the specific triplet data can include: the original input content, the positive sample input content and its corresponding response content, the negative sample input content and its corresponding response content. Among them, the response contents corresponding to the positive and negative samples can be generated by the first AI generation model in the aforementioned data collection process. By adding the response contents corresponding to the positive and negative samples to the training data, during the training process, the second AI generation model can not only rewrite the input content, but also generate the corresponding response content, and the goal is to make the response content corresponding to the rewritten input content tend to the response content corresponding to the positive sample input content, and away from the response content corresponding to the negative sample input content.

[0159] After completing the training of the second AI generation model, the second AI generation model can be used to rewrite the input query, and then assembled into the prompt template to generate a specific prompt, and then the prompt is input into the inference model to generate the response content.

[0160] For example, suppose the original query entered by the user is:

[0161] "Can I make marshmallows at home? My kids would be really excited if I could make them at home."

[0162] The query can be rewritten using the second AI generation model as follows:

[0163] “I’m interested in making marshmallows at home but I’m not sure if this is possible or safe for children. Can you provide more information on how to make marshmallows at home and tell me what safety precautions to take?”.

[0164] It can be seen that the valid information in the original query input by the user is "Can I make marshmallows at home?" and "Is there a child at home?" In fact, what the user needs to care about is the feasibility of making marshmallows at home, whether it is safe if there are children at home, etc. However, the user did not describe this information very specifically. However, the trained second AI generation model can guess the content that the user failed to express directly from the user's original query. Therefore, the rewritten query will describe the problem in more detail and provide more guidance on the angle from which the model should answer the user's question. Correspondingly, it is helpful to help the inference model output higher quality response content.

[0165] The second AI generation model trained above can cooperate with the reasoning models in multiple scenarios, rewrite the input content before specific reasoning, and then construct a prompt to enable the reasoning model to generate better reasoning results.

[0166] For example, one of the application scenarios can be the aforementioned auxiliary reception scenario, in which a human customer service representative needs to respond after the consumer user makes an online consultation. However, before the human customer service representative responds, the inference model needs to give a recommended response to the question input by the user for the human customer service representative to refer to. In this scenario, before the above-mentioned inference model gives the recommended response, the second AI generation model can first rewrite the question input by the user, generate a specific prompt after the rewriting is completed, and then the inference model gives the recommended response.

[0167] Alternatively, another application scenario may be the aforementioned AI polishing, that is, after the manual customer service gives a reply to the user's question, the reply content can be optimized. For example, in a cross-border scenario, the manual customer service may receive foreign users and need to use the corresponding language to reply. At this time, the reply content given by the manual customer service can be polished (that is, optimized) through the reasoning model to correct some grammatical errors, spelling errors, etc. In this scenario, before the reasoning model optimizes the reply content, the second AI model can first rewrite the reply content, and generate a prompt based on the rewritten reply content, and then the reasoning model generates the optimized reply content.

[0168] In addition, it can also be used in scenarios such as intelligent customer service or ordinary intelligent dialogue. That is, it is necessary to communicate with users directly through customer service robots, etc. At this time, it is also necessary to generate reply content for user questions through reasoning models. Before generating the reply content, the user's question content can be rewritten first, and so on.

[0169] In short, whether it is a reasoning model that performs fixed tasks or a reasoning model that supports free dialogue, etc., the solution provided by the embodiment of the present application can be used to optimize the input content before reasoning, thereby optimizing the prompt, so as to obtain better reasoning results after the input model. In this way, only one increase in reasoning cost (compared to the prompt project of aligned input, the cost is significantly reduced) can achieve better results (compared to aligned output).

[0170] Through the embodiments of the present application, when constructing training data, the first AI generation model can be used to generate corresponding response content for the original input content, and then the first AI generation model can be used to self-evaluate the response content generated by itself. If the evaluation result does not meet the preset conditions, the original input content can be rewritten by the first AI generation model, and then the response content generation and self-evaluation are re-executed based on the rewritten input content, and this cycle is repeated until the evaluation score of the response content corresponding to the rewritten input content meets the conditions or the number of cycles reaches the threshold. Afterwards, based on the multiple versions of input content obtained during the input content rewriting process and the corresponding evaluation scores of the response content, the input content belonging to the positive sample and the negative sample is determined, and triple data is constructed for training the second AI generation model so that the second AI generation model acquires the ability to rewrite the input content. Through this solution, in the process of obtaining training data, only the same AI generation model is needed to complete the generation of response content, self-evaluation and rewriting of input content to construct training data triplets. There is no need to call a third-party AI generation model or manual labeling. Therefore, it is beneficial to save resources and costs. The construction of positive and negative samples will not be affected by the different standards of labelers, making the model training results more stable and improving model performance.

[0171] In other words, the embodiment of the present application utilizes the knowledge and judgment of the first AI generation model itself to generate high-quality training data through self-evaluation, avoiding the high cost and bias problems of manual labeling and calling third-party AI generation models. Moreover, through the self-evaluation and feedback mechanism, the model can continuously optimize prompts, and only a small number of iterations are required to generate rich, high-quality preference data, which is conducive to improving the efficiency and quality of training data collection. In addition, with sophisticated process design, the solution has high computational efficiency in both training and reasoning stages.

[0172] In a preferred manner, the first AI generation model and the second AI generation model can be from the same series and / or have the same / similar generation preferences. In this way, through the self-evaluation and feedback mechanism of the first AI generation model, high-quality data consistent with the internal logic of the second AI generation model can be generated, thereby adapting to the needs of different second AI generation models and ensuring the effectiveness and pertinence of the optimization prompts.

[0173] In another preferred manner, the response contents corresponding to the positive and negative samples can also be added to the training data. In this way, during the training process, the second generation model needs to consider not only what the modified query looks like, but also what kind of response the modified query will generate. In addition, the training goal can be "making the response content corresponding to the rewritten input content tend to the response content corresponding to the positive sample input content, and away from the response content corresponding to the negative sample input content", so as to obtain a better training effect. In this way, in the mapping process from the original query to the rewritten query, the response that may be triggered by the rewritten query can be fully considered, and a more comprehensive and targeted optimization can be performed.

[0174] In other words, in the process of training the second AI generation model using training data, implicit reasoning can be incorporated based on the training method of reinforcement learning, so that the model can fully consider the actual effect when rewriting the input content, thereby improving the effectiveness and pertinence of the input content rewriting. In addition, thanks to the training method based on reinforcement learning and the paradigm of injecting implicit reasoning, the prompt optimizer trained in the solution provided in the embodiment of the present application shows good generalization ability. On data sets in multiple different fields, the prompt optimizer can generate high-quality optimization prompts that adapt to the characteristics of the field, helping the model to produce more accurate and professional responses.

[0175] In addition, the trained second AI generation model can be used as a query rewriter (or prompt optimizer, etc.), which can be applied to any original input content to automatically generate optimized input content to make it more in line with human preferences. Moreover, the query rewriter can quickly adapt to new tasks, generate high-quality, expected responses, and show good adaptability and flexibility. It provides a general and efficient framework for handling diverse tasks, while improving the ability to follow instructions in the field while taking into account general performance. This pluggable prompt optimizer only requires one additional reasoning during reasoning. Specifically, the original input content is input into the prompt optimizer to generate an optimized prompt, which is then sent to the reasoning model to produce the final response. Compared with the traditional prompt word engineering method, this method significantly reduces the need for multiple calls to large models, thereby reducing the cost of reasoning.

[0176] Embodiment 2

[0177] This second embodiment is an application of the prompt optimizer obtained through specific training in an auxiliary reception scenario or an intelligent dialogue scenario, and provides a method for assisting in generating dialogue reply content. Figure 3 , the method may include:

[0178] S301: Determine the current question content of the user in the target session;

[0179] S302: rewrite the current question content by using a second AI generation model to generate a rewritten question content; wherein the second AI generation model is trained in the following manner:

[0180] Get multiple raw input contents;

[0181] Calling a first artificial intelligence AI generation model according to the original input content to generate response content corresponding to the original input content;

[0182] Calling the first AI generation model to perform self-evaluation on the response content according to the original input content, the response content, and a preset first evaluation rule;

[0183] If the evaluation score does not meet the conditions, the first AI generation model is called to rewrite the original input content, and the response content generation and self-evaluation are re-executed according to the rewritten input content, until the evaluation score of the response content corresponding to the rewritten input content meets the conditions or the number of cycles reaches the threshold;

[0184] Determine the input content belonging to the positive sample and the negative sample according to the multiple versions of the input content obtained in the input content rewriting process and the evaluation scores of the corresponding response content, and construct triple data according to the original input content, the positive sample input content, and the negative sample input content for training the second AI generation model;

[0185] S303: Generate prompt information based on the question content rewritten by the second AI generation model, and input it into the reasoning model to generate reply content or generate reply content for recommendation.

[0186] Embodiment 3

[0187] This embodiment 3 is an application of the prompt optimizer obtained through specific training in the "AI polishing" scenario, and provides a method for optimizing the content of dialogue replies. Figure 4 , the method may include:

[0188] S401: Determine the answer content to the user's question content in the target session;

[0189] S402: rewrite the reply content by using a second AI generation model to generate rewritten reply content; wherein the second AI generation model is trained in the following manner:

[0190] Get multiple raw input contents;

[0191] Calling a first artificial intelligence AI generation model according to the original input content to generate response content corresponding to the original input content;

[0192] Calling the first AI generation model to perform self-evaluation on the response content according to the original input content, the response content, and a preset first evaluation rule;

[0193] If the evaluation score does not meet the conditions, the first AI generation model is called to rewrite the original input content, and the response content generation and self-evaluation are re-executed according to the rewritten input content, until the evaluation score of the response content corresponding to the rewritten input content meets the conditions or the number of cycles reaches the threshold;

[0194] Determine the input content belonging to the positive sample and the negative sample according to the multiple versions of the input content obtained in the input content rewriting process and the evaluation scores of the corresponding response content, and construct triple data according to the original input content, the positive sample input content, and the negative sample input content for training the second AI generation model;

[0195] S404: Generate prompt information based on the reply content rewritten by the second AI generation model, and input it into the reasoning model to optimize the reply content.

[0196] For the parts not described in detail in the above-mentioned embodiments 2 and 3, please refer to the records in the above-mentioned embodiment 1 or other parts of this specification, and will not be repeated here.

[0197] It should be noted that the embodiments of the present application may involve the use of user data. In actual applications, user-specific personal data can be used in the scheme described herein within the scope permitted by applicable laws and regulations, subject to the requirements of applicable laws and regulations of the country where the user is located (for example, with the user's explicit consent, effective notification to the user, etc.).

[0198] Corresponding to the first embodiment, the embodiment of the present application further provides a device for generating model training data, which may include:

[0199] An original input content acquisition unit, used to acquire a plurality of original input contents;

[0200] A response content generating unit, configured to call a first artificial intelligence AI generating model according to the original input content to generate response content corresponding to the original input content;

[0201] a self-evaluation unit, configured to call the first AI generation model to perform a self-evaluation on the response content according to the original input content, the response content, and a preset first evaluation rule;

[0202] A rewriting unit, configured to call the first AI generation model to rewrite the original input content if the evaluation score does not meet the condition, and re-execute the generation of response content and self-evaluation according to the rewritten input content until the evaluation score of the response content corresponding to the rewritten input content meets the condition or the number of cycles reaches a threshold;

[0203] A training data construction unit is used to determine the input content belonging to positive samples and negative samples based on multiple versions of input content obtained during the input content rewriting process and the evaluation scores of the corresponding response content, and to construct triple data based on the original input content, positive sample input content, and negative sample input content to be used as model training data for the second AI generation model.

[0204] The first AI generation model and the second AI generation model are models of the same series and / or have the same / similar content generation preferences.

[0205] Specifically, the performance of the first AI generated model is better than that of the second AI generated model.

[0206] Among them, the triplet data includes the original input content, the positive sample input content and its corresponding response content, the negative sample input content and its corresponding response content, so that during the training process, the second AI generation model rewrites the input content and generates corresponding response content, and makes the response content corresponding to the rewritten input content tend to the response content corresponding to the positive sample input content, and away from the response content corresponding to the negative sample input content.

[0207] In a specific implementation, the device may further include:

[0208] An improvement suggestion providing unit is used to analyze defects in the response content through the first AI generation model and provide improvement suggestions for the response content when the response content is evaluated through the first AI generation model and the evaluation score does not meet the conditions, so that the first AI generation model can rewrite the input content according to the improvement suggestions.

[0209] Furthermore, the device may further include:

[0210] An input content evaluation unit is used to call the first AI generation model to evaluate the changes of the rewritten input content relative to the original input content based on the original input content and the rewritten input content. If the evaluation fails, it triggers the rewriting of the input content.

[0211] Corresponding to the second embodiment, the embodiment of the present application further provides a device for generating dialogue reply content, which may include:

[0212] A question content determination unit, used to determine the current question content of the user in the target session;

[0213] A rewriting unit, configured to rewrite the current question content by using a second AI generation model to generate a rewritten question content; wherein the second AI generation model is trained in the following manner:

[0214] Get multiple raw input contents;

[0215] Calling a first artificial intelligence AI generation model according to the original input content to generate response content corresponding to the original input content;

[0216] Calling the first AI generation model to perform self-evaluation on the response content according to the original input content, the response content, and a preset first evaluation rule;

[0217] If the evaluation score does not meet the conditions, the first AI generation model is called to rewrite the original input content, and the response content generation and self-evaluation are re-executed according to the rewritten input content, until the evaluation score of the response content corresponding to the rewritten input content meets the conditions or the number of cycles reaches the threshold;

[0218] Determine the input content belonging to the positive sample and the negative sample according to the multiple versions of the input content obtained in the input content rewriting process and the evaluation scores of the corresponding response content, and construct triple data according to the original input content, the positive sample input content, and the negative sample input content for training the second AI generation model;

[0219] A reply content generation or recommendation unit is used to generate prompt information based on the question content rewritten by the second AI generation model, and input it into the reasoning model to generate reply content or generate reply content for recommendation.

[0220] Corresponding to the third embodiment, the embodiment of the present application further provides a device for optimizing the content of a conversation reply, which may include:

[0221] A reply content determination unit, used to determine the reply content to the user's question content in the target session;

[0222] A rewriting unit, configured to rewrite the reply content by using a second AI generation model to generate rewritten reply content; wherein the second AI generation model is trained in the following manner:

[0223] Get multiple raw input contents;

[0224] Calling a first artificial intelligence AI generation model according to the original input content to generate response content corresponding to the original input content;

[0225] Calling the first AI generation model to perform self-evaluation on the response content according to the original input content, the response content, and a preset first evaluation rule;

[0226] If the evaluation score does not meet the conditions, the first AI generation model is called to rewrite the original input content, and the response content generation and self-evaluation are re-executed according to the rewritten input content, until the evaluation score of the response content corresponding to the rewritten input content meets the conditions or the number of cycles reaches the threshold;

[0227] Determine the input content belonging to the positive sample and the negative sample according to the multiple versions of the input content obtained in the input content rewriting process and the evaluation scores of the corresponding response content, and construct triple data according to the original input content, the positive sample input content, and the negative sample input content for training the second AI generation model;

[0228] The reply content optimization unit is used to generate prompt information based on the reply content rewritten by the second AI generation model, and input it into the reasoning model to optimize the reply content.

[0229] In addition, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the steps of any one of the methods in the aforementioned method embodiments are implemented.

[0230] And an electronic device, comprising:

[0231] one or more processors; and

[0232] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of the aforementioned method embodiments.

[0233] A computer program product includes a computer program / computer executable instructions, which, when executed by a processor in an electronic device, implement the steps of the method described in the aforementioned method embodiment.

[0234] in, Figure 5 The architecture of the electronic device is shown as an example, which may include a processor 510, a video display adapter 511, a disk drive 512, an input / output interface 513, a network interface 514, and a memory 520. The processor 510, the video display adapter 511, the disk drive 512, the input / output interface 513, the network interface 514, and the memory 520 may be communicatively connected via a communication bus 530.

[0235] Among them, the processor 510 can be implemented by a general-purpose CPU (Central Processing Unit, processor), a microprocessor, an application specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, etc., to execute relevant programs to implement the technical solution provided in this application.

[0236] The memory 520 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 520 can store an operating system 521 for controlling the operation of the electronic device 500, and a basic input and output system (BIOS) for controlling the low-level operation of the electronic device 500. In addition, a web browser 523, a data storage management system 524, and a model training data processing system 525, etc. can also be stored. The above-mentioned model training data processing system 525 can be an application program that specifically implements the operations of the aforementioned steps in the embodiment of the present application. In short, when the technical solution provided in the present application is implemented by software or firmware, the relevant program code is stored in the memory 520 and is called and executed by the processor 510.

[0237] The input / output interface 513 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0238] The network interface 514 is used to connect to a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).

[0239] The bus 530 comprises a pathway for transmitting information between the various components of the device (eg, the processor 510, the video display adapter 511, the disk drive 512, the input / output interface 513, the network interface 514, and the memory 520).

[0240] It should be noted that, although the above device only shows a processor 510, a video display adapter 511, a disk drive 512, an input / output interface 513, a network interface 514, a memory 520, a bus 530, etc., in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include components necessary for implementing the solution of the present application, and does not necessarily include all the components shown in the figure.

[0241] It can be known from the description of the above implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be essentially or partly contributed to the prior art in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application or certain parts of the embodiments.

[0242] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.

[0243] The above is a detailed introduction to the method for generating model training data and the electronic device provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. At the same time, for those skilled in the art, according to the ideas of the present application, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present application.

Claims

1. A method for generating model training data, characterized in that: include: Get multiple raw input contents; Calling a first artificial intelligence AI generation model according to the original input content to generate response content corresponding to the original input content; Calling the first AI generation model to perform self-evaluation on the response content according to the original input content, the response content, and a preset first evaluation rule; If the evaluation score does not meet the conditions, the first AI generation model is called to rewrite the original input content, and the response content generation and self-evaluation are re-executed according to the rewritten input content, until the evaluation score of the response content corresponding to the rewritten input content meets the conditions or the number of cycles reaches the threshold; Based on the multiple versions of input content obtained during the input content rewriting process and the evaluation scores of the corresponding response content, the input content belonging to the positive sample and the negative sample is determined, and triple data is constructed based on the original input content, the positive sample input content, and the negative sample input content to be used as model training data for the second AI generation model.

2. The method according to claim 1, characterized in that The first AI generation model and the second AI generation model are models of the same series and / or have the same / similar content generation preferences.

3. The method according to claim 1, characterized in that The performance of the first AI generated model is better than that of the second AI generated model.

4. The method according to claim 1, characterized in that: The triplet data includes the original input content, the positive sample input content and its corresponding response content, the negative sample input content and its corresponding response content, so that during the training process, the second AI generation model rewrites the input content and generates corresponding response content, and makes the response content corresponding to the rewritten input content tend to the response content corresponding to the positive sample input content, and away from the response content corresponding to the negative sample input content.

5. The method according to claim 1, characterized in that Also includes: When the response content is evaluated by the first AI generation model and the evaluation score does not meet the requirements, the first AI generation model is used to analyze the defects in the response content and provide improvement suggestions for the response content, so that the first AI generation model can rewrite the input content according to the improvement suggestions.

6. The method according to claim 1, characterized in that Also includes: Based on the original input content and the rewritten input content, the first AI generation model is called to evaluate the changes in the rewritten input content relative to the original input content. If the evaluation fails, the rewriting of the input content is triggered.

7. A method for generating dialogue reply content, characterized in that: include: Determine the current question asked by the user in the target session; The current question content is rewritten by a second AI generation model to generate a rewritten question content; wherein the second AI generation model is trained in the following manner: Get multiple raw input contents; Calling a first artificial intelligence AI generation model according to the original input content to generate response content corresponding to the original input content; Calling the first AI generation model to perform self-evaluation on the response content according to the original input content, the response content, and a preset first evaluation rule; If the evaluation score does not meet the conditions, the first AI generation model is called to rewrite the original input content, and the response content generation and self-evaluation are re-executed according to the rewritten input content, until the evaluation score of the response content corresponding to the rewritten input content meets the conditions or the number of cycles reaches the threshold; Based on multiple versions of input content obtained during the input content rewriting process and the evaluation scores of the corresponding response content, determine the input content belonging to positive samples and negative samples, and construct triple data based on the original input content, positive sample input content, and negative sample input content for training the second AI generation model; generate prompt information based on the question content rewritten by the second AI generation model, and input it into the inference model to generate reply content or generate reply content for recommendation.

8. A method for optimizing dialogue reply content, characterized in that: include: Determine the answer content to the user's question in the target session; The reply content is rewritten by a second AI generation model to generate a rewritten reply content; wherein the second AI generation model is trained in the following manner: Get multiple raw input contents; Calling a first artificial intelligence AI generation model according to the original input content to generate response content corresponding to the original input content; Calling the first AI generation model to perform self-evaluation on the response content according to the original input content, the response content, and a preset first evaluation rule; If the evaluation score does not meet the conditions, the first AI generation model is called to rewrite the original input content, and the response content generation and self-evaluation are re-executed according to the rewritten input content, until the evaluation score of the response content corresponding to the rewritten input content meets the conditions or the number of cycles reaches the threshold; Based on the multiple versions of input content obtained during the input content rewriting process and the evaluation scores of the corresponding response content, the input content belonging to the positive sample and the negative sample is determined, and triple data is constructed based on the original input content, the positive sample input content, and the negative sample input content for training the second AI generation model; prompt information is generated based on the reply content rewritten by the second AI generation model, and input into the inference model for optimizing the reply content.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 8 are implemented.

10. An electronic device, characterized in that: include: one or more processors; as well as A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of claims 1 to 8.

11. A computer program product comprising a computer program / computer executable instructions, characterized in that: When the computer program / computer executable instructions are executed by a processor in an electronic device, the steps of the method according to any one of claims 1 to 8 are implemented.