Task processing model training method, role playing model training method and task processing method

By comparing and analyzing the response indicators of multiple sample responses, the instability problem of traditional single-sample evaluation is solved, and the efficiency and accuracy of task processing model training are improved, especially in open tasks.

CN120705532AActive Publication Date: 2025-09-26ZHEJIANG ALIBABA ROBOT CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511215565.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-09-26
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

Traditional task processing model training methods rely on single-sample evaluation, which leads to unstable evaluation results, affecting training efficiency and accuracy, especially making it difficult to achieve effective reinforcement learning in open tasks.

Method used

We obtain multiple sample reply contents based on sample conversation data, obtain reply indicators through comparative analysis, use reply indicators to train the task processing model, introduce group strategy optimization and content comparison models for quality assessment, and improve the stability and accuracy of the assessment.

Benefits of technology

Through comparative evaluation, the ambiguity of evaluation criteria is reduced, the accuracy and stability of response indicators are improved, the error amplification problem is reduced, and the efficiency and accuracy of task processing model training are significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705532A_ABST
    Figure CN120705532A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a task processing model training method, a role playing model training method and a task processing method, and the task processing model training method comprises the steps: obtaining a plurality of sample reply contents based on sample dialogue data through a task processing model; the multiple sample reply contents are compared and analyzed, reply indexes corresponding to the multiple sample reply contents are obtained, and the reply indexes are used for measuring the quality of the corresponding sample reply contents; and training the task processing model according to the reply index to obtain a trained task processing model. The reply indexes with higher distinction degree and stability are generated by simultaneously generating and comparing the reply contents of the multiple samples, so that judgment deviation caused by standard fuzziness during single-sample scoring is avoided, a more objective and accurate basis is provided for the training process, and the efficiency and accuracy of the task processing model training process are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of artificial intelligence technology, and in particular to a task processing model training method, a role-playing model training method, and a task processing method. Background Art

[0002] With the continuous advancement of artificial intelligence technology, reinforcement learning, as an important method for enabling autonomous decision-making by intelligent agents, has been widely used in natural language processing, dialogue systems, text generation, and other fields. In reinforcement learning, quality assessment of text generated by task processing models can provide learning signals for the training process of task processing models.

[0003] Currently, traditional evaluation modeling methods usually rely on large models to evaluate single samples, resulting in unstable evaluation results and seriously affecting the efficiency and accuracy of task processing model training. Summary of the Invention

[0004] In view of this, embodiments of this specification provide a task processing model training method. One or more embodiments of this specification also relate to a role-playing model training method, a task processing method, a request processing method based on a task processing model, a task platform, a task processing model training device, a role-playing model training device, a task processing device, a request processing device based on a task processing model, a computing device, an electronic device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.

[0005] According to a first aspect of an embodiment of this specification, a task processing model training method is provided, comprising: Using the task processing model, we obtain multiple sample responses based on the sample conversation data. Comparative analysis is performed on multiple sample response contents to obtain response indicators corresponding to the multiple sample response contents, wherein the response indicators are used to measure the quality of the corresponding sample response contents; The task processing model is trained according to the response indicator to obtain a trained task processing model.

[0006] One embodiment of this specification provides a task processing model training method, comprising: utilizing a task processing model to obtain multiple sample response contents based on sample conversation data; performing comparative analysis on the multiple sample response contents to obtain response indicators corresponding to the multiple sample response contents, wherein the response indicators are used to measure the quality of the corresponding sample response contents; and training the task processing model based on the response indicators to obtain a trained task processing model. By simultaneously generating and comparing multiple sample response contents, the differences in quality between responses can be accurately identified in relative comparison, thereby generating a more discriminatory and stable response indicator. This avoids judgment bias caused by fuzzy standards when scoring single samples, provides a more objective and accurate basis for the training process, and significantly improves the efficiency and accuracy of the task processing model training process. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 This is a flowchart of a task processing model training method provided by one embodiment of this specification; Figure 2 This is a schematic diagram of a processing process of a task processing model training method provided by an embodiment of this specification; Figure 3 This is a flow chart of a role-playing model training method provided by one embodiment of this specification; Figure 4 This is a flowchart of a task processing method provided by one embodiment of this specification; Figure 5 This is an architectural diagram of a task processing system provided by one embodiment of this specification; Figure 6 This is a flowchart of a request processing method based on a task processing model provided by an embodiment of this specification; Figure 7 This is a schematic diagram of the structure of a task platform provided by one embodiment of this specification; Figure 8 This is a structural diagram of a task processing model training device provided by one embodiment of this specification; Figure 9 This is a structural diagram of a role-playing model training device provided by one embodiment of this specification; Figure 10 This is a structural diagram of a task processing device provided by one embodiment of this specification; Figure 11 This is a schematic diagram of the structure of a request processing device based on a task processing model provided by an embodiment of this specification; Figure 12 This is a structural block diagram of a computing device provided by one embodiment of this specification; Figure 13This is a structural block diagram of an electronic device provided by an embodiment of this specification. DETAILED DESCRIPTION

[0008] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0009] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items. The term "at least one" in one or more embodiments of this application refers to "one or more" and "a plurality" refers to "two or more". The term "including" is an open description and should be understood as "including but not limited to", and may include other content on the basis of the content already described.

[0010] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0011] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0012] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a foundation model. It is pre-trained on a large amount of unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks and has good generalization capabilities, such as a large language model (LLM) and a multi-modal pre-training model.

[0013] When large models are used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. Large models can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image description (IC, Image Caption), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.

[0014] First, the terms involved in one or more embodiments of this specification are explained.

[0015] Supervised fine-tuning (SFT) is a method for further training a pre-trained model. In this method, the model is trained on a dataset containing input and desired output pairs, allowing it to learn how to generate responses closer to human-level performance. Supervised fine-tuning is often used to adapt a model to data specific to a task or domain, thereby improving its performance on those tasks.

[0016] Reinforcement learning (RL) is a machine learning paradigm whose core idea is to enable an agent to learn how to take optimal actions in specific situations through continuous interaction with its environment, thereby maximizing a cumulative reward signal. Utilizing a reward model, reinforcement learning can often fine-tune the LLM based on an evaluation model (also called a reward model) during LLM training, making the text generated by the LLM more consistent with human preferences.

[0017] Group Relatively Policy Optimization (GRPO): is a reinforcement learning algorithm that uses the normalized rewards of multiple responses to calculate the advantage of each response and then update the model parameters.

[0018] Proximal Policy Optimization (PPO): is a reinforcement learning algorithm that performs policy optimization through advantage function estimation and KL divergence constraints.

[0019] KL Divergence (Kullback-Leibler Divergence): Used to measure the difference between two probability distributions. Specifically, KL divergence measures the amount of information lost when using one probability distribution to approximate another probability distribution.

[0020] Reward Hacking: A model may manipulate the evaluation model or scoring mechanism by generating lengthy, repetitive, irrelevant, but seemingly "reasonable" content to inflate its evaluation score, rather than truly improving the output quality.

[0021] The clip function is a common function in programming, mathematics, and deep learning. It is used to limit values ​​to a specified range. If a number exceeds the set maximum or minimum value, it will be "clipped" to the boundary value.

[0022] Comparative Policy Optimization (CPO) is a method for improving policies within reinforcement learning. It optimizes the decision-making process by comparing the performance of multiple candidate actions or policies. This method is particularly suitable for tasks that require selecting the best option from a set of possible actions, such as multi-turn dialogue management in dialogue systems and item recommendations in recommender systems.

[0023] The LLM optimization process typically consists of two phases: the first phase involves performing SFT on role-playing dialogue corpus; the second phase involves reinforcement learning using an evaluation model. For open-ended tasks such as role-playing, for example, due to unclear evaluation criteria, constructing an evaluation model is challenging, making effective reinforcement learning difficult. Furthermore, traditional evaluation model schemes typically calculate a score for each model response individually. This approach often faces the following issues in open-ended tasks, hindering effective reinforcement learning training: Ambiguous evaluation criteria: It is difficult to implement a clear and consistent scoring rule for responses in open-ended tasks. Scoring instability: Evaluators based on single-shot scoring are sensitive to cue fluctuations, often generating unstable and indiscriminate scores. In certain circumstances, scoring collapse can occur, with most outputs falling within a narrow range. Error amplification: For reinforcement learning algorithms that use group normalization (such as the GRPO algorithm), rewards are normalized during optimization. When the model generates content containing factual errors or logical contradictions, single-shot scoring may result in higher scores due to local fluency, amplifying the error and misleading the policy learning process.

[0024] In order to solve the problems of reward ambiguity and scoring instability in open tasks such as role-playing, the embodiments of this specification propose a reinforcement learning training scheme based on CPO. Specifically, a task processing model is used to obtain multiple sample reply contents based on sample dialogue data; a comparative analysis is performed on the multiple sample reply contents to obtain reply indicators corresponding to the multiple sample reply contents, wherein the reply indicators are used to measure the quality of the corresponding sample reply contents; according to the reply indicators, the task processing model is trained to obtain a trained task processing model. This scheme transforms the traditional single-sample evaluation into a comparative sample evaluation. Through comparative evaluation, the ambiguity of the evaluation criteria can be reduced, the response quality can be evaluated more accurately, and the accuracy of the reply indicator can be improved. At the same time, the fluctuation of the reply indicator can be reduced, making the reply indicator more discriminative and stable, and reducing the error amplification problem caused by single-sample evaluation.

[0025] In this specification, a task processing model training method is provided. This specification also involves a role-playing model training method, a task processing method, a request processing method based on a task processing model, a task platform, a task processing model training device, a role-playing model training device, a task processing device, a request processing device based on a task processing model, a computing device, an electronic device, a computer-readable storage medium and a computer program product, which are described in detail one by one in the following embodiments.

[0026] See also Figure 1 , Figure 1A flowchart of a task processing model training method provided by an embodiment of this specification is shown, which specifically includes the following steps: Step 102: Utilize the task processing model to obtain multiple sample reply contents based on the sample conversation data.

[0027] It's important to note that a task processing model refers to a machine learning model used to complete specific tasks (such as role-playing, generating dialogue responses, and question-answering). A task processing model can be a general-purpose pre-trained large model or a domain-specific large model fine-tuned for a specific task. A task processing model receives input (such as user conversation history or questions) and generates corresponding output (such as a response). During training, the model parameters of the task processing model can be adjusted through methods such as reinforcement learning to improve task performance.

[0028] Sample conversation data refers to a set of real or constructed conversation data used to train task processing models. Sample conversation data can be question-and-answer data with a conversation goal or chat data without a conversation goal. Sample conversation data includes user input (such as questions and instructions), as well as character settings and conversation context. Sample conversation data can be in various modalities, such as text, voice, video, and images. This data is provided as input to the task processing model to generate corresponding sample responses. Multiple sample conversation data sets can be used. During each training iteration of the task processing model, a batch of data can be extracted to update the model parameters. After multiple iterations, reinforcement learning training is completed.

[0029] Multiple sample responses refer to the different responses generated by the task processing model for a sample conversation data point through multiple sampling (e.g., using different generation strategies or random sampling). These responses may differ in expression, information completeness, or quality, providing diverse comparisons for subsequent comparative analysis.

[0030] In practical applications, there are various ways to utilize a task processing model to obtain multiple sample responses based on sample conversation data. The specific method to be used depends on the specific circumstances and is not limited in this specification. In one possible implementation, sample conversation data sent by a user via a client can be received and input into the task processing model to obtain multiple sample responses. In another possible implementation, multiple sample responses, pre-processed by the task processing model from the sample conversation data, can be retrieved from another data acquisition device or database.

[0031] Step 104: Compare and analyze the multiple sample reply contents to obtain reply indicators corresponding to the multiple sample reply contents, wherein the reply indicators are used to measure the quality of the corresponding sample reply contents.

[0032] It should be noted that comparative analysis refers to the process of comparing and evaluating multiple sample responses. This comparative analysis typically identifies the strengths and weaknesses of each sample response based on pre-defined quality assessment dimensions (such as relevance, fluency, factual accuracy, and security), thereby assigning a corresponding response metric to each sample response.

[0033] Response metrics are used to assess the relative quality of a sample response. They reflect the relative quality of a sample response relative to other sample responses and can serve as reward signals for subsequent task processing model training. Response metrics can be a single-dimensional evaluation (such as fluency) or a weighted composite of multiple dimensions. Response metrics are typically generated based on comparative analysis strategies or content comparison models.

[0034] In actual applications, there are multiple ways to compare and analyze multiple sample response contents and obtain response indicators corresponding to the multiple sample response contents. The specific methods to be selected depend on the actual situation. In one possible implementation of this specification, for any sample response content, the sample response content can be compared with all sample response contents other than the sample response content. For example, assuming there are three sample response contents, namely sample response content A, sample response content B, and sample response content C. For sample response content A, sample response content A is compared with sample response content B and sample response content C to obtain the response indicator of the sample response content. In another possible implementation of this specification, for any sample response content, at least one sample response content can be selected (randomly or based on screening rules such as similarity) from sample response contents other than the sample response content to compare with the sample response content. For example, assuming there are three sample response contents, namely sample response content A, sample response content B, and sample response content C. For sample response content A, select sample response content C from sample response content B and sample response content C, compare sample response content A with sample response content C, and obtain the response index of the sample response content.

[0035] In an optional embodiment of the present specification, the comparative analysis of the plurality of sample reply contents to obtain the reply indicators corresponding to the plurality of sample reply contents may include the following steps: For the first sample reply content, at least one sample comparison content is selected from the second sample reply content, wherein the first sample reply content is any one of the multiple sample reply contents, and the second sample reply content is a sample reply content other than the first sample reply content among the multiple sample reply contents; Compare the first sample reply content with the sample comparison content to obtain the reply index corresponding to the first sample reply content.

[0036] It should be noted that the first sample response refers to a sample response selected arbitrarily from multiple sample responses and serves as the subject of the current evaluation. The second sample response refers to the remaining responses from the multiple sample responses, excluding the first sample response, and serves as a comparative reference for the first sample response.

[0037] Comparison sample content refers to one or more sample responses selected from the second sample response content for quality comparison with the first sample response content. Comparison sample content can be selected randomly or based on topic, semantic, or structural similarity to select comparable sample responses.

[0038] In actual applications, there are multiple ways to compare the first sample reply content with the sample comparison content to obtain the reply index corresponding to the first sample reply content. The specific selection is based on the actual situation, and the embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, the first sample reply content and the sample comparison content can be directly compared to obtain the reply index. In another possible implementation of this specification, in order to improve the quality of the reply index, sample conversation data can be introduced, and the correlation between the sample conversation data and the first sample reply content and the sample comparison content can be considered in the comparison process to obtain the reply index.

[0039] Applying the solutions in the embodiments of this specification, sample comparison content is dynamically selected for each sample response and relative quality judgments are made, achieving more refined and stable response metric modeling. Response metrics based on pairwise or intra-group comparisons can effectively reduce subjective bias and enhance discrimination, especially identifying responses that appear reasonable but contain subtle flaws. This guides the task processing model to more accurately optimize generation strategies in reinforcement learning, improving the stability and reliability of output quality.

[0040] In an optional embodiment of the present specification, the comparison of the first sample reply content with the sample comparison content to obtain the reply indicator corresponding to the first sample reply content may include the following steps: Based on the sample conversation data, the first sample reply content and the sample comparison content are compared to obtain a reply indicator corresponding to the first sample reply content.

[0041] In actual applications, there are multiple ways to compare the first sample reply content and the sample comparison content based on the sample conversation data to obtain the reply index corresponding to the first sample reply content. The specific selection is based on the actual situation, and the embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, the first sample reply content and the sample comparison content can be compared based on the content comparison strategy and the sample conversation data to obtain the reply index corresponding to the first sample reply content. In another possible implementation of this specification, the sample conversation data, the first sample reply content and the sample comparison content can be input into the content comparison model, and the content comparison model can compare the first sample reply content and the sample comparison content based on the sample conversation data to obtain the reply index corresponding to the first sample reply content.

[0042] By applying the solution of the embodiments of this specification, the contextual intent is understood by combining the sample conversation data, so that the comparative analysis between the first sample reply content and the sample comparison content can be placed in the same input context for horizontal comparison, making the generation of reply indicators more objective and accurate, and effectively improving the training efficiency of the task processing model and the reliability of the generation results.

[0043] In an optional embodiment of the present specification, the comparison of the first sample reply content with the sample comparison content to obtain the reply indicator corresponding to the first sample reply content may include the following steps: The comparison prompt information, the first sample reply content and the sample comparison content are input into the content comparison model to obtain the reply index corresponding to the first sample reply content, wherein the comparison prompt information is used to guide the content comparison model to compare the first sample reply content and the sample comparison content to obtain the reply index.

[0044] It should be noted that comparison prompt information refers to pre-designed instructional text (Prompt) that is used to clearly guide the content comparison model to perform specific comparison tasks, such as "Please compare the following two responses in terms of factual accuracy and clarity of expression." The comparison prompt information can standardize the comparison analysis standards and ensure the consistency and explainability of the comparison process.

[0045] A content comparison model is a model used to evaluate and compare the quality of different content. It receives comparison prompt information, a first sample response, and sample comparison content as input, and outputs a response indicator for the first sample response. This model can be a pre-trained large model or a discriminant model trained based on training sample responses, training comparison content, and training response indicators.

[0046] Using the solutions in the embodiments of this specification, a content comparison model simultaneously compares and analyzes multiple sample responses, generating response metrics that improve the stability and consistency of these metrics. By introducing structured "comparison prompts" to guide the content comparison model in performing standardized, controllable quality assessments, this effectively focuses on evaluation dimensions, reduces subjective bias, and generates reliable and interpretable response metrics even in complex contexts, improving the controllability of the task processing model training process.

[0047] In an optional embodiment of the present specification, the comparison of the first sample reply content with the sample comparison content to obtain the reply indicator corresponding to the first sample reply content may include the following steps: Based on a content comparison strategy, the first sample reply content and the sample comparison content are compared to obtain a reply index corresponding to the first sample reply content, wherein the content comparison strategy is constructed based on at least one quality assessment dimension.

[0048] It should be noted that the quality assessment dimensions refer to specific standards or aspects for measuring the quality of responses, such as relevance, fluency, factual accuracy, security, semantic accuracy, logical coherence, naturalness of language, whether key information points are included, etc.

[0049] A content comparison strategy refers to a method or system of rules used to compare the content of a first sample response with the sample comparison content. It is typically constructed based on one or more quality assessment dimensions and determines how to quantify or determine the differences between the two. A content comparison strategy can be a natural language comparison strategy constructed based on quality assessment dimensions, or it can be an automated comparison code constructed based on quality assessment dimensions. The specific choice is based on the actual situation and is not limited in this specification.

[0050] By applying the solution of the embodiment of this specification, the first sample reply content and the sample comparison content are compared item by item using a content comparison strategy, thereby achieving a multi-dimensional and systematic quality evaluation of the first sample reply content.

[0051] Step 106: Train the task processing model according to the response indicator to obtain a trained task processing model.

[0052] It should be noted that training refers to the process of updating the internal parameters of the task processing model through optimization algorithms (such as GRPO and PPO) based on the reward signal fed back by the response index. The goal of training is to enable the task processing model to generate higher-quality responses, that is, to obtain content with higher response indexes.

[0053] A trained task processing model is one that has been trained through reinforcement learning based on the response metrics derived from comparative analysis. This trained task processing model possesses enhanced generation capabilities, enabling it to generate responses that better meet expected quality standards based on input conversation data, making it suitable for practical deployment and application.

[0054] By applying the solution of the embodiment of this description, by simultaneously generating and comparing the content of multiple sample replies, the differences in quality between replies can be accurately identified in relative comparison, thereby generating more discriminatory and stable reply indicators, avoiding judgment bias caused by vague standards when scoring single samples, providing a more objective and accurate basis for the training process, and significantly improving the efficiency and accuracy of the task processing model training process.

[0055] In an optional embodiment of the present specification, the training process of the task processing model is described by taking the GRPO algorithm as an example. That is, the task processing model is trained according to the response indicator to obtain the trained task processing model, which may include the following steps: Calculate the policy loss based on the response metric, where the policy loss is used to measure the difference between the model performance of the task processing model and the expected model performance; Calculate the model difference index based on the responses of multiple samples. The model difference index is used to measure the difference between the task processing model and the historical task processing model. The historical task processing model is the model version of the task processing model at the previous time point. The task processing model is trained according to the strategy loss and model difference indicators to obtain a trained task processing model.

[0056] It's important to note that policy loss refers to the loss function calculated based on the response metric during reinforcement learning optimization. Policy loss can be used to measure the gap between the current task processing model's performance (current behavior / policy) and the expected model performance (expected behavior / policy), encouraging the model to take more optimal actions. A larger policy loss indicates that the quality of the task processing model's current output deviates further from the expected target. Policy loss can be calculated using a clipping mechanism.

[0057] Model divergence metrics can be used to measure the difference between the output behavior / strategy or parameter distribution of a task processing model and historical task processing models. A common form of model divergence metrics is the KL divergence, and therefore, the model divergence metric can also be called a KL divergence penalty. This metric can reflect the "step size" or degree of change in model policy adjustments, preventing overly large parameter updates from leading to training instability or catastrophic forgetting.

[0058] In practical applications, the task processing model can be trained according to the policy loss and model difference index using the following formulas (1) to (5) to obtain the trained task processing model: (1) in, Represents sample conversation data; The model parameters are Task processing model; Indicates that the sample conversation data is input into the task processing model, and the obtained Sample responses; Represents the response content of the i-th sample.

[0059] (2) in, Represents a content contrast model; Indicates that Sample response content is input into the content comparison model, and G response indicators are obtained; Indicates the response content of the i-th sample The corresponding i-th response indicator.

[0060] (3) in, represents the advantage value of the tth token in the i-th sample reply content; represents the average value of G response indicators; represents the standard deviation of the G response indicators.

[0061] (4) in, represents the strategy loss; Indicates taking the minimum value of two expressions as ,prevent the task processing model from overestimating or underestimating the probability of certain actions, resulting in unstable training; Indicates a given and the output token sequence In the case of task processing model Output probability; Indicates a given and In the case of historical task processing model Output probability; and The ratio of represents the change in the probability that the task processing model selects a specific output relative to the historical task processing model; represents a pruning operation, which is used to limit the range of changes in the strategy ratio of the task processing model to the historical task processing model; and Represents the preset thresholds used to control the maximum and minimum ratio changes allowed. The clipping results ensure that the strategies of the task processing model and the historical task processing model do not differ too much, thereby maintaining the stability of the reinforcement learning process.

[0062] (5) in, represents the function for training the task processing model, with the goal of maximizing the expected value; Express and The expected value of , which means that all input and output situations need to be considered to ensure that the model performs well in various situations; Indicates that all sample responses are averaged to ensure that each sample response contributes equally to the final result; represents the total strategy loss; Represents the model difference indicator, also known as the KL divergence penalty, which aims to prevent excessive deviation between the task processing model and the historical task processing model, thereby maintaining the stability of the training process; Is a hyperparameter used to control the weight of the KL divergence term; Represents the KL divergence of the task processing model relative to the historical task processing model; express The content length of the

[0063] By applying the solutions of the embodiments of this specification, efficient and stable model optimization is achieved through training with a combination of policy loss and model variance. Policy loss ensures that the task processing model learns toward higher response quality, while the model variance constrains the update amplitude of the task processing model, avoiding drastic policy fluctuations or performance degradation caused by reward signal noise. The synergy between policy loss and model variance improves the robustness and convergence of the model training process, ultimately resulting in a trained task processing model with high generation quality, stable behavior, and strong generalization capabilities.

[0064] In an optional embodiment of the present specification, in order to avoid the reward manipulation problem caused by the response (sample reply content) of the task processing model being too long, the embodiment of the present specification introduces a response length soft penalty mechanism, which constrains the reply index based on the content length of the sample reply content. That is, before training the task processing model based on the reply index and obtaining the trained task processing model, the following steps may be further included: According to the content length of the sample response content, the response index is constrained to obtain the target response index; Training the task processing model based on the response indicator to obtain a trained task processing model may include the following steps: According to the target response indicator, the task processing model is trained to obtain a trained task processing model.

[0065] It should be noted that content length refers to the length of each sample response generated by the task processing model. Content length is typically measured in characters or words. Content length can reflect the level of detail and complexity of the sample response.

[0066] Constraining response metrics means imposing restrictions on response metrics based on the length of sample response content, ensuring that response metrics consider not only the quality of sample response content but also the length of the content, thus avoiding reward manipulation issues.

[0067] The target response index refers to the response index after length constraint adjustment, which aims to balance the quality and length of the sample response content and provide a more comprehensive and reasonable optimization direction for the task processing model.

[0068] In actual applications, the reply index is constrained according to the content length of the sample reply content. There are many ways to obtain the target reply index, and the specific selection is based on the actual situation. The embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, when the length of the sample reply content exceeds the preset length, the reply index can be directly reduced by a certain proportion or set to a lower value to obtain the target reply index, thereby encouraging the task processing model to generate concise and clear answers. In another possible implementation of this specification, a penalty term related to the reply length can be introduced. The penalty term can be a linear, exponential or other form of function that increases with the increase of the reply length. Then, the penalty term is subtracted from the reply index to obtain the target reply index after length adjustment, so as to flexibly balance the relationship between reply quality and length.

[0069] Furthermore, the implementation method of “training the task processing model according to the target response index to obtain the trained task processing model” can refer to the above method of “training the task processing model according to the response index to obtain the trained task processing model”. For example, The target response indicator is modified and will not be described in detail in the embodiments of this specification.

[0070] By applying the solution of the embodiments of this specification, a constraint mechanism for content length is introduced to adjust the reply index, avoiding the reward manipulation problem, so that the training process can not only improve the overall quality of the reply, but also effectively control the length of the reply, and ultimately obtain a trained task processing model that can accurately respond to user needs and maintain a good user experience.

[0071] In an optional embodiment of the present specification, constraining the response index based on the content length of the sample response content to obtain the target response index may include the following steps: Determine content indicators corresponding to the multiple sample response contents according to the content length of the sample response contents; Fusing the content index and the response index to obtain an intermediate response index; The intermediate response indicators are standardized to obtain the target response indicators.

[0072] It's important to note that the content metric refers to a length penalty calculated based on the length of the sample response. For example, a sample response that's too short or too long might receive a higher content metric, while a sample response within the ideal length range might receive a lower score. The intermediate response metric is the intermediate score obtained by combining the content metric with the response metric (e.g., weighted summation or multiplication).

[0073] Normalization involves mapping intermediate response metrics to a uniform numerical range (e.g., [0, 1]) through mathematical transformations (e.g., min-max normalization or zero-mean normalization). Normalization can eliminate the effects of response metric scale. For example, in some states, the sample responses generated by the task processing model are generally high-quality, with response metrics concentrated in the range [8, 9]. In other states, the sample responses are of lower quality, with response metrics concentrated in the range [2, 3]. Without normalization, directly using the response metric as the advantage will cause the task processing model to favor updating states with high response metrics while ignoring actions with relatively better performance. After normalization, the target response metric is normalized to a mean of 0 and a standard deviation of 1, making different states comparable. Furthermore, in reinforcement learning, the direction and magnitude of parameter updates are influenced by the advantage value. Excessive fluctuations in the advantage value (e.g., sometimes +100 and sometimes -50) can lead to explosive or oscillatory parameter updates. Normalization can smooth parameter updates, helping the task processing model converge quickly.

[0074] In practical applications, the target response index can be calculated using the following formula (6): (6) in, represents the target response indicator, Indicates content indicators, represents the intermediate response indicator.

[0075] By applying the solution of the embodiments of this specification, a content indicator based on content length is introduced and integrated with the reply indicator to generate an intermediate reply indicator, thereby achieving joint modeling of reply quality and length. Then, a target reply indicator of a unified scale is obtained through standardization processing, which effectively avoids scoring bias caused by length differences. This can guide the task processing model to generate output that meets the expected length while ensuring high reply quality, thereby improving user experience and system controllability.

[0076] In an optional embodiment of the present specification, determining the content indicators corresponding to the plurality of sample reply contents respectively based on the content length of the sample reply contents may include the following steps: Obtain a first length parameter and a second length parameter, wherein the first length parameter is greater than the second length parameter; According to the first length parameter, the second length parameter and the content length of the sample reply content, content indicators corresponding to the plurality of sample reply contents are determined.

[0077] It should be noted that the first length parameter refers to the maximum allowable length of the sample response content (i.e., the length upper limit constraint). Sample response content exceeding the first length parameter will be considered too long and will be penalized (e.g., -1 point). The second length parameter refers to the length of the "soft penalty interval," which defines a transition region from "no penalty" to "strong penalty" to prevent the task processing model from being suddenly penalized strongly when approaching the length upper limit, thereby improving training stability. The first and second length parameters are selected based on actual conditions and are not limited in this specification.

[0078] In practical applications, the content index can be calculated using the following formula (7): (7) Among them, when the content length Greater than the first length parameter When, content indicators -1; when the content length Less than or equal to the first length parameter and a second length parameter The length difference When, content indicators 0; when the reply length Greater than or equal to , but less than or equal to When, according to 、 as well as Calculating content metrics .

[0079] By applying the solutions of the embodiments of this specification, by introducing first and second length parameters and combining them with the length of the sample responses, we can flexibly assess the length adaptability of the sample responses, thereby effectively controlling the length of the model responses and avoiding overly lengthy responses. Furthermore, by setting a reasonable transition zone, we provide flexibility to the task processing model, preventing a significant drop in response metrics due to even a slight excess of the length limit.

[0080] See also Figure 2 , Figure 2 A schematic diagram of the processing process of a task processing model training method provided by an embodiment of the present specification is shown. The task processing model training process includes: obtaining a query, inputting the query into the task processing model, and obtaining multiple replies (reply 1 to reply G); using a content comparison model, comparing and analyzing the multiple replies to obtain multiple reward values ​​(reward value 1 to reward value G); calculating multiple advantage values ​​(advantage value 1 to advantage value G) based on the multiple reward values; calculating the strategy loss based on the multiple advantage values; using an auxiliary model, calculating the KL divergence penalty based on the content of multiple sample replies; training the task processing model based on the strategy loss and the KL divergence penalty to obtain a trained task processing model.

[0081] It is worth mentioning that Figure 2 The snowflake symbol in the figure indicates that the model parameters of the content comparison model and the auxiliary model remain fixed during the training of the task processing model. The sparkle indicates that the model parameters of the task processing model are adjusted. The auxiliary model is a baseline model whose output probability distribution can be compared with the probability distribution generated by the task processing model to calculate the KL divergence between them.

[0082] Figure 2 Also shown are the traditional reward value modeling scheme and the reward value modeling scheme based on comparison strategy optimization proposed in the embodiment of this specification. In the traditional reward value modeling scheme, the reward value is calculated for each reply separately using the reward model, such as the reward value 1 corresponding to reply 1 is 0.6, the reward value 2 corresponding to reply 2 is 0.6, and the reward value G corresponding to reply G is 0.7. In the reward value modeling scheme based on comparison strategy optimization proposed in the embodiment of this specification, the content comparison model is used to compare multiple replies to obtain the reward value of each reply, such as the reward value 1 corresponding to reply 1 is 0.54, the reward value 2 corresponding to reply 2 is 0.62, and the reward value G corresponding to reply G is 0.69.

[0083] The following combined Figure 3Taking the application of the task processing model training method provided in this specification in a role-playing scenario as an example, the task processing model training method is further explained. Figure 3 A flowchart of a role-playing model training method provided by one embodiment of this specification is shown, which specifically includes the following steps: Step 302: Utilize the role-playing model to obtain multiple sample response contents based on the sample conversation data.

[0084] Step 304: Compare and analyze the multiple sample reply contents to obtain reply indicators corresponding to the multiple sample reply contents, wherein the reply indicators are used to measure the quality of the corresponding sample reply contents.

[0085] Step 306: Train the role-playing model according to the response indicator to obtain a trained role-playing model.

[0086] It should be noted that the role-playing model allows users to create virtual characters with their own personalities, memories, and behavioral logic based on their needs. These virtual characters can be used in a variety of scenarios, including gaming, emotional companionship, virtual assistants, education and training, and psychological counseling. The sample conversation data input to the role-playing model includes, but is not limited to, user information, character information, and conversation context. User information refers to background information or status information about the user, which helps the role-playing model better understand the conversation context and personalized needs. Character information defines the role that the role-playing model should play and its related attributes. It determines the role-playing model's behavior, tone, knowledge boundaries, personality settings, and more. Conversation context refers to the historical content of the current conversation, including multiple rounds of conversation between the user and the character.

[0087] In actual applications, the implementation of steps 302 to 306 is the same as that of steps 102 to 106, and will not be described in detail in this embodiment of the specification.

[0088] By applying the solution of the embodiments of this specification, by simultaneously generating and comparing multiple sample response contents, the differences in quality between responses can be accurately identified in relative comparison, thereby generating more discriminatory and stable response indicators, avoiding judgment bias caused by vague standards when scoring single samples, providing a more objective and accurate basis for the training process, and significantly improving the efficiency and accuracy of the role-playing model training process.

[0089] See also Figure 4 , Figure 4 A flowchart of a task processing method provided by an embodiment of this specification is shown, which specifically includes the following steps: Step 402: Obtain task data of the target task.

[0090] Step 404: Input the task data into the task processing model to obtain the task processing result of the target task, wherein the task processing model is trained based on the task processing model training method.

[0091] It should be noted that the target task refers to the specific task to be completed, such as question-answering, reasoning, classification, generation, and conversation. The target task can be tasks in different scenarios, such as multi-turn conversation tasks in role-playing scenarios and consulting tasks in customer service scenarios. Task data refers to the input data of the target task, which is used to input into the task processing model to generate the task processing results. Task data includes at least one of text, images, video, audio, and tables.

[0092] A task processing model refers to a deep learning model that can generate task processing results based on input task data. These include, but are not limited to, language models, classification models, and inference models. Because task processing models are trained using task processing model training methods, they possess high stability and precision, generating highly accurate task processing results. A task processing result is the output generated by the task processing model based on the input task data. It represents the solution or prediction of the target task by the task processing model.

[0093] By applying the solution of the embodiments of this specification, high-quality task processing results can be obtained by obtaining task data of the target task and inputting it into the task processing model trained by the task processing model training method.

[0094] Considering that the model parameters of the task processing model are relatively large and the computing resources of the client are limited, the task processing method proposed in the embodiment of this specification can be applied to the following: Figure 5 The task processing system shown is, but not limited to, Figure 5 , Figure 5 An architecture diagram of a task processing system provided by an embodiment of the present specification is shown. The task processing system may include a client 502 and a server 504; The client 502 is used to send task data of the target task to the server 504; The server 504 is configured to input the task data into the task processing model to obtain the task processing result of the target task, wherein the task processing model is trained based on the task processing model training method; and send the task processing result to the client 502; The client 502 is also used to receive the task processing result sent by the server 504.

[0095] like Figure 5As shown, the task processing model is deployed on a server 504. Server 504 can connect to one or more clients 502 via a local area network (LAN), a wide area network (WAN), the Internet, or other types of data networks. Clients 502 may include, but are not limited to, smartphones, tablet computers, laptops, PDAs, personal computers (PCs), smart home devices, and in-vehicle devices. Clients 502 can also interact with users through a graphical user interface (GUI) to invoke the task processing model and implement the task processing methods provided in the embodiments of this specification.

[0096] It is worth noting that the task processing methods provided in the embodiments of this specification are generally executed by the server. However, in other embodiments of this specification, if the client's operating resources can meet the deployment and operating conditions of the task processing model, the client can also have similar functions to the server and thus execute the task processing methods provided in the embodiments of this specification. In other embodiments, the task processing methods provided in the embodiments of this specification can also be executed jointly by the client and the server.

[0097] See also Figure 6 , Figure 6 A flowchart of a request processing method based on a task processing model provided in one embodiment of this specification is shown. The method is applied to a task platform and specifically includes the following steps: Step 602: Receive a model request sent by a terminal device.

[0098] Step 604: Based on the model request, a target task processing model is determined from a plurality of task processing models, wherein the plurality of task processing models are trained based on a task processing model training method.

[0099] It should be noted that multiple task processing models can have different model specifications and parameters, adapted to different scenarios. Because the task processing models are trained using the task processing model training method, they exhibit high stability and accuracy. The target task processing model is a task processing model suitable for the target scenario. The target scenario can be a variety of scenarios, such as role-playing scenarios, multi-round emotional dialogue scenarios, and so on. The model request includes at least one of the following: a scenario identifier for the target scenario, scenario input data for the target scenario, and model specification parameters.

[0100] In actual applications, there are many ways to determine the target task processing model from multiple task processing models based on the model request. The specific selection is based on the actual situation. The embodiments of this specification do not impose any restrictions on this. In the first possible implementation method of this specification, the corresponding target task processing model can be found from the task processing models included in the first model library based on the scenario identifier included in the model request; in the second possible implementation method of this specification, the target task processing model can be trained based on the scenario input data included in the model request; in the third possible implementation method of this specification, the corresponding target task processing model can be found from the task processing models included in the second model library based on the model specification parameters included in the model request.

[0101] For example, based on the scenario identification of the target scenario, at least one pre-trained task processing model can be searched from the first model library, and then based on the model specification parameters, an intermediate task processing model can be screened from at least one task processing model. Then, based on the scenario input data of the target scenario, the screened intermediate task processing model can be trained to obtain a target task processing model suitable for user needs.

[0102] The solution of the embodiment of this specification is applied to obtain the target task processing model according to user needs, realize personalized model service, provide users with an efficient, flexible and easy-to-use model service method, and improve user experience.

[0103] In an optional embodiment of the present specification, the above-mentioned determination of a target task processing model from multiple task processing models based on the model request may include the following steps: In a case where the model request includes a scenario identifier of a target scenario, searching a target task processing model adapted to the target scenario from a first model library based on the scenario identifier, wherein the first model library stores a plurality of task processing models adapted to different scenarios; When the model request includes scene input data of a target scene, determining a to-be-trained task processing model adapted to the target scene from a plurality of task processing models, and training the to-be-trained task processing model based on the scene input data to obtain a target task processing model; In the case where the model request includes model specification parameters, a target task processing model corresponding to the model specification parameters is searched from the second model library, wherein the second model library stores a plurality of task processing models with different model specification parameters.

[0104] It should be noted that the scenario identifier refers to a unique or specific label used to distinguish different scenarios. The first model library is a database for storing and managing various pre-trained deep learning models. Multiple task processing models adapted to different scenarios cover different application scenarios and needs. The first model library allows users to select appropriate task processing models according to their needs, or directly call the appropriate task processing model through the application programming interface to perform the target task. Multiple task processing models adapted to different scenarios are a variety of models suitable for different scenarios stored in the first model library, and each task processing model is optimized for a specific application environment. For example, based on the scenario identifier "role-playing" of the target scenario, the target task processing model adapted to the role-playing scenario can be searched from the first model library.

[0105] Each task processing model is suitable for different scenarios. For example, Task Processing Model 1 is suitable for Scenario 1 and Scenario 2, while Task Processing Model 2 is suitable for Scenario 2 and Scenario 3. The task processing model to be trained refers to the model among multiple task processing models that is suitable for the target scenario but whose performance can be further optimized. If the target scenario is Scenario 1, then the task processing model to be trained is Task Processing Model 1, which is suitable for Scenario 1. The task processing model to be trained may be applicable not only to the target scenario but also to other scenarios, making it a general task processing model that can be applied to different scenarios. The task processing model to be trained can perform the target task, but the performance may not be very good. In this case, the task processing model to be trained can be optimized based on the scenario input data of the target scenario. For example, by optimizing the task processing model based on the scenario input data of a role-playing scenario, a target task processing model suitable for the role-playing scenario can be obtained. The scenario input data of the target scenario can be understood as the sample set (sample conversation data and the corresponding multiple sample responses) used to train the task processing model in the target scenario.

[0106] Model specification parameters refer to the various parameters that define the model's structure and behavior. These parameters can be broadly divided into two categories: model parameters (learnable parameters) and hyperparameters. Model parameters are parameters that are automatically adjusted during model training via the backpropagation algorithm, including but not limited to weights and biases. For example, in a simple fully connected layer, the weight matrix is ​​a two-dimensional tensor that connects neurons in the input and output layers; the bias is a one-dimensional vector that provides an additional offset value for each output neuron. Hyperparameters are parameters set before model training begins to control the model's learning process and architecture. Hyperparameters include but are not limited to the learning rate and the number of neurons per layer, and should be selected based on actual circumstances.

[0107] By applying the solution of the embodiments of this specification, based on the scenario requirements, the target task processing model adapted to the corresponding scenario is accurately found through the scenario identification, so that the processing process of the target task is more accurate and more in line with the target scenario; based on the scenario requirements, the general task processing model to be trained is further trained through the scenario input data to obtain the target task processing model adapted to the target scenario, so that the target task processing model is more in line with the target scenario, thereby improving the user experience and the processing quality of the target task; based on the model specification parameters, the corresponding target task processing model can be accurately found, ensuring the efficient and stable operation of the target task processing model and improving the user experience.

[0108] In an optional embodiment of the present specification, after determining the target task processing model from multiple task processing models based on the model request, the following steps may also be included: Deploy the target task processing model, and build a task processing interface based on the target task processing model, so that the terminal device can schedule the target task processing model to execute the target task through the task processing interface.

[0109] It should be noted that the task processing interface is an interactive programming interface for the terminal device to schedule the target task processing model to process the target task, which is usually provided in the form of an application programming interface. Through the task processing interface, the user can input the task data of the target task to process the target task.

[0110] In actual applications, there are many ways to deploy the target task processing model, which can be selected based on actual conditions. The embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, the target task processing model can be deployed on a cloud-side device using the infrastructure provided by a cloud service provider. In another possible implementation of this specification, a lightweight framework can be used to deploy the target task processing model on an edge device. For example, the target task processing model can be deployed on a distributed system, and based on the target task processing model, a task processing interface can be constructed and provided to the terminal device, so that the terminal device can schedule the target task processing model to perform the target task.

[0111] By applying the solutions of the embodiments of this specification, deploying the target task processing model, and building a task processing interface, the terminal device can efficiently call the target task processing model, thereby improving the processing quality and response speed of the target task.

[0112] See also Figure 7 , Figure 7 FIG2 is a schematic diagram showing the structure of a task platform provided in one embodiment of the present specification. The task platform 700 includes a request interface 702 and a response unit 704; The request interface 702 is configured to receive a model request sent by a terminal device, wherein the model request includes at least one of a scene identifier of a target scene, scene input data of the target scene, and model specification parameters; The response unit 704 is configured to determine a target task processing model from a plurality of task processing models based on the model request, wherein the plurality of task processing models are trained based on a task processing model training method.

[0113] In an optional embodiment of the present specification, the task platform further includes a task processing interface, and the task processing interface is constructed based on the target task processing model; Task processing interface, used for terminal devices to schedule and execute target tasks.

[0114] The above is a schematic scheme of a task platform of this embodiment. It should be noted that the technical scheme of this task platform and the technical scheme of the request processing method based on the task processing model described above are based on the same concept. For details not described in detail in the technical scheme of the task platform, please refer to the description of the technical scheme of the request processing method based on the task processing model described above.

[0115] Corresponding to the above-mentioned task processing model training method embodiment, this specification also provides a task processing model training device embodiment, Figure 8 FIG1 shows a structural diagram of a task processing model training device provided by an embodiment of this specification. Figure 8 As shown, the device includes: A first acquisition module 802 is configured to utilize the task processing model to acquire a plurality of sample reply contents based on the sample conversation data; The first analysis module 804 is configured to perform comparative analysis on the plurality of sample reply contents to obtain reply indicators corresponding to the plurality of sample reply contents, wherein the reply indicators are used to measure the quality of the corresponding sample reply contents; The first training module 806 is configured to train the task processing model according to the response indicator to obtain a trained task processing model.

[0116] Optionally, the first analysis module 804 is further configured to filter out at least one sample comparison content from the second sample reply content for the first sample reply content, wherein the first sample reply content is any one of the multiple sample reply contents, and the second sample reply content is the sample reply content among the multiple sample reply contents except the first sample reply content; compare the first sample reply content and the sample comparison content to obtain the reply indicator corresponding to the first sample reply content.

[0117] Optionally, the first analysis module 804 is further configured to compare the first sample reply content with the sample comparison content based on the sample conversation data to obtain a reply indicator corresponding to the first sample reply content.

[0118] Optionally, the first analysis module 804 is further configured to input the comparison prompt information, the first sample reply content and the sample comparison content into the content comparison model to obtain the reply index corresponding to the first sample reply content, wherein the comparison prompt information is used to guide the content comparison model to compare the first sample reply content and the sample comparison content to obtain the reply index.

[0119] Optionally, the first analysis module 804 is further configured to compare the first sample reply content and the sample comparison content based on a content comparison strategy to obtain a reply index corresponding to the first sample reply content, wherein the content comparison strategy is constructed based on at least one quality assessment dimension.

[0120] Optionally, the first training module 806 is further configured to calculate the strategy loss based on the response index, wherein the strategy loss is used to measure the difference between the model performance of the task processing model and the expected model performance; calculate the model difference index based on the response content of multiple samples, wherein the model difference index is used to measure the difference between the task processing model and the historical task processing model, and the historical task processing model is the model version of the task processing model at the previous time point; train the task processing model based on the strategy loss and the model difference index to obtain a trained task processing model.

[0121] Optionally, the device also includes: a constraint module, configured to constrain the reply indicator according to the content length of the sample reply content to obtain a target reply indicator; a first training module 806, further configured to train the task processing model according to the target reply indicator to obtain a trained task processing model.

[0122] Optionally, the constraint module is further configured to determine the content indicators corresponding to multiple sample reply contents according to the content length of the sample reply contents; fuse the content indicators and reply indicators to obtain intermediate reply indicators; and standardize the intermediate reply indicators to obtain target reply indicators.

[0123] Optionally, the constraint module is further configured to obtain a first length parameter and a second length parameter, wherein the first length parameter is greater than the second length parameter; and determine the content indicators corresponding to the multiple sample reply contents based on the first length parameter, the second length parameter and the content length of the sample reply content.

[0124] By applying the solution of the embodiments of this specification, by simultaneously generating and comparing the content of multiple sample replies, the differences in quality between replies can be accurately identified in relative comparison, thereby generating more discriminatory and stable reply indicators, avoiding judgment bias caused by vague standards when scoring single samples, providing a more objective and accurate basis for the training process, and significantly improving the efficiency and accuracy of the task processing model training process.

[0125] The above is a schematic diagram of a task processing model training device according to this embodiment. It should be noted that the technical solution of the task processing model training device and the technical solution of the task processing model training method described above are based on the same concept. For details not described in detail in the technical solution of the task processing model training device, please refer to the description of the technical solution of the task processing model training method described above.

[0126] Corresponding to the above-mentioned role-playing model training method embodiment, this specification also provides a role-playing model training device embodiment, Figure 9 FIG1 shows a schematic diagram of a role-playing model training device provided by an embodiment of this specification. Figure 9 As shown, the device includes: The second acquisition module 902 is configured to use the role-playing model to acquire a plurality of sample reply contents based on the sample conversation data; The second analysis module 904 is configured to perform comparative analysis on the multiple sample reply contents to obtain reply indicators corresponding to the multiple sample reply contents, wherein the reply indicators are used to measure the quality of the corresponding sample reply contents; The second training module 906 is configured to train the role-playing model according to the response indicator to obtain a trained role-playing model.

[0127] By applying the solution of the embodiments of this specification, by simultaneously generating and comparing multiple sample response contents, the differences in quality between responses can be accurately identified in relative comparison, thereby generating more discriminatory and stable response indicators, avoiding judgment bias caused by vague standards when scoring single samples, providing a more objective and accurate basis for the training process, and significantly improving the efficiency and accuracy of the role-playing model training process.

[0128] The above is a schematic diagram of a role-playing model training device according to this embodiment. It should be noted that the technical solution of the role-playing model training device and the technical solution of the role-playing model training method described above are based on the same concept. For details not described in detail in the technical solution of the role-playing model training device, please refer to the description of the technical solution of the role-playing model training method described above.

[0129] Corresponding to the above-mentioned task processing method embodiment, this specification also provides a task processing device embodiment, Figure 10 FIG1 shows a schematic diagram of the structure of a task processing device provided by an embodiment of this specification. Figure 10 As shown, the device includes: The third acquisition module 1002 is configured to acquire task data of the target task; The input module 1004 is configured to input task data into a task processing model to obtain a task processing result of a target task, wherein the task processing model is trained based on a task processing model training method.

[0130] By applying the solution of the embodiments of this specification, high-quality task processing results can be obtained by obtaining task data of the target task and inputting it into the task processing model trained by the task processing model training method.

[0131] The above is a schematic scheme of a task processing device of this embodiment. It should be noted that the technical scheme of the task processing device and the technical scheme of the task processing method described above are of the same concept. For details not described in detail in the technical scheme of the task processing device, please refer to the description of the technical scheme of the task processing method described above.

[0132] Corresponding to the above-mentioned request processing method embodiment based on the task processing model, this specification also provides a request processing device embodiment based on the task processing model. Figure 11 FIG1 shows a schematic diagram of a request processing device based on a task processing model provided by an embodiment of this specification. Figure 11 As shown, the device is applied to a mission platform and includes: The receiving module 1102 is configured to receive a model request sent by a terminal device; The determination module 1104 is configured to determine a target task processing model from a plurality of task processing models based on the model request, wherein the plurality of task processing models are trained based on a task processing model training method.

[0133] Optionally, the determination module 1104 is further configured to, when the model request includes a scene identifier of the target scene, search for a target task processing model suitable for the target scene from the first model library based on the scene identifier, wherein the first model library stores multiple task processing models suitable for different scenarios; when the model request includes scene input data of the target scene, determine the task processing model to be trained that is suitable for the target scene from multiple task processing models, and train the task processing model to be trained based on the scene input data to obtain the target task processing model; when the model request includes model specification parameters, search for the target task processing model corresponding to the model specification parameters from the second model library, wherein the second model library stores multiple task processing models with different model specification parameters.

[0134] Optionally, the apparatus further includes: a deployment module configured to deploy a target task processing model and construct a task processing interface based on the target task processing model, so that the terminal device schedules the target task processing model to execute the target task through the task processing interface.

[0135] The solution of the embodiment of this specification is applied to obtain the target task processing model according to user needs, realize personalized model service, provide users with an efficient, flexible and easy-to-use model service method, and improve user experience.

[0136] The above is a schematic diagram of a request processing device based on a task processing model according to this embodiment. It should be noted that the technical solution of this request processing device based on a task processing model and the technical solution of the request processing method based on a task processing model described above share the same concept. For details not described in detail in the technical solution of the request processing device based on a task processing model, please refer to the description of the technical solution of the request processing method based on a task processing model described above.

[0137] Figure 12 The following is a block diagram of a computing device 1200 according to one embodiment of the present disclosure. The computing device 1200 includes a memory 1210 and a processor 1220. The memory 1210 is configured to store computer programs / instructions, and the processor 1220 is configured to execute the computer programs / instructions. When executed by the processor 1220, the computer programs / instructions implement the steps of the aforementioned task processing model training method, role-playing model training method, task processing method, or task processing model-based request processing method.

[0138] In one or more embodiments of this specification, the computing device can be understood as an integrated intelligent terminal, including but not limited to a server, a desktop computer, a personal computer, an all-in-one model machine, a mobile phone, a tablet computer or other portable intelligent terminal, etc., and the computing device can be pre-installed with the model described in the above embodiments of this application.

[0139] Specifically, the computing device can pre-install multiple types of models, including but not limited to models in the fields of natural language processing, visual processing, speech processing, code processing, and multimodal task processing, thereby providing a diverse selection of models. In different product forms, the computing device can support one or more model usage methods, including but not limited to model training, model calling, model fine-tuning, model deployment, model reasoning, and application. In some product forms, the computing device also supports model management, including but not limited to multi-type model management (supporting the management of multiple types of models such as discriminants and genesis), model version control (supporting the control of different model versions), and model evaluation (evaluating the performance and effectiveness of models based on model evaluation tools). In other product forms, the computing device can also create applications based on models and provide application programming interface (API) calling capabilities. Models can be called into created applications through the API interface, and application management tools are also provided to enable management and monitoring of applications.

[0140] Furthermore, the computing device can also include data management (supporting the creation and management of model tuning data sets), a training center (providing rich training resources to help users learn and master artificial intelligence technology), and basic management and control capabilities (providing enterprise-level basic management and control capabilities to ensure the security and efficient operation of the system). Through the above functions, a comprehensive, integrated artificial intelligence development, training, deployment and application device is provided.

[0141] Figure 13 FIG2 shows a structural block diagram of an electronic device 1300 provided according to an embodiment of the present specification.

[0142] The memory 1310 and the processor 1320 are connected via a bus 1330 ; The memory 1310 is used to store computer programs / instructions, and the processor 1320 is used to execute computer programs / instructions. When the computer program / instructions are executed by the processor 1320, the steps of the above-mentioned task processing model training method or role-playing model training method or task processing method or request processing method based on the task processing model are implemented.

[0143] Specifically, the components of the electronic device 1300 include but are not limited to a memory 1310 and a processor 1320. The processor 1320 and the memory 1310 may be connected via a bus 1330.

[0144] The electronic device 1300 may further include an access device 1340 that enables the electronic device 1300 to communicate with a database 1350 storing data via one or more networks 1360. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1340 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a World Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.

[0145] In one embodiment of the present specification, the above components of the electronic device 1300 and Figure 13 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 13 The electronic device structure block diagram shown is for illustrative purposes only and does not limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0146] Electronic device 1300 may be any type of stationary or mobile electronic device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable electronic device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary electronic device such as a desktop computer or PC. Electronic device 1300 may also be a mobile or stationary electronic device.

[0147] The above is a schematic scheme of an electronic device of this embodiment. It should be noted that the technical scheme of the electronic device and the technical schemes of the task processing model training method, role-playing model training method, task processing method, and request processing method based on the task processing model are of the same concept. For details not described in detail in the technical scheme of the electronic device, please refer to the description of the technical scheme of the task processing model training method, role-playing model training method, task processing method, or request processing method based on the task processing model.

[0148] One embodiment of this specification also provides a computer-readable storage medium, which stores a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned task processing model training method or role-playing model training method or task processing method or request processing method based on the task processing model.

[0149] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of this storage medium and the technical schemes of the aforementioned task processing model training method, role-playing model training method, task processing method, and request processing method based on the task processing model are of the same concept. For details not described in detail in the technical scheme of the storage medium, please refer to the description of the technical schemes of the aforementioned task processing model training method, role-playing model training method, task processing method, or request processing method based on the task processing model.

[0150] One embodiment of this specification also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned task processing model training method or role-playing model training method or task processing method or request processing method based on the task processing model.

[0151] The above is a schematic diagram of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product is based on the same concept as the technical solutions of the aforementioned task processing model training method, role-playing model training method, task processing method, and request processing method based on the task processing model. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solutions of the aforementioned task processing model training method, role-playing model training method, task processing method, or request processing method based on the task processing model.

[0152] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0153] Computer instructions include computer program code, which may be in source code, object code, executable files, or some intermediate form. Computer-readable media may include any entity or device capable of carrying computer program code, recording media, USB flash drives, removable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunications signals, and software distribution media. It should be noted that the content of computer-readable media may be appropriately expanded or reduced based on the requirements of patent practice. For example, in some jurisdictions, according to patent practice, computer-readable media does not include electric carrier signals or telecommunications signals.

[0154] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0155] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0156] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A task processing model training method, comprising: Using the task processing model, we obtain multiple sample responses based on the sample conversation data. Comparatively analyzing the multiple sample reply contents to obtain reply indicators corresponding to the multiple sample reply contents, wherein the reply indicators are used to measure the quality of the corresponding sample reply contents; The task processing model is trained according to the response indicator to obtain a trained task processing model.

2. The method according to claim 1, wherein the comparative analysis of the plurality of sample response contents to obtain response indicators corresponding to the plurality of sample response contents comprises: For the first sample reply content, at least one sample comparison content is screened out from the second sample reply content, wherein the first sample reply content is any one of the multiple sample reply contents, and the second sample reply content is a sample reply content among the multiple sample reply contents except the first sample reply content; Compare the first sample reply content with the sample comparison content to obtain the reply indicator corresponding to the first sample reply content.

3. The method according to claim 2, wherein comparing the first sample reply content with the sample comparison content to obtain the reply indicator corresponding to the first sample reply content comprises: Based on the sample conversation data, the first sample reply content and the sample comparison content are compared to obtain a reply index corresponding to the first sample reply content.

4. The method according to claim 2, wherein comparing the first sample reply content with the sample comparison content to obtain the reply indicator corresponding to the first sample reply content comprises: The comparison prompt information, the first sample reply content and the sample comparison content are input into the content comparison model to obtain the reply index corresponding to the first sample reply content, wherein the comparison prompt information is used to guide the content comparison model to compare the first sample reply content and the sample comparison content to obtain the reply index.

5. The method according to claim 2, wherein comparing the first sample reply content with the sample comparison content to obtain the reply indicator corresponding to the first sample reply content comprises: Based on a content comparison strategy, the first sample reply content and the sample comparison content are compared to obtain a reply index corresponding to the first sample reply content, wherein the content comparison strategy is constructed based on at least one quality assessment dimension.

6. The method according to any one of claims 1 to 5, wherein before training the task processing model based on the response indicator to obtain the trained task processing model, the method further comprises: Constraining the response index according to the content length of the sample response content to obtain a target response index; The step of training the task processing model according to the response indicator to obtain a trained task processing model includes: The task processing model is trained according to the target response indicator to obtain a trained task processing model.

7. The method according to claim 6, wherein constraining the response index based on the content length of the sample response content to obtain the target response index comprises: Determining content indicators corresponding to the plurality of sample response contents respectively according to the content length of the sample response contents; fusing the content indicator and the response indicator to obtain an intermediate response indicator; The intermediate recovery index is standardized to obtain the target recovery index.

8. The method according to claim 7, wherein determining the content indicators corresponding to the plurality of sample response contents respectively based on the content length of the sample response contents comprises: Obtaining a first length parameter and a second length parameter, wherein the first length parameter is greater than the second length parameter; The content indicators corresponding to the plurality of sample reply contents are determined according to the first length parameter, the second length parameter and the content length of the sample reply content.

9. The method according to claim 1, wherein the step of training the task processing model according to the response indicator to obtain a trained task processing model comprises: Calculating a policy loss based on the response indicator, wherein the policy loss is used to measure the difference between the model performance of the task processing model and the expected model performance; Calculating a model difference index based on the multiple sample response contents, wherein the model difference index is used to measure the difference between the task processing model and a historical task processing model, where the historical task processing model is a model version of the task processing model at a previous time point; The task processing model is trained according to the strategy loss and the model difference index to obtain a trained task processing model.

10. A role-playing model training method comprising: Using the role-playing model, multiple sample responses are obtained based on sample conversation data. Comparatively analyzing the multiple sample reply contents to obtain reply indicators corresponding to the multiple sample reply contents, wherein the reply indicators are used to measure the quality of the corresponding sample reply contents; The role-playing model is trained according to the response indicator to obtain a trained role-playing model.

11. A task processing method, comprising: Get the task data of the target task; The task data is input into a task processing model to obtain a task processing result of the target task, wherein the task processing model is trained based on the method according to any one of claims 1 to 8.

12. A request processing method based on a task processing model, applied to a task platform, comprising: Receive a model request sent by a terminal device; Based on the model request, a target task processing model is determined from a plurality of task processing models, wherein the plurality of task processing models are trained based on the method according to any one of claims 1 to 8.

13. The method according to claim 12, wherein determining a target task processing model from a plurality of task processing models based on the model request comprises: In a case where the model request includes a scenario identifier of a target scenario, searching a target task processing model adapted to the target scenario from a first model library based on the scenario identifier, wherein the first model library stores a plurality of task processing models adapted to different scenarios; In a case where the model request includes the scene input data of the target scene, determining a to-be-trained task processing model adapted to the target scene from the plurality of task processing models, and training the to-be-trained task processing model based on the scene input data to obtain a target task processing model; In a case where the model request includes model specification parameters, a target task processing model corresponding to the model specification parameters is searched from a second model library, wherein the second model library stores a plurality of task processing models with different model specification parameters.

14. The method according to claim 12, after determining a target task processing model from a plurality of task processing models based on the model request, further comprising: The target task processing model is deployed, and a task processing interface is constructed based on the target task processing model, so that the terminal device schedules the target task processing model to execute the target task through the task processing interface.

15. A task platform comprising a request interface and a response unit; The request interface is used to receive a model request sent by a terminal device, wherein: The model request includes a scene identifier of a target scene, scene input data of the target scene, and at least one of a model specification parameter; The response unit is configured to determine a target task processing model from a plurality of task processing models based on the model request, wherein the plurality of task processing models are trained based on the method according to any one of claims 1 to 8.

16. The task platform according to claim 15, further comprising a task processing interface, wherein the task processing interface is constructed based on the target task processing model; The task processing interface is used for the terminal device to schedule and execute the target task.

17. A computing device comprising: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 14 are implemented.

18. An electronic device comprising: a memory and a processor, wherein the memory and the processor are connected via a bus; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 14 are implemented.

19. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction is executed by a processor to implement the steps of the method according to any one of claims 1 to 14.

20. A computer program product comprising a computer program / instructions, which implement the steps of the method according to any one of claims 1 to 14 when executed by a processor.

Citation Information

Patent Citations

  • Role dialogue model training method, dialogue generation method, device and equipment

    CN117633198A

  • Reward model training method and device, electronic equipment and storage medium

    CN118656607A

  • Training method, application method, device and equipment of interactive data generation model

    CN118861262A

  • Task processing method, automatic question answering method and task processing model training method

    CN120523896A

  • Efficient multi-turn generative AI model suggested message generation

    US11947902B1