Model training method, device and equipment for interaction, medium and program product

Through multi-model interactive training, the behavior data of the agent is generated and updated, and the problem of decision accuracy and robustness of the agent in multi-role interaction scenarios is solved, achieving more efficient interactive behavior training.

CN120067674APending Publication Date: 2025-05-30BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510066037.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively train an agent to demonstrate accurate and robust behavioral decisions in multi-role interaction scenarios, especially in a changing interactive environment.

Method used

By utilizing multiple machine learning models, the corresponding behavior data of multiple roles in the target interaction process are generated, and the target model is updated based on the target behavior data and evaluation results, thereby introducing interactive training data of multiple agents or multiple models.

Benefits of technology

The accuracy and robustness of the target model in interactive scenarios are improved, allowing it to deal with complex multi-role interaction scenarios more effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067674A_ABST
    Figure CN120067674A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training method and device for interaction, equipment, a medium and a program product. The method comprises the following steps: generating corresponding behavior data of a plurality of roles in a target interaction process by using a plurality of machine learning models, the plurality of machine learning models comprising a to-be-trained target model, a machine learning model in the plurality of machine learning models is allocated to at least one role in the plurality of roles and is used for generating behavior data of the allocated at least one role; at least obtaining an evaluation result of target behavior data, wherein the target behavior data comprises behavior data of at least one role allocated by the target model; and updating the target model based on the target behavior data and the evaluation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to a method, apparatus, device, medium, and program product for training an interaction model. Background Art

[0002] With the development of computer technology, people can interact with various types of objects, such as other users, processing entities (also known as agents) based on machine learning models, etc. Generally, an agent can only interact with a single object (e.g., a single user) for conversations and the like. Currently, it has been proposed to apply agents to interaction scenarios or interaction processes including multiple roles. Summary of the Invention

[0003] In a first aspect of the present disclosure, there is provided a method for training an interaction model. The method includes: generating, by using a plurality of machine learning models, respective behavior data of multiple roles in a target interaction process, the plurality of machine learning models including a target model to be trained, and the machine learning models in the plurality of machine learning models being assigned to at least one role of the multiple roles and being used to generate behavior data of the at least one assigned role; obtaining at least an evaluation result of target behavior data, the target behavior data including behavior data of at least one role assigned to the target model; and updating the target model based on the target behavior data and the evaluation result.

[0004] In a second aspect of the present disclosure, there is provided a training apparatus for an interaction model. The apparatus includes: a behavior data generation module configured to generate, by using a plurality of machine learning models, respective behavior data of multiple roles in a target interaction process, the plurality of machine learning models including a target model to be trained, and the machine learning models in the plurality of machine learning models being assigned to at least one role of the multiple roles and being used to generate behavior data of the at least one assigned role; an evaluation result obtaining module configured to obtain at least an evaluation result of target behavior data, the target behavior data including behavior data of at least one role assigned to the target model; and a model update module configured to update the target model based on the target behavior data and the evaluation result.

[0005] In a third aspect of the present disclosure, there is provided an electronic device. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. The instructions, when executed by the at least one processor, cause the device to execute the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. Computer-executable instructions are stored on the computer-readable storage medium and can be executed by a processor to implement the method of the first aspect.

[0007] In a fifth aspect of the present disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to execute the method according to the first aspect of the present disclosure.

[0008] It should be understood that the content described in this part is not intended to define the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0010] Figure 1 A schematic diagram showing an example environment in which embodiments according to the present disclosure can be implemented is shown;

[0011] Figure 2 A schematic flow chart showing a model training architecture for interaction according to some embodiments of the present disclosure is shown;

[0012] Figure 3 A schematic diagram showing a scenario of a social reasoning game according to some embodiments of the present disclosure is shown;

[0013] Figure 4 A flowchart showing a model training process for interaction according to some embodiments of the present disclosure is shown;

[0014] Figure 5 A schematic structural block diagram showing a model training device for interaction according to some embodiments of the present disclosure is shown; and

[0015] Figure 6 A block diagram showing an electronic device capable of implementing multiple embodiments of the present disclosure is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0017] In the description of the embodiments of the present disclosure, the term "including" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter.

[0018] In this article, unless expressly stated, performing a step "in response to A" does not mean that the step is immediately performed after "A", but may include one or more intermediate steps.

[0019] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.

[0020] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner according to the relevant laws and regulations.

[0021] For example, when a user's active request is received, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require the acquisition and use of the user's personal information, so that the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server or a storage medium that performs the operations of the technical solution of the present disclosure according to the prompt message.

[0022] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may, for example, be in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0023] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0024] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training, for a given input, the corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" can also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.

[0025] As mentioned above, agents have been applied to interactive scenarios or interactive processes. In such an interactive scenario, for example, in a social reasoning game (SDG), not only does the agent need to have the ability of language dialogue, but also the ability to make decisions. That is to say, a machine learning model that can interact with other objects as an agent needs to be able to generate interactive strategies and also be able to dialogue. In addition, in such a scenario, there may be multiple agents. Different from the single-agent scenario (for example, question and answer), the environment faced by the agent is constructed by other agents and is often changeable, so this poses higher requirements for the robustness of the model in the face of changeable scenarios.

[0026] Currently, the following solutions have been proposed for the construction of this type of machine learning model. The first solution is simply based on prompt engineering, that is, designing prompts for the machine learning model. This method cannot consider the multi-agent scenario. The second solution is carried out by combining reinforcement learning (RL) and large language model (LLM). Specifically, it is mainly through the following two methods.

[0027] One way is to abstract the dialogue into structured data. Based on the structured data, first train the RL model. Then, for the part that needs to have a dialogue, use the LLM to expand the structured data to obtain the expression form for dialogue. For this method, since the structured data is a compressed expression, a lot of information in the dialogue is lost. The policy obtained from the RL model trained based on this lossy structured data is inaccurate. In addition, the robustness of this model is poor.

[0028] Another way is to first use the LLM to generate candidate behaviors (interactive actions, speeches, decisions, etc.). Then use the policy network RL trained based on embedding to select the best from the candidate behaviors. For this method, since the candidate behaviors generated by the LLM are inaccurate themselves, the best behaviors selected by the RL from them are also not accurate enough.

[0029] In view of this, embodiments of the present disclosure propose a training scheme for an interaction model. According to this scheme, multiple machine learning models are used to generate corresponding behavior data of multiple characters in a target interaction process. The multiple machine learning models include a target model to be trained. Each machine learning model among the multiple machine learning models is assigned to at least one character among the multiple characters and is used to generate the behavior data of the at least one assigned character; at least obtain an evaluation result of the target behavior data, where the target behavior data includes the behavior data of the at least one character assigned to the target model; and update the target model based on the target behavior data and the evaluation result.

[0030] In an embodiment of the present disclosure, corresponding behavior data is generated based on the interaction between multiple machine learning models. Then, the target model is updated using the target behavior data generated by the target model and the evaluation result of the target behavior data. That is, the interaction of multiple agents or multiple models is introduced in the training of the target model. This makes the training data of the target model closer to the actual interaction scenario. In this way, the performance of the target model can be improved, and the accuracy of the behavior decision of the target model in the interaction scenario can be improved. In addition, the robustness of the target model can also be improved so that the target model can handle complex interaction scenarios.

[0031] The following further describes various example implementations of this scheme in detail with reference to the accompanying drawings. Example environment

[0032] Figure 1 FIG. shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In Figure 1 the environment 100, it is desired to train and use such a target model 140, which is configured for multiple application environments. For example, the target model 140 can generate target behavior data 150, etc. in an interaction with a user or other models.

[0033] As Figure 1 shown, Figure 1 the upper half of FIG. shows the model training stage, and the lower half shows the model application stage. Before training, the parameter values of the initial machine learning model 130 can have initial values, or can have pre-trained parameter values obtained through a pre-training process.

[0034] In some embodiments, the environment 100 may include a model training system 120. During the model training phase, an initial machine learning model 130 may be trained based on the training dataset 110 and by using the model training system 120 to obtain the target model 140. Specifically, the training dataset 110 may include sample data and target behavior data generated by the target model during the interaction process, etc. During the training process, the parameter values of the initial machine learning model 130 may be iteratively updated and adjusted. After the initial machine learning model 130 is trained, the target model 140 is obtained.

[0035] In some embodiments, the environment 100 may include a model application system 160. During the model application phase, the target model 140 may be used to interact with a user or other models to generate target behavior data 150.

[0036] In Figure 1 it, the model training system 120 and the model application system 160 may include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. The terminal device may refer to any type of mobile terminal, fixed terminal or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. Servers include but are not limited to mainframes, edge computing nodes, computing devices in cloud environments, and so on.

[0037] It should be understood that Figure 1 the components and arrangements in the illustrated environment 100 are merely examples, and the computing systems suitable for implementing the exemplary implementations described in the present disclosure may include one or more different components, other components, and / or different arrangements. The implementations of the present disclosure are not limited in this regard.

[0038] Some example embodiments of the present disclosure will be further described below with reference to the accompanying drawings. Model fine-tuning stage

[0039] Figure 2 A schematic diagram of a process 200 for training a model for interaction according to some embodiments of the present disclosure is shown. In Figure 2 it, there are three stages, namely the model fine-tuning stage, the multi-model interaction stage, and the model update stage. These three stages can be regarded as Figure 1 at least a part of the model training stage shown in

[0040] The target model trained according to the method of the present disclosure is expected to be used in a target interaction process. Examples of the target interaction process may include social reasoning games, debates, negotiations, etc. For the convenience of describing the solution of the present disclosure, first, reference will be made to Figure 3 Describe a social reasoning game as an example. Figure 3 FIG. shows a schematic diagram of a scene 300 of a social reasoning game according to some embodiments of the present disclosure.

[0041] As Figure 3 shown, the players of this social reasoning game are divided into two camps, namely the pro camp 310 and the con camp 320. The pro camp 310 can be further divided into the first pro camp 311 and the second pro camp 312. Each player in the first pro camp 311 can have different special skills, and the players in the second pro camp 312 do not have special skills. In this game, except that the multiple players in the con camp 320 know each other's identities, the identities of the remaining players are invisible. In the game, each player needs to reason out the identities of other players and eliminate the players in the opposite camp. If all the players in the con camp 320 are eliminated, the pro camp 310 wins. If all the players in the first pro camp 311 or the second pro camp 312 in the pro camp 310 are eliminated, the con camp 320 wins. Of course, this game rule is only an example, and the method described in the present disclosure based on this game can also be applied to other interaction processes.

[0042] As Figure 3 shown, the game includes multiple rounds of interactions, and each round of interaction can include multiple stages, such as the first stage 330 and the second stage 340. In the first stage 330, each player makes different interaction actions based on the game rules and strategies of the corresponding role. For example, the players in the con camp 320 can jointly decide to eliminate a player. The player with the identity of "seer" in the first pro camp 311 can learn the identity of another player. The player with the identity of "guard" in the first pro camp 311 can guard another player from being eliminated by the players in the con camp 320. In the second stage 340, each player can speak, and the speech content includes but is not limited to explaining their own identity, speculating on the identities of other players, etc. After the speech ends, all players can jointly make a decision (such as voting) to select and eliminate a player. In multiple interaction rounds, the first stage 330 and the second stage 340 alternate until all the players in either camp are eliminated.

[0043] Next, reference will continue to be made to Figure 2 , and the model fine-tuning stage will be described in combination with Figure 3 the social reasoning game in

[0044] In some embodiments, the model training system 120 may utilize at least one type of dataset 210 to train an initial machine learning model to obtain a target model. The model training at this stage can be referred to as behavior cloning, that is, enabling the initial machine learning model to learn and imitate the behavior patterns of another entity (e.g., a human or other intelligent entity). Taking the above-mentioned social reasoning game as an example, the initial machine model can learn the behaviors, rules, strategies, etc. of players with different roles in the game during the model fine-tuning stage.

[0045] In some embodiments, at least one type of dataset 210 may at least include a behavior example dataset 211. The behavior example dataset 211 may indicate the corresponding behavior examples of different types of roles included in the target interaction process, that is, indicate what behaviors different types of roles should perform in the target interaction process. In some embodiments, each piece of behavior example data in the behavior example dataset 211 may correspond to one type of role among different types of roles. Alternatively or additionally, each piece of behavior example data in the behavior example dataset 211 may correspond to the behavior of one type of role in a certain round or a certain stage.

[0046] If the fine-tuning of the initial machine learning model is Supervised Fine–Tuning (SFT), then a labeled dataset is required to train the model to tell the model what the correct answer should be. In this case, the behavior example dataset 211 contains labeled data, that is, the behavior result data that the model should output is labeled.

[0047] In some embodiments, in order to enable the model to learn the logic behind behaviors such as interaction actions, speeches, and decisions, the behavior example dataset 211 may further include behavior analysis data. That is to say, the behavior data may include the behavior analysis data and behavior result data of the corresponding type of role. For example, the behavior result data may be the selection behavior of eliminating a certain player, and the behavior analysis data may be the thinking process for this selection behavior. In some embodiments, for the format of the behavior data, the behavior analysis data may be before the behavior result data. Alternatively or additionally, in some embodiments, the behavior analysis data may also be after the behavior result data.

[0048] The following uses a specific embodiment to explain the process of obtaining the behavior example dataset 211.

[0049] In a scenario where the target interaction process is a social reasoning game, in order for the initial machine learning model to simulate the role of players in the social reasoning game, a large amount of game data is required to train the initial machine learning model. The game data can include the data generated by real users when playing the game. For example, the game data can include data such as interaction actions, speeches, and decisions. Since game data usually lacks an analysis process, making it difficult to be used for model training, behavioral analysis data such as the reasons for behaviors and speech outlines can also be collected and supplemented. For example, for an interaction action, the principle and intention behind the interaction action can be annotated. The resulting dataset of behavioral examples can be used for model training in a more standardized form and more complete logic.

[0050] In some embodiments, at least one type of dataset 210 may further include a terminology dataset 212. The terminology dataset 212 can be used to explain the terms of the target interaction process. Learning these terms helps the model understand the behaviors or speeches of other interaction entities in the interaction process. For example, for Figure 3 the described game, the term "gold water" can mean that a player is confirmed by the "seer" as a good person, that is, the player belongs to the positive camp 310.

[0051] In some embodiments, at least one type of dataset 210 may further include a strategy dataset 213. The strategy dataset 213 can indicate the strategies of different types of roles in the target interaction process. Learning these strategies (such as guides, tips, etc.) helps the model make different behaviors for different roles assigned to it. For example, for Figure 3 the described game, when the model is assigned to a role belonging to the negative camp 320, it can pretend to be a role belonging to the positive camp 310 to avoid being eliminated by other players.

[0052] In summary, in the model fine-tuning stage, the model training system 120 uses at least one type of dataset 210 to train the initial machine learning model 130 to obtain the target model 140. The target model 140 can simulate different types of roles in the target interaction process and interact with other entities.

[0053] In some embodiments, in the subsequent multi-model interaction stage, other models 260 in the model pool 240 can also be obtained through the same or similar model fine-tuning stage. Multi-model interaction stage

[0054] Continue to refer to the following Figure 2 and describe the multi-model interaction stage in combination with the Figure 3 social reasoning game in it.

[0055] In the multi-model interaction stage, multiple machine learning models (e.g., the models in model pool 240) are utilized to generate corresponding behavior data of multiple characters in the target interaction process. The multiple machine learning models include the target model 140 and at least one other model 260. In some embodiments, the other model 260 may be a model obtained through supervised fine-tuning. Alternatively or additionally, in some embodiments, the other model 260 may be a model obtained in other ways and capable of being used in the target interaction process. For example, the model pool 240 may include various models such as models obtained through supervised fine-tuning, models obtained based on RL, and open-source models. The machine learning models among the multiple machine learning models are assigned to at least one of the multiple characters and are used to generate the behavior data of the at least one assigned character.

[0056] In some embodiments, to obtain a large amount of behavior data, multiple target interaction processes need to be carried out. In each target interaction process, the models in the model pool 240 are randomly assigned to at least one of the multiple characters. In particular, in each interaction process, the target model 140 is assigned at least one character.

[0057] Through multi-model interaction, the target model can generate rich target behavior data. Using this target behavior data for model update in the subsequent stage can improve the robustness of the target model and enable the target model to handle more complex interaction scenarios. Model update stage

[0058] Continue to refer to the following Figure 2 and combine with the Figure 3 social reasoning game in to describe the model update stage.

[0059] As Figure 2 shown, in the model update stage, at least obtain the evaluation result 280 of the target behavior data 150 generated by the target model 140, and update the target model 150 based on the target behavior data 150 and the evaluation result 280. For example, the target model 150 can be updated using the Kahneman-Tversky Optimization (KTO) loss function. KTO does not require paired data (e.g., including input, selected output, rejected output), and only requires a binary evaluation of good / bad for the target behavior data generated by the target model. The process of evaluating the target behavior data 150 to obtain the evaluation result 280 will be introduced first below.

[0060] Since the target interaction process is usually multi-round or multi-stage, and the final result of the target interaction process is jointly caused by the behaviors of multiple roles participating in the target interaction process, it is inappropriate to directly evaluate multiple behaviors in the target interaction process using the final result of the target interaction process. In some embodiments of the present disclosure, a phased evaluation method can be adopted to evaluate each behavior included in the target behavior data 150.

[0061] In some embodiments, the target interaction process may include at least one round of interaction, and each round of interaction in the at least one round of interaction may include multiple stages. The target behavior data may include multiple behaviors in different stages. For a given behavior among the multiple behaviors, the target role that makes the given behavior and the type of the given behavior can be determined. The type of the behavior may be related to the stage in which the behavior occurs. For example, the target role may be a "seer", a "guard" in the first positive camp 311, a "villager" in the second positive camp 312, etc. The type of the given behavior may include interaction actions (such as the prediction action of the "seer", the guarding action of the "guard"), speeches, decisions, etc. After determining the target role that makes the given behavior and the type of the given behavior, an evaluation of the given behavior can be determined based on the target role and the type of the given behavior. The evaluation of the given behavior may be positive or negative, that is, the above-mentioned good / bad binary evaluation. The evaluation method will be described in detail below for different target roles and different behavior types.

[0062] In some embodiments, if the given behavior includes an interaction action made by the target role to at least another role among the multiple roles, determine the behavior rules and behavior strategies corresponding to the role type of the target role, and based on the behavior rules and behavior strategies, determine the behavior evaluation of the interaction action, that is, determine whether the interaction action is positive or negative. For example, the given behavior may be the identity verification action of the "seer" to other players. When the player being verified by the "seer" belongs to the anti-camp 320, based on the behavior rules and behavior strategies of the "seer" role, it can be determined that the interaction action is positive.

[0063] In some embodiments, if the given behavior includes a speech of the target role, it can be determined whether the speech is positive or negative based on whether the decision result generated after the speech is beneficial to the target role. In one example, as Figure 3 shown, for the speech of a player in the second stage 340, if the player is voted out by other players in the subsequent player voting stage, it can be determined that the speech is negative.

[0064] Alternatively or additionally, in some embodiments, if a given action includes a statement by the target role, it can be determined whether the statement is positive or negative based on the consistency between the content of the statement and the facts that occurred during the target interaction process. For example, if the "seer" verified in the first stage 330 that player A belongs to the positive camp 310, but then stated in the second stage 340 that player A belongs to the negative camp 320, it can be determined that the statement is negative.

[0065] In some embodiments, if a given action includes a decision made by the target role, it can be determined whether the vote is positive or negative based on the relationship between the target role and the role the decision is directed at. For example, the decision can include a vote against other players during the elimination round. When the target role casts an elimination vote for a player in the same camp, it can be determined that the decision is negative. When the target role casts an elimination vote for a player in another camp, it can be determined that the decision is positive.

[0066] It can be understood that the evaluation methods for different actions are not limited to those described above. Other evaluation methods can also be determined based on the rules and strategies of the target interaction process. In summary, based on the evaluation of multiple actions, an evaluation result 280 for the target behavior data 150 can be obtained.

[0067] In some embodiments, based on the target behavior data 150 and the evaluation result 280, the KTO loss function can be used to update the target model 150. For example, for (x, y) ∈ D, KTO can optimize the policy π based on the following loss θ : In the above formula, λ D and λ U are hyperparameters for the ideal loss and the non - ideal loss respectively. y desirable represents the ideal state of y, and y undesirable represents the non - ideal state of y. λ y represents λ D when y is in the ideal case, and represents λ U .

[0068] In some embodiments, the update of the target model can be iteratively performed in multiple rounds. In each round of the multiple rounds, the model training system 120 can utilize multiple machine learning models to generate corresponding behavior data of multiple characters in the target interaction process. The multiple machine learning models can include the target model updated in the previous round. Then, at least obtain the evaluation result of the target behavior data generated in this round. Based on the target behavior data obtained in this round and the evaluation result, update the target model. That is to say, the above multi-model interaction stage and model update stage can be iteratively carried out. In each iteration process, based on the target behavior data generated in the multi-model interaction stage and the corresponding evaluation result, the parameters of the target model can be continuously updated.

[0069] In the above description, SDG is mainly used as an example of the target interaction scenario or target interaction process, but it should be understood that this is only for the purpose of description and is not intended to be any limitation. The embodiments of the present disclosure can be applied to any other suitable type of interaction scenario or interaction process. The target model trained by this solution has many advantages through experiments. First, the accuracy of decision-making is improved, that is, it can distinguish players in our camp and the opponent's camp and make correct decisions. Second, in the interaction process with other models, the winning rate is relatively high. Third, in the interaction process with multiple models, the winning rate is relatively high. Fourth, in the interaction process with users, the winning rate is relatively high. Example process

[0070] Figure 4 The flowchart of a model training process 400 for interaction according to some embodiments of the present disclosure is shown. The process 400 can be implemented at the model training system 120. The following refers to Figure 1 Describe the process 400.

[0071] In block 410, utilize multiple machine learning models to generate corresponding behavior data of multiple characters in the target interaction process. The multiple machine learning models include the target model to be trained. The machine learning models in the multiple machine learning models are assigned to at least one of the multiple characters and are used to generate the behavior data of the at least one assigned character.

[0072] In block 420, at least obtain the evaluation result of the target behavior data. The target behavior data includes the behavior data of at least one character assigned to the target model.

[0073] In block 430, update the target model based on the target behavior data and the evaluation result.

[0074] In some embodiments, at least the target model can be obtained by: using at least one type of dataset to train an initial machine learning model to obtain the target model, where the at least one type of dataset at least includes a behavior example dataset, and the behavior example dataset indicates corresponding behavior examples of different types of roles included in the target interaction process.

[0075] In some embodiments, each piece of behavior example data in the behavior example dataset can correspond to one type of role among different types of roles, and can include behavior analysis data and behavior result data of the corresponding type of role.

[0076] In some embodiments, the at least one type of dataset can further include at least one of the following: a term dataset for explaining terms of the target interaction process, or a strategy dataset indicating strategies of different types of roles in the target interaction process.

[0077] In some embodiments, the target behavior data can indicate multiple behaviors, and obtaining an evaluation result of the target behavior data can include: for a given behavior among the multiple behaviors, determining a target role among the multiple roles that makes the given behavior and the type of the given behavior; and based on the target role and the type of the given behavior, determining a behavior evaluation of the given behavior to obtain an evaluation result of the target behavior data, where the behavior evaluation indicates whether the given behavior is positive or negative.

[0078] In some embodiments, determining the behavior evaluation of the given behavior can include: in response to determining that the given behavior includes an interaction action made by the target role to at least another role among the multiple roles, determining a behavior rule and a behavior strategy corresponding to the role type of the target role; and based on the behavior rule and the behavior strategy, determining an evaluation of whether the interaction action is positive or negative.

[0079] In some embodiments, determining the behavior evaluation of the given behavior can include: in response to determining that the given behavior includes a speech of the target role, determining an evaluation of whether the speech is positive or negative based on at least one of the following: whether the decision result generated after the speech is beneficial to the target role, or the consistency between the content of the speech and the facts that occurred in the target interaction process.

[0080] In some embodiments, determining the behavior evaluation of the given behavior can include: in response to determining that the given behavior includes a decision made by the target role, determining an evaluation of whether the decision is positive or negative based on the relationship between the target role and the role targeted by the decision.

[0081] In some embodiments, the target interaction process can be one of multiple target interaction processes, and in each target interaction process among the multiple target interaction processes, multiple machine learning models can be randomly assigned to at least one role among the multiple roles.

[0082] In some embodiments, the update of the target model is iteratively performed in multiple rounds, and in each round of the multiple rounds, the method may include: generating, by using multiple machine learning models, corresponding behavior data of multiple roles in a target interaction process, where the multiple machine learning models include the target model updated in the previous round; obtaining at least an evaluation result of the target behavior data generated in the current round; and updating the target model based on the target behavior data obtained in the current round and the evaluation result. Example device and equipment

[0083] Figure 5 FIG. shows a schematic structural block diagram of a model training apparatus 500 for interaction according to some embodiments of the present disclosure. The apparatus 500 may be implemented as or included in a model training system 120. Each module / component in the apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.

[0084] As shown in the figure, the apparatus 500 includes a behavior data generation module 510, configured to generate, by using multiple machine learning models, corresponding behavior data of multiple roles in a target interaction process, where the multiple machine learning models include a target model to be trained, and the machine learning models in the multiple machine learning models are assigned to at least one of the multiple roles and are used to generate behavior data of the at least one assigned role; an evaluation result obtaining module 520, configured to obtain at least an evaluation result of the target behavior data, where the target behavior data includes behavior data of at least one role assigned to the target model; and a model update module 530, configured to update the target model based on the target behavior data and the evaluation result.

[0085] In some embodiments, the apparatus 500 further includes a target model obtaining module, configured to train an initial machine learning model by using at least one type of data set to obtain a target model, where the at least one type of data set at least includes a behavior example data set, and the behavior example data set indicates corresponding behavior examples of different types of roles included in the target interaction process.

[0086] In some embodiments, each piece of behavior example data in the behavior example data set may correspond to one type of role among different types of roles, and may include behavior analysis data and behavior result data of the corresponding type of role.

[0087] In some embodiments, the at least one type of data set may further include at least one of the following: a term data set for explaining terms of the target interaction process, or a strategy data set indicating strategies of different types of roles in the target interaction process.

[0088] In some embodiments, the target behavior data may indicate multiple behaviors, and the evaluation result obtaining module 520 may further be configured to: for a given behavior among the multiple behaviors, determine the target role that performs the given behavior and the type of the given behavior among the multiple roles; and based on the target role and the type of the given behavior, determine the behavior evaluation of the given behavior to obtain the evaluation result of the target behavior data, where the behavior evaluation indicates whether the given behavior is positive or negative.

[0089] In some embodiments, the evaluation result obtaining module 520 may further be configured to: in response to determining that the given behavior includes an interaction action performed by the target role on at least another role among the multiple roles, determine the behavior rules and behavior strategies corresponding to the role type of the target role; and based on the behavior rules and behavior strategies, determine the evaluation of whether the interaction action is positive or negative.

[0090] In some embodiments, the evaluation result obtaining module 520 may further be configured to: in response to determining that the given behavior includes the speech of the target role, determine the evaluation of whether the speech is positive or negative based on at least one of the following: whether the decision result generated after the speech is beneficial to the target role, or the consistency between the content of the speech and the facts that occurred in the target interaction process.

[0091] In some embodiments, the evaluation result obtaining module 520 may further be configured to: in response to determining that the given behavior includes a decision made by the target role, determine the evaluation of whether the decision is positive or negative based on the relationship between the target role and the role targeted by the decision.

[0092] In some embodiments, the target interaction process may be one of multiple target interaction processes, and in each of the multiple target interaction processes, multiple machine learning models may be randomly assigned to at least one of the multiple roles.

[0093] In some embodiments, the update of the target model is iteratively performed in multiple rounds, and in each of the multiple rounds, the behavior data generation module 510 may be configured to use multiple machine learning models to generate the corresponding behavior data of the multiple roles in the target interaction process, where the multiple machine learning models include the target model updated in the previous round; the evaluation result obtaining module 520 may be configured to at least obtain the evaluation result of the target behavior data generated in this round; and the model update module 530 may be configured to update the target model based on the target behavior data and the evaluation result obtained in this round.

[0094] Figure 6 A block diagram of an electronic device 600 is shown in which one or more embodiments of the present disclosure may be implemented. It should be understood that Figure 6The illustrated electronic device 600 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. Figure 6 The illustrated electronic device 600 may include or be implemented as Figure 1 a model training system 120 or a model application system 160.

[0095] As Figure 6 illustrated, the electronic device 600 is in the form of a general-purpose electronic device. The components of the electronic device 600 may include, but are not limited to, at least one processor 610 or processing unit, a memory 620, a storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processor 610 may be an actual or virtual processor and be capable of performing various processes according to the programs stored in the memory 620. In a multi-processor system, multiple processors execute computer-executable instructions in parallel to enhance the parallel processing ability of the electronic device 600.

[0096] The electronic device 600 generally includes multiple computer storage media. Such media may be any accessible media that can be obtained by the electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 620 may be volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 630 may be removable or non-removable media and may include machine-readable media, such as a flash drive, a magnetic disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 600.

[0097] The electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 6 it, a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 620 may include a computer program product 625 having one or more program modules configured to perform the various methods or actions of the various embodiments of the present disclosure.

[0098] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 600 may be implemented by a single computing cluster or multiple computing machines that are capable of communicating via a communication connection. Thus, the electronic device 600 may operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.

[0099] The input device 650 may be one or more input devices such as a mouse, keyboard, trackball, etc. The output device 660 may be one or more output devices such as a display, speaker, printer, etc. The electronic device 600 may also communicate with one or more external devices (not shown) as needed via the communication unit 640, the external devices such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the electronic device 600, or communicate with any device that enables the electronic device 600 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0100] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the method described above.

[0101] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0102] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, thereby producing a machine such that when the instructions are executed by the processor of the computer or other programmable data processing apparatus, a device is created that implements the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium, which instructions cause a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, so that the computer-readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0103] Computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0104] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, and the module, segment of a program, or part of an instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur out of the order noted in the figures. For example, two consecutive boxes may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box of the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.

[0105] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art in the field without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of technology in the market, or to enable other ordinary skilled artisans in the field to understand the various implementations disclosed herein.

Claims

1. A model training method for interaction, comprising: Generate corresponding behavior data of multiple roles in a target interaction process using a plurality of machine learning models, the plurality of machine learning models including a target model to be trained, a machine learning model in the plurality of machine learning models being assigned to at least one role among the plurality of roles and used to generate the behavior data of the at least one assigned role; obtaining at least an evaluation result of target behavior data, wherein the target behavior data includes behavior data of the at least one role assigned to the target model; as well as The target model is updated based on the target behavior data and the evaluation result.

2. The method according to claim 1, wherein at least the target model is obtained by: An initial machine learning model is trained using at least one type of data set to obtain the target model, wherein the at least one type of data set includes at least a behavior example data set, and the behavior example data set indicates corresponding behavior examples of different types of roles included in the target interaction process.

3. The method according to claim 2, wherein each item of behavior example data in the behavior example data set corresponds to a type of role among the different types of roles, and includes behavior analysis data and behavior result data of the corresponding type of role.

4. The method according to claim 2, wherein the at least one type of data set further comprises at least one of the following: a terminology dataset that explains the terminology of the target interaction process, or A strategy data set indicates strategies of different types of roles in the target interaction process.

5. The method according to claim 1, wherein the target behavior data indicates a plurality of behaviors, and obtaining an evaluation result of the target behavior data comprises: For a given behavior among the multiple behaviors, determining a target role among the multiple roles that performs the given behavior and a type of the given behavior; as well as Based on the target role and the type of the given behavior, a behavior evaluation of the given behavior is determined to obtain the evaluation result of the target behavior data, wherein the behavior evaluation indicates whether the given behavior is positive or negative.

6. The method of claim 5, wherein determining the behavior evaluation for the given behavior comprises: In response to determining that the given behavior includes an interactive action performed by the target character on at least another character among the plurality of characters, determining a behavior rule and a behavior strategy corresponding to a character type of the target character; as well as Based on the behavior rules and behavior strategies, an evaluation as to whether the interactive action is positive or negative is determined.

7. The method of claim 5, wherein determining the behavior evaluation for the given behavior comprises: In response to determining that the given behavior includes the speech of the target character, determining whether the evaluation of the speech is positive or negative based on at least one of the following: Whether the decision result after the speech is beneficial to the target role, or The consistency between the content of the speech and the facts that occurred during the target interaction process.

8. The method of claim 5, wherein determining the behavior evaluation for the given behavior comprises: In response to determining that the given behavior includes a decision made by the target character, determining an evaluation as to whether the decision is positive or negative based on a relationship between the target character and the character to which the decision is directed.

9. The method according to claim 1, wherein the target interaction process is one of a plurality of target interaction processes, and in each of the plurality of target interaction processes, the plurality of machine learning models are randomly assigned to at least one of the plurality of roles.

10. The method according to claim 1, wherein the updating of the target model is iteratively performed in a plurality of rounds, and in each of the plurality of rounds, the method comprises: Generate the corresponding behavior data of the multiple characters in the target interaction process by using the multiple machine learning models, wherein the multiple machine learning models include the target model after the last round of update; At least obtaining the evaluation result of the target behavior data generated in this round; as well as Based on the target behavior data and the evaluation results obtained in this round, the target model is updated.

11. A model training device for interaction, comprising: a behavior data generating module configured to generate corresponding behavior data of a plurality of roles in a target interaction process using a plurality of machine learning models, the plurality of machine learning models including a target model to be trained, a machine learning model in the plurality of machine learning models being assigned to at least one role among the plurality of roles and used to generate the behavior data of the at least one assigned role; An evaluation result obtaining module, configured to obtain at least an evaluation result of target behavior data, wherein the target behavior data includes behavior data of the at least one role assigned to the target model; as well as The model updating module is configured to update the target model based on the target behavior data and the evaluation result.

12. An electronic device comprising: at least one processor; as well as At least one memory, the at least one memory is coupled to the at least one processor and stores instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 10 when executed by the at least one processor.

13. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions can be executed by a processor to implement the method according to any one of claims 1 to 10.

14. A computer program product tangibly stored in a computer storage medium and comprising computer executable instructions which, when executed by a device, cause the device to perform the method according to any one of claims 1 to 10.