Human-computer interaction training method and device based on reinforcement learning strategy
Through the human-computer interaction training method based on reinforcement learning strategies, the interaction process of chat robots in the real estate field is optimized, and the problem of chat robots in the existing technology is difficult to effectively guide users to transfer commissions, achieving more efficient user to transfer commissions and goal achievement.
Patent Information
- Application Number
- CN202111521730.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-12-13
AI Technical Summary
It is difficult for existing chatbots to effectively guide users to transfer commissions in the real estate field. The pipeline model is not trained for achieving specific purposes, resulting in a low probability of users to transfer commissions.
Using a human-computer interaction training method based on reinforcement learning strategy, a second model is constructed by obtaining the first model trained by the target sample set, and in the simulation of the instant communication interaction, the response content output from the second model is adjusted to the degree of impact of the evaluation indicators of the interaction process, and the parameters of the second model are optimized to achieve a specific goal.
It increases the probability of users transferring commissions in the real estate field by chat robots, achieving more efficient user interaction and goal achievement.
Smart Images

Figure CN114417086B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of human-computer interaction, and in particular to a human-computer interaction training method and device based on a reinforcement learning strategy. Background Art
[0002] In order to improve the service quality for users and reduce the cost of manual services, the platform sets up chatbots before providing manual services to users. Chatbots can provide users with necessary basic services and solve some of the users' problems. Chatbots will only switch to manual services when they cannot solve the problems raised by users or have completed the current stage of communication and need to switch to the next stage of communication.
[0003] In the related art, most chatbots use a task-based pipeline model, which can solve the questions raised by users and ask users about the questions to obtain the necessary information to solve the questions. However, in the real estate field, chatbots are required to guide users to delegate. The pipeline model is not a model trained to achieve a specific purpose, and it cannot increase the probability of user delegation. Therefore, for chatbots used in the real estate field to achieve specific purposes, the task-based pipeline model is not well applicable. Summary of the invention
[0004] The purpose of this application is to provide a human-computer interaction training method and device based on reinforcement learning strategy, which is used to generate a chat robot used to achieve a specific goal.
[0005] The present application provides a human-computer interaction training method based on a reinforcement learning strategy, comprising:
[0006] Acquire a first model trained with a target sample set as a training sample; the target sample set includes interaction contents of multiple interaction processes; construct a second model, and use the second model and the first model to simulate an instant messaging interaction process; during the interaction process between the second model and the first model, the second model outputs reply content, and adjusts the parameters of the second model based on the degree of influence of the reply content output by the second model on the evaluation index of the interaction process; determine the second model after parameter optimization as the target model; wherein the evaluation index is used to indicate the probability that the interaction process can achieve a preset goal.
[0007] Optionally, obtaining the first model trained with the target sample set as training samples includes: using the target sample set as training samples to train a first GPT model and obtain the first model; wherein each sample in the training samples of the first GPT model includes category information; and the category information is used to classify the interactive content of the samples.
[0008] Optionally, constructing the second model and using the second model and the first model to simulate the instant communication interaction process includes: constructing the second model, and guiding the second model and the first model to start interacting based on the initial interaction content through the initial interaction content; using the reply content output by the second model as the input of the first model, and using the content output by the first model as the input of the second model, to realize the simulated instant communication interaction between the second model and the first model.
[0009] Optionally, the second model is a sorting model; in the interaction process between the second model and the first model, the second model outputs reply content, and based on the degree of influence of the reply content output by the second model on the evaluation index of the interaction process, the parameters of the second model are adjusted, including: in the simulated instant communication interaction process between the sorting model and the first model, the sorting model screens out the first reply content with the highest contextual relevance to the first content from the candidate reply set based on the first content output by the first model; screens out the second interaction content whose similarity with the first interaction content corresponding to the current interaction process meets the preset similarity from the retrieval library; the first interaction content includes the first reply content; feature extraction is performed on each interaction content in the third interaction content, and the feature vectors of each interaction content obtained are spliced to obtain the feature value of the third interaction content; the third interaction content includes: the first interaction content and the second interaction content; the feature value is determined as the first reward function value of the first reward function, and the parameters of the sorting model are adjusted based on the degree of influence of the content output by the sorting model on the evaluation index indicated by the first reward function value; wherein the first reward function is the reward function used by the reinforcement learning strategy constructed based on the sorting model.
[0010] Optionally, constructing the second model includes: using the target sample set as training samples to pre-train the second GPT model and obtain the second model; wherein each sample in the training samples of the second GPT model includes first object information and scene information; the first object information is used to indicate the first object corresponding to the interactive content of the sample; the scene information is used to indicate the application scenario to which the interactive content of the sample belongs.
[0011] Optionally, during the interaction between the second model and the first model, the second model outputs reply content, and based on the degree of influence of the reply content output by the second model on the evaluation index of the interaction process, the parameters of the second model are adjusted, including: during the simulated instant messaging interaction between the second GPT model and the first model, the content output by the first model is used as the input of the second GPT model, and the second reply content generated by the second GPT model is obtained; wherein the second reply content includes second object information; the second object information is used to indicate that the second reply content conforms to the language characteristics of the second object.
[0012] Optionally, during the interaction between the second model and the first model, the second model outputs reply content, and adjusts the parameters of the second model based on the degree of influence of the reply content output by the second model on the evaluation index of the interaction process, including: calculating the degree of influence of the content output by the second model on the evaluation index according to a second reward function, and adjusting the parameters of the second model based on the degree of influence; wherein the second reward function is a reward function used by the reinforcement learning strategy constructed based on the second GPT model.
[0013] The present application also provides a human-computer interaction training device based on a reinforcement learning strategy, comprising:
[0014] An acquisition module is used to acquire a first model trained with a target sample set as a training sample; the target sample set includes interaction contents of multiple interaction processes; a construction module is used to construct a second model, and use the second model and the first model to simulate an instant messaging interaction process; an adjustment module is used to adjust the parameters of the second model based on the degree of influence of the reply content output by the second model on the evaluation index of the interaction process when the second model outputs reply content during the interaction process between the second model and the first model; a determination module is used to determine the second model after parameter optimization as the target model; wherein the evaluation index is used to indicate the probability that the interaction process can achieve a preset goal.
[0015] Optionally, the acquisition module is specifically used to use the target sample set as training samples to train a first GPT model and obtain the first model; wherein each sample in the training samples of the first GPT model includes category information; and the category information is used to classify the interactive content of the sample.
[0016] Optionally, the device also includes: an interaction module; the construction module is specifically used to construct the second model, and to guide the second model to start interacting with the first model based on the initial interaction content through the initial interaction content; the interaction module is used to use the reply content output by the second model as the input of the first model, and use the content output by the first model as the input of the second model, so as to realize the simulated instant communication interaction between the second model and the first model.
[0017] Optionally, the device also includes: a screening module; the second model is a sorting model; the screening module is used to, during the simulated instant communication interaction process between the sorting model and the first model, the sorting model, based on the first content output by the first model, screens out the first reply content with the highest contextual relevance to the first content from the candidate reply set; the screening module is also used to screen out the second interactive content whose similarity with the first interactive content corresponding to the current interactive process meets a preset similarity from the retrieval library; the first interactive content includes the first reply content; the adjustment module is specifically used to extract features from each interactive content in the third interactive content, and to concatenate the feature vectors of each interactive content obtained to obtain a feature value of the third interactive content; the third interactive content includes: the first interactive content and the second interactive content; the adjustment module is specifically used to determine the feature value as the first reward function value of the first reward function, and adjust the parameters of the sorting model based on the degree of influence of the content output by the sorting model on the evaluation index indicated by the first reward function value; wherein the first reward function is a reward function used by a reinforcement learning strategy constructed based on the sorting model.
[0018] Optionally, the construction module is specifically used to pre-train the second GPT model using the target sample set as training samples and obtain the second model; wherein each sample in the training samples of the second GPT model includes first object information and scene information; the first object information is used to indicate the first object corresponding to the interactive content of the sample; the scene information is used to indicate the application scenario to which the interactive content of the sample belongs.
[0019] Optionally, the interaction module is also used to use the content output by the first model as the input of the second GPT model during the simulated instant messaging interaction between the second GPT model and the first model, and obtain second reply content generated by the second GPT model; wherein the second reply content includes second object information; the second object information is used to indicate that the second reply content conforms to the language characteristics of the second object.
[0020] Optionally, the adjustment module is specifically used to calculate the degree of influence of the content output by the second model on the evaluation index based on a second reward function, and adjust the parameters of the second model based on the degree of influence; wherein, the second reward function is a reward function used by the reinforcement learning strategy constructed based on the second GPT model.
[0021] The present application also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of any of the above-mentioned human-computer interaction training methods based on reinforcement learning strategies.
[0022] The present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of any one of the above-mentioned human-computer interaction training methods based on reinforcement learning strategies are implemented.
[0023] The present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described human-computer interaction training methods based on reinforcement learning strategies.
[0024] The human-computer interaction training method and device based on the reinforcement learning strategy provided by the present application first obtain a first model trained with a target sample set as a training sample, then construct a second model, and use the second model and the first model to simulate the instant messaging interaction process. Finally, in the process of interaction with the first model, based on the degree of influence of the interaction content of the second model on the evaluation index of the interaction process, the parameters of the second model are adjusted, and then a target model suitable for achieving a specific purpose is obtained, and a chat robot is created based on the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the present application or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0026] Figure 1 It is a flowchart of the human-computer interaction training method based on the reinforcement learning strategy provided by this application;
[0027] Figure 2 It is a structural schematic diagram of a sorting model provided by this application;
[0028] Figure 3 It is a schematic diagram of a model structure for judging the probability of delegation provided by this application;
[0029] Figure 4 It is a structural schematic diagram of a human-computer interaction training device based on a reinforcement learning strategy provided by the present application;
[0030] Figure 5 It is a structural schematic diagram of the electronic device provided by this application. DETAILED DESCRIPTION
[0031] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in this application will be clearly and completely described below in conjunction with the drawings in this application. Obviously, the described embodiments are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0032] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.
[0033] In order to facilitate the understanding of the technical solutions in the embodiments of the present application, the terms involved in the embodiments of the present application are explained below:
[0034] Reinforcement Learning: Reinforcement Learning, RL, a learning method, also known as reinforcement learning, evaluation learning or enhanced learning, is one of the paradigms and methodologies of machine learning. It is used to describe and solve the problem of how an agent can maximize rewards or achieve specific goals through learning strategies during its interaction with the environment.
[0035] The common model of reinforcement learning is the standard Markov Decision Process (MDP). According to the given conditions, reinforcement learning can be divided into model-based reinforcement learning (model-based RL) and model-free reinforcement learning (model-free RL), as well as active reinforcement learning (active RL) and passive reinforcement learning (passive RL). The algorithms used to solve reinforcement learning problems can be divided into two categories: strategy search algorithms and value function algorithms. Deep learning models can be used in reinforcement learning to form deep reinforcement learning.
[0036] Reinforcement learning theory is inspired by behaviorist psychology, focusing on online learning and trying to maintain a balance between exploration and exploitation. Unlike supervised learning and unsupervised learning, reinforcement learning does not require any data to be given in advance, but obtains learning information and updates model parameters by receiving rewards (feedback) from the environment for actions.
[0037] In order to obtain a chat robot that can be designed to achieve specific goals, the embodiments of the present application conceive that an interaction model can be obtained through a reinforcement learning strategy to achieve a preset purpose in the process of continuous interaction with the user.
[0038] It is understandable that since reinforcement learning strategies are used to achieve the problem of maximizing returns or achieving specific goals, in the real estate field, in order to achieve the purpose of delegation in communication with users, a chat model that can achieve this goal can be obtained through reinforcement learning strategies, and the chat model can be applied to the chat robot.
[0039] The following, in conjunction with the accompanying drawings, describes in detail the human-computer interaction training method based on the reinforcement learning strategy provided by the embodiment of the present application through specific embodiments and their application scenarios.
[0040] like Figure 1 As shown, an embodiment of the present application provides a human-computer interaction training method based on a reinforcement learning strategy, which may include the following steps 101 to 104:
[0041] Step 101: Obtain a first model trained using a target sample set as training samples.
[0042] The target sample set includes interaction contents of multiple interaction processes.
[0043] Exemplarily, the reinforcement learning core mainly includes actions, states, rewards, and the environment. Since it is impossible to use the online environment for model training, in order to train the model through the reinforcement learning strategy, it is necessary to build a simulation environment to simulate the real instant messaging (IM) interaction process.
[0044] Step 102, construct a second model, and use the second model and the first model to simulate the instant messaging interaction process.
[0045] Exemplarily, in the embodiments of the present application, the above first model can be used as a user simulator to simulate user communication. The above second model can be used as a chatbot to reply to the content output by the above user simulator.
[0046] Exemplarily, in order to enable the above first model to simulate user communication, a generation model can be trained with a large amount of interaction data collected during the interaction between the chatbot on the platform and the user.
[0047] Exemplarily, the generation model can be a GPT (Generate Pre-Training) model. The GPT model is trained with the above collected large amount of interaction data to obtain the above first model.
[0048] Specifically, the above step 101 may include the following step 101a:
[0049] Step 101a, use the target sample set as the training sample to train the first GPT model, and obtain the first model.
[0050] Among them, each sample in the training sample of the first GPT model includes category information; the category information is used to classify the interaction content of the sample.
[0051] Exemplarily, the above target sample set is obtained from a large amount of interaction data collected during the interaction between the chatbot on the platform and the user.
[0052] Exemplarily, considering the diversity of users, it is necessary to impose restrictions on the interaction content of the samples. These restrictions mainly include the city and category information. The category information can include types such as just-needed and investment. The data format for training GPT is as follows:
[0053] [CLS]How much house do you want?[SEP][city_110000][type_just-needed]I want a house worth five million yuan.[SEP]
[0054] It is understandable that each sample for training GPT can include multiple sentences, and each sentence is separated by the [SEP] tag. And the sample needs to include information such as city and category.
[0055] The platform's chatbot collects a large amount of interaction data during its interactions with users, including the user's city and the type of chat content. The above type information is used to further divide users, and the finer the division, the better.
[0056] It should be noted that after obtaining the above-mentioned first model, the parameters of the first model are no longer adjusted during the entire reinforcement learning process.
[0057] For example, after obtaining the above-mentioned first model, it is also necessary to construct a second model that can interact with the first model, and guide the second model to start interacting with the first model through a preset initial interaction content.
[0058] Specifically, the above step 102 may include the following steps 102a1 and 102a2:
[0059] Step 102a1: construct the second model, and guide the second model to start interacting with the first model based on the initial interaction content through the initial interaction content.
[0060] Step 102a2: Use the reply content output by the second model as the input of the first model, and use the content output by the first model as the input of the second model, so as to realize the simulated instant communication interaction between the second model and the first model.
[0061] It is understandable that, usually, the chatbot needs to reply based on the content input by the user. Therefore, the above initial interaction content can be used as the input of the second model. The second model outputs the reply content based on the initial interaction content. After that, the first model further outputs the subsequent interaction content based on the reply content, thereby realizing the interaction process between the second model and the first model.
[0062] Step 103: During the interaction between the second model and the first model, the second model outputs reply content, and adjusts the parameters of the second model based on the degree of influence of the reply content output by the second model on the evaluation index of the interaction process.
[0063] The evaluation index is used to indicate the probability that the interaction process can achieve a preset goal.
[0064] It is understandable that reinforcement learning is a learning method for training a model in order to achieve a specific goal. Therefore, it is necessary to set a goal for the second model, that is, the preset goal. All training for the second model is to achieve the preset goal. In the real estate field, the preset goal can be to enable users participating in the conversation to transfer.
[0065] Exemplarily, during the training process of the second model, it is necessary to continuously calculate the reply content output by the second model to obtain the probability that the entire interaction process can achieve the preset goal, and based on the probability of achieving the preset goal, provide feedback to the second model and adjust the parameters of the second model, so that the reply content output by the second model can maintain or improve the probability of achieving the preset goal in the interaction process.
[0066] For example, in the real estate field, the main purpose of the platform's chatbot is to transfer users. Therefore, the above evaluation indicators can be used to indicate the probability that the reply content output by the second model can achieve user transfer.
[0067] Step 104: determine the second model after parameter optimization as the target model.
[0068] Exemplarily, after the entire reinforcement learning strategy is executed, a target model is obtained, and then a chatbot can be created based on the target model to increase the probability of user transfer in the real estate field.
[0069] It should be noted that, as the four cores of reinforcement learning: environment, state, action and reward, in the embodiment of the present application, the simulated dialogue scene of the first model and the second model is used as the environment of reinforcement learning; the dialogue content of the context is used as the state, and the dialogue content output by the second model is used as the action; the dialogue content output by the second model that can achieve the preset goal of this dialogue is used as the reward.
[0070] During the reinforcement learning process, the reward function is used to calculate in real time the impact of the dialogue content output by the second model on the realization of the above-mentioned preset goals each time, so as to determine whether the dialogue content output by the second model can have a positive impact on the realization of the above-mentioned preset goals, and appropriately adjust the parameters of the second model so that the dialogue content it outputs can increase the probability of achieving the preset goals.
[0071] In this way, firstly, a first model trained with the target sample set as training samples is obtained, then a second model is constructed, and the instant messaging interaction process is simulated using the second model and the first model. Finally, in the process of interaction with the first model, the parameters of the second model are adjusted based on the degree of influence of the interaction content of the second model on the evaluation index of the interaction process, thereby obtaining a target model suitable for achieving a specific purpose, and creating a chatbot based on the model.
[0072] Optionally, in an embodiment of the present application, the above-mentioned second model can be a sorting model or a generation model, that is, the reply content output by the second model can be the reply content screened by the second model from the candidate replies, or it can be the reply content generated according to the content output by the first model.
[0073] Exemplarily, the second model may include the following two examples: Example 1 and Example 2.
[0074] Example 1:
[0075] In Example 1, the second model may be a sorting model.
[0076] Exemplarily, the above step 103 may include the following steps 103a1 to 103a4:
[0077] Step 103a1: During the simulated instant messaging interaction between the sorting model and the first model, the sorting model selects the first reply content with the highest contextual relevance to the first content from the candidate reply set based on the first content output by the first model.
[0078] Step 103a2: Filter out from the search library second interaction content whose similarity with the first interaction content corresponding to the current interaction process meets a preset similarity.
[0079] The first interactive content includes the first reply content.
[0080] Step 103a3: extract features from each interactive content in the third interactive content, and concatenate the obtained feature vectors of each interactive content to obtain a feature value of the third interactive content.
[0081] The third interactive content includes: the first interactive content and the second interactive content.
[0082] Step 103a4: determine the characteristic value as the first reward function value of the first reward function, and adjust the parameters of the ranking model based on the degree of influence of the content output by the ranking model on the evaluation index indicated by the first reward function value.
[0083] Among them, the first reward function is a reward function used by the reinforcement learning strategy constructed based on the sorting model.
[0084] Exemplarily, for the reinforcement learning strategy based on the ranking model, in the embodiment of the present application, the first reward function can be used to extract the feature vector of each interactive content in the third interactive content, and the obtained multiple feature vectors are spliced in a fully connected manner to obtain the reward function value of the first reward function, that is, the above-mentioned feature value. The feature value is used to determine the degree of influence of the first reply content output by the ranking model on the evaluation index, and then adjust the parameters of the ranking model.
[0085] It is understandable that the ranking model is used to select the most matching candidate from multiple candidate items. Since the chatbot of the platform collects a large amount of interaction data during the interaction with the user. Therefore, a search library can be created based on the above collected interaction records. When receiving the first content output by the first model, the ranking model can select the content to be replied with the highest degree of correlation with the first content from the above search library as the first reply content.
[0086] For example, for the training of the sorting model, in order to improve the contextual relevance between the reply content output by the sorting model and the content output by the first model, the first interactive content including the reply content screened by the sorting model and the second interactive content screened in the above retrieval library with a high similarity to the first interactive content can be obtained, and the third interactive content consisting of the first interactive content and the second interactive content is used as the second model input to extract the features of each interactive content in the third interactive content. Finally, the feature vectors of all interactive contents are spliced through the fully connected layer to obtain a feature value.
[0087] Exemplarily, after obtaining the above-mentioned characteristic value, the context relevance of the reply content output by the second model in the entire interaction process can be determined based on the characteristic value, and then the parameters of the second model can be adjusted.
[0088] For example, Figure 2 The figure is a schematic diagram of the structure of the ranking model provided in the embodiment of the present application. First, R represents the reply content selected by the ranking model from multiple candidate replies, s1 to s n It represents the conversation that occurs during the current interaction, that is, the simulated conversation context, where s1 and s3 are the content output by the first model, that is, the first content mentioned above, and s2 is the reply content output by the second model, that is, when the current conversation content is s1, the first reply content obtained by the second model is specifically the content of the reply to s1. s'1 to s' n represents the real conversation context of the real interaction content with the highest similarity to the above interaction content from the search library. Before the sorting model, a pre-training model Bert is also included, which is used to convert the above s1 to s n , s'1 to s' nThe conversation in R is converted into a vector through an embedding, so that a vector representation of the message level can be obtained. After that, the multiple vectors output by the above pre-trained model Bert are input into the sorting model. The sorting model mainly integrates four types of features:
[0089] Feature 1: Since the recent interactive content has a higher degree of attention, a two-layer transformer is used for feature extraction to determine s1 to s n The relationship between each conversation content, and then select the most recent interaction content S n The relationship with the context and extract the feature vector. Among them, q, k, and v represent query, keyword key, and value respectively. q = Sn means that the conversation content Sn is integrated into the above simulated conversation context and feature extraction is performed in combination with the context.
[0090] Feature 2: Select to integrate the current response into the above simulated conversation context to extract a feature vector, which can represent the appropriateness of the current response in the current conversation.
[0091] Feature 3: Evaluate the response content output by the sorting model independently without considering the simulated conversation context.
[0092] Feature 4: measure the relationship between each candidate response in the candidate response content and the actual conversation context (i.e., the above s'1 to s' n ) in the appropriateness.
[0093] Finally, the features extracted in the above steps are fused. Specifically, the final feature vector can be obtained by concatenating the feature vectors.
[0094] Example 2:
[0095] In Example 2, the second model may be a generative model.
[0096] Specifically, the step of constructing the second model in the above step 102 may include the following steps 102b:
[0097] Step 102b: Use the target sample set as training samples to pre-train the second GPT model and obtain the second model.
[0098] Among them, each sample in the training samples of the second GPT model includes first object information and scene information; the first object information is used to indicate the first object corresponding to the interactive content of the sample; the scene information is used to indicate the application scenario to which the interactive content of the sample belongs.
[0099] Exemplarily, the second GPT model may generate reply content in the following format:
[0100] [CLS][city_110000][agentId_12095][action_price]This one is five hundred thousand.
[0101] [SEP]
[0102] The above city indicates that the reply content is based on the city data indicated by the city, agentId is the agent ID, and action_price indicates that the reply content is price-related.
[0103] Exemplarily, it can be understood that since the samples used for pre-training the above-mentioned second GPT model include the second object information and the scene information, the reply content generated by the second GPT model can include the second object information and the scene information, so that we can generate reply content according to the language style of a certain broker (i.e., the above-mentioned second object), and at the same time, the action field is also included when generating the reply content, which field can limit the scope of the generated reply content. For example, the above-mentioned action_price indicates that the generated reply content is price-related content.
[0104] Furthermore, in the embodiment of the present application, the interactive content training method based on the reinforcement learning strategy needs to adjust the entire learning process through a reward function.
[0105] Exemplarily, during the training process of the second model, the parameters of the second model can be adjusted based on the reward function of the reinforcement learning strategy.
[0106] Specifically, the above step 103 may include the following step 103b:
[0107] Step 103b: Calculate the degree of influence of the content output by the second model on the evaluation index according to the second reward function, and adjust the parameters of the second model based on the degree of influence.
[0108] Among them, the second reward function is a reward function used by the reinforcement learning strategy constructed based on the second GPT model.
[0109] Exemplarily, for the reinforcement learning strategy constructed based on the above-mentioned second GPT model, a reward function different from that of the reinforcement learning strategy constructed based on the above-mentioned ranking model can be adopted.
[0110] For example, the second reward function in the embodiment of the present application may be the following formula:
[0111] R=λ1f(PPL)+λ2f(MMI)+λ3f(repetition)+λ4f(good)
[0112] +λ5turn+λ6f(delegation)+λ7delegation
[0113]
[0114]
[0115] Among them, PPL: perplexity is used to measure how well a probability distribution or probability model predicts a sample.
[0116] MMI: Mutual Information, logP(Sn|R)P, indicates that there is a reply generation s n This is mainly to prevent the generation of meaningless replies.
[0117] Repetition: Considering that the generated model is prone to repeating historical conversation information, a repetition penalty term is given here.
[0118] Turn: The number of conversation turns. In addition to handling delegation, we also hope to increase the number of conversation turns. This can be directly counted during the training process.
[0119] Delegation: It is used to detect whether the user's words are delegated and identify them at the end of the conversation.
[0120] f(delegation): used to determine the probability of delegation, which is calculated by training the third model through the interactive content in the above retrieval library. It is used to provide feedback information in the middle of training the second model. The third model is designed as follows Figure 3 shown.
[0121] The loss function of the third model is as follows:
[0122] loss=-[ylog(sigmoid(logits))+(1-y)log(1-sigmoid(logits))]
[0123] Among them, the calculation formula of logits in the above formula is as follows:
[0124]
[0125] like Figure 3 As shown, first, the above simulated conversation context (i.e., s1 to s nThe first model divides the conversation content into words (w1 to wm) and inputs them into the Bert model. The Bert model outputs a sentence vector SV (for example, SV1) corresponding to each sentence. Where T represents the source of the corresponding sentence (the sentence output by the first model or the second model).
[0126] The sentence vector of each sentence is input into the Long Short-Term Memory (LSTM) network to obtain the context vector E of each sentence. After that, the context vector E (for example, E1) of each sentence is respectively calculated with the random vector Del Imp or Uer Will corresponding to the model that outputs the sentence to obtain the probability of sub-delegation corresponding to each sentence. Among them, Del Imp is used to calculate the impact of the content output by the second model on sub-delegation; Uer Will is used to calculate the impact of the content output by the first model on sub-delegation. a represents the agent, that is, the above-mentioned second model, and u represents the user, that is, the above-mentioned first model. Dn,a represents the impact of the agent's conversation on sub-delegation; Dn,u represents the impact of the user's conversation on sub-delegation.
[0127] Then, the above effects are added to sigmoid to calculate the cross entropy loss, and the impact of each sentence on the delegation is obtained.
[0128] To determine whether a conversation is delegated, the entire conversation can be classified, as in the third model design above. Delimp is used to calculate the impact of the broker's message on the delegation of the conversation. User will is used to calculate the strength of the user's willingness to delegate based on the user's message, and finally the number of messages sent by the user and the broker is added. Then this impact is added to the sigmoid for cross entropy loss, and finally we can calculate the impact of each sentence on the delegation.
[0129] The human-computer interaction training method based on reinforcement learning strategy provided in the embodiment of the present application first builds a first model for simulating user interaction and a second model for simulating robot interaction. Then, the parameters of the second model are adjusted based on the reinforcement learning strategy, so that the second model becomes an interaction model for achieving specific goals, and then a chat robot is created based on the interaction model to increase the probability of achieving the above-mentioned specific goals when interacting with users.
[0130] It should be noted that the human-computer interaction training method based on reinforcement learning strategy provided in the embodiment of the present application can be executed by a human-computer interaction training device based on reinforcement learning strategy, or a control module in the human-computer interaction training device based on reinforcement learning strategy for executing the human-computer interaction training method based on reinforcement learning strategy. In the embodiment of the present application, the human-computer interaction training device based on reinforcement learning strategy is taken as an example to execute the human-computer interaction training method based on reinforcement learning strategy, to illustrate the human-computer interaction training device based on reinforcement learning strategy provided in the embodiment of the present application.
[0131] It should be noted that in the embodiments of the present application, the above-mentioned methods are shown in the accompanying drawings. The human-computer interaction training methods based on the reinforcement learning strategy are all illustrated by combining an accompanying drawing in the embodiments of the present application as an example. In specific implementation, the human-computer interaction training methods based on the reinforcement learning strategy shown in the accompanying drawings of the above-mentioned methods can also be implemented in combination with any other accompanying drawings that can be combined as shown in the above-mentioned embodiments, which will not be repeated here.
[0132] The human-computer interaction training device based on the reinforcement learning strategy provided by the present application is described below. The human-computer interaction training method based on the reinforcement learning strategy described above and described below can be referenced to each other.
[0133] Figure 4 A schematic diagram of the structure of a human-computer interaction training device based on a reinforcement learning strategy provided in an embodiment of the present application is shown in FIG. Figure 4 As shown, specifically including:
[0134] An acquisition module 401 is used to acquire a first model trained with a target sample set as a training sample; the target sample set includes interaction contents of multiple interaction processes; a construction module 402 is used to construct a second model, and use the second model and the first model to simulate an instant messaging interaction process; an adjustment module 403 is used to adjust the parameters of the second model based on the degree of influence of the reply content output by the second model on the evaluation index of the interaction process when the second model outputs reply content during the interaction process between the second model and the first model; a determination module 404 is used to determine the second model after parameter optimization as the target model; wherein the evaluation index is used to indicate the probability that the interaction process can achieve a preset goal.
[0135] Optionally, the acquisition module 401 is specifically used to use the target sample set as a training sample to train a first GPT model and obtain the first model; wherein each sample in the training sample of the first GPT model includes category information; and the category information is used to classify the interactive content of the sample.
[0136] Optionally, the device also includes: an interaction module 405; the construction module 402 is specifically used to construct the second model, and guide the second model to start interacting with the first model based on the initial interaction content through the initial interaction content; the interaction module 405 is used to use the reply content output by the second model as the input of the first model, and use the content output by the first model as the input of the second model, to realize simulated instant communication interaction between the second model and the first model.
[0137] Optionally, the device also includes: a screening module 406; the second model is a sorting model; the screening module 406 is used for, during the simulated instant communication interaction process between the sorting model and the first model, the sorting model, based on the first content output by the first model, screens out the first reply content with the highest contextual relevance to the first content from the candidate reply set; the screening module 406 is also used to screen out the second interactive content whose similarity with the first interactive content corresponding to the current interactive process meets the preset similarity from the retrieval library; the first interactive content includes the first reply content; the adjustment module 403 is specifically used to extract features from each interactive content in the third interactive content, and concatenate the feature vectors of each interactive content obtained to obtain the feature value of the third interactive content; the third interactive content includes: the first interactive content and the second interactive content; the adjustment module 403 is also specifically used to determine the feature value as the first reward function value of the first reward function, and adjust the parameters of the sorting model based on the degree of influence of the content output by the sorting model on the evaluation index indicated by the first reward function value; wherein the first reward function is the reward function used by the reinforcement learning strategy constructed based on the sorting model.
[0138] Optionally, the construction module 402 is specifically used to pre-train the second GPT model using the target sample set as training samples and obtain the second model; wherein each sample in the training samples of the second GPT model includes first object information and scene information; the first object information is used to indicate the first object corresponding to the interactive content of the sample; the scene information is used to indicate the application scenario to which the interactive content of the sample belongs.
[0139] Optionally, the interaction module 405 is also used to use the content output by the first model as the input of the second GPT model during the simulated instant messaging interaction between the second GPT model and the first model, and obtain second reply content generated by the second GPT model; wherein the second reply content includes second object information; the second object information is used to indicate that the second reply content conforms to the language characteristics of the second object.
[0140] Optionally, the adjustment module 403 is specifically used to calculate the degree of influence of the content output by the second model on the evaluation index according to a second reward function, and adjust the parameters of the second model based on the degree of influence; wherein, the second reward function is a reward function used by the reinforcement learning strategy constructed based on the second GPT model.
[0141] The human-computer interaction training device based on reinforcement learning strategy provided in the present application first builds a first model for simulating user interaction and a second model for simulating robot interaction. Then, the parameters of the second model are adjusted based on the reinforcement learning strategy, so that the second model becomes an interaction model for achieving specific goals, and then a chat robot is created based on the interaction model to increase the probability of achieving the above-mentioned specific goals when interacting with users.
[0142] Figure 5 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communication interface 520 and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute the human-computer interaction training method based on the reinforcement learning strategy, the method comprising: obtaining a first model trained with a target sample set as a training sample; the target sample set includes the interaction content of multiple interaction processes; constructing a second model, using the second model and the first model to simulate the instant communication interaction process; in the interaction process between the second model and the first model, the second model outputs the reply content, and based on the influence of the reply content output by the second model on the evaluation index of the interaction process, the parameters of the second model are adjusted; the second model after the parameter optimization is determined as the target model; wherein the evaluation index is used to indicate the probability that the interaction process can achieve the preset goal.
[0143] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art, and the computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0144] On the other hand, the present application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the human-computer interaction training method based on the reinforcement learning strategy provided by the above-mentioned methods, and the method includes: obtaining a first model trained with a target sample set as a training sample; the target sample set includes interaction content of multiple interaction processes; constructing a second model, and using the second model and the first model to simulate an instant communication interaction process; during the interaction process between the second model and the first model, the second model outputs reply content, and adjusts the parameters of the second model based on the degree of influence of the reply content output by the second model on the evaluation index of the interaction process; the second model after parameter optimization is determined as the target model; wherein the evaluation index is used to indicate the probability that the interaction process can achieve a preset goal.
[0145] On the other hand, the present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the above-mentioned human-computer interaction training methods based on reinforcement learning strategies, the methods comprising: obtaining a first model trained with a target sample set as a training sample; the target sample set includes interaction content of multiple interaction processes; constructing a second model, and using the second model and the first model to simulate an instant messaging interaction process; during the interaction process between the second model and the first model, the second model outputs reply content, and adjusts the parameters of the second model based on the degree of influence of the reply content output by the second model on the evaluation index of the interaction process; determining the second model after parameter optimization as the target model; wherein the evaluation index is used to indicate the probability that the interaction process can achieve a preset goal.
[0146] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0147] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0148] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A human-computer interaction training method based on reinforcement learning strategy, characterized in that: include: Obtain a first model trained using the target sample set as training samples; The target sample set includes interaction contents of multiple interaction processes; Constructing a second model, and using the second model and the first model to simulate an instant messaging interaction process; During the interaction between the second model and the first model, the second model outputs reply content, and based on the influence of the reply content output by the second model on the evaluation index of the interaction process, the parameters of the second model are adjusted; Determine the second model after parameter optimization as the target model; The evaluation index is used to indicate the probability that the interaction process can achieve user delegation; The adjusting the parameters of the second model based on the influence of the reply content output by the second model on the evaluation index of the interaction process includes: In the case where the second model is a generative model, the influence of the content output by the second model on the evaluation index is calculated according to the second reward function, and the parameters of the second model are adjusted based on the influence; the second reward function includes: a perplexity for measuring the quality of a probability distribution or a probability model predicting a sample, a mutual information for measuring the generation of meaningful replies, a repetition penalty item for reducing the model generating repeated historical dialogue information, a parameter for increasing the number of dialogue rounds, a parameter for detecting whether the user's words are transferred, and a parameter for calculating the user's transfer probability; the parameter of the user's transfer probability is calculated by the third model; The third model calculates the probability of user delegation based on the following steps: Segmenting the conversation content generated by the simulated instant messaging interaction and inputting it into the Bert model to obtain a sentence vector for each sentence; Input the sentence vector of each sentence into the long short-term memory network to obtain the context vector of each sentence; The context vector of each sentence and the influence of each sentence on the probability of user delegation are comprehensively calculated to obtain the probability of delegation corresponding to each sentence; Among them, the influence of each sentence on the probability of user delegation is: obtained based on the content output by the sentence corresponding model; the sentence corresponding model includes: the first model, or the second model.
2. The method according to claim 1, characterized in that The step of obtaining a first model obtained by training with the target sample set as training samples includes: Using the target sample set as training samples to train a first GPT model, and obtaining the first model; Among them, each sample in the training samples of the first GPT model includes category information; the category information is used to classify the interactive content of the sample.
3. The method according to claim 1, characterized in that The step of constructing the second model and using the second model and the first model to simulate the instant messaging interaction process includes: Constructing the second model, and guiding the second model to start interacting with the first model based on the initial interaction content through the initial interaction content; The reply content output by the second model is used as the input of the first model, and the content output by the first model is used as the input of the second model, so as to realize the simulated instant communication interaction between the second model and the first model.
4. The method according to claim 3, characterized in that: The constructing of the second model comprises: Pre-training a second GPT model using the target sample set as training samples to obtain the second model; Among them, each sample in the training samples of the second GPT model includes first object information and scene information; the first object information is used to indicate the first object corresponding to the interactive content of the sample; the scene information is used to indicate the application scenario to which the interactive content of the sample belongs.
5. The method according to claim 4, characterized in that During the interaction between the second model and the first model, the second model outputs reply content, and based on the degree of influence of the reply content output by the second model on the evaluation index of the interaction process, adjusting the parameters of the second model includes: In the simulated instant messaging interaction process between the second GPT model and the first model, the content output by the first model is used as the input of the second GPT model, and the second reply content generated by the second GPT model is obtained; The second reply content includes second object information; the second object information is used to indicate that the second reply content conforms to the language characteristics of the second object.
6. A human-computer interaction training device based on reinforcement learning strategy, characterized in that: The device comprises: An acquisition module, used to acquire a first model trained with a target sample set as a training sample; the target sample set includes interaction contents of multiple interaction processes; A construction module, used to construct a second model, and use the second model and the first model to simulate the instant messaging interaction process; an adjustment module, configured to adjust parameters of the second model based on the influence of the reply content output by the second model on the evaluation index of the interaction process when the second model outputs reply content during the interaction between the second model and the first model; A determination module, used for determining the second model after parameter optimization as the target model; The evaluation index is used to indicate the probability that the interaction process can achieve user delegation; The adjustment module is specifically used to calculate the influence of the content output by the second model on the evaluation index according to the second reward function when the second model is a generation model, and adjust the parameters of the second model based on the influence; the second reward function includes: a perplexity for measuring the quality of a probability distribution or a probability model predicting a sample, a mutual information for measuring the generation of meaningful replies, a repetition penalty item for reducing the model generating repeated historical dialogue information, a parameter for increasing the number of dialogue rounds, a parameter for detecting whether the user's words are transferred, and a parameter for calculating the user's transfer probability; the parameter of the user's transfer probability is calculated by the third model; The adjustment module is further specifically used to segment the conversation content generated by the simulated instant messaging interaction and input it into the Bert model to obtain a sentence vector for each sentence; input the sentence vector of each sentence into a long short-term memory network to obtain a context vector for each sentence; comprehensively calculate the context vector of each sentence and the impact of each sentence on the user's delegation probability to obtain the delegation probability corresponding to each sentence; wherein the impact of each sentence on the user's delegation probability is: obtained based on the content output by the sentence corresponding model; the sentence corresponding model includes: the first model, or, the second model.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the human-computer interaction training method based on the reinforcement learning strategy as described in any one of claims 1 to 5 are implemented.
8. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the human-computer interaction training method based on reinforcement learning strategy described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Model training method, dialogue generation method and device, equipment and medium
CN110188182A
Multi-model training method and device, electronic equipment and storage medium
CN112541570A