A Task-Oriented Dialogue Strategy Learning Method Incorporating User Satisfaction

Through a task-oriented dialogue strategy learning method that integrates user satisfaction, combined with deep reinforcement learning and emotional strategy modules, the problem of ignoring user experience in the dialogue strategy optimization process in the existing technology is solved, and a more efficient and user-friendly dialogue agent is achieved.

CN115344667BActive Publication Date: 2025-05-27SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210738419.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-27
Publication Date
2025-05-27
Estimated Expiration
2042-06-27

AI Technical Summary

Technical Problem

The existing dialogue strategy model based on deep reinforcement learning ignores the user's user experience and emotions in the process of optimizing the dialogue completion efficiency, resulting in dialogue agents that may ignore user requests.

Method used

A task-oriented dialogue strategy learning method that integrates user satisfaction is proposed. By collecting and preprocessing human-computer dialogue data, extracting intention, slot value and emotional state information, constructing a dialogue strategy module and emotional strategy module aimed at task completion efficiency and user satisfaction, and comprehensively assessing the action value and emotional value of candidate response actions through the weighted fusion module.

Benefits of technology

It effectively improves the user experience of the conversation agent, ensures that the conversation strategy takes into account efficiency and user satisfaction, avoids the problem of a single strategy, and improves the user stickiness of the conversation agent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115344667B_ABST
    Figure CN115344667B_ABST
Patent Text Reader

Abstract

The present invention discloses a task-oriented dialogue strategy learning method integrating user satisfaction. The method comprises the following steps: collecting human-computer dialogue data and performing data cleaning in combination with task scenarios; extracting the intention, slot value and emotional state information in the dialogue and performing vectorized representation; constructing a dialogue strategy module; constructing an emotional strategy module; constructing a weighted fusion module to obtain the total score of the aggregated action value and the action emotional value corresponding to the candidate response action, and predicting the response action based on the total score; obtaining the dialogue state, reward and user's real emotional state information after the predicted response action, and optimizing the network parameters of the dialogue strategy module and the emotional strategy module. The present invention fully considers the characteristics of dialogue and emotional state, and by integrating deep reinforcement learning and supervised learning technology, takes into account both dialogue efficiency and user satisfaction goals, and improves the effect of the dialogue strategy model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of human-computer dialogue, and particularly to a task-oriented dialogue strategy learning method integrating user satisfaction. Background Art

[0002] With the rapid development of technologies such as big data, cloud computing, and artificial intelligence, digitization, networking, and intelligence penetrate all aspects of the economic society and promote the transformation and upgrading of related industries. Under the strategic background of building a digital China, the research on natural language processing technology has advanced by leaps and bounds and has been widely applied in actual production and life. Among them, the dialogue system, as an important interface for human-computer interaction, occupies a key position in the research of the field of natural language processing and is an indispensable link for realizing digital and intelligent life and services. As a basic technology that has a significant impact on human production and life in the information age, the intelligent human-computer dialogue system can assist users in solving common problems in the form of dialogue, improve the convenience of services, and reduce service costs. Therefore, the strategy model of the task-oriented dialogue agent has relatively wide practical application value.

[0003] The goal of the new generation of intelligent human-computer dialogue agents is to make human-computer dialogue as efficient, convenient, and natural as human-human dialogue. Therefore, such systems must have a certain cognitive ability to identify and mine the emotions and preferences of users from the dialogue and provide more personalized responses and services. Establishing an emotional bond between humans and machines through natural communication with users helps to improve the user stickiness of the dialogue agent. However, the current dialogue strategy models based on deep reinforcement learning (such as DQN, DRQN, etc.) usually focus on the completion efficiency of the dialogue and do not consider the user experience (such as user emotions). This often leads to the situation that the dialogue agent is prone to learning some shortcuts in task-oriented dialogue and ignoring some requests of users during the dialogue.

[0004] In the prior art, a dialogue strategy method for a task-oriented dialogue system has the following specific defects:

[0005] The current dialogue strategy models based on deep reinforcement learning (such as DQN, DRQN, etc.) usually focus on the completion efficiency of the dialogue and do not consider the user experience (such as user emotions). This often leads to the situation that the dialogue agent is prone to learning some shortcuts in task-oriented dialogue and ignoring some requests of users during the dialogue. Summary of the Invention

[0006] The purpose of the present invention is to overcome the defects and deficiencies of the prior art and provide a task-oriented dialogue strategy learning method integrating user satisfaction.

[0007] The purpose of the present invention is achieved by at least one of the following technical solutions.

[0008] A task-oriented dialogue strategy learning method integrating user satisfaction, comprising the following steps:

[0009] S1. Collect human-machine dialogue data and perform data cleaning in combination with the task scenario;

[0010] S2. Preprocess the human-machine dialogue data after data cleaning, extract the intention, slot value and emotion state information in the dialogue, and perform vectorized representation;

[0011] S3. Construct a dialogue strategy module with the task completion efficiency as the goal, and evaluate the action value of candidate response actions;

[0012] S4. Construct an emotion strategy module with the user satisfaction as the goal, and evaluate the emotion value of candidate response actions;

[0013] S5. Construct a weighted fusion module, obtain the total score of the aggregated action value and the action emotion value corresponding to the candidate response action, and predict the response action according to the total score;

[0014] S6. Obtain the dialogue state, reward and user's true emotion state information after executing the response action predicted in step S5, and optimize the network parameters of the dialogue strategy module and the emotion strategy module.

[0015] Furthermore, step S1 specifically includes the following steps:

[0016] S1.1. Collect human-machine dialogue data from existing human-machine dialogue systems and publicly available task-oriented dialogue datasets;

[0017] S1.2. Clean the collected human-machine dialogue data according to the task scenario, and filter out the human-machine dialogue data samples with information missing and noise.

[0018] Furthermore, step S2 specifically includes the following steps:

[0019] S2.1. Adopt natural language processing tools to identify the user's intention in each round of dialogue from the cleaned human-machine dialogue data through deep learning-based semantic classification, and extract the corresponding dialogue semantic slots through deep learning-based named entity recognition technology;

[0020] S2.2. Identify the category and intensity of the user's emotion state in each round of dialogue from the cleaned human-machine dialogue data through a deep learning-based text emotion classification algorithm;

[0021] S2.3. Obtain the current dialogue state information through dialogue state tracking according to the user's intention and the slot value of the dialogue semantic slot obtained in step S2.1 And map the current dialogue state information and the category and intensity of the user's emotional state obtained in step S2.2 into a vectorized representation through the Lookup-Table to obtain the current dialogue state vector.

[0022] Furthermore, step S3 specifically includes the following steps:

[0023] S3.1. Define the reinforcement learning reward function r of the dialogue policy module as follows:

[0024]

[0025] Where L is the maximum dialogue length allowed by the human-machine dialogue system; the dialogue length is the maximum number of dialogue turns between the machine and the user; the reward function r will be used for the training of policy learning.

[0026] S3.2. Use the current dialogue state vector obtained in step S2.3 as the input of the dialogue policy module; among them, the dialogue policy module performs feature extraction and feature transformation through a linear network with a tanh activation function.

[0027] S3.3. The dialogue policy module predicts the action value of each candidate response action, and selects the K candidate response actions with the highest action value to form the candidate response action set A candidate , that is, the dialogue policy module outputs the K candidate response actions with the highest action value in the existing response action set according to the input current dialogue state vector to form the candidate response action set A candidate ; the candidate response actions are the set of response actions predicted by the system agent according to the dialogue, and it is a sorted set.

[0028] Furthermore, step S4 specifically includes the following steps:

[0029] S4.1. Map through the Lookup-Table to obtain each candidate response action in the candidate action set A candidate and the emotional state of the user at the current moment t corresponding candidate response action vector and the current emotional state vector, input the candidate response action vector, the current emotional state vector and the current dialogue state vector into the emotional policy module, and the emotional policy module predicts the user's emotional state at the next moment after executing the candidate response action

[0030] S4.2. Use the emotional utility function U to estimate the emotional value of each candidate response action, and the calculation is as follows:

[0031]

[0032] Among them, the Z(·) function is expressed as the difference between positive sentiment and negative sentiment, and the calculation method is as follows:

[0033]

[0034]

[0035] Among them, and are respectively the total intensities of the user's positive and negative emotions at time t. The subscript t represents the time information included in the state, and the superscripts pos and neg correspond to positive and negative sentiments respectively; and are respectively the total intensities of the predicted positive and negative emotions of the user at time t + 1, and are obtained according to the sentiment classification of the text.

[0036] Furthermore, in step S4.2, the sentiment classification method of the text includes text sentiment classification based on a sentiment dictionary.

[0037] Furthermore, step S5 specifically includes the following steps:

[0038] S5.1. For each candidate response action in the candidate response action set A candidate , calculate the total score Score of the candidate response action through a weighted fusion formula, specifically as follows:

[0039] Score = α × Q(s, a) + (1 - α) × U(s, a)

[0040] where α is a weight control parameter; s is the sentiment state, a is the candidate response action, Q(s, a) is the action value of adopting the candidate response action a in the sentiment state s, and U(s, a) is the sentiment value of adopting the candidate response action a in the sentiment state s;

[0041] S5.2. Compare the total scores of each candidate response action in the candidate response action set A candidate , and select the candidate response action with the highest total score as the output of the weighted fusion module, that is, the response action predicted by the weighted fusion module.

[0042] Furthermore, step S6 specifically includes the following steps:

[0043] S6.1. In each round of interaction, execute the response action predicted by the weighted fusion module, obtain the dialogue state information and the sentiment state data, that is, the true value of the sentiment state and use the dialogue state information after executing the response action and the true value of the sentiment state Stored in the experience replay pool;

[0044] S6.2. According to the dialogue state information after the response action and the true value of the emotional state Optimize the dialogue policy module and the emotion policy module, and continue the next round of dialogue after the update until the dialogue ends or exceeds the maximum number of dialogue rounds L set in S3.1.

[0045] Furthermore, in step S6.2, the network parameters of the dialogue policy module are updated by the Q-Learning algorithm.

[0046] Furthermore, in step S6.2, the parameters of the emotion policy module are updated by minimizing the mean squared error loss as follows:

[0047]

[0048] where and respectively represent the predicted value and the true value of the emotional state corresponding to the i-th data in the experience replay pool; N is the number of data in the experience replay pool, that is, the total number of samples.

[0049] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0050] 1. The task-oriented dialogue policy learning method integrating user satisfaction proposed by the present invention takes into account both the dialogue efficiency and the user satisfaction, overcomes the problem of single policy caused by the excessive optimization of the efficiency goal in the existing methods, and can effectively improve the usage experience of the dialogue agent.

[0051] 2. The emotional utility function proposed by the present invention can reasonably evaluate and quantify the emotional value of the dialogue agent's actions, and effectively support the application requirements of the human-computer dialogue system integrating user satisfaction.

[0052] 3. The present invention combines deep reinforcement learning and supervised regression techniques to divide and conquer the two goals of improving dialogue efficiency and improving user satisfaction, thereby reducing the state space dimension of the policy learning model, improving the learning effect and enhancing the efficiency of policy learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a flowchart of a task-oriented dialogue policy learning method integrating user satisfaction in an embodiment of the present invention;

[0054] Figure 2 is an overall framework diagram of a task-oriented dialogue policy learning method integrating user satisfaction in an embodiment of the present invention;

[0055] Figure 3This is a flowchart of the steps for preprocessing the human-machine dialogue data after data cleaning in the embodiments of the present invention. Detailed implementation manners

[0056] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present invention will be further described in detail below with reference to the drawings and embodiments.

[0057] Embodiment 1:

[0058] A task-oriented dialogue strategy learning method integrating user satisfaction, as Figure 1 and Figure 2 shown, includes the following steps:

[0059] S1. Collect human-machine dialogue data and perform data cleaning in combination with the task scenario, specifically including the following steps:

[0060] S1.1. Collect human-machine dialogue data from existing human-machine dialogue systems and publicly available task-oriented dialogue data sets;

[0061] S1.2. Clean the collected human-machine dialogue data according to the task scenario, and filter out human-machine dialogue data samples with missing information and noise. In this embodiment, the human-machine dialogue system data includes data such as ordering meals and booking cars, and the filtered data includes duplicate data and abnormal data.

[0062] S2. Preprocess the human-machine dialogue data after data cleaning, extract the intention, slot value, and emotional state information in the dialogue, and perform vectorization representation, as Figure 3 shown, specifically including the following steps:

[0063] S2.1. In this embodiment, a natural language processing tool is used to identify the intention of the user in each round of dialogue from the cleaned human-machine dialogue data through deep learning-based semantic classification, and extract the corresponding dialogue semantic slots through deep learning-based named entity recognition technology;

[0064] S2.2. Through a deep learning-based text sentiment classification algorithm, identify the category and intensity of the user's emotional state in each round of dialogue from the cleaned human-machine dialogue data;

[0065] In this embodiment, the categories of the user's emotional state include happy, sad, angry, surprised, satisfied, grateful, dissatisfied, and no obvious emotion; the intensity range of the emotional state is set to [0, 1], and normalization is performed so that the sum of the emotional intensities is 1;

[0066] S2.3. According to the intention of the user and the slot value of the dialogue semantic slot obtained in step S2.1, obtain the current dialogue state information through dialogue state tracking And map the current dialogue state information and the category and intensity of the user's emotional state obtained in step S2.2 into a vectorized representation through a Lookup-Table to obtain the current dialogue state vector.

[0067] S3. Construct a dialogue policy module aiming at task completion efficiency to evaluate the action value of candidate response actions, which specifically includes the following steps:

[0068] S3.1. Define the reinforcement learning reward function r of the dialogue policy module as follows:

[0069]

[0070] Where L is the maximum dialogue length allowed by the human-machine dialogue system; the dialogue length is the maximum number of dialogue turns between the machine and the user; the reward function r will be used for policy learning training;

[0071] S3.2. As Figure 2 shown, use the current dialogue state vector obtained in step S2.3 as the input of the dialogue policy module; among them, the dialogue policy module performs feature extraction and feature transformation through a linear network with a tanh activation function;

[0072] S3.3. In this embodiment, the dialogue policy module predicts the action value of each candidate response action, and selects the 3 candidate response actions with the highest action value to form a candidate response action set A candidate , that is, the dialogue policy module outputs the 3 candidate response actions with the highest action value in the existing response action set according to the input current dialogue state vector to form a candidate response action set A candidate ; the candidate response action is a set of response actions predicted by the system agent according to the dialogue and is a sorted set.

[0073] In this embodiment, the dialogue policy module includes a linear network with a tanh activation function and an output layer network. The number of neurons in the linear network is set to 80, and the number of neurons in the output layer network is equal to the number of legal actions of the policy model. The dialogue policy module calculates the value of each legal action a in state s, denoted as Q(s, a).

[0074] S4. Construct an emotion policy module aiming at user satisfaction to evaluate the emotion value of candidate response actions, which specifically includes the following steps:

[0075] S4.1. As Figure 2 shown, obtain each candidate response action in the candidate action set A candidate and the user's emotional state at the current moment t through Lookup-Table mapping The corresponding candidate response action vector and the current emotional state vector are input into the emotion policy module together with the current dialogue state vector. The emotion policy module predicts the user's emotional state at the next moment after executing the candidate response action.

[0076] In this embodiment, the emotion policy module is composed of a fully connected network, and the number of its neurons is set to 80.

[0077] S4.2. Use the emotion utility function U to estimate the emotional value of each candidate response action, and the calculation is as follows:

[0078]

[0079] Among them, the Z(·) function represents the difference between positive emotion and negative emotion, and the calculation method is as follows:

[0080]

[0081]

[0082] Among them, and respectively represent the total intensities of the user's positive and negative emotions at time t. The subscript t represents the time information included in the state, and the superscripts pos and neg correspond to positive and negative emotions respectively. and respectively represent the total intensities of the predicted positive and negative emotions of the user at time t + 1. and are obtained according to the emotional classification of the text.

[0083] The emotional classification method of the text includes text emotional classification based on an emotion dictionary.

[0084] S5. Construct a weighted fusion module to obtain the total score of the aggregated action value and the action emotional value corresponding to the candidate response action, and predict the response action according to the total score. Specifically, it includes the following steps:

[0085] S5.1. For each candidate response action in the candidate response action set A candidate calculate the total score Score of the candidate response action through the weighted fusion formula as follows:

[0086] Score = α × Q(s, a) + (1 - α) × U(s, a)

[0087] Among them, in this embodiment, α is a weight control parameter with a value of 0.7; s is the emotional state, a is the candidate response action, Q(s, a) is the action value of adopting the candidate response action a in the emotional state s, and U(s, a) is the emotional value of adopting the candidate response action a in the emotional state s;

[0088] S5.2. Compare the total scores of each candidate response action in the candidate response action set A candidate and select the candidate response action with the highest total score as the output of the weighted fusion module, that is, the response action predicted by the weighted fusion module.

[0089] S6. Obtain the dialogue state, reward, and user's true emotional state information after executing the response action predicted in step S5, and optimize the network parameters of the dialogue policy module and the emotional policy module, which specifically include the following steps:

[0090] S6.1. In each round of interaction, execute the response action predicted by the weighted fusion module to obtain the dialogue state information and the emotional state data, that is, the true value of the emotional state and store the dialogue state information after executing the response action and the true value of the emotional state in the experience replay pool;

[0091] S6.2. According to the dialogue state information after the response action and the true value of the emotional state, optimize the dialogue policy module and the emotional policy module, and continue the next round of dialogue after updating until the dialogue ends or exceeds the maximum number of dialogue rounds L set in S3.1;

[0092] Update the network parameters of the dialogue policy module through the Q-Learning algorithm.

[0093] Update the parameters of the emotional policy module by minimizing the mean square error loss as follows:

[0094]

[0095] Among them, and respectively represent the predicted value and the true value of the emotional state corresponding to the i-th data in the experience replay pool; N is the number of data in the experience replay pool, that is, the total number of samples.

[0096] Embodiment 2:

[0097] In this embodiment, the differences from Embodiment 1 are specifically as follows:

[0098] In step S1, collect human-machine dialogue data from real application scenarios, such as the customer service center data of a telecommunications company, and filter out human-machine dialogue data samples with missing information and noise.

[0099] In step S2, the categories of the user's emotional state include positive, negative, and no obvious emotion; the intensity range of the emotional state is set to [0, 1] and normalized;

[0100] In step S3.3, select the 5 actions with the highest action value to form the candidate response action set A candidate 。

[0101] The dialogue policy module includes a linear network with a tanh activation function and an output layer network. The number of neurons in the linear network is set to 100, and the number of neurons in the output layer network is equal to the number of legal actions of the policy model. Therefore, the dialogue policy module calculates the value of each legal action a in state s, denoted as Q(s, a).

[0102] In step S4, the emotion policy module is composed of a fully connected network, and the number of its neurons is set to 100;

[0103] In step S5, the value of the weight control parameter α is 0.8.

[0104] Embodiment 3:

[0105] In this embodiment, the differences from Embodiment 1 are as follows:

[0106] In step S1, collect human-machine dialogue data from real application scenarios, such as the customer service center data of a bank; filter out duplicate, missing, and abnormal data samples, including the duplication of punctuation marks.

[0107] In step S2, identify the user's intention in each round of dialogue through deep learning-based semantic classification from the cleaned human-machine dialogue data, and extract the corresponding dialogue semantic slots through IOB (In-out-Begin) technology;

[0108] The categories of the user's emotional state include positive, negative, and no obvious emotion; the intensity range of the emotional state is set to [0, 1] and normalized;

[0109] In step S3.3, select the 10 actions with the highest action value to form the candidate response action set A candidate 。

[0110] The dialogue policy module includes a linear network with a tanh activation function and an output layer network. The number of neurons in the linear network is set to 60, and the number of neurons in the output layer network is equal to the number of legal actions of the policy model. Therefore, the dialogue policy module calculates the value of each legal action a in state s, denoted as Q(s, a).

[0111] In step S4, the emotion policy module is composed of a fully connected network, and the number of its neurons is set to 60;

[0112] In step S5, the value of the weight control parameter α is 0.9.

[0113] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various equivalent transformations, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalent scope.

Claims

1. A task-oriented dialogue strategy learning method integrating user satisfaction, characterized in that, it includes the following steps: S1. Collect human-machine dialogue data and perform data cleaning in combination with the task scenario; S2. Preprocess the human-machine dialogue data after data cleaning, extract the intent, slot value, and emotional state information in the dialogue, and perform vectorization representation; S3. Construct a dialogue strategy module with the goal of task completion efficiency, and evaluate the action value of candidate response actions; S4. Construct an emotional strategy module with the goal of user satisfaction, and evaluate the emotional value of candidate response actions, specifically including the following steps: S4.

1. Map through the Lookup-Table to obtain the candidate action set A candidate for each candidate response action in it and the user's emotional state at the current moment t the corresponding candidate response action vector and the current emotional state vector, input the candidate response action vector, the current emotional state vector, and the current dialogue state vector into the emotion policy module, and the emotion policy module predicts the user's emotional state at the next moment after executing this candidate response action S4.

2. Use the emotional utility function U to estimate the emotional value of each candidate response action; S5. Construct a weighted fusion module, obtain the total score of the aggregated action value and the action emotional value corresponding to the candidate response action, and predict the response action according to the total score, specifically including the following steps: S5.

1. For each candidate response action in the candidate response action set A candidate calculate the total score Score of the candidate response action through the weighted fusion formula; S5.

2. Compare the total scores of each candidate response action in candidate response action set A candidate and select the candidate response action with the highest total score as the output of the weighted fusion module, that is, the response action predicted by the weighted fusion module; S6. Obtain the dialogue state, reward, and user's true emotional state information after executing the response action predicted in step S5, and optimize the network parameters of the dialogue strategy module and the emotional strategy module, including the following steps: S6.

1. In each round of interaction, execute the response action predicted by the weighted fusion module, and obtain the dialogue state information after executing the response action and the emotional state data, i.e., the true value of the emotional state and store the dialogue state information after executing the response action and the true value of the emotional state in the experience replay pool; S6.

2. Optimize the dialogue strategy module and the emotion strategy module according to the dialogue state information after the response action and the true value of the emotion state Then continue the next round of dialogue after the update until the dialogue ends or exceeds the maximum number of dialogue rounds L set in S3.1 2. The task-oriented dialogue strategy learning method integrating user satisfaction according to claim 1, characterized in that, step S1 specifically includes the following steps: S1.

1. Collect human-machine dialogue data from existing human-machine dialogue systems and publicly available task-oriented dialogue datasets; S1.

2. Clean the collected human-machine dialogue data according to the task scenario, and filter out the human-machine dialogue data samples with missing information and noise.

3. The task-oriented dialogue strategy learning method integrating user satisfaction according to claim 1, characterized in that, step S2 specifically includes the following steps: S2.

1. Use natural language processing tools to identify the intent of the user in each round of dialogue from the cleaned human-machine dialogue data through deep learning-based semantic classification, and extract the corresponding dialogue semantic slots through deep learning-based named entity recognition technology; S2.

2. Through a deep learning-based text sentiment classification algorithm, identify the category and intensity of the user's emotional state in each round of dialogue from the cleaned human-machine dialogue data. S2.

3. Based on the user's intention and the slot values of the dialogue semantic slots obtained in step S2.1, obtain the current dialogue state information through dialogue state tracking And map the current dialogue state information and the category and intensity of the user's emotional state obtained in step S2.2 into a vectorized representation through a Lookup-Table to obtain the current dialogue state vector.

4. The task-oriented dialogue strategy learning method integrating user satisfaction according to claim 3, characterized in that, step S3 specifically includes the following steps: S3.

1. Define the reinforcement learning reward function r of the dialogue strategy module, specifically as follows: where L is the maximum dialogue length allowed by the human-machine dialogue system; the dialogue length is the maximum number of dialogue rounds between the machine and the user; the reward function r will be used for the training of policy learning; S3.

2. Use the current dialogue state vector obtained in step S2.3 as the input of the dialogue policy module; among them, the dialogue policy module performs feature extraction and feature transformation through a linear network with a tanh activation function; S3.

3. Predict the action value of each candidate response action by the dialogue policy module, and select the top K candidate response actions with the highest action value to form the candidate response action set A candidate , that is, the dialogue policy module outputs the top K candidate response actions with the highest action value in the existing response action set according to the input current dialogue state vector to form the candidate response action set A candidate ; The candidate response action is the set of response actions predicted by the system agent according to the dialogue, and it is a sorted set.

5. The task-oriented dialogue strategy learning method integrating user satisfaction according to claim 1, characterized in that, in step S4.2, the text sentiment classification method includes text sentiment classification based on sentiment dictionaries.

6. The task-oriented dialogue strategy learning method integrating user satisfaction according to claim 1, characterized in that, in step S4.2, use the emotional utility function U to estimate the emotional value of each candidate response action, and the calculation is: Among them, the Z(·) function is expressed as the difference between positive sentiment and negative sentiment, and the calculation method is as follows: wherein, and are the total intensities of the user's positive and negative emotions at time t, respectively. The subscript t represents the time information included in the state, and the superscripts pos and neg correspond to positive and negative emotions, respectively; and are the total intensities of the predicted positive and negative emotions of the user at time t + 1, respectively, and are obtained according to the sentiment classification of the text.

7. The task-oriented dialogue strategy learning method integrating user satisfaction according to claim 1, characterized in that In step S5.1, for each candidate response action in the candidate response action set A candidate calculate the total score Score of the candidate response action through the weighted fusion formula as follows: Score = α × Q(s,a) + (1 - α) × U(s,a) where α is the weight control parameter; s is the sentiment state, a is the candidate response action, Q(s,a) is the action value of adopting the candidate response action a in the sentiment state s, and U(s,a) is the sentiment value of adopting the candidate response action a in the sentiment state s.

8. The task-oriented dialogue strategy learning method integrating user satisfaction according to claim 1, characterized in that In step S6.2, the network parameters of the dialogue strategy module are updated through the Q-Learning algorithm.

9. The task-oriented dialogue strategy learning method integrating user satisfaction according to claim 1, characterized in that In step S6.2, by minimizing the mean square error loss Update the parameters of the sentiment policy module as follows: Among them, and respectively represent the predicted value and the true value of the emotional state corresponding to the i-th data in the experience replay pool; N is the number of data in the experience replay pool, that is, the total number of samples.

Citation Information

Patent Citations

  • Task-oriented dialogue strategy learning method for user personality perception

    CN114611527A

  • Reinforcement Learning Techniques for Dialogue Management

    US20220108080A1