Target guidance dialogue generation method based on joint strategy driving
By introducing a joint strategy-driven goal guidance method in the active dialogue technology, combining the directional policy network and the guidance policy network, and updating network parameters using reinforcement learning algorithms, the problem of low goal orientation and achievement efficiency in the dialogue system in the existing technology is solved, and the directional and fine-grained guidance of dialogue is achieved.
Patent Information
- Application Number
- CN202510186031.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-20
AI Technical Summary
Existing active dialogue technologies are difficult to generalize in new situations and lack fine-grained guidance in specific contexts, resulting in low goal orientation and achievement efficiency of dialogue systems.
A target-guided dialogue generation method based on joint strategy is adopted, and the directional policy network and the guidance policy network are combined, and the reinforcement learning algorithm is used to update network parameters to achieve directional and fine-grained guidance of dialogue.
It improves the goal orientation and achievement efficiency of the dialogue system, ensures directionality and fine-grained guidance of the dialogue, and is suitable for a variety of dialogue tasks.
Smart Images

Figure CN120124642A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of active dialogue, and more specifically, to a goal-guided dialogue generation method based on joint strategy drive. Background Art
[0002] Currently, in active dialogue tasks, in order to simulate the behavior of human experts, traditional methods usually adopt corpus-based fine-tuning to predict dialogue strategies. However, these methods rely on a large amount of labeled data and are difficult to generalize to new situations. Recently, some studies have prompted large language models to perform self-thinking for the next round of strategies, but have ignored long-term dialogue goals. Fu et al. adopted self-play to simulate iterative optimization of strategy planning, but this method is only applicable to individual cases. The PPDPP method introduced a learnable language model plugin, which can improve the strategy planning ability by fine-tuning the plugin and assist the large language model to achieve dialogue goals, but it lacks effective guidance for the fine-grained expected behavior.
[0003] Therefore, how to provide a dialogue generation method that can not only ensure the directionality of the dialogue, but also pay attention to the fine-grained guidance in a specific context environment, and improve the goal orientation and achievement efficiency of the dialogue system is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a goal-guided dialogue generation method based on joint strategy drive.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions:
[0006] A goal-guided dialogue generation method based on joint strategy drive, comprising the following steps:
[0007] S1: Input the historical dialogue s t into the direction strategy network to obtain the direction strategy v t ; wherein, the historical dialogue s t = {(o 1 , u 1 ), (o 2 , u 2 ),...,(o t-1 , u t-1 ); t≥2; o t-1 represents the output corpus of the system dialogue model LLM sys in the (t - 1)-th round of dialogue; u t-1 represents the output corpus of the user role model LLM user in the (t - 1)-th round of dialogue; wherein, the direction strategy network Denote the direction policy base model after supervised fine-tuning;
[0008] S2: Input the historical dialogue s t and the direction policy v t into the guiding policy network to obtain the guiding keyword w t ; where the guiding policy network denotes the guiding policy base model after supervised fine-tuning;
[0009] S3: Input the historical dialogue s t , the direction policy v t and the guiding keyword w t into the system dialogue model LLM sys to obtain the output corpus o sys of the system dialogue model LLM t in the t-th round of dialogue;
[0010] S4: Input the historical dialogue s t and the output corpus o t into the user role model LLM user to obtain the output corpus u user of the user role model LLM t in the t-th round of dialogue;
[0011] S5: Update the network parameters of the direction policy network and the guiding policy network using the reinforcement learning algorithm to obtain the direction policy network and the guiding policy network
[0012] S6: Continuously repeat the above steps until the preset termination condition is met.
[0013] Preferably, the direction policy network is obtained based on the following steps:
[0014] Obtain the direction policy dialogue dataset D', where D' = {(s i , v i )}; i = 1, 2,.., N'; (s i , v i ) represents the i-th direction policy data pair in the direction policy dialogue dataset D'; s i represents the historical dialogue included in the i-th direction policy data pair; v i represents the direction policy v i corresponding to the historical dialogue s i; N' represents the number of direction policy data pairs included in the direction policy dialogue dataset D'.
[0015] Train the direction policy base model using the direction policy dialogue dataset D', and adjust the network parameters of the direction policy base model by maximizing the first log-likelihood function to obtain the direction policy network
[0016] Preferably, the expression of the first log-likelihood function is:
[0017]
[0018] where p P (v i |s i ) represents inputting the historical dialogue s i into the direction policy base model, and the direction policy base model outputs the probability of the direction policy v i ; log represents taking the logarithm;
[0019] represents taking the average of N' log p P (v i |s i ).
[0020] Preferably, the guidance policy network is obtained based on the following steps:
[0021] Obtain the guidance policy dialogue dataset D”; where, represents the j-th guidance policy data pair in the guidance policy dialogue dataset D”; s j represents the historical dialogue included in the j-th guidance policy data pair; v j represents the direction policy corresponding to the historical dialogue s j ; represents the guidance keyword corresponding to the historical dialogue s j and the direction policy v j ; N” represents the number of guidance policy data pairs included in the guidance policy dialogue dataset D”.
[0022] Train the guidance policy base model using the guidance policy dialogue dataset D”, and adjust the network parameters of the guidance policy base model by maximizing the second log-likelihood function to obtain the guidance policy network
[0023] Preferably, the guidance keyword is obtained based on the following steps:
[0024] Obtain the historical dialogue s j The user response corpus in the last turn of the dialogue;
[0025] Use the vlt5-base-keywords model to extract keywords from the user response corpus to obtain the guiding keywords
[0026] Preferably, the expression of the second log-likelihood function is:
[0027]
[0028] wherein, represents inputting the historical dialogue s j and the direction strategy v j into the guiding strategy base model, and the guiding strategy base model outputs the probability of the guiding keywords ; log represents taking the logarithm; E represents taking the average of N” ; p(k m |k 1 ,...,k m-1 ; s j ,v j ) represents the probability that the guiding strategy base model generates the mth token under the condition of the known historical dialogue s j , the direction strategy v j and the first m - 1 tokens, and k 1 ,...,k m-1 ,k m represent tokens, which together constitute the guiding keywords n represents the number of tokens included in the guiding keywords in.
[0029] Preferably, S5 specifically includes the following steps:
[0030] S51: Input the output corpus o t and the output corpus u t into the evaluation reward model LLM critic to obtain the reward value r t of the t-th turn of the dialogue;
[0031] S52: Maximize the reward to update the network parameters of the direction strategy network and the guiding strategy network to obtain the direction strategy network and the guiding strategy network wherein, represents finding the guiding strategy network and the guiding policy network ; β represents a penalty coefficient.
[0032] Preferably, the following steps are repeatedly executed M times: inputting the output corpus o t and the output corpus u t into the evaluation reward model LLM critic to obtain M reward values;
[0033] Calculate the average value of the M reward values to obtain the reward value r of the t-th round of dialogue t .
[0034] Preferably, the direction policy base model is RoBERTa; the guiding policy base model is GPT-2.
[0035] Preferably, the user role model LLM user , the system dialogue model LLM sys and the evaluation reward model LLM critic are all simulated using ChatGPT.
[0036] Through the above technical solutions, it can be seen that compared with the prior art, the present invention discloses a target-guided dialogue generation method based on joint policy drive, which can not only ensure the directionality of the dialogue, but also pay attention to the fine-grained guidance in a specific context, improving the goal orientation and achievement efficiency of the dialogue system. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0038] Figure 1 is a flowchart of a target-guided dialogue generation method based on joint policy drive provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0040] As Figure 1As shown in the figure, an embodiment of the present invention discloses a method for generating a target-guided dialogue based on a joint strategy drive, including the following steps:
[0041] S1: Input the historical dialogue s t into the direction policy network to obtain the direction policy v t ; where the historical dialogue s t ={(o 1 , u 1 ), (o 2 , u 2 ),...,(o t-1 , u t-1 ); t≥2; o t-1 represents the output corpus of the system dialogue model LLM sys in the (t - 1)-th round of dialogue; u t-1 represents the output corpus of the user role model LLM user in the (t - 1)-th round of dialogue; where the direction policy network represents the direction policy base model after supervised fine-tuning;
[0042] It can be understood that: the direction policy network represents the direction policy network obtained in the previous S5;
[0043] In one embodiment, the direction policy base model adopts the Large version of RoBERTa;
[0044] The present invention uses ChatGPT (gpt-3.5-turbo) to simulate the two roles of the user role model LLM user and the system dialogue model LLM sys .
[0045] In one embodiment, the direction policy network is obtained based on the following steps:
[0046] Obtain the direction policy dialogue dataset D', where D' ={(s i , v i )}; i = 1, 2,.., N'; (s i , v i ) represents the i-th direction policy data pair in the direction policy dialogue dataset D'; s i represents the historical dialogue included in the i-th direction policy data pair; v i represents the direction policy v i corresponding to the historical dialogue s i ; N' represents the number of direction policy data pairs included in the direction policy dialogue dataset D';
[0047] Training the direction policy base model using the direction policy dialogue dataset D', and adjusting the network parameters of the direction policy base model by maximizing the first log-likelihood function to obtain the direction policy network
[0048] It can be understood that: for different target tasks, the specific direction policy dialogue dataset D' used is also different. If the dialogue task is emotional support, the training data and validation data in the ESConv dataset (specifically refer to Liu S, Zheng C, Demasi O, et al. Towards Emotional Support Dialog
[0049] Systems[C] / / Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021: 3469-3483.
[0050] ) can be used as the direction policy dialogue dataset D'; if the dialogue task is buying and selling dialogue negotiation, the improved CraiglistBargain dataset (specifically refer to Yang R, Chen J, Narasimhan K. Improving Dialog Systems for Negotiation with Personality Modeling[C] / / Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021: 681-693) can be used
[0051] The training data and validation data therein (in the present invention, the improved CraiglistBargain dataset is divided into training data, validation data, and test data in a ratio of 6:2:2) are used as the direction policy dialogue dataset D'.
[0052] In one embodiment, the expression of the first log-likelihood function is:
[0053]
[0054] where p P (v i |s i ) represents the probability that the historical dialogue s i is input into the direction policy base model, and the direction policy base model outputs the direction policy v i ; log represents taking the logarithm;
[0055] represents taking the average of N' logp P (v i |s i ).
[0056] S2: Input the historical dialogue s t and the direction policy v t into the guiding policy network to obtain the guiding keyword w t ; where the guiding policy network represents the guiding policy base model after supervised fine-tuning;
[0057] It can be understood that: the guiding policy network represents the guiding policy network obtained in the previous round of S5;
[0058] In one embodiment, the guiding policy base model uses GPT-2;
[0059] In one embodiment, the guiding policy network is obtained based on the following steps:
[0060] Obtain the guiding policy dialogue dataset D”; where represents the jth guiding policy data pair in the guiding policy dialogue dataset D”; s j represents the historical dialogue included in the jth guiding policy data pair; v j represents the direction policy corresponding to the historical dialogue s j ; represents the historical dialogue s jGuiding keywords corresponding to the direction strategy vj; "N" represents the number of guiding strategy data pairs included in the guiding strategy dialogue dataset D".
[0061] Use the guiding strategy dialogue dataset D" to train the guiding strategy base model, and maximize the second log-likelihood function to adjust the network parameters of the guiding strategy base model to obtain the guiding strategy network
[0062] In a certain embodiment, the guiding keywords are obtained based on the following steps:
[0063] Obtain the user response corpus in the last turn of the historical dialogue s j ;
[0064] Use the vlt5-base-keywords model to extract keywords from the user response corpus to obtain the guiding keywords
[0065] It can be understood that: for different target tasks, the specific guiding strategy dialogue dataset D" used is also different. If the dialogue task is emotional support, keyword extraction can be performed based on the data in the ESConv dataset to construct the guiding strategy dialogue dataset D"; if the dialogue task is a buying and selling dialogue negotiation, keyword extraction can be performed based on the data in the improved CraiglistBargain dataset to construct the direction strategy dialogue dataset D'.
[0066] In a certain embodiment, the expression of the second log-likelihood function is:
[0067]
[0068] where represents inputting the historical dialogue s j and the direction strategy v j into the guiding strategy base model, and the guiding strategy base model outputs the probability of the guiding keyword ; log represents taking the logarithm; E represents taking the average of N" ; p(k m |k 1 ,...,k m-1 ; s j ,v j ) represents the probability that the guiding strategy base model generates the m-th token under the condition of the known historical dialogue s j , the direction strategy v j and the previous m - 1 tokens, k 1 ,...,km-1 , k m represents tokens that together constitute the guiding keyword n represents the number of tokens included in the guiding keyword .
[0069] S3: Input the historical dialogue s t , the direction strategy v t and the guiding keyword w t into the system dialogue model LLM sys to obtain the output corpus o of the system dialogue model LLM sys in the t-th round of dialogue t ;
[0070] S4: Input the historical dialogue s t and the output corpus o t into the user role model LLM user to obtain the output corpus u of the user role model LLM user in the t-th round of dialogue t ;
[0071] S5: Update the network parameters of the direction policy network and the guiding policy network using the reinforcement learning algorithm to obtain the direction policy network and the guiding policy network
[0072] In one embodiment, S5 specifically includes the following steps:
[0073] S51: Input the output corpus o t and the output corpus u t into the evaluation reward model LLM critic to obtain the reward value r of the t-th round of dialogue t ;
[0074] In one embodiment, the evaluation reward model LLM critic is simulated using ChatGPT
[0075] Specifically: By asking a selection question to the evaluation reward model LLM critic to generate an evaluation feedback result for a specific target. The specific evaluation feedback results include three types: making progress (score of 1), maintaining the status quo (score of 0), and regressing (score of -1). Further, through the mapping function fmap(·), the evaluation feedback result is mapped to a scalar as the final reward value
[0076] S52: Maximize the reward to update the direction policy network and the guidance policy network of the network parameters to obtain the direction policy network and the guidance policy network wherein, represents the divergence between the guidance policy network and the guidance policy network ; β represents the penalty coefficient.
[0077] In one embodiment, the following steps are repeatedly executed M times: input the output corpus o t and the output corpus u t into the evaluation reward model LLM critic to obtain M reward values;
[0078] Calculate the average value of the M reward values to obtain the reward value r of the t-th round of dialogue t .
[0079] S6: Continuously repeat the above steps until the preset termination condition is met.
[0080] It can be understood that the preset termination condition is to complete the dialogue goal or reach the preset maximum number of dialogue rounds.
[0081] To verify the effectiveness of the present invention, the present invention (named JSD) was compared with 6 baseline models, namely DiagoGPT, Vanilla, Proactive, ProCoT, AnE, ICL-AIF, and PPDPP, in terms of performance in two tasks: emotional support dialogue and buying and selling negotiation dialogue;
[0082] Task 1: Emotional support dialogue task
[0083] This task aims to reduce personal emotional pain and help people understand and solve the challenges they face. This task uses the ESConv dataset, which involves 1300 dialogues under 10 topic questions, and 8 support strategies (corresponding to the direction strategy of the present invention) are included in the dialogue process. Each dialogue is marked with question categories, emotional categories, and the basic situation of the applicant, etc. This task divides ESConv into a training set / validation set / test set according to a ratio of 6:2:2. During the emotional support dialogue, if the average value of the M reward values output by the current round of dialogue evaluation reward model LLM critic is greater than 0.5, it is considered that the current dialogue goal is completed.
[0084] Task 2: Buying and selling negotiation dialogue task
[0085] To further verify the effectiveness of the model in non - collaborative conversations (when the goals of the user and the system conflict), the present invention uses an improved CraiglistBargain dataset. This dataset is based on negotiation conversations about real - world goods and covers the negotiation process between buyers and sellers regarding the price of the goods. Each negotiation sample includes the category of the goods, the description of the goods, the buyer's target price, and the seller's target price. In this task, 11 marked negotiation strategies (corresponding to the directional strategies of the present invention) are selected, and the dataset is divided into a training set / validation set / test set in a ratio of 6:2:2.
[0086] All experiments and model training of the present invention are completed on a Ubuntu server equipped with 2 Nvidia A800 GPUs. The main environments used include Python 3.10, PyTorch 2.4, and CUDA 12.1. To ensure the determinacy of the output of the large - language model, the temperature is set to 0 in the experiment, and the evaluation sampling number M of the evaluation reward model LLM critic for each conversation is set to 10. During the reinforcement learning training process, the learning rate of the policy model is set to 2e - 5, the upper limit of the number of dialogue rounds is 7, and a total of 5 epochs are trained.
[0087] Table 1 shows the performance of the JSD of the present invention and 6 baseline models in two tasks: emotional support conversations and buying - selling negotiation conversations.
[0088] Table 1 Performance of the JSD of the present invention and 6 baseline models in two tasks: emotional support conversations and buying - selling negotiation conversations
[0089]
[0090] In Table 1:
[0091] 1) Boldface indicates the best result in each column, and underlining indicates the second - best result.
[0092] 2)
[0093] Among them, N represents the total number of conversations, t i represents the number of rounds used to achieve the goal in the i - th conversation, or the maximum number of rounds reached when the goal is not achieved; AT measures the efficiency of goal completion by calculating the average number of rounds to reach the goal, and the smaller the value, the better.
[0094] 3) SR% = SR@T = S T / N;
[0095] Among them, S Tdenotes the number of dialogue goals achieved before or equal to T rounds, where N is the total number of dialogues; SR@T measures the effectiveness of goal completion by calculating the success rate of achieving the goal within the predefined maximum number of rounds T, and the larger its value, the better.
[0096] 4)
[0097] where d p denotes the deal price at which the transaction is concluded, s p denotes the seller's target price, b p denotes the buyer's target price; SL% represents the proportion of the buyer's gain relative to the price expectations of both parties in the transaction. If the transaction fails to be concluded, SL% is set to 0.
[0098] As can be seen from Table 1:
[0099] On the ESConv dataset for the emotional support dialogue task, the JSD of the present invention performs best with a success rate of 93.85%, significantly higher than other models, indicating that it has the highest goal achievement rate. Compared with the second-best performing model PDPDP, the JSD of the present invention has improved by 9.23 percentage points, most effectively guiding the dialogue to reach the expected goal. In terms of the average number of rounds to reach the dialogue goal, the JSD of the present invention reaches the predefined goal with an average of 3.35 rounds. Compared with PDPDP, it reduces by 1.21 rounds, indicating that the JSD of the present invention reaches the goal most efficiently.
[0100] On the CraigslistBargain dataset for the buying and selling dialogue negotiation task, the JSD of the present invention also performs excellently, with a success rate of 82.45% significantly higher than other models, which indicates that the JSD of the present invention has the highest goal achievement rate in the dialogue negotiation task and successfully guides the dialogue to reach the expected goal. Compared with the other best model PDPDP, the JSD of the present invention has improved by 21.28 percentage points. In terms of the average number of rounds, the JSD of the present invention reaches the goal with an average of 4.00 rounds. Compared with PDPDP, the JSD of the present invention reduces by 1.62 rounds, also indicating that it reaches the goal more efficiently in multiple tasks. Although the JSD of the present invention (38.88%) is not the highest in the SL% index, its overall performance is still outstanding.
[0101] In addition, the present invention also conducts ablation experiments on the JSD provided by the present invention on ESConv and CraigslistBargain, and the analysis results are shown in Table 2;
[0102] Table 2 Analysis Results of the Ablation Experiment of the Present Invention
[0103]
[0104] As can be seen from Table 2: On the ESConv dataset, when the direction strategy module is removed (w / o SP), the performance of the variant model drops from 93.85% to 90.58% in terms of the goal achievement rate, and increases to 4.23 in terms of the average number of dialogue turns. This indicates that the direction strategy module makes a significant contribution to improving the goal achievement rate and dialogue efficiency. When the guiding strategy module is removed (w / o SG), the performance of the variant model drops to 84.62% in terms of the goal achievement rate, and increases to 4.56 in terms of the average number of dialogue turns, showing that the guiding strategy plays a key role in the fine-grained control of dialogue content. When the reinforcement learning module is removed (w / o RL), the performance of the variant model drops to 88.46% in terms of the goal achievement rate, and increases to 4.35 in terms of the average number of dialogue turns, indicating that reinforcement learning plays an important role in optimizing the overall strategy.
[0105] On the CraigslistBargain dataset, when the direction strategy module is removed, the goal achievement rate of the variant model drops from 82.45% to 58.51%, and the average number of dialogue turns increases to 5.17, which indicates the importance of the direction strategy in the negotiation of buying and selling conversations. When the guiding strategy module is removed, the goal achievement rate of the variant model drops to 79.79%, and the average number of dialogue turns slightly increases to 4.05, indicating that the guiding strategy also plays an important role in controlling the direction of the dialogue. When the reinforcement learning module is removed, the goal achievement rate of the variant model rises to 86.17%, the average number of dialogue turns decreases to 3.61, and the SL% reaches 42.79%. This result shows that in the absence of reinforcement learning, the model may tend to adopt more conservative and simplified strategies, avoiding complex negotiation processes and directly achieving the transaction goal.
[0106] Finally, the present invention demonstrates the effectiveness of the JSD of the present invention in guiding goal-oriented dialogues through cases from two aspects: the direction strategy and the guiding strategy.
[0107] 1) Case analysis of the direction strategy
[0108] To better analyze the capabilities of the JSD of the present invention in the dialogue guidance task, the present invention presents a complete case of an emotional support dialogue. As shown in Table 3, based on the dual-process theory, JSD provides decision-making support for the system assistant (i.e., the system dialogue model described in the present invention) in terms of countermeasure direction and guiding clues through joint strategy prompts, forming a path plan of dynamic programming. This plan gradually guides the dialogue towards the established goal, driving the model to actively advance in the target direction. In a specific case, the user (a patient with psychological problems) has negative emotions such as anxiety due to "personal health problems", and the system assistant (also known as the therapist) provides emotional treatment for the patient with psychological problems under the guidance of the joint strategy. Among them, the yellow highlight indicates the dialogue direction strategy and the corresponding generated content, and the red highlight indicates the specific guiding strategy and the corresponding generated content. In the presented case, with the assistance of the joint strategy, the model guides the user from an anxious state to a more positive emotional state through three conversations. In each dialogue turn, the system selects different strategies according to the patient's current state. In the first round, the therapist selects the strategy of "Affirmation and Reassurance", by affirming the patient's feelings and emphasizing the importance of health and well-being, the system guides the patient to turn their attention to their own health priorities and gradually soothe negative emotions. In the second round, the therapist adopts the strategy of "Reflection of Feelings". The therapist expresses understanding of the patient's sense of frustration and at the same time proposes that they can work together to find a solution, helping the patient see hope from the emotional dilemma. In the last round of the dialogue, the therapist adopts the strategy of "Providing Suggestions", promising to support the patient's every move and further guiding the patient to find more possibilities and a sense of hope when facing challenges.
[0109] Case Analysis of Direction Strategy Guiding Dialogue
[0110]
[0111] 2) Case Analysis of Guiding Strategy
[0112] In order to better analyze the ability of the guidance strategy network, the present invention presents a series of typical cases. As shown in Table 4, the JSD of the present invention can generate reasonable guidance keywords for each round of dialogue through joint strategy prompts. The model will generate specific dialogue details based on these keywords to guide the dialogue. Under the guidance of the pre-planned dialogue path, the system assistant (i.e., the system dialogue model of the present invention) can better plan the dialogue details and actively promote the dialogue according to the goal. In three specific cases, different users have negative emotions such as anxiety due to "remote work", "unemployment" and "emotional companionship". The guidance strategy is responsible for optimizing the specific details of the dialogue, so that the system assistant can more efficiently alleviate the patient's distress. Among them, the yellow highlight indicates the direction strategy of the dialogue and its corresponding generated content, and the red highlight indicates the specific guidance strategy and its corresponding generated content. In the displayed case, JSD uses different guidance keywords to enable the system to generate the content of the corresponding user dialogue in the current round. In the first dialogue case, the patient expressed negative feelings about remote work and believed that "remote work is disadvantageous". The therapist adopted the "Self-Disclosure" strategy and used "remote work" and "functional work" as guidance keywords. This strategy allows the therapist to express understanding of the patient's feelings, for example, "Feeling disconnected from work affects your work performance." This facilitation strategy helps the patient focus on how to maintain work performance, especially in a remote work environment, and provides constructive suggestions. In the second conversation case, the patient was experiencing ongoing anxiety and frustration due to unemployment. The therapist used the "Affirmation and Reassurance" strategy, using "job position" and "despair" as facilitating keywords. The therapist guided the patient to recognize the root of the emotions, such as "You feel desperate and overwhelmed by the lack of job opportunities." This strategy helps the patient express his or her emotions and seek support and advice through empathetic guidance. The third conversation case explored the loneliness and stress felt by the patient after the loss of his or her partner. The patient faced psychological difficulties due to the lack of emotional companionship. The therapist again used the Affirmation and Reassurance strategy, using the facilitating keywords "difficult" and "lonely." By using the facilitating phrases "It is normal to feel confused and overwhelmed," the therapist helped the patient recognize that his or her emotions were understandable and reminded the patient that he or she did not have to deal with difficulties alone. This strategy alleviated the patient's loneliness and provided emotional comfort through emotional support.
[0113] Table 4 Case analysis of facilitation strategies to guide dialogue
[0114]
[0115] In the present specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0116] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A goal-guided dialogue generation method based on joint strategy driving, characterized in that: The following steps are involved: S1: Historical dialogues t Input to the direction policy network Get direction strategy v t ; Wherein, the historical dialogue s t ={(o1,u1),(o2,u2),...,(o t-1 ,u t-1 );t≥2;o t-1 Represents the system dialogue model LLM in the t-1th round of dialogue sys The output corpus of u t-1 Represents the user role model LLM in the t-1th round of dialogue user The output corpus of the direction strategy network represents the basic model of the direction strategy after supervised fine-tuning; S2: The historical dialogue s t and the direction strategy v t Input to the guided policy network Get the guide keyword w t ; Among them, the guiding strategy network represents the basic model of the guidance strategy after supervised fine-tuning; S3: The historical conversation s t 、The direction strategy v t and the guiding keyword w t Input to the System Dialogue Model LLM sys , obtain the system dialogue model LLM in the tth round of dialogue sys The output corpus o t ; S4: The historical dialogue s t and the output corpus o t Input to the User Role Model LLM user , obtain the user role model LLM in the tth round of dialogue user The output corpus u t ; S5: Using reinforcement learning algorithm to update the direction strategy network and the boot strategy network The network parameters of the direction strategy network are obtained and guide strategy network S6: Repeat the above steps until the preset termination condition is met.
2. The method for generating a target-guided dialogue based on joint strategy drive according to claim 1, characterized in that: The Direction Strategy Network Based on the following steps: Get the direction strategy dialogue dataset D', where D' = {(s i ,v i )};i=1.2..N';(s i ,v i ) represents the i-th direction strategy data pair in the direction strategy dialogue data set D'; s i represents the historical dialogue included in the i-th direction strategy data pair; v i Indicates historical conversations i The corresponding direction strategy v i ; N' represents the number of direction strategy data pairs included in the direction strategy dialogue data set D'; The directional strategy dialogue data set D' is used to train the directional strategy basic model, and the network parameters of the directional strategy basic model are adjusted by maximizing the first log-likelihood function to obtain the directional strategy network 3. The method for generating a target-guided dialogue based on joint strategy drive according to claim 2, characterized in that: The expression of the first log-likelihood function is: Among them, p P (v i |s i ) means to convert historical dialogues i The direction strategy basic model outputs the direction strategy v i The probability of ; log means finding the logarithm; Indicates finding N' logp P (v i |s i )’s average value.
4. The method for generating a target-guided dialogue based on joint strategy drive according to claim 1, characterized in that: The guided strategy network Based on the following steps: Get the guidance strategy dialogue dataset D"; where, represents the jth guidance strategy data pair in the guidance strategy dialogue data set D”; s j represents the historical dialogue included in the j-th guidance strategy data pair; v j Indicates historical conversations j Corresponding direction strategy; Indicates historical conversations j and direction strategyv j corresponding guiding keywords; N" represents the number of guiding strategy data pairs included in the guiding strategy dialogue data set D"; The guiding strategy dialogue data set D" is used to train the guiding strategy basic model, and the network parameters of the guiding strategy basic model are adjusted by maximizing the second log-likelihood function to obtain the guiding strategy network 5. The method for generating a target-guided dialogue based on joint strategy drive according to claim 4, characterized in that: The guiding keywords Based on the following steps: Get the historical conversations j The user response corpus in the last round of dialogue; The vlt5-base-keywords model is used to extract keywords from the user response corpus to obtain the guide keywords 6. The method for generating a target-guided dialogue based on joint strategy drive according to claim 4, characterized in that: The expression of the second log-likelihood function is: in, Indicates that historical dialogues j and direction strategyv j Input to the guidance strategy basic model, the guidance strategy basic model outputs guidance keywords The probability of ; log means finding the logarithm; E means finding N” The average value of p(k m |k1,...,k m-1 ;s j ,v j ) indicates that in the known history conversations j 、Direction strategyv j The probability of the boot strategy base model generating the mth token under the condition of the first m-1 tokens, k1,...,k m-1 ,k m Represents a token, which together constitutes the guiding keyword n represents the guiding keyword The number of tokens included in .
7. The method for generating a target-guided dialogue based on joint strategy drive according to claim 5, characterized in that: S5 specifically includes the following steps: S51: Output the corpus o t and the output corpus u t Input to the Evaluation Reward Model LLM critic , get the reward value r of the tth round of dialogue t ; S52: Maximizing Rewards To update the direction strategy network and the boot strategy network The network parameters of the direction strategy network are obtained and the boot strategy network in, Represents the guidance strategy network and guide strategy network The divergence between them; β represents the penalty coefficient.
8. The method for generating a target-guided dialogue based on joint strategy drive according to claim 7, characterized in that: Repeat the following steps M times: t and the output corpus u t Input to the Evaluation Reward Model LLM critic , obtain M reward values; Calculate the average of the M reward values to obtain the reward value r of the tth round of dialogue t .
9. The method for generating a target-guided dialogue based on joint strategy drive according to claim 1, characterized in that: The direction strategy basic model is RoBERTa; the guidance strategy basic model is GPT-2.
10. The method for generating a target-guided dialogue based on joint strategy drive according to claim 7, characterized in that: The User Role Model LLM user 、The system dialogue model LLM sys And the evaluation reward model LLM critic ChatGPT is used for simulation.
Citation Information
Patent Citations
Emotion support dialogue system based on variational Bayesian inverse reinforcement learning strategy
CN119293181A
Systems and methods for safe policy improvement for task oriented dialogues
US20220036884A1
Dynamic goal-oriented dialogue with virtual agents
US20230063131A1
Deep reinforcement learning-based adaptive game algorithm
WO2020024097A1