A positive emotion encouragement-based emotional support conversation method and component

CN117010410BActive Publication Date: 2026-09-29TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310761790.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-26
Publication Date
2026-09-29
Estimated Expiration
2043-06-26

AI Technical Summary

Technical Problem

然而,共情对话在提供情绪支持方面存在天生的不足:(1)缺乏多轮对话考虑

Benefits of technology

[0015]本发明提供的一种基于正向情绪激励的情绪支持对话方法及组件,该方法包括设计多任务混合专家;多任务混合专家包括正向和负向的情绪专家和关键词专家;基于多任务混合专家设计强化学习框架;强化学习框架包括情绪支持奖励和对话连贯奖励;对强化学习框架进行优化。该方法将情绪支持对话形式化为一个正向情绪激励过程,在回复中不仅能够激励正向的情绪同时能够维持对话的连贯性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117010410B_ABST
    Figure CN117010410B_ABST
Patent Text Reader

Abstract

The application provides an emotion support dialogue method and component based on positive emotion stimulation, which comprises designing a multi-task hybrid expert; the multi-task hybrid expert comprises positive and negative emotion experts and keyword experts; a reinforcement learning framework is designed based on the multi-task hybrid expert; the reinforcement learning framework comprises emotion support rewards and dialogue coherence rewards; and the reinforcement learning framework is optimized. The method formalizes the emotion support dialogue into a positive emotion stimulation process, and can stimulate positive emotions and maintain the coherence of the dialogue in the reply.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of natural language processing and dialogue systems, and in particular to an emotion-supported dialogue method and components based on positive emotion incentives. Background Technology

[0002] Emotional support aims to understand and comfort a person to help them recover from emotional stress and improve their mental state. It is a manifestation of emotional intelligence in social interactions and plays a crucial role in maintaining close relationships. Infusing social dialogue systems with emotional support capabilities to construct warm, helpful, and trustworthy dialogue agents has become a popular trend and has attracted increasing attention.

[0003] To achieve this goal, a typical practice is to model empathy, a key component of providing successful support, which aims to identify and understand the experiences and feelings of others. However, empathic dialogues have inherent shortcomings in providing emotional support: (1) lack of consideration for multi-turn dialogues. Empathic dialogues typically focus on making empathetic responses in a single dialogue turn, ignoring the user's feedback and changes in mental state during multiple interactions. (2) lack of emotional incentive strategies. Empathic dialogues often focus on generating emotional resonance, making it difficult to help users escape negative mental states. Although introducing emotional support dialogue tasks is expected to make up for the above two shortcomings, existing working systems on emotional support dialogue tasks still only focus on fitting grounded response and reply strategies (e.g., inquiry strategies), ignoring their effects on emotional support. They do not adequately model the essential working mechanism of emotional support dialogues and lack explicit goals to guide the user's emotions to shift towards positive emotions during multi-turn dialogues. Therefore, the existing working systems mentioned above are insufficient to lay out the entire process of emotional support dialogues and cannot effectively improve a person's mental state. Summary of the Invention

[0004] This invention provides an emotion support dialogue method and components based on positive emotion incentives, which addresses the problem of how to construct an emotion support dialogue system that can incentivize users' emotions to shift towards a positive direction and improve their mental state in multi-turn dialogues. This method formalizes emotion support dialogue as a positive emotion incentive process, which can not only incentivize positive emotions in responses but also maintain the coherence of the dialogue.

[0005] This invention provides an emotion-support dialogue method based on positive emotion incentives, comprising: designing a multi-task hybrid expert; the multi-task hybrid expert including positive and negative emotion experts and keyword experts; designing a reinforcement learning framework based on the multi-task hybrid expert; the reinforcement learning framework including emotion support rewards and dialogue coherence rewards; and optimizing the reinforcement learning framework.

[0006] According to the present invention, an emotion-supported dialogue method based on positive emotion incentives is provided. The method involves designing a multi-task hybrid expert, comprising: encoding the state of the input dialogue context to obtain an encoding result; optimizing the contextual emotion loss of the user's utterances in the context using positive and negative contextual emotion experts based on the encoding result; optimizing the future emotion loss of the user's future utterances using positive and negative future emotion experts based on the encoding result; optimizing the joint loss of emotion experts based on the contextual emotion loss and the future emotion loss; optimizing the contextual keyword loss of the keywords in the user's utterances in the context using positive and negative contextual keyword experts based on the encoding result and a preset bidirectional emotion keyword graph; optimizing the future keyword loss of the keywords in the user's future utterances in the preset bidirectional emotion keyword graph using positive and negative future emotion experts based on the encoding result and the preset bidirectional emotion keyword graph; optimizing the future keyword loss of the keywords in the user's future utterances in the preset bidirectional emotion keyword graph using positive and negative future emotion experts based on the encoding result and the preset bidirectional emotion keyword graph; optimizing the joint loss of keyword experts based on the contextual keyword loss and the future keyword loss; and optimizing the joint loss of multi-task hybrid experts based on the joint loss of emotion experts, the joint loss of keyword experts, and a preset mean squared error loss.

[0007] According to the present invention, an emotion-supporting dialogue method based on positive emotion incentives is provided. The method, based on the multi-task hybrid expert design reinforcement learning framework, includes: acquiring a state sequence of the context; determining expert actions from the action space according to a preset expert selection policy network based on the state sequence of the context and the multi-task hybrid expert; determining a cue token sequence that can trigger a state update according to the expert actions and the dialogue decoder, and generating a response according to the cue token sequence and the response decoder; and designing the emotion-supporting reward and the dialogue coherence reward to optimize the generation of the response.

[0008] According to the present invention, an emotion support dialogue method based on positive emotion incentives is provided, wherein the preset expert selection strategy network includes an action network that selects appropriate expert action learning expert search strategies based on the current state and action space, and a value network that measures the value of the current state.

[0009] According to the present invention, an emotion-supporting dialogue method based on positive emotion incentives optimizes the generation of the response by designing emotion support rewards and dialogue coherence rewards. This includes: designing a session-level emotion support reward designed to dynamically adjust the intensity of positive emotion incentives as the conversation progresses; designing a round-level emotion support reward designed to capture the user's next round of emotional feedback; designing a contextual dialogue coherence reward designed to measure the coherence of the generated response with the dialogue context at the keyword and sentence levels; designing a future dialogue coherence reward designed to measure the coherence of the generated response with the user's future utterances; and optimizing the generation of the response based on the session-level emotion support reward, the round-level emotion support reward, the contextual dialogue coherence reward, and the future dialogue coherence reward.

[0010] According to the present invention, an emotion-supporting dialogue method based on positive emotion incentives is provided, wherein optimizing the reinforcement learning framework includes: performing proxy loss optimization and decoding loss optimization on the response; performing warm-start loss optimization based on the decoding loss optimization and the multi-task hybrid expert joint loss optimization; and performing joint training based on the proxy loss optimization and decoding loss optimization.

[0011] The present invention also provides an emotion support dialogue system based on positive emotion incentives, comprising: a hybrid expert design module for designing multi-task hybrid experts; the multi-task hybrid experts include emotion experts and keyword experts; a reinforcement learning framework design module for designing a reinforcement learning framework based on the multi-task hybrid experts; the reinforcement learning framework includes emotion support rewards and dialogue coherence rewards; and an optimization module for optimizing the reinforcement learning framework.

[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the emotion support dialogue method based on positive emotion incentives as described above.

[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the emotion-supported dialogue method based on positive emotion incentives as described above.

[0014] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the emotion support dialogue method based on positive emotion incentives as described above.

[0015] This invention provides an emotion-supporting dialogue method and components based on positive emotion incentives. The method includes designing a multi-task hybrid expert; the multi-task hybrid expert includes positive and negative emotion experts and keyword experts; designing a reinforcement learning framework based on the multi-task hybrid expert; the reinforcement learning framework includes emotion support rewards and dialogue coherence rewards; and optimizing the reinforcement learning framework. This method formalizes emotion-supporting dialogue as a positive emotion incentive process, which not only incentivizes positive emotions during responses but also maintains dialogue coherence. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating an emotion support dialogue method based on positive emotion incentives provided by the present invention.

[0018] Figure 2 This is a schematic diagram illustrating the principle of an emotion support dialogue method based on positive emotion incentives provided by the present invention.

[0019] Figure 3 This is a schematic diagram of the structure of an emotion support dialogue system based on positive emotion incentives provided by the present invention;

[0020] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0022] The following is combined with Figures 1-4 This invention describes an emotion support dialogue method and components based on positive emotion incentives.

[0023] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating an emotion support dialogue method based on positive emotion incentives provided by the present invention.

[0024] Please refer to Figure 2 , Figure 2 This is a schematic diagram illustrating the principle of an emotion support dialogue method based on positive emotion incentives provided by the present invention.

[0025] This invention proposes an emotion-support dialogue method based on positive emotion incentives. It designs a hybrid expert framework and employs reinforcement learning to generate responses in multi-turn dialogues to improve the user's mental state. The designed hybrid experts include heuristic experts associated with a specific task, who learn diverse semantics by characterizing the dialogue context. Specifically: (1) all experts are designed as either positive or negative experts to address potential fluctuations in the user's emotions during the ongoing dialogue; (2) the emotion expert among the hybrid experts is used to predict possible shifts in the user's emotions, thereby inspiring emotion-support capabilities in generated responses; and (3) the keyword expert among the hybrid experts is used to predict keywords that maintain dialogue coherence, thereby inspiring coherent expression in generated responses. Using these experts as candidates, a reinforcement learning agent learns a dialogue semantic encoding strategy and uses an expert selection strategy to purposefully select experts for response generation. In order to achieve the goals of positive emotional incentives in generating responses and maintaining dialogue coherence, we construct emotional support rewards and dialogue coherence rewards to optimize strategy learning: (1) emotional support rewards take into account the dialogue process to dynamically adjust the intensity of positive emotional incentives; (2) dialogue coherence rewards involve keyword-level and sentence-level guidance to finely maintain dialogue coherence.

[0026] This invention provides an emotion support dialogue method based on positive emotion incentives, comprising:

[0027] 101: Design a multi-task hybrid expert; multi-task hybrid experts include positive and negative emotion experts and keyword experts;

[0028] As a preferred embodiment, a multi-task hybrid expert is designed, including:

[0029] The state of the input dialogue context is encoded to obtain the encoding result;

[0030] Specifically, for each input dialogue context, the input tokens are concatenated and an [CLS] token is appended. This is then encoded using a dialogue encoder based on BlenderBot (an open-source chatbot) to obtain the hidden state H. The representation of the [CLS] token is then used as the representation h of the dialogue context sequence.

[0031] Based on the encoding results, contextual sentiment experts with positive and negative biases are used to optimize the emotional response of user utterances in the context; future sentiment experts with positive and negative biases are used to optimize the emotional response of user utterances in the future based on the encoding results; and joint loss optimization of sentiment experts is performed based on contextual sentiment loss and future sentiment loss.

[0032] Specifically, COMET is used beforehand to extract the emotional response of each utterance in the corpus, represented by emotion words, and VAD is used to identify the positive or negative emotional polarity of each emotional response. Based on the A1 encoding results, two MLPs (Multi-Layer Perceptrons) are used as positive and negative contextual emotion experts, respectively. By performing semantic space transformation on the hidden states and using [CLS] representation, they are used to predict the positive and negative emotional responses of the user's last utterance in the dialogue context. and The loss is optimized to enable the context-based sentiment expert to encode states based on the user's emotional state characteristics within the context. Similarly, two other MLPs are used as future sentiment experts to predict the positive and negative emotional responses to the user's future utterances, respectively. and The loss is optimized to allow future sentiment experts to anticipate the potential shifts in user emotions. Ultimately, the sentiment expert uses the loss... Joint optimization was carried out.

[0033] Based on the encoding results and the preset bidirectional sentiment keyword graph, positive and negative context keyword experts are used to optimize the context keyword loss of keywords in the user's utterance in the preset bidirectional sentiment keyword graph. Based on the encoding results and the preset bidirectional sentiment keyword graph, positive and negative future sentiment experts are used to optimize the future keyword loss of keywords in the user's future utterance in the preset bidirectional sentiment keyword graph. Based on the context keyword loss and the future keyword loss, joint loss optimization by keyword experts is performed.

[0034] Specifically, a rule-based keyword extraction method was used beforehand to extract the keyword set for each utterance in the corpus. VAD was then used to identify the positive or negative sentiment polarity of each keyword. Considering the order of utterances, a bidirectional sentiment keyword map was constructed using the keyword set in the utterances. Based on the constructed bidirectional keyword map and the A1 encoding results, two MLPs were used as positive and negative context keyword experts, respectively. By performing semantic space transformation on the hidden states and using [CLS] representation, the one-hop positive and negative keyword neighbors of the keyword in the user's last utterance in the dialogue context were predicted in the bidirectional sentiment keyword map, representing "forward-positive" and "forward-negative" relationships. and Loss is optimized to ensure that the context keyword expert captures emotional expression in authentic responses while maintaining conversational coherence within the context. Similarly, two other MLPs are used as future keyword experts to predict the positive and negative keyword neighbors of the keywords in the user's last utterance within the two-way sentiment keyword graph, representing the "forward→forward→backward-positive" and "forward→forward→backward-negative" relationships. and The loss is optimized to ensure that future keyword experts can maintain dialogue coherence with users in future conversations. Ultimately, the keyword expert uses the loss... Joint optimization was carried out.

[0035] The multi-task hybrid expert joint loss is optimized based on the joint loss of emotion experts, the joint loss of keyword experts, and the preset mean squared error loss.

[0036] Specifically, to ensure that experts maintain the semantics of the original context without hindering their diverse expressions, we use the MSE (Mean Square Error) loss L with a minimal hyperparameter α constraint. mse The average pooling vector of all expert [CLS] representations is made close to the [CLS] representation of the original context. Finally, this is achieved by optimizing L... exp Loss joint training multi-task hybrid expert.

[0037] L exp =L emo +L kws +α·L mse

[0038] 102: A reinforcement learning framework based on multi-task hybrid expert design; the reinforcement learning framework includes emotion support rewards and dialogue coherence rewards;

[0039] As a preferred embodiment, a reinforcement learning framework based on multi-task hybrid expert design includes:

[0040] Obtain the state sequence of the context;

[0041] Specifically, the initial state is obtained by concatenating the dialogue context C and the keywords within the dialogue context. At each step, the strategy selects an expert to generate a sequence of cue tokens. This triggers a state update. The observed state s at each step... k Represented as s k ={C,ε1,…,ε k The state sequence is encoded by the dialogue context encoder to obtain the hidden state H. s,k and hiding the representation h s,k The state representation is obtained by concatenating the hidden representation of the history and the representation of the k-th step.

[0042] Based on contextual state sequences and multi-task hybrid experts, the network determines expert actions from the action space according to a preset expert selection policy.

[0043] As a preferred embodiment, a pre-defined expert selection strategy network is included, comprising an action network that selects appropriate expert action learning expert search strategies based on the current state and action space, and a value network that measures the value of the current state.

[0044] Specifically, the action space of step k The hybrid expert described in A1 takes the state sequence at step k as its input. The reinforcement learning agent learns to select an expert from the action space as the expert action 'a' at step k. k It uses a BlenderBot-based dialog decoder to generate a sequence of cue tokens to trigger state updates.

[0045] In addition to using the dialogue encoder as a semantic encoding policy network, this invention designs an expert selection policy network, which includes an actor network and a value network. The actor network is based on the current state s. k and action space Choose the appropriate expert action a k Let's learn an expert's search strategy The value network is used to measure the state s. k Value Q δ (s k Their network structures are defined as follows:

[0046] o k =η((η(s)k W1)W2)),

[0047]

[0048] Q δ (s k ) = o k W k .

[0049] Where η(·) is the ELU activation function with a dropout layer, ⊙ is the Hadamard product, φ(·) is the softmax function, and A k This is a binary vector used to prune the action space. Since the number of experts is small, it is set to all 1s here.

[0050] The expert action and dialogue decoder determine the sequence of cue tokens that can trigger a state update, and the response is generated based on the cue token sequence and the response decoder.

[0051] The design incorporates emotion-supporting rewards and dialogue coherence rewards to optimize response generation.

[0052] As a preferred embodiment, the generation of responses is optimized by designing emotion support rewards and dialogue coherence rewards, including:

[0053] The design aims to dynamically adjust the intensity of positive emotional incentives as the conversation progresses, providing conversation-level emotional support rewards.

[0054] Specifically, to guide strategy learning, this invention designs four reward-optimized agent responses that provide emotional support while maintaining dialogue coherence:

[0055] Conversation-level emotional support rewards: designed to dynamically adjust the intensity of positive emotional incentives as the conversation progresses, and defined as follows:

[0056] PED cES =f ES (y)-f ES (c t ),

[0057]

[0058] Among them, f ES (·) A sentiment classification model is used to measure the positive sentiment score of a utterance. This model achieves 66% accuracy in sentiment classification. Here, the generated response y is encouraged to correlate with the context of the user utterance c. t Positive Emotional Distance (PED) between cESa) It is non-negative, meaning that expressing empathy (equal to 0) or motivation (greater than 0) is a basic requirement; b) It increases with the number of dialogue rounds, meaning that empathy dominates in the early stages of the conversation, while motivation dominates in the later stages. MT is the maximum number of dialogue rounds set, and T is the current round.

[0059] The design aims to capture the user's next round of emotional feedback through round-level emotional support rewards;

[0060] Specifically, round-based emotional support rewards: designed to capture user emotional feedback for the next round, defined as:

[0061] PED tES =|f ES (y)-f ES (c f )|,

[0062]

[0063] Among them, PED tES Used to measure the generated response y and the user's future (i.e., next round) discourse c. f The relative positive emotional distance between them. PED is encouraged here. tES As the current round T approaches MT, the incentives become smaller, which means that in the later stages, supervision is smoothed out and tolerance for emotional fluctuations is increased.

[0064] The design aims to measure the contextual dialogue coherence reward by measuring the coherence of the generated response with the dialogue context at the keyword and sentence levels.

[0065] Specifically, the contextual dialogue coherence reward aims to constrain coherent response generation by measuring the coherence of the generated response y with the dialogue context at the keyword and sentence levels. First, an emotion-supported dialogue dataset containing coherent and incoherent context-response pairs is reconstructed. Then, a BERT-based text classification model f is designed. cDC The classification model was trained using sentence-keyword pairs as input, achieving an accuracy of 85%, with coherence probability as the coherence score. This reward was defined as follows:

[0066]

[0067] Among them, y kws Let N be the set of keywords for y. c,kws Indicates y kws The keyword in is C kws The number of one-hop neighbors of the keywords in the two-way sentiment keyword graph under the "forward" relationship.

[0068] The design aims to measure the coherence of generated responses with future conversational rewards.

[0069] Specifically, the future dialogue coherence reward aims to measure the coherence of the generated response with the user's future (i.e., next) utterances. Similarly, an emotion-supported dialogue dataset containing coherent and incoherent future utterance-response pairs is reconstructed here, and another text classification model f is trained. cDC Achieving an accuracy rate of 77%, the reward was defined as:

[0070]

[0071] Where, N f,kws Indicates y kws Keywords and c fkws The keywords in the bidirectional emotion keyword graph have a "backward" relationship.

[0072] The generation of responses is optimized based on conversation-level emotion support rewards, round-level emotion support rewards, contextual dialogue coherence rewards, and future dialogue coherence rewards.

[0073] Specifically, the total reward is defined as:

[0074] r = w cES *r cES +w tES *r tES +w cDC *r cDC +w fDC *r fDC

[0075] 103: Optimize the reinforcement learning framework.

[0076] As a preferred embodiment, the reinforcement learning framework is optimized, including:

[0077] Perform proxy loss optimization and decoding loss optimization on the response;

[0078] Specifically, this invention sets up a K-step iteration for a reinforcement learning agent, where the agent's goal is to maximize the cumulative reward. Where θ is the parameter to be learned, and γ is the discount factor. Agent optimization uses L... agent The loss, whose policy gradient is defined as:

[0079]

[0080] Where G is the cumulative reward with discounts from the initial state to the final state. Finally, using state s... K+1 Hidden state H s,K+1 The decoder generates a response via Lgen Loss optimization:

[0081]

[0082] Hot-start loss optimization is performed based on decoding loss optimization and multi-task hybrid expert joint loss optimization.

[0083] Specifically, this invention first uses a pre-trained BlenderBot to initialize the model, using the initial state as input to fine-tune the model for warm start, and optimizes it through L... warm loss:

[0084] L warm =L exp +L gen

[0085] Joint training is performed based on proxy loss optimization and decoding loss optimization.

[0086] Specifically, this invention ultimately optimizes L joint Loss-based joint training:

[0087]

[0088] In addition, the parameters can be set as follows:

[0089] The number of iterations is set to K = 2;

[0090] The reward weight is set to w cES =w cDC =0.1, w tES =w fDC =1.0;

[0091] Set the number of dialogue rounds to MT=10;

[0092] The discount factor is set to γ ​​= 0.99;

[0093] The hyperparameter α is set to 1e-5;

[0094] The learning rate was set to 2e-5, the number of warm-start training rounds was set to 5, and the number of training rounds in the joint training phase was 3.

[0095] Please refer to Figure 3 , Figure 3 This is a schematic diagram of the structure of an emotion support dialogue system based on positive emotion incentives provided by the present invention.

[0096] This invention also provides an emotion support dialogue system based on positive emotion incentives, comprising: a hybrid expert design module for designing multi-task hybrid experts; the multi-task hybrid experts include emotion experts and keyword experts; a reinforcement learning framework design module for designing a reinforcement learning framework based on the multi-task hybrid experts; the reinforcement learning framework includes emotion support rewards and dialogue coherence rewards; and an optimization module for optimizing the reinforcement learning framework.

[0097] For an introduction to the emotion support dialogue system based on positive emotion incentives provided by this invention, please refer to the above method embodiments; the invention itself will not be described in detail here.

[0098] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device may include a processor 401, a communications interface 402, a memory 403, and a communication bus 404. The processor 401, communications interface 402, and memory 403 communicate with each other via the communication bus 404. The processor 401 can call logical instructions in the memory 403 to execute an emotion-supporting dialogue method based on positive emotion incentives. This method includes: designing a multi-task hybrid expert; the multi-task hybrid expert includes positive and negative emotion experts and keyword experts; designing a reinforcement learning framework based on the multi-task hybrid expert; the reinforcement learning framework includes emotion support rewards and dialogue coherence rewards; and optimizing the reinforcement learning framework.

[0099] Furthermore, the logical instructions in the aforementioned memory 403 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0100] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the emotion-supporting dialogue method based on positive emotion incentives provided by the above methods. The method includes: designing a multi-task hybrid expert; the multi-task hybrid expert includes positive and negative emotion experts and keyword experts; designing a reinforcement learning framework based on the multi-task hybrid expert; the reinforcement learning framework includes emotion support rewards and dialogue coherence rewards; and optimizing the reinforcement learning framework.

[0101] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, is implemented to perform the emotion-supporting dialogue method based on positive emotion incentives provided by the methods described above. This method includes: designing a multi-task hybrid expert; the multi-task hybrid expert including positive and negative emotion experts and keyword experts; designing a reinforcement learning framework based on the multi-task hybrid expert; the reinforcement learning framework including emotion support rewards and dialogue coherence rewards; and optimizing the reinforcement learning framework.

[0102] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0103] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A dialogue method for emotional support based on positive emotional incentives, characterized in that, include: Design a multi-task hybrid expert; The multi-task hybrid experts include positive and negative emotion experts and keyword experts; The reinforcement learning framework is designed based on the aforementioned multi-task hybrid expert design. The reinforcement learning framework includes emotional support rewards and dialogue coherence rewards. The reinforcement learning framework is optimized. The design multi-task hybrid expert includes: The state of the input dialogue context is encoded to obtain the encoding result; Based on the encoding results, contextual sentiment loss is optimized by using positive and negative contextual sentiment experts to assess the emotional response to user utterances in the context; future sentiment loss is optimized by using positive and negative future sentiment experts to assess the emotional response to user utterances in the future based on the encoding results; and joint loss optimization by sentiment experts is performed based on the contextual sentiment loss and the future sentiment loss. Based on the encoding results and the preset bidirectional sentiment keyword graph, positive and negative context keyword experts are used to optimize the context keyword loss of keywords in the user's utterance in the context and their keyword neighbors in the preset bidirectional sentiment keyword graph; based on the encoding results and the preset bidirectional sentiment keyword graph, positive and negative future sentiment experts are used to optimize the future keyword loss of keywords in the user's future utterance and their keyword neighbors in the preset bidirectional sentiment keyword graph; and keyword expert joint loss optimization is performed based on the context keyword loss and the future keyword loss. Optimize the multi-task hybrid expert joint loss based on the joint loss of emotion experts, the joint loss of keyword experts, and the preset mean squared error loss; The reinforcement learning framework based on the multi-task hybrid expert design includes: Obtain the state sequence of the context; Based on the state sequence of the context and the multi-task hybrid expert, the expert action is determined from the action space according to the preset expert selection strategy network; The expert action and dialogue decoder determine a sequence of cue tokens that can trigger a state update, and the response is generated based on the cue token sequence and the response decoder. The design of the emotion support reward and the dialogue coherence reward optimizes the generation of the response; The design of emotion support rewards and dialogue coherence rewards optimizes the generation of the response, including: The design aims to dynamically adjust the intensity of positive emotional incentives as the conversation progresses, providing conversation-level emotional support rewards. The design aims to capture the user's next round of emotional feedback through round-level emotional support rewards; The design aims to measure the contextual dialogue coherence reward, which measures the coherence of the generated response with the dialogue context at the keyword and sentence levels. The design aims to measure the coherence of the generated response with the user's future utterances, resulting in a future conversational coherence reward. The generation of the response is optimized based on conversation-level emotion support rewards, round-level emotion support rewards, contextual dialogue coherence rewards, and future dialogue coherence rewards.

2. The emotion support dialogue method based on positive emotion incentives according to claim 1, characterized in that, The preset expert selection strategy network includes an action network that selects a suitable expert action learning expert search strategy based on the current state and action space, and a value network that measures the value of the current state.

3. The emotion support dialogue method based on positive emotion incentives according to claim 1, characterized in that, The optimization of the reinforcement learning framework includes: The response is then optimized using proxy loss and decoding loss optimization. Hot start loss optimization is performed based on the decoding loss optimization and the multi-task hybrid expert joint loss optimization. Joint training is performed based on the aforementioned proxy loss optimization and decoding loss optimization.

4. An emotion support dialogue system based on positive emotion incentives, characterized in that, include: The Hybrid Expert Design Module is used to design multi-task hybrid experts; The multi-task hybrid expert includes emotion experts and keyword experts; A reinforcement learning framework design module is used to design a reinforcement learning framework based on the multi-task hybrid expert. The reinforcement learning framework includes emotional support rewards and dialogue coherence rewards. An optimization module is used to optimize the reinforcement learning framework; The design multi-task hybrid expert includes: The state of the input dialogue context is encoded to obtain the encoding result; Based on the encoding results, contextual sentiment loss is optimized by using positive and negative contextual sentiment experts to assess the emotional response to user utterances in the context; future sentiment loss is optimized by using positive and negative future sentiment experts to assess the emotional response to user utterances in the future based on the encoding results; and joint loss optimization by sentiment experts is performed based on the contextual sentiment loss and the future sentiment loss. Based on the encoding results and the preset bidirectional sentiment keyword graph, positive and negative context keyword experts are used to optimize the context keyword loss of keywords in the user's utterance in the context and their keyword neighbors in the preset bidirectional sentiment keyword graph; based on the encoding results and the preset bidirectional sentiment keyword graph, positive and negative future sentiment experts are used to optimize the future keyword loss of keywords in the user's future utterance and their keyword neighbors in the preset bidirectional sentiment keyword graph; and keyword expert joint loss optimization is performed based on the context keyword loss and the future keyword loss. Optimize the multi-task hybrid expert joint loss based on the joint loss of emotion experts, the joint loss of keyword experts, and the preset mean squared error loss; The reinforcement learning framework based on the multi-task hybrid expert design includes: Obtain the state sequence of the context; Based on the state sequence of the context and the multi-task hybrid expert, the expert action is determined from the action space according to the preset expert selection strategy network; The expert action and dialogue decoder determine a sequence of cue tokens that can trigger a state update, and the response is generated based on the cue token sequence and the response decoder. The design of the emotion support reward and the dialogue coherence reward optimizes the generation of the response; The design of emotion support rewards and dialogue coherence rewards optimizes the generation of the response, including: The design aims to dynamically adjust the intensity of positive emotional incentives as the conversation progresses, providing conversation-level emotional support rewards. The design aims to capture the user's next round of emotional feedback through round-level emotional support rewards; The design aims to measure the contextual dialogue coherence reward, which measures the coherence of the generated response with the dialogue context at the keyword and sentence levels. The design aims to measure the coherence of the generated response with the user's future utterances, resulting in a future conversational coherence reward. The generation of the response is optimized based on conversation-level emotion support rewards, round-level emotion support rewards, contextual dialogue coherence rewards, and future dialogue coherence rewards.

5. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the emotion-supported dialogue method based on positive emotion incentives as described in any one of claims 1 to 3.

6. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the emotion-supported dialogue method based on positive emotion incentives as described in any one of claims 1 to 3.

7. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the emotion-supported dialogue method based on positive emotion incentives as described in any one of claims 1 to 3.