Adaptive Outbound Dialogue Strategy Configuration Method and Device Based on User Emotional State
By building user sentiment detection and semantic analysis modules, the Q&A strategy of the intelligent outgoing call robot is adaptively adjusted, and the problem of stiff interaction in the existing technology is solved, the quality of Q&A and the user hang-up rate is reduced.
Patent Information
- Application Number
- CN202310017555.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-06
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-01-06
AI Technical Summary
Existing intelligent outgoing call robots cannot adjust the Q&A strategy based on the user's emotional state, resulting in stiff interactions and high user hang-up rate.
By building a user emotion detection module, a semantic analysis module and a speech configuration module, based on user speech emotion recognition and semantic analysis, the Q&A strategy is adaptively adjusted, including emotion detection model training, voice to text, keyword matching and Q&A strategy processing, and the question and answer strategy processing are optimized for each round of questioning strategy.
It improves the Q&A quality of intelligent outbound call robots, reduces user hang-up rate, and enhances the flexibility and effectiveness of interaction.
Smart Images

Figure CN116386604B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent outbound calls, and in particular to a method and device for configuring an adaptive outbound call dialogue strategy based on the emotional state of a user. Background Art
[0002] With the rapid development of artificial intelligence technology, especially the continuous development of automatic speech recognition technology ASR (Automatic Speech Recognition) and natural language processing technology NLP (Natural Language Processing), more and more intelligent outbound call robots have been practically implemented in real business scenarios, and the market feedback has been good, replacing a large amount of manual labor. However, in the prior art, the question-and-answer response mode of intelligent robots is still configured based on fixed templates. Although a convenient front-end visual interface allows for the configuration of multiple templates, usually only one can be enabled in a single conversation. This results in a rigid and inflexible interaction between the robot and the user, and the robot is unable to adjust its own conversation skills or question-and-answer strategies according to the emotional state of the user, which is also the core reason for users to hang up.
[0003] The following are the related patents based on speech emotion recognition in the prior art:
[0004] Related patent document, "Speech Emotion Recognition Method and Device", CN108122552B. This patent mainly describes a speech emotion recognition method, and its main method is to match the audio feature vector of a speech segment with multiple emotion feature models to obtain the corresponding emotion classification. By using the method of recognizing the emotion of recorded speech, it solves the problem that the prior art cannot real-time monitor the emotional states of customer service and customers in a call center system.
[0005] Related patent document, "A Speech Emotion Recognition and Application System for Call Center Calls", CN109767791B. This patent mainly describes a speech emotion recognition and application system for call center calls, which can enable customer service personnel to accurately understand the emotions of customers, provide effective response solutions, and accurately assess customer service personnel.
[0006] In order to improve the ability of an intelligent outbound call robot to analyze the emotional state of a user and make corresponding responses, based on the speech emotion recognition technology, and reduce the hang-up rate of users, a real-time analysis system that can be based on the emotions and responses of users is required, and according to the analysis results, the strategy of the question-and-answer conversation skills is adaptively adjusted. The goal of this strategy should be to retain customers as much as possible and more accurately and efficiently complete all rounds of tasks of the intelligent outbound call robot. Summary of the Invention
[0007] To address the deficiencies of the prior art and implement the function of an intelligent outbound robot to adaptively adjust the corresponding conversation strategies based on user voice emotion recognition and semantic analysis, the present invention adopts the following technical solutions:
[0008] An adaptive outbound conversation strategy configuration method based on the user's emotional state, comprising the following steps:
[0009] Step S1: Build a user emotion detection module and train to generate a user emotion detection model for analyzing and detecting the voice input by the user to obtain the user's emotional state;
[0010] Through constructing a voice emotion analysis model, the analysis of the user input voice and the training of the voice emotion analysis model include the following steps:
[0011] Step S1.1: Obtain the voice data of the user and label the emotion tags to obtain the training data;
[0012] Step S1.2: Extract the Mel-frequency cepstral coefficients features of each segment of voice data, input them into a predefined voiceprint perception classification model for emotion recognition, and train the voice emotion analysis model based on the recognized emotion and the labeled emotion tags.
[0013] Step S2: Build a semantic analysis module to analyze whether the user's reply conforms to the expected answer;
[0014] Step S3: Build a conversation strategy configuration module to store and configure the conversation strategies and expected answers corresponding to each question;
[0015] Step S4: Build a question-and-answer strategy processing module to formulate the robot's next-round question strategy based on the results of each round of human-machine interaction, including the following steps:
[0016] Step S4.1: For each question, select the answer group configured for the corresponding question from the conversation strategy configuration module as the candidate question-and-answer strategy for this question;
[0017] Step S4.2: Based on the question-and-answer strategy and the user's emotional state, select the corresponding emotion probability mean values for different user groups;
[0018] Step S4.3: Based on the question-and-answer strategy and the completeness of the user feedback, select the corresponding probability mean value for obtaining the answer;
[0019] Step S4.4: Set the emotion gain and the gain for obtaining the answer, calculate the expected gain for each round based on the selected corresponding emotion probability mean value and the probability mean value for obtaining the answer, and formulate the robot's question strategy according to the expected gain.
[0020] Further, in step S4.2, based on the positive and negative emotional fluctuations and the questioning phrases, select the corresponding positive and negative probability means; in step S4.3, based on the obtained answers and unanswered questions and the questioning phrases, select the corresponding probability means of obtained answers and unanswered questions.
[0021] Further, step S4.4 includes the following steps:
[0022] Step S4.4.1: Based on the questioning phrases for different types of emotional experiences, calculate the expected benefits and select the questioning phrases.
[0023] Step S4.4.2: According to the questioning phrases, calculate the expected benefits and select the next round of Q&A strategy.
[0024] Use the user emotion detection module in step S1 to detect the user's voice reply to obtain the current user's emotional state, and use the semantic analysis module in step S2 to analyze the user's answer to determine whether the current user has answered the question.
[0025] If the user has answered the question, return to step S4.4.1. According to the probability means and benefits of the questioning phrases corresponding to the newly detected user emotional state, and the probability means and benefits of whether an answer can be obtained, calculate the expected benefits of each questioning phrase. According to the questioning phrase corresponding to the maximum expected benefit, determine the questioning phrase for the second question.
[0026] Further, in step S4.2, construct a follow-up Q&A strategy, that is, a Q&A strategy for asking questions again after the first question fails to obtain a clear answer. Based on step S1, select the corresponding follow-up emotion probability means for different user groups.
[0027] In step S4.4.2, if it is analyzed that the user has not answered the question, it is necessary to consider that continued follow-up will also cause emotional fluctuations at this time. According to the follow-up probability means and benefits corresponding to the detected user emotional state, and the probability means and benefits of whether an answer can be obtained, calculate whether to execute the follow-up dialogue strategy. If the calculated benefit is less than the set threshold, do not execute the follow-up dialogue strategy, return to step S4.4.1, and according to the probability means and benefits of the questioning phrases corresponding to the newly detected user emotional state, and the probability means and benefits of whether an answer can be obtained, calculate the expected benefits of each questioning phrase. According to the questioning phrase corresponding to the maximum expected benefit, determine the questioning phrase for the second question; if the calculated benefit is greater than the set threshold, execute the follow-up dialogue strategy.
[0028] An adaptive outbound call dialogue strategy configuration device based on the user's emotional state includes a user emotion detection module, a semantic analysis module, a phraseology configuration module, and a Q&A strategy processing module;
[0029] The user emotion detection module generates an emotion detection model for the user through training, and is used to analyze and detect the speech input by the user to obtain the user's emotion state;
[0030] The semantic analysis module analyzes whether the user's reply conforms to the expected answer;
[0031] The conversation strategy configuration module stores and configures the conversation words and expected answers corresponding to each question;
[0032] The question-answer strategy processing module formulates the robot's question strategy for the next round based on the results of the human-machine interaction in each round, including a strategy preparation module and a strategy execution module;
[0033] The strategy preparation module selects, for each question, the group of conversation words configured for the corresponding question from the conversation strategy configuration module as the candidate conversation words for the question; based on the conversation words and the user's emotion state, it selects the corresponding average emotion probability of different user groups; based on the conversation words and the completeness of the user feedback, it selects the corresponding average probability of obtaining an answer;
[0034] The strategy execution module sets the emotion gain and the gain of obtaining an answer, calculates the expected gain for each round based on the selected corresponding average emotion probability and the average probability of obtaining an answer, and formulates the robot's question strategy according to the expected gain.
[0035] Further, the user emotion detection module includes a first voice receiving module and a voice emotion analysis module, and the voice emotion analysis module includes a preprocessing module and a voice emotion analysis model;
[0036] The first voice receiving module is used to receive the user's voice signal;
[0037] The preprocessing module uses voice active endpoint detection to extract the audio data containing human voices in the received voice;
[0038] The voice emotion analysis model is a pre-trained model inference file, which can infer the corresponding emotion state according to the input voice data.
[0039] Further, the semantic analysis module includes a second voice receiving module and a semantic analysis module, and the semantic analysis module includes a voice-to-text module and a keyword matching module;
[0040] The second voice receiving module is used to receive the user's voice signal;
[0041] The speech-to-text module uses an open-source ASR (Automatic Speech Recognition) model to convert speech into text.
[0042] The keyword matching module performs regular matching on the text converted from speech with the keyword dictionary through the candidate keyword dictionary and the regular matcher, and the matching keywords are used as candidate answers.
[0043] Furthermore, the strategy preparation module selects corresponding positive and negative probability means based on positive and negative emotional fluctuations and the questioning language.
[0044] Furthermore, the strategy execution module calculates the expected benefits and selects questioning scripts based on questioning scripts of different types of emotional experiences; calculates the expected benefits and selects the next round of question-answering strategies based on the questioning scripts; utilizes the user emotion detection module to detect the user's voice response to obtain the current user's emotional state, utilizes the semantic analysis module to analyze the user's answer to determine whether the current user has answered the question; if the user has answered the question, returns to continue selecting questioning scripts, calculates the expected benefits of each questioning script based on the probability mean and benefits of the questioning scripts corresponding to the new user emotional state obtained by detection, and the probability mean and benefits of whether an answer can be obtained, and determines the questioning script for the second question based on the questioning script corresponding to the maximum expected benefit.
[0045] Furthermore, in the strategy execution module, a follow-up question-answering strategy is constructed, that is, a question-answering strategy for asking questions again after the first question does not match the expected answer. Based on the user emotion detection module, the corresponding follow-up question emotion probability means of different user groups are selected; if it is analyzed that the user has not answered the question, it is necessary to consider that continuing to ask questions will also bring about emotional fluctuations. According to the follow-up question probability mean and benefit corresponding to the user's emotional state obtained by detection, as well as the probability mean and benefit of whether an answer can be obtained, the dialogue strategy for follow-up questioning is calculated whether to execute; if the calculated benefit is less than the set threshold, the dialogue strategy for follow-up questioning is not executed, and the questioning script is returned to continue to be selected; according to the questioning script probability mean and benefit corresponding to the new user emotional state obtained by detection, as well as the probability mean and benefit of whether an answer can be obtained, the expected benefit of each questioning script is calculated, and the questioning script for the second question is determined based on the questioning script corresponding to the maximum expected benefit; if the calculated benefit is greater than the set threshold, the dialogue strategy for follow-up questioning is executed.
[0046] The advantages and beneficial effects of the present invention are:
[0047] The method and device for configuring an adaptive outbound call dialogue strategy based on the user's emotional state of the present invention, based on the emotional analysis and semantic analysis of the user's voice, calculates the expected revenue for each question-and-answer round through a pre-set probability model, and based on this, adaptively adjusts the question-and-answer conversation logic of the intelligent outbound call robot, thereby improving the quality of the outbound call reply and reducing the hang-up rate. Description of the Drawings
[0048] Figure 1 It is a flowchart of the method for configuring an adaptive outbound call dialogue strategy based on the user's emotional state in the present invention.
[0049] Figure 2 It is a schematic structural diagram of the device for configuring an adaptive outbound call dialogue strategy based on the user's emotional state in the present invention. Detailed Embodiments
[0050] The following will describe in detail the specific embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for explaining and illustrating the present invention, and are not used to limit the present invention.
[0051] As Figure 1 shown, the method for configuring an adaptive outbound call dialogue strategy based on the user's emotional state includes the following steps:
[0052] Step S1: Build a user emotion detection module;
[0053] This module is mainly responsible for analyzing and detecting the voice input by the user, and the main analysis target is the emotional state of the user.
[0054] Specifically, this module includes 2 parts. The first part is the voice reception module, which is responsible for receiving the voice signal of the user. The second part is the voice emotion analysis module, which mainly includes a voice preprocessing module and a voice emotion analysis model.
[0055] Specifically, the preprocessing module includes voice active endpoint detection, and the main purpose is to extract the audio data containing human voices in the received voice. This is a voice activity detection technology (Voice Activity Detection).
[0056] Specifically, the voice emotion analysis model is a pre-trained model inference file, which can infer the corresponding emotional state according to the input voice data.
[0057] Specifically, the training steps of the voice emotion analysis model are as follows.
[0058] Step S1.1: Obtain the voice data of the user;
[0059] The user's voice data usually refers to real-time input voice data, or can also be pre-recorded voice data manually. Preferably, this part of the data should include as much recorded data as possible under different emotions. Specifically, the collected data should at least contain three emotional states, "peaceful", "impatient", "pleased", or emotions close to these three states. The collected data is manually verified and labeled, and the training data required for model training is standardized and sorted out.
[0060] Step S1.2: Train the voice emotion analysis model;
[0061] For the training audio data obtained in step S1.1, first extract the Mel-frequency cepstral coefficients (MFCC) features of each audio segment, and then input it into a pre-defined voiceprint perception classification model for emotion recognition.
[0062] Specifically, the voiceprint perception classification model can be a traditional machine learning method, including support vector machine model, Gaussian mixture model, hidden Markov model, etc., or a deep neural network model, or a decision tree classification model.
[0063] Preferably, the model specifically adopted in the present invention is a convolutional neural network model in the deep neural network. Because it has strong feature extraction ability and is suitable for classification scenarios. The classification objective function it adopts is the cross-entropy loss function.
[0064] Specifically, the goal of model training is three classifications, namely, "peaceful", "impatient", "pleased".
[0065] Step S2: Build a semantic analysis module;
[0066] This module is mainly responsible for analyzing and detecting the user's input voice. The main analysis target is whether the user's reply is the expected answer.
[0067] Specifically, this module includes two parts. The first part is the voice receiving module, which is responsible for receiving the user's voice signal.
[0068] The second part is the semantic analysis module, which mainly includes a speech-to-text module and a keyword matching module.
[0069] Specifically, the speech-to-text module will adopt an open-source ASR (Automatic Speech Recognition) recognition model interface.
[0070] Specifically, the keyword matching module mainly includes a candidate keyword dictionary and a regular matcher. This module will match the text converted from the audio with the keyword dictionary, and the keywords in the match are the candidate answers.
[0071] Step S3: Build a conversation configuration module;
[0072] This module is mainly responsible for storing the configuration templates for each question and should be designed for visual configuration.
[0073] Based on this template, humans can configure different conversation scripts for each question and set the corresponding expected answers.
[0074] Specifically, one of the storage formats is as follows:
[0075] {"Question 1": {"Question Content": "x",
[0076] "Expected Answer": ["xxx", "xxx",...],
[0077] "Inquiry Script": ["xxx", "xxx", "xxx"]
[0078] "Question 2":...}
[0079] Step S4: Build a question - answering strategy processing module;
[0080] This module is mainly responsible for formulating the robot response strategy for the next round based on the results of each round of human - machine interaction.
[0081] Specifically, taking "Follow - up of discharged patients" as a scenario case for elaboration.
[0082] Suppose this scenario case involves 2 questions, as shown below:
[0083] Q1: Are you satisfied with the rehabilitation effect?
[0084] Q2: How did you get admitted to the hospital?
[0085] Next, the specific question - answering response strategies for these two questions will be elaborated in detail.
[0086] Step S4.1: Set 3 types of question scripts for each question;
[0087] Specifically, this scenario case is mainly divided into the following 3 types of question scripts. Different question scripts take different lengths of time and bring different emotional experiences to users. For impatient users, simple scripts are often sufficient. On the contrary, for calm users, polite scripts are used. Moreover, different scripts have different probabilities of obtaining clear answers for different user groups.
[0088] Specifically, the questioning techniques corresponding to each question are shown in the following table.
[0089]
[0090] Step S4.2: For each questioning technique, select the corresponding emotional probability mean of different user groups.
[0091] Specifically, each questioning technique will produce different emotional fluctuations in different user groups. The main emotional fluctuations are positive and negative. After collecting and organizing the past questionnaire dataset, the corresponding probability means are calculated according to the user emotional states defined in step S1: impatience, calmness, and joy.
[0092] Specifically, one of the peaceful probability statistics is shown in Table 1:
[0093] Table 1 Mean probability of Pinghe-questioning techniques
[0094] Direct question Direct question with answer Complete question with answer Positive <![CDATA[P qt1_pos = 0.8]]> <![CDATA[P qt2_pos = 0.7]]> <![CDATA[P qt3_pos = 0.6]]> Negative <![CDATA[P qt1_neg = 0.2]]> <![CDATA[P qt2_neg = 0.3]]> <![CDATA[P qt3_neg = 0.4]]>
[0095] Specifically, there's a special type of questioning: follow-up questions. This strategy involves asking again after a clear answer hasn't been obtained after the initial question. While this strategy can increase the probability of obtaining an answer, it can also inflame the user's emotions. After collecting and organizing past survey data, the corresponding probability means are analyzed and statistically analyzed based on the user's emotional states defined in step S1: impatience, calmness, and patience.
[0096] Specifically, the probability statistics of one type of impatience are shown in Table 2:
[0097] Table 2 Mean probability of impatience-questioning
[0098] Follow-up question Positive <![CDATA[P plus_pos = 0.2]]> Negative <![CDATA[P plus_pos = 0.8]]>
[0099] Step S4.3: For each questioning technique, select the corresponding mean probability of obtaining the answer.
[0100] Specifically, user responses to each question are bound to vary. For more complete question formulations, users tend to be more likely to give the expected answer. After collecting and organizing past questionnaire datasets, we calculated the corresponding probability means.
[0101] Specifically, the mean probability of the answer corresponding to one of the speech techniques is shown in Table 3:
[0102] Table 3 Mean probability of answers corresponding to the speech
[0103]
[0104] Step S4.4: Calculate the expected revenue for each round and formulate the robot's conversation strategy.
[0105] For the convenience of description, some mathematical symbols and corresponding default values are now defined:
[0106] E_pos: Revenue of regular positive emotion, default value is 1
[0107] E_neg: Revenue of regular negative emotion, default value is -1
[0108] E_pos_plus: Revenue of pursuing positive emotion, default value is 0.8
[0109] E_neg_plus: Revenue of pursuing negative emotion, default value is -1.2
[0110] E_yes: Revenue of getting an answer, default value is 1
[0111] E_no: Revenue of not getting an answer, default value is -1
[0112] E_total: Total revenue
[0113] In one of the preferred strategies, as the number of conversation rounds increases, the revenue of positive / negative emotions can also change linearly. For the convenience of elaboration, fixed values are adopted in this specific case.
[0114] In each round, the robot has a corresponding conversation strategy, and its final selection criterion depends on maximizing the expected revenue.
[0115] Specifically, the formula for calculating the final expected revenue of each questioning conversation is as follows:
[0116] E total =E pos *P pos +E neg *P neg +E yes *P yes +E no *P no
[0117] Step S4.4.1: Calculate the expected revenue and select the first-round questioning conversation.
[0118] Specifically, when just dialing the user's phone number, due to the ambiguity of the user portrait information, the default probability value of "peaceful user" will be directly adopted for strategy selection.
[0119] Specifically, the revenue calculation of each questioning conversation is as follows:
[0120] E total_直接提问= 0.8 * 1 - 0.2 * 1 + 0.3 * 1 - 0.7 * 1 = 0.2
[0121] E total_直接提问含答案 = 0.7 * 1 - 0.3 * 1 + 0.6 * 1 - 0.4 * 1 = 0.6
[0122] E total_完整提问含答案 = 0.6 * 1 - 0.4 * 1 + 0.9 * 1 - 0.1 * 1 = 1.0
[0123] Obviously, the final maximum benefit is to adopt the conversation strategy of "complete question with answer", and the expected benefit is 1.0.
[0124] Step S4.4.2: Calculate the expected benefit and select the secondary round Q&A strategy.
[0125] Specifically, according to Step S4.4.1, adopt the conversation strategy of "complete question with answer" to ask the first question. After obtaining the user's voice reply, formulate the secondary round Q&A strategy.
[0126] Specifically, use the emotion detection module in Step S1 to detect the user's voice reply and detect which emotional state the current user is in.
[0127] Specifically, use the semantic analysis module in Step S2 to analyze the user's answer and analyze whether the current user's voice has answered the question. There will be two situations at this time. One situation is to obtain the answer and complete the question.
[0128] Step S4.4.2.1: If Question 1 is completed, only repeat Step S4.4.1. According to the newly detected user emotional state, use the probability mean of the question-asking conversation strategy corresponding to this emotion to calculate the benefit of each question-asking conversation strategy to determine the question-asking conversation strategy for the second question.
[0129] Specifically, assume that the currently detected user emotional state is "impatient", and the probability mean of the corresponding question-asking conversation strategy is shown in Table 4:
[0130] Table 4 Impatient - Probability Mean of Question - Asking Conversation Strategy
[0131] Direct question Direct question with answer Complete question with answer Positive <![CDATA[P qt1_pos = 0.7]]> <![CDATA[P qt2_pos = 0.5]]> <![CDATA[P qt3_pos = 0.1]]> Negative <![CDATA[P qt1_neg = 0.3]]> <![CDATA[P qt2_neg = 0.5]]> <![CDATA[P qt3_neg = 0.9]]>
[0132] Then the expected benefit calculation for the corresponding second question is as follows:
[0133] E total_直接提问 = 0.7 * 1 - 0.3 * 1 + 0.3 * 1 - 0.7 * 1 = 0.0
[0134] E total_直接提问含答案 = 0.5 * 1 - 0.5 * 1 + 0.6 * 1 - 0.4 * 1 = 0.2
[0135] E total_完整提问含答案 = 0.1 * 1 - 0.9 * 1 + 0.9 * 1 - 0.1 * 1 = 0.0
[0136] As above, the maximum benefit of the second question is 0.2, and the corresponding strategy is to directly ask questions with answers. At this time, the total expected benefit statistics of the two questions are shown in Table 5:
[0137] Expected return Corresponding strategy Round 1 1.0 Question 1: Complete question with answer Round 2 0.2 Question 2: Direct question with answer Overall 1.2 -
[0138] In step S4.4.2.2, if question one is not completed, it is necessary to consider that continuing to ask questions will also cause fluctuations in emotions at this time. Specifically, the calculation formula for the expected benefit at this time is changed as follows,
[0139] E plus = E yes * P plus_yes + E no * P plus_no + E pos_plus * P plus_pos + E neg_plus * P plus_neg
[0140] If the expected benefit of asking questions is greater than 0, execute the dialogue strategy of asking questions; otherwise, do not execute, skip this question, and repeat step S4.4.1. According to the newly detected user emotion state, use the probability mean of the question-asking words corresponding to this emotion to calculate the benefit of each question-asking word, and determine the question-asking word for the second question.
[0141] Specifically, it is still assumed that the currently detected user emotion state is "impatient". For the first question, the benefit of the strategy of asking questions is:
[0142] E plus = 1 * 0.7 - 1 * 0.3 + 0.2 * 0.8 - 0.8 * 1.2 = -0.4
[0143] The benefit is less than 0, so do not execute the strategy of asking questions. Repeat step S4.3.2.1.
[0144] As Figure 2 shown, the adaptive outbound dialogue strategy configuration device based on the user emotion state includes a user emotion detection module, a semantic analysis module, a dialogue strategy configuration module, and a question-answer strategy processing module;
[0145] The user emotion detection module generates a user emotion detection model through training, and is used to analyze and detect the voice input by the user to obtain the user's emotion state; it includes a first voice reception module and a voice emotion analysis module, and the voice emotion analysis module includes a preprocessing module and a voice emotion analysis model;
[0146] The first voice receiving module is used to receive the user's voice signal;
[0147] The preprocessing module uses voice endpoint detection of human voices to extract the audio data containing human voices in the received voice;
[0148] The voice emotion analysis model is a pre-trained model inference file that can infer the corresponding emotional state based on the input voice data.
[0149] The semantic analysis module analyzes whether the user's reply conforms to the expected answer; it includes the second voice receiving module and the semantic analysis module, and the semantic analysis module includes a voice-to-text module and a keyword matching module;
[0150] The second voice receiving module is used to receive the user's voice signal;
[0151] The voice-to-text module uses an open-source ASR (Automatic Speech Recognition) automatic speech recognition model to convert speech into text;
[0152] The keyword matching module uses a candidate keyword dictionary and a regular matcher to perform regular matching between the text converted from speech and the keyword dictionary, and the matched keywords are used as candidate answers.
[0153] The conversation script configuration module stores and configures the conversation scripts and expected answers corresponding to each question;
[0154] The Q&A strategy processing module formulates the next round of robot questioning strategies based on the results of each round of human-machine interaction, including a strategy preparation module and a strategy execution module;
[0155] The strategy preparation module, for each question, selects the corresponding group of conversation scripts configured for the question from the conversation script configuration module as the candidate questioning conversation scripts for this question; based on the questioning conversation scripts and the user's emotional state, selects the corresponding average emotional probability of different user groups; based on the questioning conversation scripts, selects the corresponding average probability of obtaining answers;
[0156] Specifically, the average emotional probability is selected based on the positive and negative emotional fluctuations and the questioning conversation scripts, and the corresponding positive and negative average probabilities are selected.
[0157] The strategy execution module sets the emotional gain and the gain of obtaining answers, calculates the expected gain of each round based on the selected corresponding average emotional probability and the average probability of obtaining answers, and formulates the robot's questioning strategy according to the expected gain; at the same time, constructs a follow-up Q&A strategy, that is, the Q&A strategy for asking questions again after the first question fails to match the expected answer, and selects the corresponding average follow-up emotional probability of different user groups based on the user emotion detection module.
[0158] Specifically, the policy execution module calculates the expected benefits based on the questioning phrases for different types of emotional experiences, selects the questioning phrases; calculates the expected benefits according to the questioning phrases, and selects the next round of Q&A strategy; uses the user emotion detection module to detect the user's speech reply to obtain the current user's emotional state, and uses the semantic analysis module to analyze the user's answer to determine whether the current user has answered the question; if the user has answered the question, it returns to continue to select the questioning phrases, calculates the expected benefits of each questioning phrase according to the average probability and benefit of the questioning phrases corresponding to the newly detected user emotional state, and the average probability and benefit of whether an answer can be obtained, and determines the questioning phrase for the second question according to the questioning phrase corresponding to the maximum expected benefit; if it is analyzed that the user has not answered the question, it is necessary to consider that continuous questioning will also cause emotional fluctuations at this time, calculates the average probability and benefit of continuous questioning corresponding to the detected user emotional state, and the average probability and benefit of whether an answer can be obtained, and calculates the dialogue strategy of whether to execute continuous questioning. If the calculated benefit is less than the set threshold, the dialogue strategy of not executing continuous questioning is not executed, and it returns to continue to select the questioning phrases, calculates the expected benefits of each questioning phrase according to the average probability and benefit of the questioning phrases corresponding to the newly detected user emotional state, and the average probability and benefit of whether an answer can be obtained, and determines the questioning phrase for the second question according to the questioning phrase corresponding to the maximum expected benefit; if the calculated benefit is greater than the set threshold, the dialogue strategy of executing continuous questioning is executed... and so on until the selection of the Q&A strategy is completed.
[0159] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An adaptive outbound call dialogue strategy configuration method based on the user's emotional state, characterized in that It includes the following steps: Step S1: Build a user emotion detection module and train to generate a user emotion detection model for analyzing and detecting the speech input by the user to obtain the user's emotion state; Step S2: Build a semantic analysis module to analyze whether the user's reply conforms to the expected answer; Step S3: Build a conversation strategy configuration module to store and configure the conversation strategies and expected answers corresponding to each question; Step S4: Build a question-and-answer strategy processing module to formulate the robot's question strategy for the next round based on the results of each round of human-machine interaction, including the following steps: Step S4.1: For each question, select the answer group configured for the corresponding question from the conversation strategy configuration module as the candidate question strategy for this question; Step S4.2: Based on the question strategy and the user's emotion state, select the corresponding average emotion probability of different user groups; Step S4.3: Based on the question strategy and the completeness of the user feedback, select the corresponding average probability of obtaining an answer; Step S4.4: Set the emotion gain and the gain of obtaining an answer, calculate the expected gain of each round based on the selected corresponding average emotion probability and the average probability of obtaining an answer, and formulate the robot's question strategy according to the expected gain.
2. The method for configuring an adaptive outbound call dialogue strategy based on the user's emotional state according to claim 1, characterized in that: In step S4.2, based on the positive and negative emotion fluctuations and the question strategy, select the corresponding positive and negative average probabilities; in step S4.3, based on whether an answer is obtained or not and the question strategy, select the corresponding average probabilities of obtaining an answer and not obtaining an answer.
3. The adaptive outbound call dialogue strategy configuration method based on the user's emotional state according to claim 1, wherein: Step S4.4 includes the following steps: Step S4.4.1: Calculate the expected gain based on the question strategies of different types of emotion experiences, and select the question strategy; Step S4.4.2: Calculate the expected gain according to the question strategy and select the question-and-answer strategy for the next round; Use the user emotion detection module in step S1 to detect the user's speech reply to obtain the current user's emotion state, and use the semantic analysis module in step S2 to analyze the user's answer to determine whether the current user has answered the question; If the user has answered the question, return to step S4.4.1, calculate the expected gain of each question strategy based on the average probability and gain of the question strategy corresponding to the new user emotion state detected, and the average probability and gain of whether an answer can be obtained, and determine the question strategy of the second question according to the question strategy corresponding to the maximum expected gain.
4. The adaptive outbound call conversation strategy configuration method based on the user emotion state according to claim 3, characterized in that: In step S4.2, construct a follow-up question-and-answer strategy, that is, a question-and-answer strategy for asking questions again after the first question fails to obtain a clear answer, and based on step S1, select the corresponding average follow-up emotion probability of different user groups; In step S4.4.2, if it is analyzed that the user has not answered the question, calculate the dialogue strategy of whether to execute a follow-up question based on the average follow-up probability and benefit corresponding to the detected user emotional state, as well as the average probability and benefit of whether an answer can be obtained. If the calculated benefit is less than the set threshold, do not execute the dialogue strategy of the follow-up question, return to step S4.4.1, and calculate the expected benefit of each question formula based on the average probability and benefit of the question formula corresponding to the newly detected user emotional state, as well as the average probability and benefit of whether an answer can be obtained. Determine the question formula of the second question according to the question formula corresponding to the maximum expected benefit; if the calculated benefit is greater than the set threshold, execute the dialogue strategy of the follow-up question.
5. An adaptive outbound call dialogue strategy configuration device based on the user emotional state, including a user emotion detection module, a semantic analysis module, a formula configuration module, and a question and answer strategy processing module, characterized in that: The user emotion detection module generates an emotion detection model of the user through training, and is used to analyze and detect the voice input by the user to obtain the user's emotional state; The semantic analysis module analyzes whether the user's reply conforms to the expected answer; The formula configuration module stores and configures the formula and expected answer corresponding to each question; The question and answer strategy processing module formulates the robot's question strategy for the next round based on the human-computer interaction results of each round, including a strategy preparation module and a strategy execution module; The strategy preparation module, for each question, selects the formula group configured for the corresponding question from the formula configuration module as the candidate question formula for the question; based on the question formula and the user's emotional state, selects the corresponding average emotion probability of different user groups; based on the question formula, selects the corresponding average probability of obtaining an answer; The strategy execution module sets the emotion benefit and the benefit of obtaining an answer, calculates the expected benefit of each round based on the selected corresponding average emotion probability and the average probability of obtaining an answer, and formulates the robot's question strategy according to the expected benefit.
6. The adaptive outbound call dialogue strategy configuration device based on the user's emotional state according to claim 5, characterized in that: The user emotion detection module includes a first voice receiving module and a voice emotion analysis module, and the voice emotion analysis module includes a preprocessing module and a voice emotion analysis model; The first voice receiving module is used to receive the user's voice signal; The preprocessing module uses voice voice endpoint detection to extract the audio data containing human voices in the received voice; The voice emotion analysis model is a pre-trained model inference file, and can infer the corresponding emotional state according to the input voice data.
7. The adaptive outbound call dialogue strategy configuration device based on the user's emotional state according to claim 5, characterized in that: The semantic analysis module includes a second voice receiving module and a semantic analysis module, and the semantic analysis module includes a voice-to-text module and a keyword matching module; The second voice receiving module is used to receive the user's voice signal; The voice-to-text module uses an automatic speech recognition model to convert speech into text; The keyword matching module uses a candidate keyword dictionary and a regular matcher to perform regular matching on the text converted from the speech with the keyword dictionary, and the matched keywords are used as candidate answers.
8. The adaptive outbound call dialogue policy configuration device based on the user's emotional state according to claim 5, characterized in that: The strategy preparation module selects the corresponding positive and negative probability means based on positive and negative mood fluctuations and the questioning phrases.
9. The adaptive outbound call dialogue policy configuration device based on the user's emotional state according to claim 5, characterized in that: The strategy execution module calculates the expected return based on the questioning phrases for different types of emotional experiences, selects the questioning phrase, calculates the expected return according to the questioning phrase, and selects the next round of Q&A strategy. Using the user emotion detection module to detect the user's voice reply to obtain the current user's emotion state, and using the semantic analysis module to analyze the user's answer to determine whether the current user has answered the question; if the user has answered the question, return to continue selecting the questioning phrase, and calculate the expected return of each questioning phrase according to the probability mean and return of the questioning phrase corresponding to the newly detected user emotion state, as well as the probability mean and return of whether an answer can be obtained, and determine the questioning phrase of the second question according to the questioning phrase corresponding to the maximum expected return.
10. The adaptive outbound dialogue strategy configuration device based on the user emotion state according to claim 9, characterized in that: In the strategy execution module, a follow-up Q&A strategy is constructed, that is, a Q&A strategy for asking questions again after the first question fails to match the expected answer. Based on the user emotion detection module, the corresponding follow-up emotion probability means of different user groups are selected; if it is analyzed that the user has not answered the question, calculate whether to execute the follow-up dialogue strategy according to the follow-up probability mean and return of the user emotion state corresponding to the detection, as well as the probability mean and return of whether an answer can be obtained. If the calculated return is less than the set threshold, do not execute the follow-up dialogue strategy, return to continue selecting the questioning phrase, and calculate the expected return of each questioning phrase according to the probability mean and return of the questioning phrase corresponding to the newly detected user emotion state, as well as the probability mean and return of whether an answer can be obtained, and determine the questioning phrase of the second question according to the questioning phrase corresponding to the maximum expected return; if the calculated return is greater than the set threshold, execute the follow-up dialogue strategy.
Citation Information
Patent Citations
Voice emotion recognition method and device
CN108122552B
A voice emotion recognition and application system for call center calls
CN109767791B
Customer service robot dialogue method based on reinforcement learning and related components thereof
CN112507094A
Question and answer data processing method and device based on question and answer equipment
CN113051375A