Intelligent dialogue control method based on multi-modal emotion perception and related device
Through multi-modal emotion perception technology, the intelligent outgoing robot all-in-one machine processes user voice and audio data, recognizes emotions and semantics, and generates timely response audio data, solving the problem of inability to accurately respond to complex dialogues in the existing technology and improving user experience.
Patent Information
- Application Number
- CN202510486867.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-08-26
AI Technical Summary
The existing intelligent outgoing robot all-in-one machine cannot respond in a timely and accurate manner when facing emotional conversations, resulting in poor user experience.
Through multimodal emotion perception technology, voice audio data is collected using microphones, tone feature matching, speech emotion recognition and semantic analysis are performed, and response audio data matching the user's emotional state is generated.
It realizes timely capture and respond to user emotions and improves user conversation experience.
Smart Images

Figure CN120544555A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an intelligent dialogue control method based on multimodal emotion perception and related devices. Background Art
[0002] Some all-in-one intelligent outbound call robots may not only have the outbound call function integrated, but may also have the conversation and chat function integrated; however, even if the conversation and chat function is integrated, it is generally some simple conversation and chat functions integrated, which can only realize some simple chat conversations. It cannot respond to some emotional conversations or more complex conversations in a timely manner, or cannot respond accurately, so it cannot bring a good user experience to the target users. Summary of the Invention
[0003] The purpose of the present invention is to overcome the shortcomings of the existing technology. The present invention provides an intelligent dialogue control method and related devices based on multimodal emotion perception, which can timely capture the emotional state data of the target user, and timely generate response audio data based on the emotional state data and semantic information data to respond, thereby achieving timely comfort for the target user.
[0004] In order to solve the above technical problems, an embodiment of the present invention provides an intelligent dialogue control method based on multimodal emotion perception, which is applied to an intelligent outbound call robot all-in-one machine. The method includes:
[0005] The intelligent outbound call robot all-in-one enters the intelligent conversation monitoring mode, performs conversation monitoring processing on the voice and audio data of the target user in the current environment, and obtains the conversation monitoring result;
[0006] Entering a conversation mode based on the conversation monitoring result, and performing speech emotion recognition processing on the speech audio data of the target user to obtain speech emotion data of the target user in the speech audio data;
[0007] Performing semantic analysis on the speech and audio data of the target user to obtain semantic information data of the target user in the speech and audio data;
[0008] Obtaining emotional state data of the target user based on the speech emotion data and the semantic information data analysis;
[0009] A response strategy corresponding to the emotional state data is indexed based on the emotional state data, and response audio data of the target user is generated based on the response strategy.
[0010] Optionally, performing conversation monitoring processing on the voice and audio data of the target user in the current environment to obtain a conversation monitoring result includes:
[0011] When the intelligent outbound call robot all-in-one is in the intelligent dialogue monitoring mode, the microphone device carried by the intelligent outbound call robot starts to collect and process the audio data in the current environment to obtain voice audio data;
[0012] Extracting timbre feature data from the speech audio data, and matching the timbre feature data with target timbre feature data corresponding to the target user to obtain a matching result;
[0013] Performing conversation monitoring processing on the voice audio data based on the matching result to obtain a conversation monitoring result.
[0014] Optionally, the entering into a conversation mode based on the conversation monitoring result and performing speech emotion recognition processing on the speech audio data of the target user to obtain speech emotion data of the target user in the speech audio data includes:
[0015] When the conversation monitoring result indicates that audio data corresponding to target timbre feature data exists in the speech audio data, extracting target audio data corresponding to the target timbre feature data from the speech audio data;
[0016] The target audio data is subjected to speech emotion recognition processing based on a discrete speech emotion recognition model to obtain speech emotion data in the target audio data.
[0017] Optionally, performing speech emotion recognition processing on the target audio data based on a discrete speech emotion recognition model to obtain speech emotion data in the target audio data includes:
[0018] Performing audio frame processing on the target audio data, and performing audio feature extraction processing on the formed target audio sequence frame data in units of frames to obtain initial audio feature data corresponding to the target audio data, wherein the initial audio feature data includes energy feature data, fundamental frequency feature data, and formant feature data;
[0019] Performing short-time Fourier transform processing on the initial audio feature data to form transformed audio feature data;
[0020] The transformed audio feature data is windowed using any short-time window as the target window to form global speech feature data;
[0021] Performing feature combination update processing on the global audio feature data and the emotion dimension information to form updated global audio feature data, wherein the emotion dimension information is obtained by the emotion dimension recognition model using the global audio feature data process prediction;
[0022] The updated global speech feature data is input into the discrete speech emotion recognition model for speech emotion recognition processing to obtain the speech emotion data in the target audio data.
[0023] Optionally, performing semantic analysis on the speech and audio data of the target user to obtain semantic information data of the target user in the speech and audio data includes:
[0024] Converting the speech and audio data into text using an automatic speech recognition algorithm to obtain text data corresponding to the speech and audio data;
[0025] Segmenting the text data into a plurality of segmentation character strings according to a segmentation strategy, and performing recognition processing on semantic sequence information in the plurality of segmentation character strings to obtain semantic sequence information corresponding to the plurality of segmentation character strings;
[0026] Using each of the plurality of segmentation character strings, perform candidate expression search processing in a preset expression set to obtain a candidate expression corresponding to each segmentation character string;
[0027] The semantic information data is generated based on a combination of a candidate expression corresponding to each segmentation character string and semantic sequence information corresponding to the plurality of segmentation character strings.
[0028] Optionally, obtaining the target user's emotional state data based on the speech emotion data and the semantic information data analysis includes:
[0029] Performing text preprocessing on the semantic information data to form semantic text data;
[0030] Performing feature conversion processing on the semantic text data to form a semantic text feature vector;
[0031] Performing sentiment decision processing on the semantic text feature vector based on a decision tree model to obtain a sentiment decision result corresponding to the semantic information data;
[0032] The speech emotion data and the emotion decision result are subjected to emotion multimodal weighted fusion analysis processing to obtain the emotional state data of the target user.
[0033] Optionally, indexing a response strategy corresponding to the emotional state data based on the emotional state data, and generating response audio data of the target user based on the response strategy includes:
[0034] Using the emotional state data, indexing a response strategy corresponding to the emotional state data in a response strategy database;
[0035] Based on the response strategy, response text data is generated according to the semantic information data, and the response text data is converted into response audio data of the target user according to the voice emotion data.
[0036] In addition, an embodiment of the present invention further provides an intelligent dialogue control device based on multimodal emotion perception, which is applied to an intelligent outbound call robot all-in-one machine, and the device includes:
[0037] Conversation monitoring module: used for the intelligent outbound call robot all-in-one machine to enter the intelligent conversation monitoring mode, conduct conversation monitoring processing on the voice and audio data of the target user in the current environment, and obtain conversation monitoring results;
[0038] Emotion recognition module: used to enter the conversation mode based on the conversation monitoring result, and perform speech emotion recognition processing on the speech audio data of the target user to obtain speech emotion data of the target user in the speech audio data;
[0039] Semantic analysis module: used to perform semantic analysis on the voice and audio data of the target user to obtain semantic information data of the target user in the voice and audio data;
[0040] Emotion analysis module: used to obtain the emotional state data of the target user based on the speech emotion data and the semantic information data analysis;
[0041] Response dialogue module: used to index the response strategy corresponding to the emotional state data based on the emotional state data, and generate the response audio data of the target user based on the response strategy.
[0042] In addition, an embodiment of the present invention further provides an electronic device, including a processor and a memory, wherein the processor runs a computer program or code stored in the memory to implement the adaptive dialogue generation method as described in any one of the above.
[0043] In addition, an embodiment of the present invention further provides a computer-readable storage medium for storing a computer program or code. When the computer program or code is executed by a processor, the adaptive dialogue generation method as described in any one of the above is implemented.
[0044] In an embodiment of the present invention, after the intelligent outbound call robot all-in-one machine enters the dialogue mode, voice emotion data is obtained by performing voice emotion recognition processing on the voice audio data; semantic analysis processing is performed on the voice audio data to obtain semantic information data; emotional state data is obtained based on the analysis of the voice emotion data and semantic information data; a response strategy corresponding to the emotional state data is indexed, and response audio data is generated based on the response strategy; the emotional state data of the target user can be captured in a timely manner, and response audio data can be generated in a timely manner according to the emotional state data and semantic information data for response, and the generated response audio data has a soothing effect on the emotional state data of the target user, thereby achieving timely soothing of the target user and improving the dialogue experience of the target user. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0046] Figure 1 is a flow chart of an intelligent dialogue control method based on multimodal emotion perception in an embodiment of the present invention;
[0047] Figure 2 is a flow chart of an intelligent dialogue control method based on multimodal emotion perception in another embodiment of the present invention;
[0048] Figure 3 Schematic diagram of the structure of an intelligent dialogue control device based on multimodal emotion perception in an embodiment of the present invention;
[0049] Figure 4 This is a schematic diagram of the structure of the electronic device in the embodiment of the present invention. DETAILED DESCRIPTION
[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0051] For example 1, please refer to Figure 1 , Figure 1 This is a flow chart of an intelligent dialogue control method based on multimodal emotion perception in an embodiment of the present invention.
[0052] like Figure 1 As shown, an intelligent dialogue control method based on multimodal emotion perception is applied to an intelligent outbound call robot all-in-one machine, the method comprising:
[0053] S101: The intelligent outbound call robot all-in-one enters an intelligent conversation monitoring mode, performs conversation monitoring processing on the voice and audio data of the target user in the current environment, and obtains a conversation monitoring result;
[0054] During the specific implementation of the present invention, the voice audio data of the target user in the current environment is subjected to conversation monitoring processing to obtain a conversation monitoring result, including: when the intelligent outbound call robot all-in-one is in the intelligent conversation monitoring mode, starting the microphone device it carries to collect and process the audio data in the current environment to obtain voice audio data; extracting the timbre feature data in the voice audio data, and using the timbre feature data to match the target timbre feature data corresponding to the target user to obtain a matching result; and performing conversation monitoring processing on the voice audio data based on the matching result to obtain a conversation monitoring result.
[0055] Specifically, the implementation carrier is an intelligent outbound call robot all-in-one machine, that is, when the intelligent outbound call robot all-in-one machine is in the power-on state, it will automatically be in the intelligent dialogue monitoring mode. In this mode, the intelligent outbound call robot all-in-one machine will monitor the voice and audio data of the target user in the current environment; that is, by starting the microphone device (pickup device) carried by the intelligent outbound call robot all-in-one machine, all audio data in the current environment are collected and processed, thereby obtaining the voice and audio data in the current environment; at this time, it is necessary to confirm whether there is the target user's speaking audio in the voice and audio data.
[0056] Therefore, it is necessary to extract the timbre feature data from the voice audio data, and then use the timbre feature data to match the target timbre feature data corresponding to the target user to obtain a matching result. Before using the intelligent outbound call robot all-in-one machine, it is necessary to register with the relevant information of the target user, and then process the target timbre feature data corresponding to the target user on the intelligent outbound call robot all-in-one machine.
[0057] Then, the matching results are used to perform conversation monitoring processing on the voice audio data to obtain the conversation monitoring results; that is, the conversation monitoring processing is performed on the voice audio data by matching whether there is audio data generated by the target user's speech in the matching results to obtain the conversation monitoring results.
[0058] Among them, conversation monitoring is to identify whether there is a conversation audio of the target user in the voice audio data. When there is a conversation audio of the target user, it can be confirmed that voice conversation processing is required. Otherwise, it can be confirmed that voice conversation processing is not required, thereby obtaining a conversation monitoring result.
[0059] S102: Entering a conversation mode based on the conversation monitoring result, and performing speech emotion recognition processing on the speech audio data of the target user to obtain speech emotion data of the target user in the speech audio data;
[0060] In the specific implementation process of the present invention, the method enters a conversation mode based on the conversation monitoring result, performs speech emotion recognition processing on the speech audio data of the target user, and obtains the speech emotion data of the target user in the speech audio data, including: when the conversation monitoring result is that there is audio data corresponding to target timbre feature data in the speech audio data, extracting target audio data corresponding to the target timbre feature data in the speech audio data; and performs speech emotion recognition processing on the target audio data based on a discrete speech emotion recognition model to obtain the speech emotion data in the target audio data.
[0061] Furthermore, the target audio data is subjected to speech emotion recognition processing based on a discrete speech emotion recognition model to obtain speech emotion data in the target audio data, including: performing audio frame processing on the target audio data, and performing audio feature extraction processing on the formed target audio sequence frame data in units of frames to obtain initial audio feature data corresponding to the target audio data, wherein the initial audio feature data includes energy feature data, fundamental frequency feature data and resonance peak feature data; performing short-time Fourier transform processing on the initial audio feature data to form transformed audio feature data; performing windowing processing on the transformed audio feature data using an arbitrary short-time window as a target window to form global speech feature data; performing feature combination update processing on the global audio feature data and emotion dimension information to form updated global audio feature data, wherein the emotion dimension information is obtained by prediction by the emotion dimension recognition model using the global audio feature data process; inputting the updated global speech feature data into the discrete speech emotion recognition model for speech emotion recognition processing to obtain speech emotion data in the target audio data.
[0062] Specifically, when the conversation monitoring result is that there is audio data corresponding to the target timbre feature data in the speech audio data, that is, when there is a conversation audio of the target user in the speech audio data, it is necessary to extract the target audio data corresponding to the target timbre feature data in the speech audio data; this step is mainly to extract the target audio data in the speech audio data by using the target timbre feature data as a template; and then use the discrete speech emotion recognition model to identify the speech emotion data in the target audio data, so as to obtain the speech emotion data in the target audio data.
[0063] When using a discrete speech emotion recognition model to identify speech emotion data in target audio data, it is first necessary to perform audio framing processing on the target audio data in the form of audio frames. The audio framing here is mainly implemented according to the audio acquisition frequency; after completing the audio framing, it is necessary to perform audio feature extraction processing on the formed target audio sequence frame data; at this time, it is necessary to implement audio feature extraction processing in frames to obtain the initial audio feature data corresponding to the target audio data. It can be understood that this initial audio feature data is basic acoustic feature data, and mainly includes energy feature data, fundamental frequency feature data and resonance peak feature data.
[0064] Taking into account that the frequency change in the speech signal is a parameter that is more likely to reflect the intensity of the speech, and Fourier transform is a common means of analyzing the frequency domain characteristics of the signal; taking into account that the emotion, speaking speed and pitch in the speech signal may change greatly in a short period of time, the short-time Fourier transform should be used when performing Fourier transform; that is, the short-time Fourier transform function transforms and analyzes the initial audio feature data to form the transformed audio feature data; thereby, the audio feature data can be segmented and the frequency characteristics of the audio feature data in a short period of time can be captured, so as to better reflect the time-varying characteristics of the signal, so that the changes in the frequency characteristics of intense emotions can be more accurately identified.
[0065] Then, in order to convert the transformed audio feature data into global audio feature data, windowing processing is required. In this technical solution, an arbitrary short-time window is used as the target window, and the target window is used to perform windowing processing on the transformed audio feature data to form global speech feature data; that is, the target window is used to perform statistical regression on the transformed audio feature data to obtain global speech feature data of fixed dimension.
[0066] First, it is necessary to obtain the emotional dimension information contained in the global speech feature data. At this time, it is necessary to use the emotional dimension recognition model to achieve this, that is, input the global audio feature data into the emotional dimension recognition model, and identify the emotional dimension information of the global audio feature data in the emotional dimension recognition model, wherein the emotional dimension recognition model is a random forest algorithm; wherein the random forest algorithm is a machine learning algorithm, which is an algorithm that can make random changes in the use of variables and data, generate a large number of classification trees, and then summarize the results of the classification trees to obtain prediction results; in this technical solution, the model can be used to well predict the emotional dimension information contained in the global speech feature data; after obtaining the emotional dimension information, it is necessary to add the emotional dimension information to the global audio feature data, so as to update the global audio feature data and obtain updated global audio feature data.
[0067] Finally, the updated global speech feature data is input into a discrete speech emotion recognition model for speech emotion recognition processing, thereby obtaining speech emotion data in the target audio data; in this technical solution, the discrete speech emotion recognition model can be a logistic regression model.
[0068] S103: Performing semantic analysis on the voice and audio data of the target user to obtain semantic information data of the target user in the voice and audio data;
[0069] In the specific implementation process of the present invention, the semantic analysis processing is performed on the voice and audio data of the target user to obtain the semantic information data of the target user in the voice and audio data, including: performing text conversion processing on the voice and audio data using an automatic speech recognition algorithm to obtain text data corresponding to the voice and audio data; dividing the text data into a number of word segmentation strings according to a word segmentation strategy, and identifying the semantic sequence information in the number of word segmentation strings to obtain the semantic sequence information corresponding to the number of word segmentation strings; using each of the number of word segmentation strings to perform candidate expression search processing in a preset expression set to obtain a candidate expression corresponding to each word segmentation string; generating the semantic information data based on a combination of the candidate expression corresponding to each word segmentation string and the semantic sequence information corresponding to the number of word segmentation strings.
[0070] Specifically, first, the voice audio data needs to be converted into text data. At this time, an automatic speech recognition algorithm needs to be used, that is, the voice audio data is converted into text using an automatic speech recognition algorithm to obtain text data corresponding to the voice audio data; then the text data needs to be divided into several word segmentation strings according to the word segmentation strategy, and the semantic sequence information in the several word segmentation strings is identified and processed to obtain the semantic sequence information corresponding to the several word segmentation strings; for example, the sentence "I want to go shopping" can be divided into "I", "going", "shopping", and the order can be recorded.
[0071] At this time, it is necessary to use each of the several segmentation strings to perform candidate expression search processing in the preset expression set, so as to obtain the candidate expression corresponding to each segmentation string; finally, the candidate expression corresponding to each segmentation string is combined with the semantic sequence information corresponding to the several segmentation strings, so as to generate semantic information data.
[0072] S104: Obtaining emotional state data of the target user based on the analysis of the voice emotion data and the semantic information data;
[0073] In the specific implementation process of the present invention, the emotional state data of the target user is obtained based on the analysis of the voice emotion data and the semantic information data, including: performing text preprocessing on the semantic information data to form semantic text data; performing feature conversion processing on the semantic text data to form a semantic text feature vector; performing emotion decision processing on the semantic text feature vector based on a decision tree model to obtain the emotion decision result corresponding to the semantic information data; performing emotion multimodal weighted fusion analysis processing on the voice emotion data and the emotion decision result to obtain the emotional state data of the target user.
[0074] Specifically, the semantic information needs to be processed accordingly, which is text preprocessing here, that is, preprocessing the semantic information into text to form semantic text data; the semantic text data is subjected to corresponding feature conversion processing, and the N-Gram feature in the field of natural speech processing is adopted here, that is, the content in the semantic text data is operated with an active window of size N according to the word to form a word fragment of length N. Each of these fragments becomes a Gram, and the occurrence frequency of all Grams is statistically analyzed accordingly, and filtered according to the preset threshold to form a key Gram list. This Gram list can be a vector feature space, and each Gram in the list is a feature vector dimension, thus forming a semantic text feature vector; the semantic text feature vector is subjected to emotional decision processing through the decision tree model, and the emotional decision result corresponding to the semantic information data can be obtained; finally, the speech emotional data and the emotional decision result can be subjected to emotional multimodal weighted fusion analysis and processing to obtain the emotional state data of the target user.
[0075] S105: Indexing a response strategy corresponding to the emotional state data based on the emotional state data, and generating response audio data of the target user based on the response strategy.
[0076] In the specific real-time process of the present invention, the response strategy corresponding to the emotional state data is indexed based on the emotional state data, and the response audio data of the target user is generated based on the response strategy, including: using the emotional state data to index the response strategy corresponding to the emotional state data in a response strategy database; generating response text data according to the semantic information data based on the response strategy, and converting the response text data into the response audio data of the target user according to the voice emotion data.
[0077] Specifically, first, a response strategy needs to be obtained, and response text data is generated through the response strategy, and finally response audio data is generated through the response text data; that is, first, the emotional state data needs to be searched in the response strategy database, and the response strategy corresponding to the emotional state data is indexed through the retrieval; wherein the response strategy database stores the corresponding response strategies under various emotional state data; then the indexed response strategy is used to generate the corresponding response text data according to the semantic information data, and the response text data is converted into the response audio data of the target user according to the voice emotion data, and then played to the target user through the playback device carried by the intelligent outbound call robot all-in-one machine.
[0078] In an embodiment of the present invention, after the intelligent outbound call robot all-in-one machine enters the dialogue mode, voice emotion data is obtained by performing voice emotion recognition processing on the voice audio data; semantic analysis processing is performed on the voice audio data to obtain semantic information data; emotional state data is obtained based on the analysis of the voice emotion data and semantic information data; a response strategy corresponding to the emotional state data is indexed, and response audio data is generated based on the response strategy; the emotional state data of the target user can be captured in a timely manner, and response audio data can be generated in a timely manner according to the emotional state data and semantic information data for response, and the generated response audio data has a soothing effect on the emotional state data of the target user, thereby achieving timely soothing of the target user and improving the dialogue experience of the target user.
[0079] For example 2, please refer to Figure 2 , Figure 2 It is a flowchart of an intelligent dialogue control method based on multimodal emotion perception in another embodiment of the present invention.
[0080] like Figure 2 As shown, an intelligent dialogue control method based on multimodal emotion perception is applied to an intelligent outbound call robot all-in-one machine, the method comprising:
[0081] S201: The intelligent outbound call robot all-in-one enters the intelligent conversation monitoring mode, performs conversation monitoring processing on the voice and audio data of the target user in the current environment, and obtains the conversation monitoring result;
[0082] S202: Entering a conversation mode based on the conversation monitoring result, and when the conversation monitoring result indicates that audio data corresponding to target timbre feature data exists in the speech audio data, extracting target audio data corresponding to the target timbre feature data from the speech audio data;
[0083] S203: performing audio frame processing on the target audio data, and performing audio feature extraction processing on the formed target audio sequence frame data in units of frames to obtain initial audio feature data corresponding to the target audio data, wherein the initial audio feature data includes energy feature data, fundamental frequency feature data, and formant feature data;
[0084] S204: Performing short-time Fourier transform processing on the initial audio feature data to form transformed audio feature data;
[0085] S205: performing windowing processing on the transformed audio feature data using any short-time window as a target window to form global speech feature data;
[0086] S206: performing feature combination update processing on the global audio feature data and the emotion dimension information to form updated global audio feature data, wherein the emotion dimension information is obtained by prediction by the emotion dimension recognition model using the global audio feature data;
[0087] S207: Inputting the updated global speech feature data into the discrete speech emotion recognition model to perform speech emotion recognition processing to obtain speech emotion data in the target audio data;
[0088] S208: Performing semantic analysis on the voice and audio data of the target user to obtain semantic information data of the target user in the voice and audio data;
[0089] S209: Obtaining emotional state data of the target user based on the speech emotion data and the semantic information data analysis;
[0090] S210: Indexing a response strategy corresponding to the emotional state data based on the emotional state data, and generating response audio data of the target user based on the response strategy.
[0091] The specific implementation of the second embodiment can be found in the above embodiments, which will not be described in detail here.
[0092] For example three, please refer to Figure 3 , Figure 3 This is a schematic diagram of the structural composition of an intelligent dialogue control device based on multimodal emotion perception in an embodiment of the present invention.
[0093] like Figure 3 As shown, an intelligent dialogue control device based on multimodal emotion perception is applied to an intelligent outbound call robot all-in-one machine, and the device includes:
[0094] Conversation monitoring module 301: used for the intelligent outbound call robot all-in-one machine to enter the intelligent conversation monitoring mode, perform conversation monitoring processing on the voice and audio data of the target user in the current environment, and obtain conversation monitoring results;
[0095] During the specific implementation of the present invention, the voice audio data of the target user in the current environment is subjected to conversation monitoring processing to obtain a conversation monitoring result, including: when the intelligent outbound call robot all-in-one is in the intelligent conversation monitoring mode, starting the microphone device it carries to collect and process the audio data in the current environment to obtain voice audio data; extracting the timbre feature data in the voice audio data, and using the timbre feature data to match the target timbre feature data corresponding to the target user to obtain a matching result; and performing conversation monitoring processing on the voice audio data based on the matching result to obtain a conversation monitoring result.
[0096] Specifically, the implementation carrier is an intelligent outbound call robot all-in-one machine, that is, when the intelligent outbound call robot all-in-one machine is in the power-on state, it will automatically be in the intelligent dialogue monitoring mode. In this mode, the intelligent outbound call robot all-in-one machine will monitor the voice and audio data of the target user in the current environment; that is, by starting the microphone device (pickup device) carried by the intelligent outbound call robot all-in-one machine, all audio data in the current environment are collected and processed, thereby obtaining the voice and audio data in the current environment; at this time, it is necessary to confirm whether there is the target user's speaking audio in the voice and audio data.
[0097] Therefore, it is necessary to extract the timbre feature data from the voice audio data, and then use the timbre feature data to match the target timbre feature data corresponding to the target user to obtain a matching result. Before using the intelligent outbound call robot all-in-one machine, it is necessary to register with the relevant information of the target user, and then process the target timbre feature data corresponding to the target user on the intelligent outbound call robot all-in-one machine.
[0098] Then, the matching results are used to perform conversation monitoring processing on the voice audio data to obtain the conversation monitoring results; that is, the conversation monitoring processing is performed on the voice audio data by matching whether there is audio data generated by the target user's speech in the matching results to obtain the conversation monitoring results.
[0099] Among them, conversation monitoring is to identify whether there is a conversation audio of the target user in the voice audio data. When there is a conversation audio of the target user, it can be confirmed that voice conversation processing is required. Otherwise, it can be confirmed that voice conversation processing is not required, thereby obtaining a conversation monitoring result.
[0100] Emotion recognition module 302: used to enter the conversation mode based on the conversation monitoring result, and perform speech emotion recognition processing on the speech audio data of the target user to obtain speech emotion data of the target user in the speech audio data;
[0101] In the specific implementation process of the present invention, the method enters a conversation mode based on the conversation monitoring result, performs speech emotion recognition processing on the speech audio data of the target user, and obtains the speech emotion data of the target user in the speech audio data, including: when the conversation monitoring result is that there is audio data corresponding to target timbre feature data in the speech audio data, extracting target audio data corresponding to the target timbre feature data in the speech audio data; and performs speech emotion recognition processing on the target audio data based on a discrete speech emotion recognition model to obtain the speech emotion data in the target audio data.
[0102] Furthermore, the target audio data is subjected to speech emotion recognition processing based on a discrete speech emotion recognition model to obtain speech emotion data in the target audio data, including: performing audio frame processing on the target audio data, and performing audio feature extraction processing on the formed target audio sequence frame data in units of frames to obtain initial audio feature data corresponding to the target audio data, wherein the initial audio feature data includes energy feature data, fundamental frequency feature data and resonance peak feature data; performing short-time Fourier transform processing on the initial audio feature data to form transformed audio feature data; performing windowing processing on the transformed audio feature data using an arbitrary short-time window as a target window to form global speech feature data; performing feature combination update processing on the global audio feature data and emotion dimension information to form updated global audio feature data, wherein the emotion dimension information is obtained by prediction by the emotion dimension recognition model using the global audio feature data process; inputting the updated global speech feature data into the discrete speech emotion recognition model for speech emotion recognition processing to obtain speech emotion data in the target audio data.
[0103] Specifically, when the conversation monitoring result is that there is audio data corresponding to the target timbre feature data in the speech audio data, that is, when there is a conversation audio of the target user in the speech audio data, it is necessary to extract the target audio data corresponding to the target timbre feature data in the speech audio data; this step is mainly to extract the target audio data in the speech audio data by using the target timbre feature data as a template; and then use the discrete speech emotion recognition model to identify the speech emotion data in the target audio data, so as to obtain the speech emotion data in the target audio data.
[0104] When using a discrete speech emotion recognition model to identify speech emotion data in target audio data, it is first necessary to perform audio framing processing on the target audio data in the form of audio frames. The audio framing here is mainly implemented according to the audio acquisition frequency; after completing the audio framing, it is necessary to perform audio feature extraction processing on the formed target audio sequence frame data; at this time, it is necessary to implement audio feature extraction processing in frames to obtain the initial audio feature data corresponding to the target audio data. It can be understood that this initial audio feature data is basic acoustic feature data, and mainly includes energy feature data, fundamental frequency feature data and resonance peak feature data.
[0105] Taking into account that the frequency change in the speech signal is a parameter that is more likely to reflect the intensity of the speech, and Fourier transform is a common means of analyzing the frequency domain characteristics of the signal; taking into account that the emotion, speaking speed and pitch in the speech signal may change greatly in a short period of time, the short-time Fourier transform should be used when performing Fourier transform; that is, the short-time Fourier transform function transforms and analyzes the initial audio feature data to form the transformed audio feature data; thereby, the audio feature data can be segmented and the frequency characteristics of the audio feature data in a short period of time can be captured, so as to better reflect the time-varying characteristics of the signal, so that the changes in the frequency characteristics of intense emotions can be more accurately identified.
[0106] Then, in order to convert the transformed audio feature data into global audio feature data, windowing processing is required. In this technical solution, an arbitrary short-time window is used as the target window, and the target window is used to perform windowing processing on the transformed audio feature data to form global speech feature data; that is, the target window is used to perform statistical regression on the transformed audio feature data to obtain global speech feature data of fixed dimension.
[0107] First, it is necessary to obtain the emotional dimension information contained in the global speech feature data. At this time, it is necessary to use the emotional dimension recognition model to achieve this, that is, input the global audio feature data into the emotional dimension recognition model, and identify the emotional dimension information of the global audio feature data in the emotional dimension recognition model, wherein the emotional dimension recognition model is a random forest algorithm; wherein the random forest algorithm is a machine learning algorithm, which is an algorithm that can make random changes in the use of variables and data, generate a large number of classification trees, and then summarize the results of the classification trees to obtain prediction results; in this technical solution, the model can be used to well predict the emotional dimension information contained in the global speech feature data; after obtaining the emotional dimension information, it is necessary to add the emotional dimension information to the global audio feature data, so as to update the global audio feature data and obtain updated global audio feature data.
[0108] Finally, the updated global speech feature data is input into a discrete speech emotion recognition model for speech emotion recognition processing, thereby obtaining speech emotion data in the target audio data; in this technical solution, the discrete speech emotion recognition model can be a logistic regression model.
[0109] Semantic analysis module 303: used to perform semantic analysis on the speech and audio data of the target user to obtain semantic information data of the target user in the speech and audio data;
[0110] In the specific implementation process of the present invention, the semantic analysis processing is performed on the voice and audio data of the target user to obtain the semantic information data of the target user in the voice and audio data, including: performing text conversion processing on the voice and audio data using an automatic speech recognition algorithm to obtain text data corresponding to the voice and audio data; dividing the text data into a number of word segmentation strings according to a word segmentation strategy, and identifying the semantic sequence information in the number of word segmentation strings to obtain the semantic sequence information corresponding to the number of word segmentation strings; using each of the number of word segmentation strings to perform candidate expression search processing in a preset expression set to obtain a candidate expression corresponding to each word segmentation string; generating the semantic information data based on a combination of the candidate expression corresponding to each word segmentation string and the semantic sequence information corresponding to the number of word segmentation strings.
[0111] Specifically, first, the voice audio data needs to be converted into text data. At this time, an automatic speech recognition algorithm needs to be used, that is, the voice audio data is converted into text using an automatic speech recognition algorithm to obtain text data corresponding to the voice audio data; then the text data needs to be divided into several word segmentation strings according to the word segmentation strategy, and the semantic sequence information in the several word segmentation strings is identified and processed to obtain the semantic sequence information corresponding to the several word segmentation strings; for example, the sentence "I want to go shopping" can be divided into "I", "going", "shopping", and the order can be recorded.
[0112] At this time, it is necessary to use each of the several segmentation strings to perform candidate expression search processing in the preset expression set, so as to obtain the candidate expression corresponding to each segmentation string; finally, the candidate expression corresponding to each segmentation string is combined with the semantic sequence information corresponding to the several segmentation strings, so as to generate semantic information data.
[0113] Emotion analysis module 304: used to obtain the emotional state data of the target user based on the speech emotion data and the semantic information data analysis;
[0114] In the specific implementation process of the present invention, the emotional state data of the target user is obtained based on the analysis of the voice emotion data and the semantic information data, including: performing text preprocessing on the semantic information data to form semantic text data; performing feature conversion processing on the semantic text data to form a semantic text feature vector; performing emotion decision processing on the semantic text feature vector based on a decision tree model to obtain the emotion decision result corresponding to the semantic information data; performing emotion multimodal weighted fusion analysis processing on the voice emotion data and the emotion decision result to obtain the emotional state data of the target user.
[0115] Specifically, the semantic information needs to be processed accordingly, which is text preprocessing here, that is, preprocessing the semantic information into text to form semantic text data; the semantic text data is subjected to corresponding feature conversion processing, and the N-Gram feature in the field of natural speech processing is adopted here, that is, the content in the semantic text data is operated with an active window of size N according to the word to form a word fragment of length N. Each of these fragments becomes a Gram, and the occurrence frequency of all Grams is statistically analyzed accordingly, and filtered according to the preset threshold to form a key Gram list. This Gram list can be a vector feature space, and each Gram in the list is a feature vector dimension, thus forming a semantic text feature vector; the semantic text feature vector is subjected to emotional decision processing through the decision tree model, and the emotional decision result corresponding to the semantic information data can be obtained; finally, the speech emotional data and the emotional decision result can be subjected to emotional multimodal weighted fusion analysis and processing to obtain the emotional state data of the target user.
[0116] The response dialogue module 305 is configured to index a response strategy corresponding to the emotional state data based on the emotional state data, and generate response audio data of the target user based on the response strategy.
[0117] In the specific real-time process of the present invention, the response strategy corresponding to the emotional state data is indexed based on the emotional state data, and the response audio data of the target user is generated based on the response strategy, including: using the emotional state data to index the response strategy corresponding to the emotional state data in a response strategy database; generating response text data according to the semantic information data based on the response strategy, and converting the response text data into the response audio data of the target user according to the voice emotion data.
[0118] Specifically, first, a response strategy needs to be obtained, and response text data is generated through the response strategy, and finally response audio data is generated through the response text data; that is, first, the emotional state data needs to be searched in the response strategy database, and the response strategy corresponding to the emotional state data is indexed through the retrieval; wherein the response strategy database stores the corresponding response strategies under various emotional state data; then the indexed response strategy is used to generate the corresponding response text data according to the semantic information data, and the response text data is converted into the response audio data of the target user according to the voice emotion data, and then played to the target user through the playback device carried by the intelligent outbound call robot all-in-one machine.
[0119] In an embodiment of the present invention, after the intelligent outbound call robot all-in-one machine enters the dialogue mode, voice emotion data is obtained by performing voice emotion recognition processing on the voice audio data; semantic analysis processing is performed on the voice audio data to obtain semantic information data; emotional state data is obtained based on the analysis of the voice emotion data and semantic information data; a response strategy corresponding to the emotional state data is indexed, and response audio data is generated based on the response strategy; the emotional state data of the target user can be captured in a timely manner, and response audio data can be generated in a timely manner according to the emotional state data and semantic information data for response, and the generated response audio data has a soothing effect on the emotional state data of the target user, thereby achieving timely soothing of the target user and improving the dialogue experience of the target user.
[0120] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the program implements the intelligent dialogue control method of any of the above-described embodiments. The computer-readable storage medium includes, but is not limited to, any type of disk (including floppy disks, hard disks, optical disks, CD-ROMs, and magneto-optical disks), ROM (Read-Only Memory), RAM (Random Access Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory, magnetic cards, or optical cards. In other words, a storage device includes any medium that can store or transmit information in a readable form by a device (e.g., a computer or mobile phone), and can be a read-only memory, a disk, or an optical disk.
[0121] An embodiment of the present invention further provides a computer application program that runs on a computer and is used to execute the intelligent dialogue control method of any one of the above embodiments.
[0122] also, Figure 4 It is a schematic diagram of the structure of an electronic device in an embodiment of the present invention.
[0123] The embodiment of the present invention further provides an electronic device, such as Figure 4 The electronic device includes a processor 402, a memory 403, an input unit 404, a display unit 405 and other components. Those skilled in the art will understand that Figure 4The structural components of the electronic device shown do not constitute a limitation on all devices, and may include more or fewer components than shown, or combine certain components. The memory 403 can be used to store the application 401 and various functional modules, and the processor 402 runs the application 401 stored in the memory 403, thereby executing various functional applications and data processing of the device. The memory can be an internal memory or an external memory, or include both internal and external memories. The internal memory may include a read-only memory (ROM), a programmable ROM (PROM), an electrically programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a flash memory, or a random access memory. The external memory may include a hard disk, a floppy disk, a ZIP disk, a USB flash drive, a magnetic tape, etc. The memory disclosed in the present invention includes but is not limited to these types of memories. The memory disclosed in the present invention is only an example and not a limitation.
[0124] The input unit 404 is used to receive input signals and keywords entered by the user. The input unit 404 may include a touch panel and other input devices. The touch panel can collect user touch operations on or near it (such as operations performed by the user using a finger, stylus, or any other suitable object or accessory on or near the touch panel) and drive the corresponding connected device according to a pre-set program; other input devices may include, but are not limited to, one or more of a physical keyboard, function keys (such as playback control keys, on / off keys, etc.), a trackball, a mouse, a joystick, etc. The display unit 405 can be used to display information entered by the user or information provided to the user, as well as various menus of the terminal device. The display unit 405 can be in the form of a liquid crystal display, an organic light-emitting diode, etc. The processor 402 is the control center of the terminal device, connecting the various parts of the entire device using various interfaces and lines. It performs various functions and processes data by running or executing software programs and / or modules stored in the memory 403 and calling data stored in the memory.
[0125] As an embodiment, the electronic device includes: one or more processors 402, a memory 403, and one or more applications 401, wherein the one or more applications 401 are stored in the memory 403 and are configured to be executed by the one or more processors 402, and the one or more applications 401 are configured to execute the corresponding intelligent dialogue control method in any one of the above embodiments.
[0126] In an embodiment of the present invention, after the intelligent outbound call robot all-in-one machine enters the dialogue mode, voice emotion data is obtained by performing voice emotion recognition processing on the voice audio data; semantic analysis processing is performed on the voice audio data to obtain semantic information data; emotional state data is obtained based on the analysis of the voice emotion data and semantic information data; a response strategy corresponding to the emotional state data is indexed, and response audio data is generated based on the response strategy; the emotional state data of the target user can be captured in a timely manner, and response audio data can be generated in a timely manner according to the emotional state data and semantic information data for response, and the generated response audio data has a soothing effect on the emotional state data of the target user, thereby achieving timely soothing of the target user and improving the dialogue experience of the target user.
[0127] In addition, the above is a detailed introduction to an intelligent dialogue control method based on multimodal emotion perception and related devices provided by an embodiment of the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. An intelligent dialogue control method based on multimodal emotion perception, characterized in that: Applied to an intelligent outbound call robot all-in-one machine, the method includes: The intelligent outbound call robot all-in-one enters the intelligent conversation monitoring mode, performs conversation monitoring processing on the voice and audio data of the target user in the current environment, and obtains the conversation monitoring result; Entering a conversation mode based on the conversation monitoring result, and performing speech emotion recognition processing on the speech audio data of the target user to obtain speech emotion data of the target user in the speech audio data; Performing semantic analysis on the speech and audio data of the target user to obtain semantic information data of the target user in the speech and audio data; Obtaining emotional state data of the target user based on the speech emotion data and the semantic information data analysis; A response strategy corresponding to the emotional state data is indexed based on the emotional state data, and response audio data of the target user is generated based on the response strategy.
2. The intelligent dialogue control method according to claim 1, characterized in that: The performing conversation monitoring processing on the voice audio data of the target user in the current environment to obtain the conversation monitoring result includes: When the intelligent outbound call robot all-in-one is in the intelligent dialogue monitoring mode, the microphone device carried by the intelligent outbound call robot starts to collect and process the audio data in the current environment to obtain voice audio data; Extracting timbre feature data from the speech audio data, and matching the timbre feature data with target timbre feature data corresponding to the target user to obtain a matching result; Performing conversation monitoring processing on the voice audio data based on the matching result to obtain a conversation monitoring result.
3. The intelligent dialogue control method according to claim 1, characterized in that: The step of entering a conversation mode based on the conversation monitoring result and performing speech emotion recognition processing on the speech audio data of the target user to obtain speech emotion data of the target user in the speech audio data includes: When the conversation monitoring result indicates that audio data corresponding to target timbre feature data exists in the speech audio data, extracting target audio data corresponding to the target timbre feature data from the speech audio data; The target audio data is subjected to speech emotion recognition processing based on a discrete speech emotion recognition model to obtain speech emotion data in the target audio data.
4. The intelligent dialogue control method according to claim 3, characterized in that: The performing speech emotion recognition processing on the target audio data based on the discrete speech emotion recognition model to obtain speech emotion data in the target audio data includes: Performing audio frame processing on the target audio data, and performing audio feature extraction processing on the formed target audio sequence frame data in units of frames to obtain initial audio feature data corresponding to the target audio data, wherein the initial audio feature data includes energy feature data, fundamental frequency feature data, and formant feature data; Performing short-time Fourier transform processing on the initial audio feature data to form transformed audio feature data; The transformed audio feature data is windowed using any short-time window as the target window to form global speech feature data; Performing feature combination update processing on the global audio feature data and the emotion dimension information to form updated global audio feature data, wherein the emotion dimension information is obtained by the emotion dimension recognition model using the global audio feature data process prediction; The updated global speech feature data is input into the discrete speech emotion recognition model for speech emotion recognition processing to obtain the speech emotion data in the target audio data.
5. The intelligent dialogue control method according to claim 1, characterized in that: The performing semantic analysis on the speech and audio data of the target user to obtain semantic information data of the target user in the speech and audio data includes: Converting the speech and audio data into text using an automatic speech recognition algorithm to obtain text data corresponding to the speech and audio data; Segmenting the text data into a plurality of segmentation character strings according to a segmentation strategy, and performing recognition processing on semantic sequence information in the plurality of segmentation character strings to obtain semantic sequence information corresponding to the plurality of segmentation character strings; Using each of the plurality of segmentation character strings, perform candidate expression search processing in a preset expression set to obtain a candidate expression corresponding to each segmentation character string; The semantic information data is generated based on a combination of a candidate expression corresponding to each segmentation character string and semantic sequence information corresponding to the plurality of segmentation character strings.
6. The intelligent dialogue control method according to claim 1, characterized in that: The obtaining of the target user's emotional state data based on the speech emotion data and the semantic information data analysis includes: Performing text preprocessing on the semantic information data to form semantic text data; Performing feature conversion processing on the semantic text data to form a semantic text feature vector; Performing sentiment decision processing on the semantic text feature vector based on a decision tree model to obtain a sentiment decision result corresponding to the semantic information data; The speech emotion data and the emotion decision result are subjected to emotion multimodal weighted fusion analysis processing to obtain the emotional state data of the target user.
7. The intelligent dialogue control method according to claim 1, characterized in that: The step of indexing a response strategy corresponding to the emotional state data based on the emotional state data, and generating response audio data of the target user based on the response strategy, includes: Using the emotional state data, indexing a response strategy corresponding to the emotional state data in a response strategy database; Based on the response strategy, response text data is generated according to the semantic information data, and the response text data is converted into response audio data of the target user according to the voice emotion data.
8. An intelligent dialogue control device based on multimodal emotion perception, characterized in that: Applied to an intelligent outbound call robot all-in-one machine, the device includes: Conversation monitoring module: used for the intelligent outbound call robot all-in-one machine to enter the intelligent conversation monitoring mode, conduct conversation monitoring processing on the voice and audio data of the target user in the current environment, and obtain conversation monitoring results; Emotion recognition module: used to enter the conversation mode based on the conversation monitoring result, and perform speech emotion recognition processing on the speech audio data of the target user to obtain speech emotion data of the target user in the speech audio data; Semantic analysis module: used to perform semantic analysis on the voice and audio data of the target user to obtain semantic information data of the target user in the voice and audio data; Emotion analysis module: used to obtain the emotional state data of the target user based on the speech emotion data and the semantic information data analysis; Response dialogue module: used to index the response strategy corresponding to the emotional state data based on the emotional state data, and generate the response audio data of the target user based on the response strategy.
9. An electronic device comprising a processor and a memory, characterized in that: The processor runs the computer program or code stored in the memory to implement the adaptive dialogue generation method according to any one of claims 1 to 7.
10. A computer-readable storage medium for storing a computer program or code, characterized in that: When the computer program or code is executed by a processor, the adaptive dialogue generation method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Intelligent outbound robot dialogue regulation and control system fused with multi-mode emotion recognition
CN121037501A