Interaction method and device based on emotion recognition, electronic device, and storage medium

By acquiring user data through emotion recognition technology, identifying intent and emotions, and combining this with interaction strategies to generate interaction data, the problem of poor user experience and insufficient intelligence in existing AI interaction products is solved, achieving more efficient user interaction and making it applicable to multiple fields.

CN115860000BActive Publication Date: 2026-02-10ANHUI IFLYHEALTH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211600338.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-13
Publication Date
2026-02-10
Estimated Expiration
2042-12-13

AI Technical Summary

Technical Problem

Existing terminal products with AI interaction capabilities rely on sentence segmentation nodes for speech generation, resulting in a poor user experience, insufficient intelligence, and speech output dependent on existing data.

Method used

By using emotion recognition-based interaction methods, user data is acquired, intent is understood, and emotion is recognized. Interaction data is generated by combining interaction strategies, and more complete interaction data is generated by using multimodal feature fusion and cross-modal feature extraction.

Benefits of technology

It improves the reliability and accuracy of interaction, enhances the user experience, and improves intelligence. It is suitable for depression recognition and clinical diagnosis, and can be applied in fields such as banking, tourism, education, healthcare, and recruitment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115860000B_ABST
    Figure CN115860000B_ABST
Patent Text Reader

Abstract

The application provides an interactive method and device based on emotion recognition, an electronic device and a storage medium, wherein the method comprises: obtaining user data; performing intention understanding on the user data based on data features of the user data to obtain an interactive intention; generating interactive data corresponding to the user data based on the correlation between the data features and the dialogue features of the interactive strategy corresponding to the interactive intention; and the dialogue features are obtained by performing feature extraction on each sample interactive data under the interactive strategy. The method, device, electronic device and storage medium provided by the application can extract more complete semantic feature information, thereby ensuring that the generated interactive data corresponding to the user data can refer to the information provided by the complete and comprehensive user data, thereby improving the reliability and accuracy of the interaction, further improving the intelligence of the interaction, and improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-computer interaction technology, and in particular to an interaction method, device, electronic device and storage medium based on emotion recognition. Background Technology

[0002] Currently, various terminal products with AI (Artificial Intelligence) interaction capabilities have emerged, which can help people solve problems in daily life and are widely used in fields such as banking, tourism, education, healthcare, and recruitment.

[0003] In existing technologies, the speech generation of terminal products with AI interaction capabilities usually relies on sentence segmentation nodes, which inevitably results in a stiff feel and a poor user experience. Furthermore, the speech output depends on existing data and lacks sufficient intelligence. Summary of the Invention

[0004] This invention provides an interaction method, device, electronic device, and storage medium based on emotion recognition, which addresses the shortcomings of existing terminal products with AI interaction capabilities, where the generation of dialogue usually relies on sentence segmentation nodes, which inevitably results in stiffness and a poor user experience. Furthermore, the dialogue output depends on existing data and lacks sufficient intelligence.

[0005] This invention provides an interaction method based on emotion recognition, comprising:

[0006] Obtain user data;

[0007] Based on the data characteristics of the user data, the user data is used to understand the intent and obtain the interaction intent;

[0008] Based on the correlation between the data features and the verbal features of the interaction strategy corresponding to the interaction intent, interaction data corresponding to the user data is generated;

[0009] The verbal features are obtained by extracting features from the interaction data of each sample under the interaction strategy.

[0010] According to an interaction method based on emotion recognition provided by the present invention, the data features are determined based on the following steps:

[0011] Modal data of at least two modalities in the user data are respectively feature-encoded to obtain modal-encoded features of the at least two modalities;

[0012] Based on the correlation between the modality coding features of the at least two modalities, feature fusion is performed on the modality coding features of the at least two modalities to obtain the data features.

[0013] According to an interaction method based on emotion recognition provided by the present invention, the method involves fusing the modality coding features of the at least two modalities based on the correlation between the modality coding features of the at least two modalities to obtain the data features, including:

[0014] Obtain cross-modal features that align the modal coding features of the non-reference modalities in the at least two modalities to the reference modalities in the at least two modalities;

[0015] Based on the correlation between the modality coding features of the at least two modalities and the cross-modality features, feature fusion is performed on the modality coding features of the at least two modalities to obtain the data features.

[0016] According to an interaction method based on emotion recognition provided by the present invention, generating interaction data corresponding to the user data based on the correlation between the data features and the verbal features of the interaction strategy corresponding to the interaction intent includes:

[0017] Obtain the emotional characteristics of the user data;

[0018] Based on the correlation between the data features and the verbal features of the interaction strategy corresponding to the interaction intent, as well as the emotional features, interaction data corresponding to the user data is generated.

[0019] According to an interaction method based on emotion recognition provided by the present invention, the step of obtaining the emotional features of the user data includes:

[0020] Obtain the confidence level of historical sentiment features obtained from the last interaction;

[0021] When the confidence level is greater than a preset threshold, emotion recognition is performed on the user data based on the data characteristics of the user data to obtain the emotion characteristics.

[0022] According to an interaction method based on emotion recognition provided by the present invention, after obtaining the confidence level of historical emotion features obtained from the previous interaction, the method further includes:

[0023] If the confidence level is less than or equal to the preset threshold, the emotional characteristics of the user data are determined based on the preset scale corpus, and the interaction data corresponding to the user data is generated based on the preset scale corpus.

[0024] According to an interaction method based on emotion recognition provided by the present invention, the step of performing intent understanding on the user data based on the data features of the user data to obtain the interaction intent includes:

[0025] Based on the historical interaction intent obtained from the previous interaction and the data characteristics of the user data, the user data is subjected to intent understanding to obtain the interaction intent;

[0026] The interaction intent and the historical interaction intent Figure 1 This means that the interaction intent is in an interaction phase following the historical interaction intent.

[0027] The present invention also provides an interactive device based on emotion recognition, comprising:

[0028] The acquisition unit is used to acquire user data;

[0029] The intent understanding unit is used to understand the intent of the user data based on the data characteristics of the user data to obtain the interaction intent;

[0030] The generation unit is used to generate interaction data corresponding to the user data based on the correlation between the data features and the verbal features of the interaction strategy corresponding to the interaction intent;

[0031] The verbal features are obtained by extracting features from the interaction data of each sample under the interaction strategy.

[0032] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the emotion recognition-based interaction method as described above.

[0033] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the emotion-based interaction method as described above.

[0034] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the emotion recognition-based interaction method as described above.

[0035] The present invention provides an interaction method, device, electronic device, and storage medium based on emotion recognition. Based on the correlation between data features and the verbal features of the interaction strategy corresponding to the interaction intent, interaction data corresponding to user data is generated. The verbal features are obtained by feature extraction of each sample interaction data under the interaction strategy. As a result, more complete semantic feature information can be extracted, thereby ensuring that the generated interaction data corresponding to user data can refer to the complete and comprehensive information provided by the user data. This improves the reliability and accuracy of the interaction, further enhances the intelligence of the interaction, and improves the user experience. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0037] Figure 1 This is one of the flowcharts illustrating the emotion recognition-based interaction method provided by the present invention;

[0038] Figure 2 This is a flowchart illustrating the process of determining data features provided by the present invention;

[0039] Figure 3 This is a flowchart illustrating step 220 in the emotion recognition-based interaction method provided by the present invention.

[0040] Figure 4 This is a flowchart illustrating step 130 in the emotion recognition-based interaction method provided by the present invention.

[0041] Figure 5 This is a flowchart illustrating step 131 in the emotion recognition-based interaction method provided by the present invention.

[0042] Figure 6 This is the second flowchart illustrating the emotion recognition-based interaction method provided by the present invention;

[0043] Figure 7 This is a schematic diagram of the structure of the emotion recognition-based interactive device provided by the present invention;

[0044] Figure 8 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0046] Currently, various AI-enabled interactive terminal products have emerged, helping people solve problems in daily life. For example, when people encounter problems in a restaurant, they can ask the interactive terminal for information such as the restaurant's address and phone number. It can understand people's needs through dialogue and provide corresponding services accordingly. It is expected that in the future, the application of AI interaction will become increasingly widespread, and its form will become more realistic, making users feel as if they are conversing with a real person. It will be widely used in fields such as banking, tourism, education, healthcare, and recruitment.

[0047] In related technologies, the speech generation of terminal products with AI interaction capabilities usually relies on sentence segmentation nodes, which inevitably leads to awkwardness and a poor user experience. Furthermore, the speech output depends on existing data and lacks sufficient intelligence.

[0048] To address the above problems, this invention provides an interaction method based on emotion recognition. Figure 1 This is one of the flowcharts illustrating the emotion recognition-based interaction method provided by this invention, such as... Figure 1 As shown, the method includes:

[0049] Step 110: Obtain user data.

[0050] Specifically, user data can be acquired, which is the data generated by the user during the interaction process. User data may include text data and audio data, or text data and video data, or text data, audio data and video data. This embodiment of the invention does not make specific limitations in this regard.

[0051] The text data here can be directly input by the user or obtained by transcribing the collected audio data into speech. This embodiment of the invention does not specifically limit this.

[0052] The audio data here can be obtained through a sound pickup device, which can be a smartphone, tablet, or smart appliance such as a speaker, television, or air conditioner. After the sound pickup device obtains the audio data through a microphone array, it can also amplify and reduce the noise of the audio data. This embodiment of the invention does not specifically limit this.

[0053] The video data here can be pre-shot and stored, or it can be real-time video stream data. This embodiment of the invention does not specifically limit this.

[0054] It is understandable that the user data here refers to the user who is interacting with the computer. That is, user data can include audio recordings of the user's speech during human-computer interaction, recorded videos, and text input by the user.

[0055] Step 120: Based on the data characteristics of the user data, perform intent understanding on the user data to obtain the interaction intent.

[0056] Specifically, after acquiring user data, the user intent can be understood based on the data characteristics of the user data to obtain the interaction intent. The data characteristics of the user data here reflect the feature-level information of the user data. Accordingly, the data characteristics may include text features and audio features, or text features and video features, or text features, audio features, and video features. This embodiment of the invention does not specifically limit this.

[0057] Furthermore, the data features can be obtained by encoding the modal data of at least two modalities in the user data to obtain modal coding features of at least two modalities, and then fusing the modal coding features of at least two modalities. Alternatively, they can be obtained by aligning the modal coding features of the non-reference modalities in the at least two modalities to the cross-modal features of the reference modalities in the at least two modalities, and then obtaining them based on the cross-modal features and the modal coding features of the at least two modalities. This embodiment of the invention does not specifically limit the specific methods used.

[0058] Here, intent understanding of user data refers to accurately understanding the user's intent at the semantic level based on user data.

[0059] The interaction intent here refers to the user's intent during the interaction process. The interaction intent can be a scale interaction intent, an intent to seek empathy, or an intent to solicit suggestions. This embodiment of the invention does not specifically limit this.

[0060] Step 130: Based on the correlation between the data features and the verbal features of the interaction strategy corresponding to the interaction intent, generate the interaction data corresponding to the user data;

[0061] The verbal features are obtained by extracting features from the interaction data of each sample under the interaction strategy.

[0062] Specifically, in related technologies, during the speech generation process of terminal products with AI interaction capabilities, user data and the interaction strategies corresponding to the interaction intentions are not associated, resulting in insufficient intelligence and a poor user experience. To address this issue, embodiments of the present invention associate the data characteristics of user data with the speech characteristics of the interaction strategies corresponding to the interaction intentions, thereby improving the intelligence of terminal products with AI interaction capabilities during the interaction process.

[0063] Furthermore, after obtaining the interaction intent, interaction data corresponding to the user data can be generated based on the correlation between data features and the verbal features of the interaction strategy corresponding to the interaction intent. Here, the interaction strategy refers to the psychological strategy corresponding to the interaction intent. The interaction strategy can be a strategy from a pre-collected strategy knowledge base, and there can be multiple interaction strategies.

[0064] The strategy knowledge base here can be a scale strategy knowledge base, an empathy strategy knowledge base, or a suggestion strategy knowledge base. This embodiment of the invention does not specifically limit this.

[0065] The knowledge base of scale strategies here may include the PHQ9 scale strategy, or the Hamilton Depression Scale (HAMD) scale strategy, or both the PHQ9 scale strategy and the HAMD scale strategy. This embodiment of the invention does not specifically limit this.

[0066] The verbal features here are obtained by extracting features from the interaction data of each sample under the interaction strategy. The interaction data of each sample here is the interaction data under the strategy collected in the strategy knowledge base in advance.

[0067] For example, multiple interaction coding features can be obtained by extracting features from the interaction data of each sample under the interaction strategy, and then the average value of the corresponding positions in the multiple interaction coding features can be obtained. Here, the feature extraction of each sample interaction data under the interaction strategy can be performed using the BERT (Bidirectional Encoder Representations from Transformer) model, or a natural language processing model such as m-BERT that has outstanding performance in tasks such as semantic understanding. This embodiment of the invention does not make specific limitations in this regard.

[0068] Among them, the speech feature S j The formula is as follows:

[0069] S j =Self-Pooling(BERT(K) idx ))

[0070] Among them, K idx It is the interaction data of each sample under the interaction strategy, BERT(K idx ) indicates that features are extracted from the interaction data of each sample under the interaction strategy, and Self-Pooling means taking the average of the corresponding positions among multiple interaction encoded features.

[0071] Here, the dimensions of the verbal features can be 256, 128, or 64, and this embodiment of the invention does not specifically limit them.

[0072] Here, the correlation between data features and the verbal features of the interaction strategy corresponding to the interaction intent can be achieved through Cross-Attention.

[0073] For example, data features and the verbal features of the interaction strategy corresponding to the interaction intent can be input into Cross-Attention. Cross-Attention then fuses these features to output fused semantic features. In essence, these fused semantic features combine the data features and the verbal features of the interaction strategy corresponding to the interaction intent, thus providing more complete feature information for generating subsequent interaction data corresponding to user data.

[0074] After outputting the fused semantic features, these features can be input into the decoder, which decodes them to generate the interactive data corresponding to the user data. Here, the decoder can use sequence decoding methods, Beam Search, Greedy Search, etc., to decode the fused semantic features; this embodiment of the invention does not specifically limit the method used.

[0075] Furthermore, the interaction data generated corresponding to the user data also takes into account the emotional characteristics of the user data, that is, emotion recognition can be performed on the user data, further improving the reliability and accuracy of the interaction. Here, emotion recognition can output the emotion category or the emotion intensity. The emotion category can be any one of happiness, anger, resentment, surprise, disgust, fear, and neutrality; taking depressive emotion as an example, the emotion category can be analyzing whether the test subject has depressive emotions, and the emotion intensity can be divided into no depressive emotions, mild depressive emotions, moderate depressive emotions, and severe depressive emotions. This embodiment of the invention does not make specific limitations on this.

[0076] The interactive data here refers to the data generated by a terminal product with AI interaction capabilities in response to the input user data, which is used for interaction.

[0077] The method provided in this invention generates interactive data corresponding to user data based on the correlation between data features and the verbal features of the interactive strategy corresponding to the interactive intent. The verbal features are obtained by feature extraction of each sample of interactive data under the interactive strategy. As a result, more complete semantic feature information can be extracted, thereby ensuring that the interactive data corresponding to the user data can refer to the complete and comprehensive information provided by the user data. This improves the reliability and accuracy of the interaction, further enhances the intelligence of the interaction, and improves the user experience.

[0078] Based on any of the above embodiments, the emotion-recognition-based interactive method provided by the embodiments of the present invention can be applied to the identification and quantitative analysis of depressive mood. Furthermore, the severity of depressive mood obtained based on emotion recognition in the embodiments of the present invention can be applied to clinical diagnosis as a reference factor for doctors in diagnosing depression. In addition, the severity of depressive mood obtained based on emotion recognition in the embodiments of the present invention can also be applied to medical record quality inspection, comparing the severity of depressive mood obtained from automated analysis with the severity of depression diagnosed by doctors in the medical records, thereby verifying the quality of the medical records. The embodiments of the present invention do not specifically limit this application.

[0079] Based on the above embodiments, Figure 2 This is a flowchart illustrating the process of determining data features provided by the present invention, such as... Figure 2 As shown, the data features are determined based on the following steps:

[0080] Step 210: Perform feature encoding on the modal data of at least two modalities in the user data to obtain the modal encoding features of the at least two modalities;

[0081] Step 220: Based on the correlation between the modal coding features of the at least two modalities, perform feature fusion on the modal coding features of the at least two modalities to obtain the data features.

[0082] Specifically, after obtaining user data, feature encoding can be performed on the modal data of at least two modalities in the user data to obtain modal encoding features of at least two modalities.

[0083] The term "at least two modalities" here refers to two or more modalities. At least two modalities may include a text modal and an audio modal, or a text modal and a video modal, or a text modal, an audio modal, and a video modal. This embodiment of the invention does not specifically limit the specific modalities mentioned.

[0084] Accordingly, the modal data of at least two modalities may include text modal data and audio modal data, or text modal data and video modal data, or text modal data, audio modal data and video modal data. The embodiments of the present invention do not specifically limit this.

[0085] Here, text modal data can be feature encoded to obtain text modal encoded features. The text modal data can be feature encoded using the BERT-BiLSTM (Bidirectional Encoder Representations from Transformer-Bi-directional Long Short-Term Memory) model, deep neural networks (DNN), or the Transformer model, etc. The embodiments of the present invention do not specifically limit this.

[0086] Here, the text modal data T can be represented as:

[0087] T = {t1, t2, t3, ..., t} N}

[0088] The text modal data can be divided into N segments.

[0089] Here, the audio modal data can be feature-encoded to obtain audio modal encoded features. The feature encoding of the audio modal data can be performed using the Conformer model, a cascaded multilayer convolutional neural network (CNN), or a DNN network. This embodiment of the invention does not specifically limit the specific method used.

[0090] Here, the audio modal data A can be represented as:

[0091] A = {a1, a2, a3, ..., a} N}

[0092] The audio modal data can be divided into N segments.

[0093] Here, feature encoding can be performed on the video modal data to obtain video modal encoded features. The feature encoding of the video modal data can be performed using a CNN-BiLSTM model, or a combination of CNN and DNN structures, etc. This embodiment of the invention does not make specific limitations on this.

[0094] Here, the video modal data V can be represented as:

[0095] V = {v1, v2, v3, ..., v} N}

[0096] The video modal data can be divided into N segments.

[0097] Accordingly, the modal coding features of at least two modalities may include text modal coding features and audio modal coding features, or text modal coding features and video modal coding features, or text modal coding features, audio modal coding features and video modal coding features. The embodiments of the present invention do not specifically limit this.

[0098] Here, at least two modal coding features H m It can be represented as:

[0099] H m ={h1,h2,h3,...,h N} m

[0100] Where m∈{T,A,V}.

[0101] Having obtained modality coding features of at least two modalities, feature fusion can be performed on the modality coding features of at least two modalities based on the correlation between them to obtain data features.

[0102] Here, based on the correlation between the modality-encoded features of at least two modalities, feature fusion is performed on the modality-encoded features of at least two modalities. This can be achieved using a Self-Attention model, concatenating the modality-encoded features of at least two modalities, or weighting the modality-encoded features of at least two modalities using an attention mechanism before concatenation. This embodiment of the invention does not specifically limit the method used. For example, data features can be represented as Self-Attention(H... T H V H A ).

[0103] The method provided in this invention fuses the modality coding features of at least two modalities based on the correlation between them, resulting in data features that reference more complete and comprehensive user data, thereby improving the reliability and accuracy of subsequent interactions.

[0104] Based on the above embodiments, Figure 3 This is a flowchart illustrating step 220 of the emotion-based interaction method provided by the present invention, as shown below. Figure 3 As shown, step 220 includes:

[0105] Step 221: Obtain the cross-modal features that align the modal coding features of the non-reference modalities in the at least two modalities to the reference modalities in the at least two modalities;

[0106] Step 222: Based on the correlation between the modal coding features of the at least two modalities and the cross-modal features, perform feature fusion on the modal coding features of the at least two modalities to obtain the data features.

[0107] Specifically, after obtaining the modality coding features of at least two modalities, cross-modality features can be obtained to align the modality coding features of the non-reference modalities in the at least two modalities to the reference modalities in the at least two modalities. Here, the reference modalities in the at least two modalities refer to the modalities that serve as reference standards in the at least two modalities. These can be pre-defined, for example, the text modalities in the at least two modalities can be used as the reference modalities.

[0108] The non-reference mode in the at least two modes here refers to the mode other than the reference mode in the at least two modes. For example, the audio mode and the video mode in the at least two modes can be used as non-reference modes. The embodiments of the present invention do not limit this.

[0109] Here, aligning the modal coding features of the non-reference modalities in at least two modalities to the cross-modal features of the reference modalities in at least two modalities can be achieved through a cross-modal attention mechanism. Since the text modality encompasses richer user semantic information, the modal coding features of the audio modality in at least two modalities can be aligned to the text modal coding features to obtain cross-modal features, and similarly, the modal coding features of the video modality in at least two modalities can be aligned to the text modal coding features to obtain cross-modal features. This embodiment of the invention does not specifically limit the specific approach taken in this way.

[0110] For example, cross-modal features can be represented as:

[0111] C A->T =CMM(H A H T )

[0112] C V->T =CMM(H V H T )

[0113] Among them, C A->T This represents the cross-modal features obtained by aligning the modal coding features of the audio modalities from at least two modalities to the text modal coding features. C V->T Aligning the modal coding features of video modalities from at least two modalities to text modal coding features yields cross-modal features. CMM stands for Cross-Modal Attention Mechanism. A H represents the audio modality coding feature. T H represents the text modality coding features. VThis represents the video modal coding features.

[0114] After obtaining the cross-modal features, feature fusion can be performed on the modal coding features of at least two modalities based on the correlation between the modal coding features of at least two modalities and the cross-modal features to obtain data features. The formula for data feature G is as follows:

[0115] G = Concat(Self-Attention(H) T H V H A ),C A->T C V->T )

[0116] Among them, Self-Attention (H T H V H A C represents a fused semantic feature determined based on the correlation between modality-encoded features of at least two modalities. A->T C V->T This indicates cross-modal features, and Concat indicates feature fusion operation.

[0117] Here, feature fusion is performed on the modality coding features of at least two modalities. This can be achieved by concatenating the modality coding features of at least two modalities, or by using an attention mechanism to weight the modality coding features of at least two modalities before concatenation. This embodiment of the invention does not impose specific limitations on this method.

[0118] The method provided in this invention fuses the modality coding features of at least two modalities based on the correlation between modality coding features of at least two modalities and cross-modality features. The resulting data features also incorporate feature information reflected by cross-modality features, allowing reference to more complete and comprehensive user data information, thereby improving the reliability and accuracy of subsequent emotion-based interaction.

[0119] Based on the above embodiments, Figure 4 This is a flowchart illustrating step 130 of the emotion-based interaction method provided by the present invention, as shown below. Figure 4 As shown, step 130 includes:

[0120] Step 131: Obtain the emotional characteristics of the user data;

[0121] Step 132: Based on the correlation between the data features and the verbal features of the interaction strategy corresponding to the interaction intent, as well as the emotional features, generate the interaction data corresponding to the user data.

[0122] Specifically, the emotional characteristics of user data can be obtained, which can be obtained by performing emotion recognition on user data.

[0123] After obtaining the emotional characteristics of user data, interaction data corresponding to the user data can be generated based on the correlation between the data characteristics and the verbal characteristics of the interaction strategy corresponding to the interaction intent, as well as the emotional characteristics.

[0124] For example, feature fusion can be performed based on the emotional features of user data and the correlation between the data features and the verbal features of the interaction strategy corresponding to the interaction intent. Then, based on the features obtained from feature fusion, interaction data corresponding to user data can be generated.

[0125] The method provided in this invention, based on the correlation between data features and the verbal features of the interaction strategy corresponding to the interaction intent, as well as emotional features, generates user data corresponding to interaction data that also takes into account emotional features, further improving the reliability and accuracy of the interaction.

[0126] Based on the above embodiments, Figure 5 This is a flowchart illustrating step 131 of the emotion-based interaction method provided by the present invention, as shown below. Figure 5 As shown, step 131 includes:

[0127] Step 1311: Obtain the confidence level of the historical emotional features obtained from the last interaction;

[0128] Step 1312: When the confidence level is greater than a preset threshold, perform emotion recognition on the user data based on the data features of the user data to obtain the emotion features.

[0129] Specifically, the confidence level of historical sentiment features obtained from the previous interaction can be obtained. This confidence level reflects the reliability of the historical sentiment features obtained from the previous interaction. Historical sentiment features refer to the features obtained by performing sentiment recognition on user data in the previous interaction before the current interaction.

[0130] When the confidence level is greater than a preset threshold, emotion recognition can be performed on the user data based on its data characteristics to obtain emotion features. The preset threshold here can be 0.8, 0.7, 0.9, etc., and this embodiment of the invention does not specifically limit it.

[0131] For example, the data features of user data can be input into an emotion recognition model to obtain the emotion features output by the emotion recognition model.

[0132] The method provided in this embodiment of the invention performs emotion recognition on user data based on the data features of user data when the confidence level is greater than a preset threshold, thereby obtaining emotion features and improving the reliability of the obtained emotion features.

[0133] Based on the above embodiments, step 1311 further includes:

[0134] If the confidence level is less than or equal to the preset threshold, the emotional characteristics of the user data are determined based on the preset scale corpus, and the interaction data corresponding to the user data is generated based on the preset scale corpus.

[0135] Specifically, when the confidence level is less than or equal to a preset threshold, the preset scale corpus may include the PHQ9 scale corpus, the HAMD scale corpus, or both the PHQ9 scale corpus and the HAMD scale corpus, etc. The embodiments of the present invention do not specifically limit this.

[0136] It is understandable that the pre-set scale corpus covers the characteristics of users' emotions. Therefore, the emotional characteristics of user data can be determined based on the pre-set scale corpus.

[0137] After determining the emotional characteristics of the user data, interactive data corresponding to the user data can be generated based on a pre-set scale corpus. For example, a response can be generated word by word based on the preceding text, and the generated response can be matched with sentences in the pre-set scale corpus. The sentence with the highest similarity to the sentence in the pre-set scale corpus is then used as the interactive data corresponding to the user data. Here, the similarity between sentences can be calculated using methods such as cosine similarity and Pearson correlation coefficient, and this embodiment of the invention does not specifically limit this method.

[0138] The method provided in this embodiment of the invention determines the emotional characteristics of user data based on a preset scale corpus when the confidence level is less than or equal to a preset threshold, and generates interactive data corresponding to the user data based on the preset scale corpus, thereby improving the convenience of generating interactive data corresponding to the user data.

[0139] Based on the above embodiments, step 120 includes:

[0140] Based on the historical interaction intent obtained from the previous interaction and the data characteristics of the user data, the user data is subjected to intent understanding to obtain the interaction intent;

[0141] The interaction intent and the historical interaction intent Figure 1 This means that the interaction intent is in an interaction phase following the historical interaction intent.

[0142] Specifically, based on the historical interaction intent obtained from the previous interaction and the data characteristics of the user data, the user data can be used to understand the intent and obtain the interaction intent.

[0143] The historical interaction intent here refers to the intent obtained by understanding the user data based on the data characteristics of the user data in the previous interaction before the current interaction. The historical interaction intent can be a scale interaction intent, an intent to seek empathy, or an intent to solicit suggestions. This embodiment of the invention does not specifically limit it in this way.

[0144] Understandably, the three interaction stages—scale interaction, seeking empathy, and soliciting advice—occur sequentially. In the scale interaction stage, it's necessary to ask users questions to capture details and understand the underlying causes of their current emotions. In the seeking empathy stage, after understanding the causes of the user's emotions, it's necessary to deepen the user's trust in the AI-enabled product through emotional resonance and self-disclosure, making the product (e.g., the chatbot) feel more approachable and thus more receptive. In the soliciting advice stage, the groundwork laid in the seeking empathy stage ensures better results in the third stage of communication, making users more willing to accept the chatbot's suggestions and ultimately helping them.

[0145] The interaction intent here is related to the historical interaction intent. Figure 1 This indicates that the interaction intent is in the interaction phase following a previous interaction intent.

[0146] For example, when the historical interaction intent is a scale interaction intent, the interaction intent can be a scale interaction intent, an intent to seek empathy, or an intent to solicit advice; when the historical interaction intent is an intent to seek empathy, the interaction intent can be an intent to seek empathy or an intent to solicit advice; when the historical interaction intent is an intent to solicit advice, the interaction intent can be an intent to solicit advice.

[0147] The method provided in this invention, based on the historical interaction intent obtained from the previous interaction and the data characteristics of user data, performs intent understanding on user data to obtain interaction intent, gradually understanding the user's interaction intent and improving the user experience.

[0148] Based on any of the above embodiments, the present invention provides an interaction method based on emotion recognition. Figure 6 This is the second flowchart illustrating the emotion-based interaction method provided by this invention, as shown below. Figure 6 As shown, the method includes:

[0149] The first step is to obtain user data.

[0150] The second step involves understanding the user data based on the historical interaction intent obtained from the previous interaction and the data characteristics of the user data to derive the interaction intent. This interaction intent is then compared with the historical interaction intent. Figure 1 This indicates that the interaction intent is in the interaction phase following a previous interaction intent.

[0151] The data features here are determined based on the following steps:

[0152] Modal data of at least two modalities in the user data are respectively feature-encoded to obtain modal-encoded features of the at least two modalities.

[0153] Obtain the cross-modal features of the non-reference modalities in at least two modalities aligned to the reference modalities in at least two modalities.

[0154] Based on the correlation between modal coding features of at least two modalities and cross-modal features, feature fusion is performed on the modal coding features of at least two modalities to obtain data features.

[0155] The third step is to obtain the confidence level of the historical emotional features obtained from the last interaction;

[0156] When the confidence level is greater than a preset threshold, emotion recognition is performed on the user data based on the data characteristics of the user data to obtain emotion features.

[0157] The fourth step is to generate interaction data corresponding to user data based on the correlation between data features and the verbal features of the interaction strategy corresponding to the interaction intent, as well as the emotional features. Here, the verbal features are obtained by feature extraction of each sample interaction data under the interaction strategy.

[0158] The interactive device provided by the present invention will be described below. The emotion recognition-based interactive device described below and the emotion recognition-based interactive method described above can be referred to in correspondence.

[0159] Based on any of the above embodiments, the present invention provides an interactive device based on emotion recognition. Figure 7 This is a schematic diagram of the structure of the emotion recognition-based interactive device provided by the present invention, as shown below. Figure 7 As shown, the device includes:

[0160] Acquisition unit 710 is used to acquire user data;

[0161] The intent understanding unit 720 is used to understand the user data based on the data characteristics of the user data to obtain the interaction intent;

[0162] The generation unit 730 is used to generate interaction data corresponding to the user data based on the correlation between the data features and the verbal features of the interaction strategy corresponding to the interaction intent;

[0163] The verbal features are obtained by extracting features from the interaction data of each sample under the interaction strategy.

[0164] The apparatus provided in this invention generates interactive data corresponding to user data based on the correlation between data features and the verbal features of the interactive strategy corresponding to the interactive intent. The verbal features are obtained by extracting features from the interactive data of each sample under the interactive strategy. As a result, more complete semantic feature information can be extracted, thereby ensuring that the interactive data corresponding to the user data can refer to the complete and comprehensive information provided by the user data. This improves the reliability and accuracy of the interaction, further enhances the intelligence of the interaction, and improves the user experience.

[0165] Based on any of the above embodiments, the data features are determined based on the following steps:

[0166] The feature encoding unit is used to perform feature encoding on the modal data of at least two modalities in the user data to obtain the modal encoding features of the at least two modalities;

[0167] The feature fusion unit is used to perform feature fusion on the modality coding features of the at least two modalities based on the correlation between the modality coding features of the at least two modalities to obtain the data features.

[0168] Based on any of the above embodiments, the feature fusion unit is specifically used for:

[0169] Obtain cross-modal feature units, used to obtain cross-modal features that align the modal coding features of the non-reference modalities in the at least two modalities to the reference modalities in the at least two modalities;

[0170] The feature fusion subunit is used to perform feature fusion on the modality coding features of the at least two modalities based on the correlation between the modality coding features of the at least two modalities and the cross-modality features, to obtain the data features.

[0171] Based on any of the above embodiments, the generating unit is specifically used for:

[0172] An emotion feature acquisition unit is used to acquire the emotion features of the user data;

[0173] A generation subunit is used to generate interaction data corresponding to the user data based on the correlation between the data features and the verbal features of the interaction strategy corresponding to the interaction intent, as well as the emotional features.

[0174] Based on any of the above embodiments, the emotion feature unit is specifically used for:

[0175] The confidence unit is used to obtain the confidence level of historical sentiment features obtained from the last interaction.

[0176] An emotion recognition unit is used to perform emotion recognition on the user data based on the data features of the user data when the confidence level is greater than a preset threshold, so as to obtain the emotion features.

[0177] Based on any of the above embodiments, after obtaining the confidence unit, the process further includes:

[0178] If the confidence level is less than or equal to the preset threshold, the emotional characteristics of the user data are determined based on the preset scale corpus, and the interaction data corresponding to the user data is generated based on the preset scale corpus.

[0179] Based on any of the above embodiments, the intent understanding unit is specifically used for:

[0180] Based on the historical interaction intent obtained from the previous interaction and the data characteristics of the user data, the user data is subjected to intent understanding to obtain the interaction intent;

[0181] The interaction intent and the historical interaction intent Figure 1 This means that the interaction intent is in an interaction phase following the historical interaction intent.

[0182] Figure 8 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 8 As shown, the electronic device may include a processor 810, a communications interface 820, a memory 830, and a communication bus 840, wherein the processor 810, communications interface 820, and memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute an emotion-based interaction method. This method includes: acquiring user data; performing intent understanding on the user data based on data features to obtain an interaction intent; and generating interaction data corresponding to the user data based on the correlation between the data features and the verbal features of the interaction strategy corresponding to the interaction intent. The verbal features are obtained by feature extraction from each sample interaction data under the interaction strategy.

[0183] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0184] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the emotion-based interaction method provided by the above methods. The method includes: acquiring user data; performing intent understanding on the user data based on the data features of the user data to obtain an interaction intent; generating interaction data corresponding to the user data based on the correlation between the data features and the verbal features of the interaction strategy corresponding to the interaction intent; wherein the verbal features are obtained by feature extraction of each sample interaction data under the interaction strategy.

[0185] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the emotion-based interaction method provided by the above methods. The method includes: acquiring user data; performing intent understanding on the user data based on data features to obtain an interaction intent; generating interaction data corresponding to the user data based on the correlation between the data features and the verbal features of the interaction strategy corresponding to the interaction intent; wherein the verbal features are obtained by feature extraction from each sample interaction data under the interaction strategy.

[0186] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0187] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0188] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An interaction method based on emotion recognition, characterized in that, include: Obtain user data; Based on the data characteristics of the user data, the user data is used to understand the intent and obtain the interaction intent; The interaction intents include scale interaction intents, empathy-seeking intents, and suggestion-seeking intents; Based on the correlation between the data features and the verbal features of the interaction strategy corresponding to the interaction intent, interaction data corresponding to the user data is generated; The dialogue features are obtained by taking the average of the corresponding positions among the multiple interaction coding features obtained by extracting features from the interaction data of each sample under the interaction strategy. The step of generating interaction data corresponding to the user data based on the correlation between the data features and the verbal features of the interaction strategy corresponding to the interaction intent includes: The data features and the verbal features of the interaction strategy corresponding to the interaction intent are fused to obtain fused semantic features; The fused semantic features are input into the decoder for decoding to generate the interaction data corresponding to the user data.

2. The interaction method based on emotion recognition according to claim 1, characterized in that, The data features are determined based on the following steps: Modal data of at least two modalities in the user data are respectively feature-encoded to obtain modal-encoded features of the at least two modalities; Based on the correlation between the modality coding features of the at least two modalities, feature fusion is performed on the modality coding features of the at least two modalities to obtain the data features.

3. The interaction method based on emotion recognition according to claim 2, characterized in that, The data features are obtained by fusing the modality coding features of the at least two modalities based on the correlation between the modality coding features of the at least two modalities, including: Obtain cross-modal features that align the modal coding features of the non-reference modalities in the at least two modalities to the reference modalities in the at least two modalities; Based on the correlation between the modality coding features of the at least two modalities and the cross-modality features, feature fusion is performed on the modality coding features of the at least two modalities to obtain the data features.

4. The interaction method based on emotion recognition according to claim 1, characterized in that, The step of generating interaction data corresponding to the user data based on the correlation between the data features and the verbal features of the interaction strategy corresponding to the interaction intent includes: Obtain the emotional characteristics of the user data; Based on the correlation between the data features and the verbal features of the interaction strategy corresponding to the interaction intent, as well as the emotional features, interaction data corresponding to the user data is generated.

5. The interaction method based on emotion recognition according to claim 4, characterized in that, The emotional features of the user data obtained include: Obtain the confidence level of historical sentiment features obtained from the last interaction; When the confidence level is greater than a preset threshold, emotion recognition is performed on the user data based on the data characteristics of the user data to obtain the emotion characteristics.

6. The interaction method based on emotion recognition according to claim 5, characterized in that, The process of obtaining the confidence level of historical emotional features obtained from the previous interaction further includes: If the confidence level is less than or equal to the preset threshold, the emotional characteristics of the user data are determined based on the preset scale corpus, and the interaction data corresponding to the user data is generated based on the preset scale corpus.

7. The interaction method based on emotion recognition according to any one of claims 1 to 6, characterized in that, The process of understanding the user data based on its data features to obtain the interaction intent includes: Based on the historical interaction intent obtained from the previous interaction and the data characteristics of the user data, the user data is subjected to intent understanding to obtain the interaction intent; The interaction intent is consistent with the historical interaction intent, or the interaction intent is in an interaction stage following the historical interaction intent.

8. An interactive device based on emotion recognition, characterized in that, include: The acquisition unit is used to acquire user data; The intent understanding unit is used to understand the intent of the user data based on the data characteristics of the user data to obtain the interaction intent; the interaction intent includes scale interaction intent, empathy-seeking intent, and suggestion-seeking intent; The generation unit is used to generate interaction data corresponding to the user data based on the correlation between the data features and the verbal features of the interaction strategy corresponding to the interaction intent; The dialogue features are obtained by taking the average of the corresponding positions among the multiple interaction coding features obtained by extracting features from the interaction data of each sample under the interaction strategy. The generation unit is specifically used for: The data features and the verbal features of the interaction strategy corresponding to the interaction intent are fused to obtain fused semantic features; The fused semantic features are input into the decoder for decoding to generate the interaction data corresponding to the user data.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the emotion recognition-based interaction method as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the emotion-based interaction method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Man-machine interaction method and man-machine interaction device

    CN110110169A