User emotion recognition method based on AI and voice data

By introducing a three-level emotion induction mechanism into speech emotion recognition technology, combining passive data obtained from natural language dialogue and voice interaction, the problem of low recognition accuracy when users actively hide emotions is solved, and a higher emotional recognition accuracy is achieved.

CN120148561AInactive Publication Date: 2025-06-13HENAN POLYTECHNIC
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510487810.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing voice emotion recognition technology is difficult to effectively obtain the user's true emotions, especially when users actively hide their emotions or express their implicit expressions, the recognition accuracy rate has dropped significantly.

Method used

User emotion recognition methods based on AI and voice data are used to obtain passive data through natural language dialogue or voice interaction, and when the emotion recognition model detects that emotions do not meet the discriminant threshold, a three-level emotion induction mechanism is triggered, including the first level, the second level and the third level of emotion induction, gradually increasing the emotion induction data, and finally emotional recognition is performed through multimodal assisted analysis.

Benefits of technology

Through the three-level emotion induction mechanism, users can effectively supplement their active retention or implicit emotional data, significantly improving the accuracy of emotion recognition, especially in scenarios where deep emotional interaction is required.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148561A_ABST
    Figure CN120148561A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data processing, and particularly discloses a user emotion recognition method based on AI and voice data, which obtains passive data of a user through natural language dialogue or voice interaction, and dynamically analyzes an emotion state in combination with an emotion recognition model. When the model detects that the emotion does not reach a discrimination threshold value, three levels of emotion induction mechanisms are triggered in sequence: in the primary stage, basic emotion data are supplemented by adopting a general guide verbal skill; in the intermediate stage, a personalized induction strategy is adapted based on user portrait features; and in the advanced stage, multi-modal auxiliary analysis is started. According to the technology, recessive emotions are effectively mined through a progressive data enhancement strategy, the proportion of effective emotion samples in customer service, psychological counseling and other scenes is increased, the emotion recognition accuracy is improved through a three-layer data compensation mechanism especially for the scene that a user actively hides the emotions, and deep emotion interaction and accurate recognition are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a method for identifying user emotions based on AI and voice data. Background Art

[0002] Existing voice emotion recognition technologies rely on "passive data" in normal user interactions (in this application, "passive data" refers to data during the interaction between a large voice model and a user based on preset algorithms and training data, such as natural conversations or voices with intelligent assistants). However, in the process of only analyzing "passive data", there are situations where the user's emotions cannot be effectively obtained due to the retention of user initiative or randomness. The essence of this problem is the retention of user initiative or insufficient data volume. For example, users may actively hide their true emotions due to social etiquette, privacy concerns, etc. (such as suppressing anger in a customer service conversation), resulting in the voice data not reflecting the true emotional state. This kind of user is the retention of initiative. The occurrence frequency of specific emotions (such as fear, surprise) in passive data is low, and it is difficult for the model to learn complex emotional features. Especially when the user's emotional expression is implicit, the recognition accuracy drops significantly. This is caused by insufficient data volume. Therefore, the existing technologies cannot recognize the user's emotions in the face of this situation. Summary of the Invention

[0003] Aiming at the deficiencies of the existing technology, the present invention provides a method for identifying user emotions based on AI and voice data, which solves the existing problems.

[0004] To achieve the above objectives, the present invention is realized through the following technical solutions:

[0005] On the one hand, the present invention discloses a method for identifying user emotions based on AI and voice data, including:

[0006] A method for identifying user emotions based on AI and voice data includes the steps of: conducting a natural language conversation or voice interaction with the user based on AI, obtaining "passive data" during the natural language conversation or voice interaction, and identifying and obtaining the user's emotions through an emotion recognition model;

[0007] When the emotion recognition model detects that the emotion does not meet the discrimination threshold, trigger the first-level emotion induction for the user, and add first-level emotion induction data during the natural language conversation or voice interaction with the user;

[0008] If the emotion recognition model still detects that the emotion does not meet the discrimination threshold, then trigger the second-level emotion induction for the user, and add second-level emotion induction data during the natural language conversation or voice interaction with the user;

[0009] If the emotion recognition model still detects that the emotion does not meet the discrimination threshold; then, trigger the third-level emotion induction for the user, conduct multimodal auxiliary analysis, and complete the emotion recognition of the user by the emotion recognition model after active induction.

[0010] Furthermore, the discrimination threshold is calculated as follows: generate the probability distribution of candidate emotion categories through the Softmax function, obtain the maximum probability value, the second-largest probability value, and the category discrimination degree. The discrimination threshold is the comprehensive confidence of the weighted fusion probability value and the discrimination degree.

[0011] Furthermore, conduct natural language conversations or voice interactions with the user based on AI, and obtain the "passive data" in the process of natural language conversations or voice interactions, specifically including synchronously collecting the user's basic attributes, interaction metadata, and real-time behavior characteristics;

[0012] For voice interaction, segment the effective voice segments through endpoint detection, extract the Mel-frequency cepstral coefficients, fundamental frequency range, and formant frequencies; convert the voice to text through a pre-trained speech recognition model, and output the word-level timestamps at the same time;

[0013] Perform word segmentation, label the sentiment polarity words and modal particles, and calculate the text sentiment tendency score;

[0014] Statistical text entropy value, repeated word frequency, punctuation density; store the processed passive data in a distributed database.

[0015] Furthermore, identify and obtain the user's emotions through the emotion recognition model, including:

[0016] Concatenate the voice acoustic features and prosodic features after processing the passive data to form a multi-dimensional voice feature vector;

[0017] For continuous voice segments, use a sliding window to generate a time series feature sequence, and input it into a bidirectional LSTM network to capture the time-dependent relationship of emotion changes;

[0018] Combine the text sentiment tendency score, emotion trigger word density, punctuation density, and text entropy value to form a multi-dimensional text feature vector;

[0019] For multi-round conversation texts, generate context embedding vectors through the Transformer model to capture long-distance semantic dependencies;

[0020] Align the dimensions of the voice time series feature vector and the text context embedding vector through a fully connected layer, and use concatenation to generate multimodal joint features;

[0021] Use a multi-layer CNN to extract frequency-domain and time-domain features from the multimodal joint features to generate high-dimensional abstract features;

[0022] Output the probability distribution of multiple basic emotions through the Softmax function;

[0023] Generate the dominant emotion category and the category probability distribution vector of the current conversation;

[0024] Based on the dominant emotion category, extract the intensity parameter and duration parameter of the corresponding emotion from the joint features;

[0025] Output the emotion parameter vector through the fully connected layer;

[0026] Take the maximum probability value output by Softmax as the model's judgment of the dominant emotion category;

[0027] Calculate the difference between the maximum probability and the second largest probability, and fuse the probability confidence and the discrimination confidence according to a fixed weight to form a comprehensive judgment value in the range of 0-1;

[0028] Finally, output the emotion recognition result, which includes the dominant emotion category, probability distribution, emotion intensity, duration, and comprehensive confidence.

[0029] Furthermore, during the process of natural language conversation or voice interaction with the user, add the first-level emotion induction data, including:

[0030] Ask open-ended questions during the interaction, without giving limited answers, guide the user to express emotions independently, perform multi-modal scene adaptation, and adjust the interaction complexity based on user attributes.

[0031] Furthermore, trigger the second-level emotion induction for the user. During the process of natural language conversation or voice interaction with the user, add the second-level emotion induction data, including:

[0032] Adjust the second-level emotion induction based on the feedback of the first-level emotion induction; conduct interactions based on the speculation of the emotional tendency; for voice interaction, adjust the voice style according to the user's previous voice features and add sound effects for assistance;

[0033] For text scene interaction, adjust the text wording, use infectious and friendly vocabulary, and increase the frequency and diversity of the use of emojis;

[0034] Deeply match the user portrait and associate historical interaction information.

[0035] Furthermore, trigger the third-level emotion induction for the user and conduct multi-modal auxiliary analysis, including:

[0036] For voice modal interaction, extract the fundamental frequency fluctuation curve, speech rate mutation point, and silent duration ratio of the real-time speech, adjust the subsequent induction according to the "fundamental frequency fluctuation curve, speech rate mutation point, and silent duration ratio of the real-time speech", extract the prosodic features, identify the potential emotional tendency, and adjust the subsequent induction according to the prosodic features.

[0037] Furthermore, trigger the third-level emotional induction of the user, and conduct multimodal auxiliary analysis, including:

[0038] For text-modal interaction, analyze the density of emotion trigger words, the change in text entropy value, and the abnormal punctuation. Adjust the subsequent induction based on the "density of emotion trigger words, the change in text entropy value, and the abnormal punctuation", and adjust the subsequent induction according to the semantic association of the context.

[0039] Furthermore, after the active induction, complete the emotion recognition of the user by the emotion recognition model, including:

[0040] Integrate the passive data, the first-level induction data, the second-level induction data, and the third-level multimodal analysis data to form a long-range interaction feature sequence including temporal dependence;

[0041] Stitch the fundamental frequency fluctuation curve, speech rate change, and silence interval temporal features in each round of induction according to the dialogue rounds to generate a dynamic prosody feature matrix;

[0042] Extract the co-occurrence pattern of emotion words and the evolution trend of punctuation in multiple rounds of dialogue;

[0043] According to the user's basic attributes, historical interaction behaviors, and real-time behavior characteristics, perform weighted adjustment on the fused features;

[0044] Align the dimensions of the speech prosody features and the text semantic features to generate a multimodal joint feature vector. For the long text formed by multiple rounds of induction, use the Transformer model to generate the context encoding including the dialogue history. For the long speech segment, use a sliding window and bidirectional LSTM to capture the emotion change trajectory, forming multimodal temporal fusion data of multiple rounds of interaction and multi-layer processing. The multimodal temporal fusion data is input into the emotion recognition model to output the result of the emotion recognition of the user.

[0045] On the other hand, the present invention discloses an electronic device, including:

[0046] At least one processor; and

[0047] A memory communicatively connected to the at least one processor; wherein,

[0048] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned user emotion recognition method based on AI and voice data.

[0049] The beneficial effects of the present invention include but are not limited to: For the scenario where users actively hide their emotions, this application gradually mines implicit emotions through a three-layer induction mechanism, improving the accuracy of emotion recognition. The active induction mechanism significantly increases the effective data. Specifically, through a three-layer progressive induction strategy (primary general guidance → intermediate portrait adaptation → deep multi-modal analysis), it actively supplements the emotion data retained or implied by users, increasing the proportion of effective samples with clear emotions. It has excellent application prospects especially in scenarios such as customer service and psychological counseling that require in-depth emotional interaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a flowchart of the user emotion recognition method based on AI and voice data of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0052] This application discloses a user emotion recognition method based on AI and voice data, as Figure 1 , including the steps: S100. Conduct natural language conversations or voice interactions with the user based on AI, and obtain "passive data" during the natural language conversations or voice interactions; S101. Identify and obtain the user's emotions through an emotion recognition model; S102. When the emotion recognition model detects that the emotion does not meet the discrimination threshold, trigger the first-level emotion induction for the user, and add first-level emotion induction data during the natural language conversations or voice interactions with the user; S103. If the emotion recognition model still detects that the emotion does not meet the discrimination threshold, then trigger the second-level emotion induction for the user, and add second-level emotion induction data during the natural language conversations or voice interactions with the user; S104. If the emotion recognition model still detects that the emotion does not meet the discrimination threshold; then trigger the third-level emotion induction for the user, and conduct multi-modal auxiliary analysis; S105. After the active induction, complete the emotion recognition of the user by the emotion recognition model.

[0053] For S100. Conduct natural language conversations or voice interactions with the user based on AI, and the "passive data" obtained during the natural language conversations or voice interactions includes:

[0054] The user inputs voice through a microphone (such as a smart speaker, in-vehicle voice assistant, or customer service call recording), and the original audio signal is collected in real time (sampling rate ≥ 16 kHz, resolution 16 bit), while synchronously recording interaction timing features such as voice duration and silent intervals.

[0055] The user inputs text through a chat interface (such as an APP dialog box or web customer service chat), and the natural language text content is collected. At the same time, implicit emotional cues such as input speed (characters / second) and punctuation usage habits (such as the frequency of exclamation marks) are recorded.

[0056] Regardless of voice or text interaction, the following information is collected synchronously:

[0057] User basic attributes: age, gender, region, and language preference (such as dialect) provided during registration;

[0058] Interaction metadata: conversation occurrence time, device type (mobile phone / PC / in-vehicle), network environment (4G / Wi-Fi), and conversation turn number;

[0059] Real-time behavior features: speech rate (words / minute) and average fundamental frequency (Hz) in voice interaction, and sentence length and frequency of emotion trigger words (such as "hate" and "satisfied") in text interaction.

[0060] The Wiener filtering algorithm is used to remove environmental noise (such as background noise and current noise). The effective voice segments are segmented through voice activity detection (VAD), and basic acoustic features such as Mel-frequency cepstral coefficients (MFCC, 13 dimensions), fundamental frequency range (F0_min / F0_max), and formant frequencies (the first 3 orders) are extracted.

[0061] The voice is converted into text through a pre-trained speech recognition model (such as DeepSpeech), and at the same time, word-level timestamps are output (for subsequent prosody feature analysis, such as phoneme duration and pause position).

[0062] Natural language processing tools (such as NLTK and jieba) are used for word segmentation, sentiment polarity words (such as "happy" and "frustrated") and modal particles ("ah" and "ya") are annotated, and the text sentiment tendency score is calculated (based on a pre-trained sentiment dictionary).

[0063] The text entropy value (reflecting semantic complexity), frequency of repeated words (such as "very very" reflecting emotion intensity), and punctuation density (the proportion of exclamation marks reflecting excitement level) are statistically analyzed.

[0064] The processed passive data is stored in a distributed database to form a standardized data record; the user's voice / text data is anonymized, and sensitive information (such as phone numbers and addresses) is deleted; the basic features are extracted on the local device, and only the desensitized feature vectors are uploaded to the server to avoid the leakage of original data.

[0065] For S101. Identifying and obtaining the user's emotion through an emotion recognition model includes:

[0066] Concatenate the speech acoustic features (MFCC 13-dimensional, fundamental frequency range, formant frequency) and prosodic features (phoneme duration, pause position, speech rate) after passive data processing to form a 30-dimensional speech feature vector.

[0067] For continuous speech segments, use a sliding window (window length 1 second, step size 0.5 second) to generate a temporal feature sequence, and input it into a bidirectional LSTM network to capture the time-dependent relationship of emotion changes (such as the fundamental frequency mutation trajectory of angry emotion).

[0068] Combine the text sentiment tendency score (-1 to 1, negative value for negative and positive value for positive), emotion trigger word density (number of emotion words per unit number of words), punctuation mark density (ratio of exclamation marks / question marks), and text entropy value (reflecting semantic complexity) to form an 8-dimensional text feature vector.

[0069] For multi-turn dialogue text, generate context embedding vectors through a Transformer model to capture long-distance semantic dependencies (such as the impact of dissatisfaction in previous conversations on the current response).

[0070] Align the dimensions of the speech temporal feature vector and the text context embedding vector through a fully connected layer (unify to 256 dimensions), and use the concatenation method to generate a 512-dimensional multi-modal joint feature, while retaining the independence of speech and text features for subsequent hierarchical recognition.

[0071] Use a 3-layer CNN to extract frequency-domain and time-domain features from the multi-modal joint feature to generate high-dimensional abstract features (such as the high-frequency energy concentration region unique to angry emotion, the co-occurrence pattern of exclamation marks and "dissatisfaction" words in the text).

[0072] Output the probability distribution of 10 basic emotions (happy, angry, sad, surprised, frightened, disgusted, calm, anxious, disappointed, neutral) through the Softmax function, and support multi-label output (such as a mixed emotion of "angry + disappointed").

[0073] Generate the dominant emotion category (single category or multi-category combination with the highest probability) and category probability distribution vector of the current conversation.

[0074] Based on the dominant emotion category, extract the intensity parameters (such as the standard deviation of the fundamental frequency of anger, the occurrence frequency of emotion words in the text) and duration parameters (speech segment duration, number of emotion continuation rounds in multi-turn conversations) corresponding to the emotion from the joint feature.

[0075] Output the emotion parameter vector through a fully connected layer.

[0076] Take the maximum probability value of the Softmax output, which reflects the judgment certainty of the model for the dominant emotion category.

[0077] Calculate the difference between the maximum probability and the second-largest probability. The larger the difference (approaching 1), the clearer the emotion category, and vice versa (approaching 0) indicates a fuzzy category (e.g., the probabilities of "anger" and "anxiety" are close).

[0078] Fuse the probability confidence and the discrimination confidence with a weight of 8:2 to form a comprehensive judgment value in the range of 0 - 1.

[0079] For implicit users (speech rate ≤ 150 words per minute and text sentence length ≤ 10 words), if the comprehensive confidence is lower than 0.6 but a sudden increase in fundamental frequency is detected (more than 20% higher than the historical average), the confidence is forcibly increased by 10% (to avoid missing hidden emotions).

[0080] For extroverted users (speech rate > 200 words per minute and the text contains ≥ 2 exclamation marks), if the comprehensive confidence is higher than 0.7 but the emotion parameter intensity is lower than 0.3, trigger the "inconsistent emotion expression" mark to prompt that subsequent induction needs to focus on confirming the emotion intensity.

[0081] The final output result, the emotion recognition result: includes the dominant emotion category (e.g., "anger"), probability distribution (e.g., 0.72), emotion intensity (0.65), duration (12 seconds), and comprehensive confidence (0.75).

[0082] For S102. When the emotion recognition model detects that the emotion does not meet the discrimination threshold, trigger the first-level emotion induction for the user, and add first-level emotion induction data during the natural language conversation or voice interaction with the user; the first-level emotion induction refers to a preliminary guidance mechanism actively triggered when the confidence of the recognition result of the emotion recognition model based on passive data (speech / text in natural conversation) is insufficient. Its core goal is to guide the user to supplement emotion-related information through generalized and open-ended interaction words, and solve the problem of insufficient effective data caused by the user actively withholding emotions or expressing implicitly. The calculation of the discrimination threshold is as follows: generate the probability distribution of candidate emotion categories through the Softmax function, obtain the maximum probability value, the second-largest probability value, and the category discrimination (reflecting the clarity of the model's judgment of the dominant emotion), and the discrimination threshold is the comprehensive confidence of the weighted fusion probability value and the discrimination.

[0083] Specifically, adding first-level emotion induction data during the natural language conversation or voice interaction with the user includes:

[0084] During the interaction, mainly use open-ended questions, avoid restrictive answers, and guide the user to express emotions independently.

[0085] Exemplary voice interaction: "Your expression just now was very concise. Can you tell me more specifically how you feel now?"

[0086] Exemplary,text interaction: “I saw your reply and want to know more about your current emotional,state. Can you describe it in detail?”.

[0087] Adapt multimodal scenarios:

[0088] Exemplary, voice scenario: use a gentle tone (reduce the fundamental frequency mean by 10%, and slow down the speaking speed to 180 words / minute) to convey a signal of encouragement for expression; Exemplary, text scenario: use friendly punctuation (such as "~" and "?") and emoticons to reduce the user's psychological defense.

[0089] Lightweight matching of user portraits, adjusting interaction complexity based on user attributes (age, region):

[0090] Example, young user: "I feel like you seem a little different, what exactly happened?" Example, older user: "Your feedback just now is very important, can you tell me more about your thoughts?"

[0091] For S103, if the emotion recognition model detects that the emotion still does not meet the discrimination threshold, the second level of emotion induction for the user is triggered, and the second level of emotion induction data is added during the natural language dialogue or voice interaction with the user; the second level of emotion induction is a further guidance mechanism initiated when the confidence of the recognition result of the emotion recognition model is still insufficient after the first level of emotion induction. Its core goal is to more accurately stimulate the user's emotional information, and adopt a more targeted induction method based on the characteristics of different users and previous interactions, so as to break through the possible emotional reservations of users and obtain richer and more effective emotional data, thereby improving the accuracy of emotion recognition.

[0092] Specifically, triggering the second level of emotion induction for the user, adding the second level of emotion induction data during the natural language dialogue or voice interaction with the user includes:

[0093] Adjust the second level of emotion induction based on the feedback of the first level of emotion induction.

[0094] For example, the second level of emotional induction is adjusted based on the user's response content in the first level of emotional induction. If the user mentions a specific event but does not clearly express emotions, further questions can be asked around the event. For example, if the user mentions "some problems at work" in the first level of response, you can ask "How do the problems at work make you feel? Are you a little anxious or a little irritated?"

[0095] Interact based on emotional inclinations.

[0096] Exemplarily, questions are asked based on the emotional tendencies reflected in the previous passive data and the first-level induced data. If a negative emotional tendency is recognized, questions such as "It seems that you are not very happy. Have you encountered something disappointing?" can be asked; if there is a positive emotional tendency, questions such as "You seem to be in a good mood. Can you share something that makes you happy with me?" can be asked.

[0097] For voice interaction, the voice style is adjusted according to the user's previous voice characteristics.

[0098] Exemplarily, if the user has a slow speaking speed and a steady intonation, the induced voice can adopt a more soothing and gentle intonation, the average fundamental frequency is further reduced by 15%, and the speaking speed is slowed down to 160 words per minute; if the user has a fast speaking speed and large fluctuations in intonation, the induced voice can appropriately increase the volume and vitality of the intonation to better respond to the user.

[0099] For voice interaction, sound effects are added for assistance.

[0100] Exemplarily, some soft and encouraging sound effects are appropriately added, such as gentle reminder sounds, which are played before and after the questions to enhance the guiding effect.

[0101] For text scenario interaction, the text wording is adjusted, and words with infectivity and affinity are used.

[0102] Exemplarily, change "Can you talk about it in detail?" to "I really want to hear you talk about it in more detail."

[0103] For text scenario interaction, the frequency and diversity of the use of emojis are increased.

[0104] Exemplarily, when negative emotions are speculated, comforting emojis are used; when positive emotions are present, cheerful emojis are used.

[0105] Deeply match the user portrait and associate historical interaction information. In the interaction, in addition to age and region, the user's language characteristics, historical interaction behaviors, and emotional expression patterns are combined for induction, and key information in the previous conversation is mentioned, so that the user feels being cared about and understood.

[0106] Exemplarily, "You mentioned something about work before. Now thinking about it, does it still make you feel a little uncomfortable?"

[0107] For S104. If the emotion recognition model still detects that the emotion does not meet the discrimination threshold; then, trigger the third-level emotion induction for the user and conduct multimodal auxiliary analysis; the third-level emotion induction is a deep induction mechanism triggered when the confidence level of the emotion recognition model still fails to meet the standard after the first two levels of induction. Its core goal is to break through the bottleneck of the user's deep emotion retention or complex emotion expression through in-depth fusion analysis of multimodal data, and combine multi-dimensional information such as speech prosody, text semantics, and interaction history to conduct personalized emotion stimulation, ultimately achieving accurate recognition of the user's emotion.

[0108] Specifically, triggering the third-level emotion induction for the user and conducting multimodal auxiliary analysis includes:

[0109] For speech modality interaction:

[0110] Extract the fundamental frequency fluctuation curve of the real-time speech (such as a sudden increase in fundamental frequency when angry and a low fundamental frequency when sad), the speed mutation point of the speech rate (a sudden increase in the speech rate may reflect anxiety), and the proportion of the silent duration (a long pause may imply hesitation or hidden emotion), and adjust the subsequent induction based on the "fundamental frequency fluctuation curve, speed mutation point, and proportion of the silent duration of the real-time speech".

[0111] Exemplarily, if it is detected that the fundamental frequency of the user's speech suddenly increases (more than 30% above the average) at a certain word and is accompanied by an increase in the speech rate, then it is speculated that this word is an emotion trigger point (such as "complaint" "deception"), and the subsequent induction is carried out around this keyword.

[0112] Extract prosodic features (such as stress position, intonation inflection point), identify potential emotion tendencies, and adjust the subsequent induction based on the prosodic features.

[0113] Exemplarily, the user emphasizes the word "simply unusable" and the intonation drops suddenly. Combining with the text keyword "unusable", it is judged as a strong negative emotion, and the induction question focuses on the specific dissatisfaction point: "You mentioned'simply unusable'. Is it that the function cannot be realized, or did you encounter an obstacle during the use process?"

[0114] For text modality interaction:

[0115] Analyze the density of emotion trigger words (such as the appearance frequency of "collapse" "despair"), the change amount of text entropy value (a sudden drop in entropy value may indicate that the emotion tends to be clear), and the abnormal situation of punctuation marks (continuous exclamation marks / question marks reflect the intensity of emotion), and adjust the subsequent induction based on the "density of emotion trigger words, change amount of text entropy value, and abnormal situation of punctuation marks".

[0116] Exemplarily, the user replies "How can this be solved!!!". Combining with the high punctuation density and negative trigger words, the induction question directly targets the solution: "Your tone makes me feel your eagerness. Can you tell me specifically what difficulties you have encountered, and we will figure it out together?"

[0117] Adjust the subsequent induction according to the semantic association in the context (such as the complaint content in the previous conversation, the emotional consistency in multi-round interactions).

[0118] Exemplarily, if the user repeatedly mentions "poor service" in multi-round conversations but the current response is short, trigger in-depth induction: "You have mentioned service issues many times before. Is this experience particularly unsatisfactory for you? Can you be more specific about which part?"

[0119] For S105. After the active induction, the emotion recognition model's recognition of the user's emotion includes:

[0120] Integrate passive data (initial voice / text features), first-level induction data, second-level induction data, and third-level multi-modal analysis data to form a long-range interaction feature sequence with temporal dependence.

[0121] Stitch the temporal features such as the fundamental frequency fluctuation curve, speech rate change, and silence interval in each round of induction according to the conversation round to generate a dynamic prosody feature matrix.

[0122] Extract the co-occurrence pattern of emotional words in multi-round conversations (such as the high-frequency co-occurrence of "dissatisfaction" and "efficiency" reflecting anger towards the service), and the evolution trend of punctuation marks (from no punctuation to consecutive exclamation marks indicating an emotional escalation).

[0123] According to the user's basic attributes (age, region), historical interaction behaviors (such as users who frequently use negative words are more likely to be in long-term anxiety), and real-time behavior characteristics (when the current conversation device is in a vehicle scenario, prioritize attention to urgent emotions such as "anxious"), perform weighted adjustment on the fused features (such as increasing the weight of internet buzzwords for young users by 20%).

[0124] Align the dimensions of the speech prosody features (fundamental frequency mean, MFCC sequence) and text semantic features (emotional tendency score, context embedding vector) (unify to 256 dimensions), and generate a 512-dimensional multi-modal joint feature vector by stitching, capturing cross-modal associations while retaining the independence of each modality (such as the co-occurrence relationship between the high-frequency energy of angry speech and the text "!").

[0125] For the long text formed by multi-round induction, use the Transformer model to generate context encoding including the conversation history, and for long speech segments, use a sliding window combined with bidirectional LSTM to capture the emotional change trajectory (such as the fundamental frequency mutation point from "calm" to "angry").

[0126] Through the above processing, form multi-modal temporal fusion data of multi-round interaction and multi-layer processing, and input the multi-modal temporal fusion data into the emotion recognition model to output the result of the user's emotion recognition.

[0127] Obviously, the method of the present invention can be implemented by a computer program. The computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the computer program is executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0128] Therefore, it can be understood that the present invention discloses an electronic device, comprising:

[0129] at least one processor; and

[0130] a memory communicatively connected to the at least one processor; wherein,

[0131] the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the above-mentioned user emotion recognition method based on AI and voice data.

[0132] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A user emotion recognition method based on AI and voice data, characterized in that: The process includes: conducting natural language dialogue or voice interaction with users based on AI, and obtaining "passive data" in the process of natural language dialogue or voice interaction, where "passive data" refers to data in the process of interaction between the big voice model and users based on preset algorithms and training data; Identify and obtain the user's emotions through the emotion recognition model; When the emotion recognition model detects that the emotion does not meet the discrimination threshold, the first level of emotion induction for the user is triggered, and the first level of emotion induction data is added during the natural language dialogue or voice interaction with the user; If the emotion recognition model detects that the emotion still does not meet the discrimination threshold, the second level emotion induction for the user is triggered, and the second level emotion induction data is added during the natural language dialogue or voice interaction with the user; If the emotion recognition model detects that the emotion still does not meet the discrimination threshold, the third level of emotion induction for the user is triggered, and multimodal auxiliary analysis is performed to complete the emotion recognition model's recognition of the user's emotion after active induction.

2. The user emotion recognition method based on AI and voice data according to claim 1, characterized in that: The discrimination threshold is calculated by generating the probability distribution of candidate emotion categories through the Softmax function, obtaining the maximum probability value, the second largest probability value, and the category discrimination degree. The discrimination threshold is the comprehensive confidence of the weighted fusion probability value and the discrimination degree.

3. The user emotion recognition method based on AI and voice data according to claim 1, characterized in that: Based on AI, the "passive data" obtained in the process of natural language dialogue or voice interaction with users includes synchronous collection of basic attributes, interaction metadata, and real-time behavior characteristics of users; For voice interaction, endpoint detection is used to segment valid voice segments, extract Mel-frequency cepstral coefficients, fundamental frequency range, and formant frequency; the voice is converted into text through a pre-trained speech recognition model, and word-level timestamps are output at the same time; Perform word segmentation, mark sentiment polarity words and modal particles, and calculate the sentiment tendency score of the text; Statistics of text entropy, repeated word frequency, and punctuation density; Store the processed passive data into the distributed database.

4. The user emotion recognition method based on AI and voice data according to claim 1, characterized in that: Identifying and acquiring user emotions through the emotion recognition model includes: The speech acoustic features and prosodic features after passive data processing are concatenated to form a multi-dimensional speech feature vector; For continuous speech segments, a sliding window is used to generate a temporal feature sequence, which is input into a bidirectional LSTM network to capture the temporal dependency of emotional changes. The text sentiment tendency score, emotion trigger word density, punctuation density, and text entropy value are combined into a multi-dimensional text feature vector; For multi-round conversation texts, the Transformer model is used to generate context embedding vectors to capture long-distance semantic dependencies. The speech time sequence feature vector and the text context embedding vector are dimensionally aligned through a fully connected layer, and multimodal joint features are generated by splicing; Use multi-layer CNN to extract frequency-domain and time-domain features of multi-modal joint features to generate high-dimensional abstract features; Output the probability distribution of multiple categories of basic emotions through the Softmax function; Generate the dominant emotion category and category probability distribution vector of the current conversation; Based on the dominant emotion category, the intensity parameter and duration parameter of the corresponding emotion are extracted from the joint features; Output the emotion parameter vector through the fully connected layer; Take the maximum probability value of Softmax output as the model's judgment on the dominant emotion category; Calculate the difference between the maximum probability and the second-highest probability, fuse the probability confidence and the discrimination confidence according to a fixed weight, and form a comprehensive judgment value in the range of 0-1; The final output is the emotion recognition result, which includes the dominant emotion category, probability distribution, emotion intensity, duration, and comprehensive confidence.

5. The user emotion recognition method based on AI and voice data according to claim 1, characterized in that: The first level of emotion induction data added during natural language conversation or voice interaction with users includes: Ask open-ended questions during the interaction without giving restrictive answers, guide users to express their emotions independently, perform multimodal scene adaptation, and adjust the complexity of the interaction based on user attributes.

6. The user emotion recognition method based on AI and voice data according to claim 1, characterized in that: Triggering the second level of emotion induction for users, adding the second level of emotion induction data during natural language dialogue or voice interaction with users includes: Adjust the second level of emotional induction based on the first level of emotional induction feedback; interact based on emotional tendency speculation; for voice interaction, adjust the voice style and add sound effects according to the user's previous voice characteristics; For text scene interactions, adjust the wording of the text, use appealing and friendly words, and increase the frequency and diversity of emoticons; Deeply match user portraits and associate historical interaction information.

7. The user emotion recognition method based on AI and voice data according to claim 1, characterized in that: Triggering the third level of emotional induction for users and conducting multimodal auxiliary analysis includes: For voice modal interaction, extract the fundamental frequency fluctuation curve, speech rate mutation points, and silence duration ratio of real-time speech, adjust subsequent induction based on "fundamental frequency fluctuation curve, speech rate mutation points, and silence duration ratio of real-time speech", extract rhythmic features, identify potential emotional tendencies, and adjust subsequent induction based on rhythmic features.

8. The user emotion recognition method based on AI and voice data according to claim 1, characterized in that: Triggering the third level of emotional induction for users and conducting multimodal auxiliary analysis includes: For text modal interaction, analyze the density of emotional trigger words, changes in text entropy values, and abnormal punctuation marks, and adjust subsequent induction based on "density of emotional trigger words, changes in text entropy values, and abnormal punctuation marks". Adjust subsequent induction based on contextual semantic associations.

9. The user emotion recognition method based on AI and voice data according to claim 1, characterized in that: After active induction, the emotion recognition model completes the user's emotion recognition including: Integrate passive data, first-level induced data, second-level induced data, and third-level multimodal analysis data to form a long-range interactive feature sequence containing temporal dependencies; The fundamental frequency fluctuation curve, speech rate change, and silence interval time sequence features in each round of induction are spliced ​​according to the dialogue rounds to generate a dynamic rhythm feature matrix; Extract the co-occurrence patterns of sentiment words and the evolution trend of punctuation marks in multi-round conversations; According to the user's basic attributes, historical interaction behaviors and real-time behavior characteristics, the fusion features are weighted and adjusted; The speech prosodic features and text semantic features are dimensionally aligned to generate a multimodal joint feature vector. For long texts formed by multiple rounds of induction, the Transformer model is used to generate context encoding containing the conversation history. For long voice clips, sliding windows and bidirectional LSTM are used to capture the trajectory of emotional changes, forming multimodal time series fusion data of multi-round interactions and multi-layer processing. The multimodal time series fusion data is input into the emotion recognition model to output the results of user emotion recognition.

10. An electronic device, comprising: at least one processor; as well as A memory communicatively connected to the at least one processor; characterized in that The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the user emotion recognition method based on AI and voice data as described in any one of claims 1-9.

Citation Information

Cited By

  • Multi-modal digital human interaction method and system based on large language model

    CN120317879A

  • Intelligent dialogue system and method based on AI multi-mode large model

    CN120491834A

  • Voice quality inspection method and device, computer equipment and storage medium

    CN120690231A

  • A voice quality inspection method, apparatus, computer equipment, and storage medium

    CN120690231B

  • User emotion recognition method and device, electronic equipment and storage medium

    CN120951293A