Emotional intention recognition system and method based on AI recognition

By generating emotionally unintended intent speech and combining training and learning of artificial intelligence models, the problem of insufficient accuracy and robustness in traditional methods is solved, and higher accuracy and robust emotional intent recognition are achieved.

CN120472941APending Publication Date: 2025-08-12SHEN ZHEN WEI LAI SHI XU REN GONG ZHI NENG YOU XIAN GONG SI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510627744.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

Traditional emotional intention recognition methods rely on semantic extraction and vocabulary clustering, ignoring the hidden emotional characteristics in the speech, resulting in low accuracy and robustness.

Method used

User voice audio is collected through intelligent voice sensors, generate emotional intent voice, extract emotional identification voice and convert audio features, combine the first and second artificial intelligence models for training and learning, and generate emotional intent recognition results.

Benefits of technology

It improves the accuracy and robustness of emotional intention recognition, can better adapt to emotional changes in different users and situations, and enhances the accuracy and system adaptability of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472941A_ABST
    Figure CN120472941A_ABST
Patent Text Reader

Abstract

The invention provides an emotional intention recognition system and method based on AI recognition, and the method comprises the steps: carrying out the generation of a non-emotional intention sample based on audio parameter verification information and voice text data, and obtaining a non-emotional intention voice of a user; performing emotion interference extraction according to the voice audio of the user and the voice without emotional intention to obtain an emotion identification feature vector corresponding to the user; performing training learning on the voice text data based on a first artificial intelligence model to obtain an initial emotion label, and performing training learning on the emotion identification feature vector through a second artificial intelligence model to obtain a corrected emotion label; and performing emotional intention recognition of the corresponding user according to the initial emotional label and the corrected emotional label, and generating a final emotional intention recognition result of the user, so that emotional intention recognition of the user can be performed through the non-emotional intention voice of the user, and the accuracy and robustness of emotional intention recognition of the user are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of speech emotion recognition, and more specifically, to an emotion intention recognition system and method based on AI recognition. Background Art

[0002] Emotion, as a human's subjective experience and reaction to external stimuli, typically manifests as psychological states such as happiness, sadness, or shock in response to specific events. This complex and ever-changing internal experience plays a crucial role in human behavior and decision-making. Speech emotion recognition aims to accurately determine individual emotional intent by extracting the perceptual characteristics of user audio. By deeply exploring and integrating user audio features, it can more accurately understand and analyze user emotional intent, providing strong support for the development of fields such as artificial intelligence and human-computer interaction.

[0003] Traditional emotional intent recognition methods usually infer users' emotional intent by performing semantic extraction on user speech and clustering semantic words. Therefore, they focus on analyzing the vocabulary and sentence structure in the speech content, aiming to identify the user's emotional state through parsing the topic and semantic information of the conversation. However, this approach only considers the explicit emotional information contained in the speech content, while ignoring the latent emotional characteristics in the speech. Simply relying on semantic extraction and vocabulary clustering methods can easily lead to low accuracy and robustness in the emotional intent recognition process. Summary of the Invention

[0004] The present application provides an emotion intention recognition system and method based on AI recognition, which can identify the user's emotional intention through the user's non-emotional intention speech, thereby improving the accuracy and robustness of user emotional intention recognition.

[0005] On the first aspect, the present application provides an emotion intention recognition method based on AI recognition, which can be executed by a network device, or by a chip configured in the network device, and the present application does not limit this.

[0006] Specifically, the method includes:

[0007] Collect user voice audio through intelligent voice sensors;

[0008] Performing speech-to-text recognition on the user's voice audio to obtain speech-to-text data, obtaining audio parameter verification information corresponding to the user's voice audio, and generating a non-emotional intent sample based on the audio parameter verification information and the speech-to-text data to obtain the user's non-emotional intent speech;

[0009] Performing emotion interference extraction based on the user voice audio and the speech without emotion intention to obtain emotion recognition speech, and performing audio feature conversion based on the emotion recognition speech to obtain an emotion recognition feature vector corresponding to the user;

[0010] Training the speech text data based on a first artificial intelligence model to obtain an initial emotion label, and training the emotion recognition feature vector using a second artificial intelligence model to obtain a corrected emotion label;

[0011] The emotional intention of the corresponding user is identified based on the initial emotional label and the corrected emotional label to generate a final emotional intention recognition result of the user.

[0012] In combination with the first aspect, in certain implementations of the first aspect, a microphone array and a front-end audio processing chip equipped with corresponding audio processing software are used as the intelligent voice sensor.

[0013] In combination with the first aspect, in certain implementations of the first aspect, performing emotional interference extraction based on the user voice audio and the speech without emotional intention to obtain the emotion recognition speech specifically includes: obtaining the Mel-cepstrum corresponding to the user voice audio and the speech without emotional intention respectively, performing spectral differentiation based on the Mel-cepstrum corresponding to the user voice audio and the speech without emotional intention respectively, obtaining the differential spectrum corresponding to the two, and then performing spectral restoration to obtain the emotion recognition speech.

[0014] In conjunction with the first aspect, in certain implementations of the first aspect, performing audio feature conversion based on the emotion recognition speech to obtain an emotion recognition feature vector corresponding to the user specifically includes:

[0015] Acquire speech-to-text data, determine a timestamp corresponding to a start time and an end time of each semantic word in the emotion recognition speech based on the emotion recognition speech and the speech-to-text data, and determine a plurality of semantic word time windows based on the timestamp corresponding to the start time and the end time of each semantic word;

[0016] Window features corresponding to the audio information in each semantic word time window are extracted, and each window feature is normalized and then formed into the emotion recognition feature vector according to the time sequence.

[0017] In conjunction with the first aspect, in certain implementations of the first aspect, performing the emotional intent recognition of the corresponding user based on the initial emotion label and the corrected emotion label to generate a final emotional intent recognition result of the user specifically includes:

[0018] Performing emotional intention credibility detection based on the initial emotion recognition result and the corrected emotion recognition result to obtain emotional intention similarity;

[0019] When the emotion intention similarity is higher than a preset threshold, the emotion recognition initial result is corrected using the emotion recognition correction result to obtain a final emotion intention recognition result of the user;

[0020] When the emotional intention similarity is lower than a preset threshold, the initial emotional label is used as the final emotional intention recognition result of the user.

[0021] In combination with the first aspect, in certain implementations of the first aspect, a long short-term memory neural network is used as the first artificial intelligence model to train and learn the audio-text data to obtain an initial emotion label.

[0022] In combination with the first aspect, in certain implementations of the first aspect, performing speech-to-text recognition on the user voice audio to obtain the speech-to-text data further includes: removing background noise from the user voice audio.

[0023] In a second aspect, the present application provides an emotion intention recognition system based on AI recognition, which includes a speech emotion recognition unit, and the speech emotion recognition unit includes:

[0024] Voice collection module, used to collect user voice audio through intelligent voice sensor;

[0025] A speech generation module is configured to perform speech-to-text recognition on the user's speech audio to obtain speech-to-text data, acquire audio parameter verification information corresponding to the user's speech audio, generate a non-emotional intent sample based on the audio parameter verification information and the speech-to-text data, and obtain the user's non-emotional intent speech;

[0026] An emotion recognition module is configured to extract emotion interference from the user's voice audio and the speech without emotion intention to obtain emotion recognition speech, and perform audio feature conversion based on the emotion recognition speech to obtain an emotion recognition feature vector corresponding to the user;

[0027] The emotion recognition module is further configured to train and learn the speech text data based on a first artificial intelligence model to obtain an initial emotion label, and to train and learn the emotion recognition feature vector using a second artificial intelligence model to obtain a corrected emotion label;

[0028] The emotion recognition module is further configured to recognize the emotion intention of the corresponding user based on the initial emotion label and the corrected emotion label, and generate a final emotion intention recognition result of the user.

[0029] In a third aspect, the present application provides a computer terminal device, which includes a memory and a processor, wherein the memory stores a code, and the processor is configured to obtain the code and execute the above-mentioned emotion intention recognition method based on AI recognition.

[0030] In a fourth aspect, the present application provides a computer-readable storage medium, which stores at least one computer program, and the computer program is loaded and executed by a processor to implement the operations performed by the above-mentioned emotion intention recognition method based on AI recognition.

[0031] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects:

[0032] In an emotion intention recognition system and method based on AI recognition provided by the present application, user voice audio is first collected through an intelligent voice sensor; voice text recognition is performed on the user voice audio to obtain voice text data, audio parameter verification information corresponding to the user voice audio is obtained, and a sample without emotion intention is generated based on the audio parameter verification information and the voice text data to obtain the user's emotion-free intention voice; emotion interference is extracted based on the user voice audio and the emotion-free intention voice to obtain emotion recognition voice, and audio feature conversion is performed based on the emotion recognition voice to obtain the user's corresponding emotion recognition feature vector; the voice text data is trained and learned based on a first artificial intelligence model to obtain an initial emotion label, and the emotion recognition feature vector is trained and learned through a second artificial intelligence model to obtain a corrected emotion label; the corresponding user's emotion intention is recognized based on the initial emotion label and the corrected emotion label to generate the user's final emotion intention recognition result.

[0033] It can be seen that the present application generates speech without emotion intention, that is, removes emotion-related features from the user's speech, removes emotional interference in the speech, and provides a clearer and purer audio signal for subsequent emotion recognition, thereby improving the accuracy of emotion feature extraction. Then, pure emotion information is identified and extracted from the user's speech by means of emotion interference extraction, and the emotion recognition feature vector is obtained as the core data for learning the emotion recognition system, so that the system can capture subtle changes in emotion signals, improve the accuracy of emotion intention recognition, and process the speech text data through the first model to obtain the initial emotion label and provide preliminary emotion recognition results. Based on the first stage, the second model is used to train and correct the extracted emotion recognition feature vector to further optimize the accuracy of the emotion label, so that emotion recognition not only relies on semantic analysis, but also combines audio features, enhances the robustness of recognition, and can better adapt to the emotional changes of different users and different situations.

[0034] In summary, this application solves the problems of insufficient accuracy and robustness in traditional emotion intention recognition methods by generating speech without emotion intention and extracting emotion interference, combining the emotional characteristics of speech and analysis of text data, and improving the accuracy of emotion recognition and system robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 This is an exemplary flow chart of an emotional intention recognition method based on AI recognition according to some embodiments of the present application;

[0036] Figure 2 is an exemplary flow chart for generating emotion intent recognition results in some embodiments of the present application;

[0037] Figure 3 is a structural diagram of a speech emotion recognition unit according to some embodiments of the present application;

[0038] Figure 4 This is a structural diagram of a computer terminal device that implements an emotion intention recognition method based on AI recognition according to some embodiments of the present application. DETAILED DESCRIPTION

[0039] This application collects user voice audio through an intelligent voice sensor; performs voice-to-text recognition on the user voice audio to obtain voice-to-text data, obtains audio parameter verification information corresponding to the user voice audio, generates a sample without emotion intention based on the audio parameter verification information and the voice-to-text data, and obtains the user's voice without emotion intention; performs emotion interference extraction based on the user voice audio and the voice without emotion intention to obtain emotion recognition voice, and performs audio feature conversion based on the emotion recognition voice to obtain the user's corresponding emotion recognition feature vector; trains and learns the voice-to-text data based on a first artificial intelligence model to obtain an initial emotion label, and trains and learns the emotion recognition feature vector through a second artificial intelligence model to obtain a corrected emotion label; performs corresponding user emotion intention recognition based on the initial emotion label and the corrected emotion label to generate the user's final emotion intention recognition result, which can recognize the user's emotion intention through the user's voice without emotion intention, thereby improving the accuracy and robustness of user emotion intention recognition.

[0040] In order to better understand the above technical solution, the following will be combined with the accompanying drawings and specific implementation methods to describe the above technical solution in detail. Figure 1 , which is an exemplary flow chart of an emotional intention recognition method based on AI recognition according to some embodiments of the present application. The emotional intention recognition method 100 based on AI recognition mainly includes the following steps:

[0041] In step S101, user voice audio is collected through an intelligent voice sensor.

[0042] Optionally, in some embodiments, a microphone array and a front-end audio processing chip equipped with corresponding audio processing software can be used as the intelligent voice sensor. In some other embodiments, other voice sensor devices capable of audio processing can also be used. This application does not limit this.

[0043] In step S102, speech-to-text recognition is performed on the user's voice audio to obtain speech-to-text data, audio parameter verification information corresponding to the user's voice audio is obtained, and a non-emotional intention sample is generated based on the audio parameter verification information and the speech-to-text data to obtain the user's non-emotional intention speech.

[0044] Optionally, in some embodiments, before performing speech-to-text recognition on the user voice audio to obtain the speech-to-text data, the method further includes: removing background noise from the user voice audio.

[0045] It should be noted that the process of performing speech-to-text recognition on the user voice audio in the present application can be achieved by using the existing speech recognition engine equipped with the front-end audio processing chip in the intelligent voice sensor. For example, the user voice audio can be subjected to speech-to-text recognition by calling a third-party cloud service API, such as the Alibaba Cloud intelligent speech recognition engine commonly used in speech-to-text recognition in the prior art, to obtain speech-to-text data. In some other embodiments suitable for scenarios requiring private deployment and data security, a locally deployed recognition engine such as OpenAI's open source Whisper intelligent speech recognition model can also be called for speech-to-text recognition. This application does not limit this.

[0046] Optionally, in some embodiments, obtaining the audio parameter verification information corresponding to the user voice audio specifically includes: performing user identity identification based on the user voice audio to obtain user identity information, and retrieving the audio parameter verification information corresponding to the user identity information.

[0047] In a specific implementation, the process of performing user identity recognition based on the user voice audio and obtaining user identity information can be carried out by extracting the Mel-frequency cepstral coefficients in the user voice audio to form an embedding vector, and comparing it with the embedded vectors of the stored standard voice audios of different users, confirming the user identity information based on the standard voice audio with the highest similarity, and then determining the corresponding audio parameter verification information based on the user identity information, wherein the audio parameter verification information is determined based on the detailed parameters of the standard voice audio stored by the user, and is used to restore the user's timbre and generate the user's emotionless voice sample. In some embodiments, the audio parameter verification information includes: the standard voice template and audio parameters stored for the user identity corresponding to the user voice audio, wherein the audio parameters include: a feature vector composed of speech rate information, fundamental frequency range, sound energy characteristics (the overall energy response in dB can be used), vocal tract characteristics, etc.

[0048] It should be noted that the audio parameter verification information described in this application is determined by parameter extraction based on the standard voice audio template stored by the user, and is used to restore the sound quality information and emotional characteristics of the user when storing his or her own standard voice audio. Among them, the user can be guided to maintain a neutral emotion through voice prompts when the user stores the standard voice audio template, thereby reducing emotional fluctuations during the recording process. The standard voice audio template is input by the user under a standard test environment, which avoids emotional interference during vocalization to the greatest extent from the external conditions, and also reflects the natural, emotion-free vocal characteristics of the user under normal circumstances. Emotionless intention speech is performed based on the audio parameter verification information and used for emotion recognition of daily speech, which can improve the accuracy of emotional intention recognition and system robustness.

[0049] In some specific embodiments of the present application, the standard speech template of the user and its corresponding audio parameter feature vector pre-stored in the database can be queried based on the confirmed user identity ID, wherein the parameter categories in the audio parameter feature vector include: speaking rate information: the number of pronounced words per unit time (WPM); fundamental frequency range: the mean, maximum and minimum values of the main fundamental frequency in the speech; sound energy characteristics: the overall energy response in decibels (the short-time energy curve mean can be used), vocal tract characteristics: Formant resonance peak (F1, F2, F3) distribution parameters.

[0050] Optionally, in some embodiments, the process of generating a sample without emotion intention based on the audio parameter verification information and the voice text data to obtain the user's voice without emotion intention can be performed by a text-to-speech (TTS) system based on the audio parameter verification information and the voice text data to obtain a sample without emotion intention. In specific implementation, the VITS open source system with an end-to-end modeling method in the prior art can be used as a TTS voice generation algorithm, wherein the audio parameter verification information is used to restore the sound quality information of the user when storing his own standard voice audio, and the audio parameter feature vector in the audio parameter verification information can be input in a conditional coding manner to achieve audio parameter control. wherein, the present application does not need to express emotional features, such as tone fluctuations, sudden changes in speech speed, sentence stress, etc., in the conditional coding process, and the emotional features can be forcibly avoided during voice synthesis, and the voice text data reflects the semantic content of the user's voice audio.

[0051] In step S103, emotion interference extraction is performed based on the user voice audio and the speech without emotion intention to obtain emotion recognition speech, and audio feature conversion is performed based on the emotion recognition speech to obtain an emotion recognition feature vector corresponding to the user.

[0052] It should be noted that the user's voice audio contains two parts of information, namely: content information (i.e. what was said) and emotional information (i.e. how it was said, tone, pitch, rhythm changes, etc.), while speech without emotional intention only retains content information (emotional intention is close to zero) and is expressed in a stable and neutral manner. Therefore, the difference between the two is the emotional characteristics expressed by the user in the current voice. Extracting this difference as an emotion recognition voice feature can enhance the accuracy of identifying the user's emotional intention.

[0053] Optionally, in some embodiments, performing emotion interference extraction based on the user voice audio and the speech without emotion intention to obtain emotion recognition speech specifically includes:

[0054] The mel-cepstrum corresponding to the user voice audio and the speech without emotion intention is obtained respectively, and spectrum difference is performed based on the mel-cepstrum corresponding to the user voice audio and the speech without emotion intention respectively. After obtaining the differential spectrum corresponding to the two, spectrum restoration is performed to obtain the emotion recognition speech.

[0055] In specific implementation, the audio signals of the user voice audio and the speech without emotion intention can be framed and windowed, and the Mel-frequency inverse spectrum can be obtained by short-time Fourier transform. Then, the dynamic time warping (DTW) algorithm is used to align the Mel-frequency inverse spectrum of the user voice audio and the speech without emotion intention. Then, they are subtracted frame by frame to obtain the differential spectrum. Finally, the inverse Fourier transform is used to restore the spectrum, and the differential spectrum is converted into a time domain signal to obtain the emotion recognition speech corresponding to the user.

[0056] It should be noted that in this application, the spectrum is restored by inverse Fourier transform, and the differential spectrum is converted into a time domain signal to obtain the user's corresponding emotion recognition voice. Due to the signal conversion from frequency domain to time domain, there is a certain accuracy loss, especially the phase part cannot be completely reconstructed. However, for emotion recognition, it is only necessary to extract the signal change trend as the recognition feature. Therefore, a certain degree of audio detail loss does not affect the final emotion recognition result.

[0057] Preferably, in some embodiments, performing audio feature conversion based on the emotion recognition speech to obtain the emotion recognition feature vector corresponding to the user specifically includes:

[0058] Acquire speech-to-text data, determine a timestamp corresponding to a start time and an end time of each semantic word in the emotion recognition speech based on the emotion recognition speech and the speech-to-text data, and determine a plurality of semantic word time windows based on the timestamp corresponding to the start time and the end time of each semantic word;

[0059] Window features corresponding to the audio information in each semantic word time window are extracted, and each window feature is normalized and then formed into the emotion recognition feature vector according to the time sequence.

[0060] In specific implementation, the mean value of the Mel-frequency cepstral coefficient in the semantic word time window can be extracted as the window feature corresponding to the semantic word time window, and then the window features corresponding to each semantic word time window are normalized according to the maximum value of the window features corresponding to each semantic word time window, and the emotion recognition feature vector is composed according to the time sequence.

[0061] In specific implementation, other methods can also be used to determine the emotion recognition feature vector. The following is a specific embodiment of the present application for determining the emotion recognition feature vector. For any semantic word time window, first, the same type of audio features are extracted for the user voice audio and the speech audio without emotion intention respectively. For example, the real speech fundamental frequency curve-emotionless speech fundamental frequency curve is used as the pitch change feature, the real speech energy curve-emotionless speech energy curve is used as the energy change feature, the pronunciation duration difference is used as the speech speed change feature, and the maximum difference value of the Mel-frequency inverse spectrum frequency domain value is used as the spectrum change feature. Then, the statistical features of all the difference features in each semantic word time window, such as variance or standard deviation, are combined into a feature vector according to time sequence to obtain the emotion recognition feature vector.

[0062] In step S104, the speech text data is trained and learned based on the first artificial intelligence model to obtain an initial emotion label, and the emotion recognition feature vector is trained and learned using the second artificial intelligence model to obtain a corrected emotion label;

[0063] It should be noted that this application adopts a two-stage artificial intelligence model (a first artificial intelligence model and a second artificial intelligence model) for training and correction, in which the first model processes speech text data, obtains initial emotion labels, and provides preliminary emotion recognition results. The second model, based on the first stage, trains and corrects the extracted emotion recognition feature vectors to further optimize the accuracy of emotion labels. Through the dual training mechanism, emotion recognition not only relies on semantic analysis, but also combines audio features, thereby enhancing the robustness of recognition and being able to better adapt to emotional changes among different users and in different situations.

[0064] Optionally, in some embodiments, a long short-term memory neural network (LSTM) is used as the first artificial intelligence model to train and learn the audio-text data to obtain an initial emotion label, wherein the audio-text data is preprocessed, including but not limited to text segmentation, encoding and length normalization, and then input as input features into the LSTM neural network for sequence feature learning. During the training process, based on a large number of preset sample data sets, each audio-text data is labeled with a corresponding artificial emotion classification label, and a cross-entropy loss function is used for training optimization, so that the LSTM neural network can automatically learn the emotional expression features in the audio-text data, and finally generate the initial emotion label based on the output emotion prediction result.

[0065] In specific implementation, a large amount of text data generated by real conversations or speech transcription can be first collected to form a semantic sample library, and each text data can be manually classified and labeled with emotions. Emotion categories include but are not limited to: happiness, anger, sadness, surprise, disgust, calmness, etc. The preprocessing of text data can be achieved by performing Chinese word segmentation on the speech text data; mapping each word to the corresponding word vector index; unifying the length of the text sequence, padding short texts with zeros, truncating long texts, etc., and then constructing an emotion recognition neural network model based on the LSTM structure; the input layer is the Embedding layer, mapping the word index to a dense word vector; the hidden layer is one or more layers of LSTM Unit, used to extract contextual features in text sequences; the output layer is a Softmax layer, used to output the predicted probability of each emotion category, and then use manually annotated emotion classification labels as supervision signals; the parameters of the intermediate layer weights are optimized through the cross entropy loss function, and the optimization algorithm can use the Adam optimizer; and the model is trained by batch gradient descent; according to a certain number of iterative rounds (Epochs) until the model converges, it is judged that the training of the first artificial intelligence model is completed, and then the speech text data is input into the trained first artificial intelligence model, and the model can output the predicted probability corresponding to each emotion category; usually the category with the largest probability is taken as the initial emotion label.

[0066] Optionally, in some embodiments, a single hidden layer neural network is used as the second artificial intelligence model to train and learn the emotion recognition feature vector to obtain a corrected emotion label. The following is a specific embodiment of the present application using a single hidden layer neural network to train and learn the emotion recognition feature vector to obtain a corrected emotion label: a number of standard emotion recognition feature vectors and their corresponding manually labeled emotion classifications are input into the single hidden layer neural network as training samples, wherein the manually labeled emotion classification is calibrated using artificial experience, and the middle layer of the single hidden layer neural network has multiple activation functions for classifying the training samples, and then the clustering results of the emotion recognition feature vectors are output through the output layer. , and adjust the activation function parameters in the single hidden layer neural network until the standard deviation of the manually labeled emotion classification corresponding to the emotion recognition feature vector in the output clustering result is lower than the preset threshold, and the single hidden layer neural network training is determined to be completed, and then the emotion recognition feature vector is input into the trained single hidden layer neural network for clustering to obtain the clustering result corresponding to the emotion recognition feature vector. In some embodiments, the clustering result corresponding to the emotion recognition feature vector can be first obtained, and in the cluster cluster corresponding to the clustering result, the manually labeled emotion classification corresponding to all the standard emotion recognition feature vectors is used as the corrected emotion label corresponding to the emotion recognition feature vector.

[0067] In step S105 , the emotional intention of the corresponding user is identified based on the initial emotional label and the corrected emotional label to generate a final emotional intention recognition result of the user.

[0068] Optionally, in some embodiments, reference Figure 2 As shown in FIG, this figure is an exemplary flow chart of generating emotion intention recognition results in some embodiments of the present application. The emotion intention of the corresponding user is recognized based on the initial emotion label and the corrected emotion label, and generating the final emotion intention recognition result of the user specifically includes:

[0069] In step S1051, emotional intention credibility detection is performed based on the initial emotion recognition result and the corrected emotion recognition result to obtain emotional intention similarity;

[0070] In step S1052, when the emotion intention similarity is higher than a preset threshold, the emotion recognition initial result is corrected using the emotion recognition correction result to obtain a final emotion intention recognition result of the user;

[0071] In step S1053, when the emotional intention similarity is lower than a preset threshold, the initial emotional label is used as the final emotional intention recognition result of the user.

[0072] Preferably, in some embodiments, a mode of segmenting the user voice audio can be adopted to obtain multiple time periods of the user voice audio. In each time period, the initial emotion recognition result and the correction result of the emotion recognition are obtained, and corresponding emotion feature values are set for different emotion recognition results. Then, the emotion feature values corresponding to the initial emotion recognition results in different time periods are composed into a first emotion feature value sequence, and the emotion feature values corresponding to the correction results of the emotion recognition in different time periods are composed into a second emotion feature value sequence. The Pearson correlation coefficient of the first emotion feature value sequence and the second emotion feature value sequence is used as the emotion intention similarity.

[0073] In specific implementation, the initial emotion recognition result is corrected by the emotion recognition correction result to obtain the user's final emotion intention recognition result, which specifically includes: obtaining the confidence levels corresponding to the emotion recognition correction result and the initial emotion recognition result, and taking the emotion recognition result with a higher confidence level as the user's final emotion intention recognition result. In specific implementation, the predicted probability of the corresponding emotion category output by the first artificial intelligence model can be used as the confidence level of the initial emotion recognition result, and the proportion of the largest number of manually labeled emotion classifications in the clustering result corresponding to the emotion recognition feature vector can be used as the confidence level corresponding to the emotion recognition correction result.

[0074] It should be noted that this application introduces emotionless speech for training, and the system can reduce errors caused by user emotional fluctuations. For example, when a user is in an emotional state such as nervousness, anger or happiness, changes in speech speed, pitch, etc. may affect semantic understanding. The generation of emotionless speech helps the system avoid over-reliance on emotional information in the speech content during the emotion recognition process, thereby improving the system's adaptability to different emotional states. Even if the user's voice is interfered with by environmental noise, emotional changes, etc., the system can still perform more robust emotional intent recognition based on emotionless intent samples and audio features.

[0075] In addition, in another aspect of the present application, in some embodiments, the present application provides an emotion intention recognition system based on AI recognition, the device includes a speech emotion recognition unit, reference Figure 3 , which is a schematic diagram of exemplary hardware and / or software structure of a speech emotion recognition unit according to some embodiments of the present application. The speech emotion recognition unit 200 includes: a speech acquisition module 201, a speech generation module 202, and an emotion recognition module 203, which are described as follows:

[0076] The voice collection module 201 is used to collect user voice audio through an intelligent voice sensor;

[0077] The speech generation module 202 is configured to perform speech-to-text recognition on the user's speech audio to obtain speech-to-text data, obtain audio parameter verification information corresponding to the user's speech audio, generate a non-emotional intent sample based on the audio parameter verification information and the speech-to-text data, and obtain the user's non-emotional intent speech;

[0078] The emotion recognition module 203 is configured to extract emotion interference based on the user's voice audio and the speech without emotion intention to obtain emotion recognition speech, and perform audio feature conversion based on the emotion recognition speech to obtain an emotion recognition feature vector corresponding to the user;

[0079] The emotion recognition module 203 is further configured to train the speech text data based on a first artificial intelligence model to obtain an initial emotion label, and train the emotion recognition feature vector using a second artificial intelligence model to obtain a corrected emotion label;

[0080] The emotion recognition module 203 is further configured to recognize the emotion intention of the corresponding user based on the initial emotion label and the corrected emotion label, and generate a final emotion intention recognition result of the user.

[0081] The above describes in detail an example of an emotion intention recognition system and method based on AI recognition provided in an embodiment of the present application. It can be understood that in order to achieve the above functions, the corresponding device includes hardware structures and / or software modules corresponding to executing each function.

[0082] Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function in the application is executed in hardware or in a computer software-driven hardware manner depends on the specific application and design constraints of the technical solution. Therefore, professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0083] In addition, the present application also provides a computer terminal device, which includes a memory and a processor, the memory stores code, and the processor is configured to obtain the code and execute the above-mentioned emotion intention recognition method based on AI recognition.

[0084] In some embodiments, reference Figure 4 , which is a schematic diagram of the structure of a computer terminal device that implements an AI-based emotional intention recognition method according to some embodiments of the present application. In the above embodiment, an AI-based emotional intention recognition method can be Figure 4 The computer terminal device 300 shown in FIG. 1 is implemented as shown in FIG. 1 , and the computer terminal device 300 includes at least one communication bus 301 , a communication interface 302 , a processor 303 and a memory 304 .

[0085] The processor 303 can be a general-purpose central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more for controlling the execution of an emotion intention recognition method based on AI recognition in this application.

[0086] The communication bus 301 may include a path for transmitting information between the aforementioned components.

[0087] The memory 304 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, an optical disc storage (including a compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 304 may exist independently and be connected to the processor 303 via the communication bus 301. The memory 304 may also be integrated with the processor 303.

[0088] Memory 304 is used to store program code for executing the solution of the present application, and is controlled by processor 303 for execution. Processor 303 is used to execute the program code stored in memory 304. The program code may include one or more software modules. In the above embodiment, the determination of the emotion recognition feature vector can be implemented by processor 303 and one or more software modules in the program code in memory 304.

[0089] The communication interface 302 uses any transceiver or other device for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area network (WLAN), etc.

[0090] Optionally, the computer terminal device 300 may further include a power supply 305 for providing power to various devices or circuits in the real-time computer terminal device.

[0091] In a specific implementation, as an embodiment, a computer terminal device may include multiple processors, each of which may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. The processor herein may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0092] The computer terminal device can be a general-purpose computer terminal device or a dedicated computer terminal device. In a specific implementation, the computer terminal device can be a desktop computer, a portable computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device, or an embedded device. The embodiments of the present application do not limit the type of computer terminal device.

[0093] In addition, in other aspects of the present application, a computer-readable storage medium is provided, which stores at least one computer program, and the computer program is loaded and executed by a processor to implement the operations performed by the above-mentioned emotion intention recognition method based on AI recognition.

[0094] In summary, in an emotion intention recognition system and method based on AI recognition disclosed in an embodiment of the present application, first, an intelligent voice sensor collects user voice audio; voice text recognition is performed on the user voice audio to obtain voice text data, audio parameter verification information corresponding to the user voice audio is obtained, and a sample without emotion intention is generated based on the audio parameter verification information and the voice text data to obtain the user's emotionless intention voice; emotion interference is extracted based on the user voice audio and the emotionless intention voice to obtain emotion recognition voice, and audio feature conversion is performed based on the emotion recognition voice to obtain the emotion recognition feature vector corresponding to the user; the voice text data is trained and learned based on the first artificial intelligence model to obtain an initial emotion label, and the emotion recognition feature vector is trained and learned through the second artificial intelligence model to obtain a corrected emotion label; the emotion intention of the corresponding user is recognized based on the initial emotion label and the corrected emotion label to generate the user's final emotion intention recognition result, which can recognize the user's emotion intention through the user's emotionless intention voice, thereby improving the accuracy and robustness of user emotion intention recognition.

[0095] The above description is merely an embodiment of the present application. Common knowledge such as the specific technical solutions or features of the solutions is not described in detail herein. It should be noted that those skilled in the art may make various modifications and improvements without departing from the technical solution of the present application, and these modifications and improvements should also be considered within the scope of protection of the present application. These modifications and improvements will not affect the effectiveness of the implementation of the present application or the practical application of the patent.

[0096] The scope of protection claimed by this application shall be determined by the content of the claims. The specific embodiments and other descriptions in the specification may be used to interpret the content of the claims. Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of the invention. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application is intended to include such modifications and variations.

Claims

1. A method for emotional intention recognition based on AI recognition, characterized in that: include: Collect user voice audio through intelligent voice sensors; Performing speech-to-text recognition on the user's voice audio to obtain speech-to-text data, obtaining audio parameter verification information corresponding to the user's voice audio, and generating a non-emotional intent sample based on the audio parameter verification information and the speech-to-text data to obtain the user's non-emotional intent speech; Performing emotion interference extraction based on the user voice audio and the speech without emotion intention to obtain emotion recognition speech, and performing audio feature conversion based on the emotion recognition speech to obtain an emotion recognition feature vector corresponding to the user; Training the speech text data based on a first artificial intelligence model to obtain an initial emotion label, and training the emotion recognition feature vector using a second artificial intelligence model to obtain a corrected emotion label; The emotional intention of the corresponding user is identified based on the initial emotional label and the corrected emotional label to generate a final emotional intention recognition result of the user.

2. The method according to claim 1, wherein A microphone array and a front-end audio processing chip equipped with corresponding audio processing software are used as the intelligent voice sensor.

3. The method according to claim 1, wherein Emotional interference extraction is performed based on the user voice audio and the speech without emotion intention to obtain the emotion recognition speech, which specifically includes: obtaining the Mel-frequency cepstrum corresponding to the user voice audio and the speech without emotion intention respectively, performing spectral difference based on the Mel-frequency cepstrum corresponding to the user voice audio and the speech without emotion intention respectively, obtaining the corresponding differential spectrum between the two, and then performing spectral restoration to obtain the emotion recognition speech.

4. The method according to claim 1, wherein The audio feature conversion based on the emotion recognition speech is performed to obtain the emotion recognition feature vector corresponding to the user, specifically including: Acquire speech-to-text data, determine a timestamp corresponding to a start time and an end time of each semantic word in the emotion recognition speech based on the emotion recognition speech and the speech-to-text data, and determine a plurality of semantic word time windows based on the timestamp corresponding to the start time and the end time of each semantic word; Window features corresponding to the audio information in each semantic word time window are extracted, and each window feature is normalized and then formed into the emotion recognition feature vector according to the time sequence.

5. The method according to claim 1, wherein Identifying the emotional intention of the corresponding user based on the initial emotional label and the corrected emotional label to generate a final emotional intention identification result of the user specifically includes: Performing emotional intention credibility detection based on the initial emotion recognition result and the corrected emotion recognition result to obtain emotional intention similarity; When the emotion intention similarity is higher than a preset threshold, the emotion recognition initial result is corrected using the emotion recognition correction result to obtain a final emotion intention recognition result of the user; When the emotional intention similarity is lower than a preset threshold, the initial emotional label is used as the final emotional intention recognition result of the user.

6. The method according to claim 1, wherein A long short-term memory neural network is used as the first artificial intelligence model to train and learn the audio-text data to obtain initial emotion labels.

7. The method according to claim 1, wherein Before performing speech-to-text recognition on the user voice audio to obtain speech-to-text data, the method further includes: removing background noise from the user voice audio.

8. An emotion intention recognition system based on AI recognition, including a speech emotion recognition unit, characterized in that: The speech emotion recognition unit includes: Voice collection module, used to collect user voice audio through intelligent voice sensor; A speech generation module is configured to perform speech-to-text recognition on the user's speech audio to obtain speech-to-text data, acquire audio parameter verification information corresponding to the user's speech audio, generate a non-emotional intent sample based on the audio parameter verification information and the speech-to-text data, and obtain the user's non-emotional intent speech; An emotion recognition module is configured to extract emotion interference from the user's voice audio and the speech without emotion intention to obtain emotion recognition speech, and perform audio feature conversion based on the emotion recognition speech to obtain an emotion recognition feature vector corresponding to the user; The emotion recognition module is further configured to train and learn the speech text data based on a first artificial intelligence model to obtain an initial emotion label, and to train and learn the emotion recognition feature vector using a second artificial intelligence model to obtain a corrected emotion label; The emotion recognition module is further configured to recognize the emotion intention of the corresponding user based on the initial emotion label and the corrected emotion label, and generate a final emotion intention recognition result of the user.

9. A computer terminal device, characterized in that: The computer terminal device includes a memory and a processor, the memory stores a code, and the processor is configured to obtain the code and execute the emotion intention recognition method based on AI recognition as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing at least one computer program, characterized in that: The computer program is loaded and executed by a processor to implement the operations performed by the emotion intention recognition method based on AI recognition as described in any one of claims 1 to 7.