Emotion recognition method and device, equipment and medium
By determining the speech determination anchor point and multiple predictors in the speech recognition emotion recognition system, and using the promotion tree for model ensemble learning, the problems of low accuracy and poor robustness in the prior art are solved, and higher recognition accuracy and better individual differences are achieved.
Patent Information
- Application Number
- CN202510344420.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-20
AI Technical Summary
In the prior art, the accuracy of speech recognition emotions is low and the robustness of a single model is poor, so individual differences cannot be fully considered.
By determining the speech determination anchor point and multiple predictors, a training set corresponding to each predictor is constructed, and a promotion tree is used to integrate learning of multiple emotions recognition models to obtain the target emotion recognition model.
The accuracy of emotion recognition is improved, the problems of poor robustness of a single model and inadequate individual differences are avoided, and the integration of multi-dimensional information and multi-model integration are achieved.
Smart Images

Figure CN120183442A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and medical health technology, and in particular to an emotion recognition method, device, equipment and medium. Background Art
[0002] In recent years, with the continuous development of artificial intelligence, artificial intelligence technology has been widely used in various fields such as finance and medical health. How to recognize user emotions based on user voice is an important topic. For example, in the field of medical health, how to effectively identify customer emotions to better serve customers is the key to intelligent medical customer service.
[0003] In the prior art, when performing emotion recognition based on user voice, the following methods are mainly used:
[0004] (1) Use real-time signal processing technology and deep learning models to perform real-time emotion recognition on sound signals. For example, use Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN) to extract features from sound signals and perform emotion classification.
[0005] This method uses a deep learning solution, which is an end-to-end learning solution. It does not adopt effective feature engineering, resulting in a huge amount of data required for training, poor robustness, and difficulty in improving accuracy, resulting in poor actual use results.
[0006] (2) Combine sound, text and visual information for emotion recognition and use deep learning models for multimodal feature fusion.
[0007] Although this method looks good, the cost of acquiring and annotating the corpus is huge, and many scenarios (such as call center scenarios) do not have visual information, so the usage scenarios are restricted.
[0008] (3) Extract the pitch, intensity, speaking speed and other features of the sound signal and use the support vector machine (SVM) for classification.
[0009] This method belongs to traditional machine learning, usually binary classification, which does not fully reflect the emotions reflected by process changes. Therefore, it not only has poor effects, but also has poor scenario applicability.
[0010] As can be seen from the above several methods, the existing solutions for emotion recognition in speech have not fully considered the emotional changes represented by temporal variations. For example, some people are naturally fast and loud speakers, but this does not necessarily mean they are angry or impatient. Moreover, a single method easily reaches a bottleneck in accuracy and is difficult to further improve. The GAP (Global Average Processing Time) of accuracy makes it difficult to satisfy the business side in practical applications, so it is difficult to truly be put into production and generate benefits. Summary of the Invention
[0011] In view of the above, it is necessary to provide a method, device, equipment and medium for emotion recognition, aiming to solve the problem of low accuracy in emotion recognition.
[0012] A method for emotion recognition, the emotion recognition method includes:
[0013] Determine the speech judgment anchor points and determine multiple prediction factors;
[0014] Construct a training set corresponding to each prediction factor, and use the training set corresponding to each prediction factor to train an emotion recognition model respectively;
[0015] Use boosting trees to perform ensemble learning on the trained multiple emotion recognition models to obtain a target emotion recognition model;
[0016] In response to an emotion recognition instruction for a target user, collect the speech information of the target user as the speech to be recognized according to the speech judgment anchor points;
[0017] Input the speech to be recognized into the target emotion recognition model, and obtain the recognition results of each tree in the target emotion recognition model;
[0018] Determine the recognition type according to the emotion recognition instruction;
[0019] Process the recognition results of each tree according to the recognition type to obtain the emotion recognition result of the target user.
[0020] An emotion recognition device, the emotion recognition device includes:
[0021] A determination unit, configured to determine the speech judgment anchor points and determine multiple prediction factors;
[0022] A training unit, configured to construct a training set corresponding to each prediction factor, and use the training set corresponding to each prediction factor to train an emotion recognition model respectively;
[0023] A learning unit, configured to use boosting trees to perform ensemble learning on the trained multiple emotion recognition models to obtain a target emotion recognition model;
[0024] An acquisition unit, configured to respond to an emotion recognition instruction for a target user, and acquire voice information of the target user as the voice to be recognized according to the voice determination anchor point;
[0025] An input unit, configured to input the voice to be recognized into the target emotion recognition model, and obtain recognition results of each tree in the target emotion recognition model;
[0026] The determination unit is further configured to determine a recognition type according to the emotion recognition instruction;
[0027] A processing unit, configured to process the recognition results of each tree according to the recognition type to obtain an emotion recognition result of the target user.
[0028] A computer device, the computer device includes:
[0029] A memory, storing at least one instruction; and
[0030] A processor, configured to execute the instruction stored in the memory to implement the emotion recognition method.
[0031] A computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is executed by a processor in a computer device to implement the emotion recognition method.
[0032] It can be seen from the above technical solutions that the present invention can respectively train an emotion recognition model by using the training set corresponding to each prediction factor, and perform ensemble learning on the multiple trained emotion recognition models by using boosting trees to obtain a target emotion recognition model, avoiding problems such as poor robustness of a single model and inability to fully consider individual differences, and realizing a large voice emotion determination model with multi-dimensional information, multiple model boosting trees for enhancement, and comprehensive use of voice changes and semantic information; acquiring voice information of a target user as the voice to be recognized according to the voice determination anchor point, inputting the voice to be recognized into the target emotion recognition model, and processing the recognition results of each tree according to the recognition type to obtain an emotion recognition result of the target user. By configuring the voice determination anchor point, it is possible to avoid short voice judgments at the word level, thereby greatly reducing judgment noise and improving the accuracy of emotion recognition. Description of the Drawings
[0033] Figure 1 is a flowchart of a preferred embodiment of the emotion recognition method of the present invention.
[0034] Figure 2 is a functional module diagram of a preferred embodiment of the emotion recognition device of the present invention.
[0035] Figure 3It is a schematic structural diagram of a computer device which is a preferred embodiment for implementing the emotion recognition method of the present invention. Detailed implementation manners
[0036] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0037] As Figure 1 shown, it is a flowchart of a preferred embodiment of the emotion recognition method of the present invention. According to different requirements, the order of steps in this flowchart can be changed and some steps can be omitted.
[0038] The emotion recognition method is applied to one or more computer devices. The computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.
[0039] The computer device can be any electronic product that can perform human-computer interaction with a user. For example, a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an Internet protocol television (IPTV), a smart wearable device, etc.
[0040] The computer device may further include a network device and / or a user device. Among them, the network device includes, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing.
[0041] The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms.
[0042] Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0043] The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0044] The network where the computer device is located includes but is not limited to the Internet, wide area network, metropolitan area network, local area network, Virtual Private Network (VPN), etc.
[0045] S10, determine the speech determination anchor point and determine multiple prediction factors.
[0046] Among them, the prediction factors may include factors that play a key role in emotion recognition.
[0047] In this embodiment, the determining the speech determination anchor point and determining multiple prediction factors includes:
[0048] Obtain historical emotion recognition data;
[0049] Conduct validity analysis based on the historical emotion recognition data to obtain the effective speech duration and the effective word count threshold as the speech determination anchor point;
[0050] Extract features from the historical emotion recognition data to obtain multiple features associated with emotion recognition;
[0051] Conduct importance analysis on the multiple features to obtain an analysis result;
[0052] Select the multiple prediction factors from the multiple features according to the analysis result.
[0053] Among them, the effective speech duration and the effective word count threshold can be subjected to sentence segmentation through technologies such as VAD (Voice Activity Detection) and ASR (Automatic Speech Recognition) to obtain the optimal speech duration and word count threshold. For example, the effective speech duration can be configured to be 8 seconds, and the word count threshold can be configured to be 16 words. That is to say, when the speech duration is greater than or equal to 8 seconds and the input word count is greater than or equal to 16 words, the input speech can be used for emotion recognition, fully considering the characteristic that emotion is related to the speaking sentence, and avoiding the unnecessary consumption of judgment resources and judgment noise caused by too short speech. Because in principle, it is unlikely that a person's emotion will mutate at the word level, thus avoiding the judgment of too short speech at the word level and greatly reducing the judgment noise.
[0054] Among them, importance analysis can be performed on the multiple features by means of random forest or the like, and the features with higher importance can be obtained as the multiple predictors.
[0055] For example: the multiple predictors may include, but are not limited to, a combination of one or more of the following factors:
[0056] (1) Pitch. Feature: Pitch reflects the frequency of vocal cord vibration, and the pitch changes when the emotion fluctuates. Implementation: By extracting the fundamental frequency (F0) and analyzing its change range and fluctuation pattern. For example, the pitch usually increases and fluctuates greatly when angry, and decreases and is stable when sad.
[0057] Specific manifestations: 1) Pitch increase. Feature: The fundamental frequency (F0) increases, and the frequency of the sound increases; Emotions: Anger, excitement, joy, anxiety. 2) Pitch decrease. Feature: The fundamental frequency (F0) decreases, and the frequency of the sound decreases; Emotions: Sadness, depression, fatigue, calm. 3) Pitch fluctuation. Feature: The fundamental frequency (F0) is unstable, sometimes high and sometimes low; Emotions: Anxiety, tension, complex emotions.
[0058] (2) Intensity. Feature: Intensity is related to the amplitude of the sound, and emotional changes will affect the speaking volume. Implementation: Calculate the root mean square energy (RMS) of the sound signal and analyze its changes. For example, the volume usually increases when angry and decreases when sad.
[0059] Specific manifestations: 1) Intensity increase. Feature: The amplitude of the sound increases, and the volume is significantly increased; Emotions: Anger, excitement, joy, anxiety. 2) Intensity decrease. Feature: The amplitude of the sound decreases, and the volume is significantly reduced; Emotions: Sadness, depression, fatigue, calm. 3) Intensity fluctuation. Feature: The volume is unstable, sometimes large and sometimes small; Emotions: Anxiety, tension, complex emotions.
[0060] (3) Speech Rate. Feature: Speech rate refers to the number of syllables pronounced per unit time, and emotional changes can affect the speaking speed. Implementation: Calculate the pronunciation speed of syllables or words through speech recognition technology. For example, the speech rate speeds up when anxious and slows down when relaxed.
[0061] Specific manifestations: 1) Speeding up of speech rate. Feature: The intervals between syllables and words are shortened, and the overall speech rate is significantly accelerated; Emotions: Anxiety, tension, excitement, anger. 2) Slowing down of speech rate. Feature: The intervals between syllables and words are lengthened, and the overall speech rate is significantly slowed down; Emotions: Sadness, depression, fatigue, contemplation. 3) Fluctuation of speech rate. Feature: The speech rate is unstable, sometimes fast and sometimes slow; Emotions: Hesitation, uncertainty, complex emotions.
[0062] (4) Spectral Features. Feature: Spectral features such as formant frequencies and bandwidths reflect the shape changes of the vocal tract. Implementation: Use Fourier transform or Linear Predictive Coding (LPC) to extract spectral features. Under different emotions, the positions and distributions of formants will be different.
[0063] Specific manifestations: 1) Anger. Formant position: When angry, the vocal cord tension increases and the shape of the vocal tract changes, usually resulting in an increase in the frequencies of the first formant (F1) and the second formant (F2). Formant distribution: The distribution of formants may be more dispersed and fluctuate greatly, reflecting the excitement and instability of emotions. 2) Excitement and joy. Formant position: When excited and joyful, the vocal cord tension increases and the shape of the vocal tract changes, usually resulting in an increase in the frequencies of the first formant (F1) and the second formant (F2). Formant distribution: The distribution of formants is more concentrated and fluctuates more significantly, reflecting the activity and positivity of emotions. 3) Anxiety and tension. Formant position: When anxious and tense, the vocal cord tension increases and the shape of the vocal tract changes, usually resulting in an increase in the frequencies of the first formant (F1) and the second formant (F2), but there may also be unstable fluctuations. Formant distribution: The distribution of formants may be unstable, sometimes high and sometimes low, reflecting the fluctuations and uncertainties of emotions. 4) Sadness and depression. Formant position: When sad and depressed, the vocal cords relax and the shape of the vocal tract changes, usually resulting in a decrease in the frequencies of the first formant (F1) and the second formant (F2). Formant distribution: The distribution of formants is more concentrated and fluctuates less, reflecting the low and stable emotions. 5) Calm and relaxed. Formant position: When calm and relaxed, the vocal cords and vocal tract are in a natural state, and the positions of the formants are usually more stable and moderate. Formant distribution: The distribution of formants is relatively stable and fluctuates little, reflecting the stable and relaxed emotions.
[0064] (5) Voice Quality. Characteristics: Voice quality involves the smoothness, roughness, etc. of the voice, and emotional changes will affect the vibration mode of the vocal cords. Implementation: Evaluate voice quality by analyzing the Harmonic-to-Noise Ratio (HNR) or Jitter and shimmer perturbations. The voice may be rougher when tense and smoother when relaxed.
[0065] (6) Pauses and Rhythm. Characteristics: The frequency and duration of pauses and the rhythm of speech can also reflect emotions. Implementation: Detect the silent segments in the speech and analyze the frequency and duration of pauses. Pauses may decrease when anxious and increase when sad.
[0066] Specific relationships between emotions and pauses, rhythm: 1) Anger: Short and frequent pauses, fast speech rate, high voice intensity. May be accompanied by sudden pauses or rhythm changes. 2) Sadness: Long and frequent pauses, slow speech rate, low voice intensity. The rhythm may seem drawn - out or incoherent. 3) Anxiety: Frequent and irregular pauses, the speech rate may be fast or slow. The rhythm is unstable and may be accompanied by repetition or stuttering. 4) Excitement: Few pauses, fast speech rate, smooth rhythm. High voice intensity, obvious intonation changes. 5) Relaxation: Natural pauses, moderate speech rate, stable rhythm. Moderate voice intensity, gentle intonation.
[0067] (7) Text information. Implementation: Combine the text information recognized by ASR (Automatic Speech Recognition), and then combine with LLM (Large Language Model) to determine emotions, so as to provide the emotion type determined solely from the text.
[0068] S11, construct a training set corresponding to each predictor, and use the training set corresponding to each predictor to train an emotion recognition model respectively.
[0069] In this embodiment, the constructing a training set corresponding to each predictor and using the training set corresponding to each predictor to train an emotion recognition model respectively includes:
[0070] Collect user speech according to each predictor;
[0071] Mark the collected user speech, and use the marked user speech to construct a training set corresponding to each predictor;
[0072] Use the training set corresponding to each predictor to train a model with prediction function respectively, and obtain an emotion recognition model corresponding to each predictor.
[0073] Among them, the predictive factor values and corresponding emotion types in the user's speech can be marked.
[0074] Among them, the model with prediction function may include, but is not limited to, decision trees, large language models, etc.
[0075] Through the above embodiments, it is possible to analyze voice characteristics such as pitch, intensity, speech rate, spectral characteristics, voice quality, pauses, and rhythm, train corresponding models respectively, and combine text information and the determination results of the LLM large model to effectively detect the emotional changes of the speaker.
[0076] S12. Use the Boosting Tree to perform ensemble learning on multiple trained emotion recognition models to obtain the target emotion recognition model.
[0077] In this embodiment, the use of the Boosting Tree to perform ensemble learning on multiple trained emotion recognition models to obtain the target emotion recognition model includes:
[0078] Construct an initial model;
[0079] Use the initial model as the first-round model. In each round of training, add emotion recognition models one by one on the basis of the current-round model and perform iterative training;
[0080] When all emotion recognition models have completed iteration, stop training and determine the currently obtained model as the target emotion recognition model;
[0081] Among them, in each round of iteration, use the newly added emotion recognition model to fit the residuals of the current-round model.
[0082] Among them, the initial model can be a simple model such as a constant model. For example: Suppose there are 5 samples in the training data, and the corresponding emotion intensity values are -0.2, 0.1, 0.3, -0.1, 0.2 respectively. First calculate the sum of these values: (-0.2 + 0.1 + 0.3 - 0.1 + 0.2) = 0.3, and then divide by the number of samples 5 to get an average of 0.3 ÷ 5 = 0.06. Then the constant model will predict the emotion intensity of all new input data to be 0.06, that is, it is considered to be in a relatively neutral but slightly positive emotional state.
[0083] Of course, the initial prediction value of the initial model can also be the emotion recognition result of the majority type, the median, etc.
[0084] Among them, the use of the newly added emotion recognition model to fit the residuals of the current-round model includes:
[0085] For each training sample, calculate the residual between the predicted value and the true value of the current-round model to obtain the residual corresponding to each training sample;
[0086] Use a decision tree to fit the residuals corresponding to each training sample to minimize the value output by each leaf node of the decision tree.
[0087] Among them, the residual represents the difference between the true value and the predicted value of the current model. For each sample, calculate the residual between the predicted value of the current model and the true value, which represents the part that the current model fails to explain.
[0088] Among them, when using the boosting tree for ensemble learning of multiple emotion recognition models obtained by training, update the predicted value of the newly added emotion recognition model in each round of iteration to the predicted value of the model in that round, and obtain the predicted value of the model obtained after each round of training;
[0089] Among them, calculate the product of the predicted value of the newly added emotion recognition model and the configuration value in each round of iteration to obtain the update step size, and use the update step size to control the update amplitude to prevent overfitting.
[0090] Among them, a certain number of iterations can also be configured, and the training stops when the number of iterations is reached. Or the overall performance of the model can also be detected, and the training stops when the overall performance no longer improves.
[0091] Among them, the boosting tree can significantly improve the prediction accuracy by combining multiple weak models, is applicable to regression, classification, and ranking problems, and has a certain robustness to missing values and outliers.
[0092] In the above embodiment, ensemble learning is performed using a multiple-model combined boosting tree. The core idea of the boosting tree is to improve the performance of the overall model by continuously correcting the errors of the previous model, so that a high accuracy of the overall model can be achieved even when the accuracy of a single model is limited.
[0093] S13, in response to an emotion recognition instruction for a target user, collect the voice information of the target user as the voice to be recognized according to the voice determination anchor point.
[0094] In this embodiment, the emotion recognition instruction can be triggered automatically. For example: in the field of medical and health, for a smart medical customer service, when it detects that a user inputs voice, it can automatically trigger real-time emotion recognition of the customer, so as to be able to pay attention to the customer's emotional changes and provide more considerate consulting services and care. For a mental health counselor, when having a conversation with a patient, the emotional state of the patient can be monitored in real time according to the patient's voice, so as to relieve the patient's emotions more scientifically and in a timely manner. For example, patients with depression and anxiety often have unique emotional characteristics in their voices. By using voice emotion recognition technology to analyze the emotional state in the patient's voice, it can provide a reference basis for doctors' diagnosis and assist in the early screening and diagnosis of mental diseases.
[0095] Of course, in the financial field, emotion recognition can also be used to assist in better serving customers. For example, in the financial customer service scenario, with the help of voice emotion recognition technology, the emotional changes in the customer's voice can be captured in real time. When detecting negative emotions such as irritability and anger in the customer, the call can be transferred to experienced customer service staff in a timely manner to prioritize the handling of customer problems and avoid customer loss. By analyzing the emotional information in a large amount of customer service voice data, financial institutions can understand the satisfaction of customers with products and services, identify the key links of customer dissatisfaction, and optimize product design and service processes. In the loan application or financial transaction scenario, abnormal emotions may be a manifestation of fraud. Voice emotion recognition technology can be used as an auxiliary means to analyze the voice emotions of customers when applying for loans or conducting important transactions. If a customer shows excessive nervousness, anxiety or unnatural emotions, it may mean that there are potential risks, prompting financial institutions to conduct more in-depth investigations and reviews. For the sales process of high-risk investment products, voice emotion recognition can also be used to monitor the emotional state of investors, judge whether their understanding and tolerance of risks match the products they invest in, avoid investors buying unsuitable products due to impulsive decisions, and reduce investment risks.
[0096] In this embodiment, the user voice can be collected according to the duration and number of words specified by the voice determination anchor point to improve the usability of the voice, thereby improving the accuracy of emotion recognition.
[0097] S14, input the voice to be recognized into the target emotion recognition model, and obtain the recognition results of each tree in the target emotion recognition model.
[0098] Among them, the recognition result of each tree can be a specific value or an emotion type. For different types of recognition tasks, the model output will also be different.
[0099] S15, determine the recognition type according to the emotion recognition instruction.
[0100] In this embodiment, the recognition type may include a classification type and a regression type.
[0101] Among them, the classification type means that a specific emotion type needs to be determined.
[0102] Among them, the regression type means that a specific value is output, and the degree of emotion can be determined through this value.
[0103] S16, process the recognition results of each tree according to the recognition type to obtain the emotion recognition result of the target user.
[0104] In this embodiment, the process of processing the recognition results of each tree according to the recognition type to obtain the emotion recognition result of the target user includes:
[0105] Obtain the weight of each tree;
[0106] Determine the prediction value of each tree according to the recognition result of each tree;
[0107] Calculate the weighted sum according to the weight of each tree and the prediction value of each tree to obtain the target prediction value;
[0108] When the recognition type is a classification type, obtain a pre-configured classification threshold, compare the target prediction value with the classification threshold to obtain a comparison result, and determine the emotion recognition result of the target user according to the comparison result; or
[0109] When the recognition type is a regression type, determine the target prediction value as the emotion recognition result of the target user.
[0110] Among them, the weight of each tree is determined by its performance during the training process.
[0111] For example: for the classification problem, emotions are divided into a finite number of discrete categories. For example, emotions are divided into fixed categories such as "happy", "sad", "angry", "calm", etc. When training the model, these categories are used as labels to let the model learn the association between different features and various emotions. During prediction, the model will output the probability of belonging to each category according to the input features. If it is higher than a certain threshold, it means happy, and if it is lower than a certain threshold, it means not happy.
[0112] Another example: for the regression problem, the emotion intensity is quantified as a continuous value. For example, the degree of pleasure of emotions is continuously scored from -10 (extremely sad) to 10 (extremely happy). When training the model, sample data with such continuous emotion intensity values is used to let the model learn how to predict the corresponding value according to the input features. During prediction, the model will output a specific value representing the recognized emotion intensity. For example, if the predicted current emotion intensity value is 3, it means being in a relatively positive emotional state but not yet extremely happy.
[0113] In the above embodiment, the boosting tree gradually corrects the residuals of the model by iteratively training multiple decision trees, and finally obtains the decision-making result by weighted summation of the prediction results of all trees. This method can effectively improve the performance of the model, especially in dealing with complex non-linear relationships, with strong robustness and high accuracy, truly meeting the needs of actual scenarios such as medical health and finance.
[0114] As can be seen from the above technical solutions, the present invention can separately train an emotion recognition model using the training set corresponding to each prediction factor, and use a boosting tree to perform ensemble learning on the multiple trained emotion recognition models to obtain a target emotion recognition model, avoiding problems such as poor robustness of a single model and failure to fully consider individual differences, and realizing a large voice emotion determination model with multi-dimensional information, multiple model boosting trees for enhancement, and comprehensive use of voice changes and semantic information; collect the voice information of the target user as the voice to be recognized according to the voice determination anchor point, input the voice to be recognized into the target emotion recognition model, and process the recognition results of each tree according to the recognition type to obtain the emotion recognition result of the target user. By configuring the voice determination anchor point, it is possible to avoid short voice judgments at the word level, thereby greatly reducing judgment noise and improving the accuracy of emotion recognition.
[0115] As Figure 2 shown, it is a functional module diagram of a preferred embodiment of the emotion recognition device of the present invention. The emotion recognition device 11 includes a determination unit 110, a training unit 111, a learning unit 112, a collection unit 113, an input unit 114, and a processing unit 115. The module / unit referred to in the present invention means a series of computer program segments that can be executed by a processor and can complete a fixed function, and is stored in a memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0116] The determination unit 110 is used to determine a voice determination anchor point and determine multiple prediction factors.
[0117] Among them, the prediction factors may include factors that play a key role in emotion recognition.
[0118] In this embodiment, the determination unit 110 determines a voice determination anchor point and determines multiple prediction factors, including:
[0119] Obtain historical emotion recognition data;
[0120] Perform validity analysis on the historical emotion recognition data to obtain the effective voice duration and the effective word count threshold as the voice determination anchor point;
[0121] Extract features from the historical emotion recognition data to obtain multiple features associated with emotion recognition;
[0122] Perform importance analysis on the multiple features to obtain an analysis result;
[0123] Select the multiple prediction factors from the multiple features according to the analysis result.
[0124] Among them, the effective speech duration and the effective word count threshold can be subjected to sentence segmentation processing through VAD (Voice Activity Detection) and ASR (Automatic Speech Recognition) technologies to obtain the optimal speech duration and word count threshold. For example, the effective speech duration can be configured to be 8 seconds, and the word count threshold can be configured to be 16 words. That is to say, when the speech duration is greater than or equal to 8 seconds and the input word count is greater than or equal to 16 words, the input speech can be used for emotion recognition, fully considering the characteristic that emotion is related to the spoken sentence, and avoiding the meaningless consumption of judgment resources and judgment noise caused by overly short speech. Because in principle, it is unlikely that a person's emotion will mutate at the word level, thus avoiding the judgment of overly short speech at the word level and greatly reducing the judgment noise.
[0125] Among them, methods such as random forest can be used to analyze the importance of the multiple features, and obtain the features with higher importance as the multiple predictors.
[0126] For example: the multiple predictors may include, but are not limited to, one or more combinations of the following factors:
[0127] (1) Pitch. Feature: Pitch reflects the frequency of vocal cord vibration, and the pitch will change when the emotion fluctuates. Implementation: By extracting the fundamental frequency (F0) and analyzing its change range and fluctuation pattern. For example, when angry, the pitch usually rises and fluctuates greatly, while when sad, it decreases and is stable.
[0128] Specific manifestations: 1) Pitch increase. Feature: The fundamental frequency (F0) increases, and the frequency of the sound increases; Emotions: Anger, excitement, joy, anxiety. 2) Pitch decrease. Feature: The fundamental frequency (F0) decreases, and the frequency of the sound decreases; Emotions: Sadness, depression, fatigue, calmness. 3) Pitch fluctuation. Feature: The fundamental frequency (F0) is unstable, sometimes high and sometimes low; Emotions: Anxiety, tension, complex emotions.
[0129] (2) Intensity. Feature: Intensity is related to the amplitude of the sound, and emotional changes will affect the speaking volume. Implementation: Calculate the root mean square energy (RMS) of the sound signal and analyze its changes. For example, when angry, the volume usually increases, while when sad, it decreases.
[0130] Specific manifestations: 1) Intensity increase. Feature: The amplitude of the sound increases, and the volume is significantly increased; Emotions: Anger, excitement, joy, anxiety. 2) Intensity decrease. Feature: The amplitude of the sound decreases, and the volume is significantly reduced; Emotions: Sadness, depression, fatigue, calmness. 3) Intensity fluctuation. Feature: The volume is unstable, sometimes large and sometimes small; Emotions: Anxiety, tension, complex emotions.
[0131] (3) Speech Rate. Feature: Speech rate refers to the number of syllables pronounced per unit time, and emotional changes can affect the speaking speed. Implementation: Calculate the pronunciation speed of syllables or words through speech recognition technology. For example, the speech rate speeds up when anxious and slows down when relaxed.
[0132] Specific manifestations: 1) Speeding up of speech rate. Feature: The intervals between syllables and words are shortened, and the overall speech rate is significantly accelerated; Emotions: anxiety, tension, excitement, anger. 2) Slowing down of speech rate. Feature: The intervals between syllables and words are lengthened, and the overall speech rate is significantly slowed down; Emotions: sadness, depression, fatigue, contemplation. 3) Fluctuation of speech rate. Feature: The speech rate is unstable, sometimes fast and sometimes slow; Emotions: hesitation, uncertainty, complex emotions.
[0133] (4) Spectral Features. Feature: Spectral features such as formant frequencies and bandwidths reflect the shape changes of the vocal tract. Implementation: Use Fourier transform or Linear Predictive Coding (LPC) to extract spectral features. Under different emotions, the positions and distributions of formants will be different.
[0134] Specific manifestations: 1) Anger. Formant position: When angry, the vocal cord tension increases, and the shape of the vocal tract changes, usually resulting in an increase in the frequencies of the first formant (F1) and the second formant (F2). Formant distribution: The distribution of formants may be more dispersed and fluctuate more, reflecting the excitement and instability of emotions. 2) Excitement and joy. Formant position: When excited and joyful, the vocal cord tension increases, and the shape of the vocal tract changes, usually resulting in an increase in the frequencies of the first formant (F1) and the second formant (F2). Formant distribution: The distribution of formants is more concentrated and fluctuates more significantly, reflecting the activity and positivity of emotions. 3) Anxiety and tension. Formant position: When anxious and tense, the vocal cord tension increases, and the shape of the vocal tract changes, usually resulting in an increase in the frequencies of the first formant (F1) and the second formant (F2), but there may also be unstable fluctuations. Formant distribution: The distribution of formants may be unstable, sometimes high and sometimes low, reflecting the fluctuations and uncertainties of emotions. 4) Sadness and depression. Formant position: When sad and depressed, the vocal cords relax, and the shape of the vocal tract changes, usually resulting in a decrease in the frequencies of the first formant (F1) and the second formant (F2). Formant distribution: The distribution of formants is more concentrated and fluctuates less, reflecting the low and stable emotions. 5) Calm and relaxed. Formant position: When calm and relaxed, the vocal cords and vocal tract are in a natural state, and the positions of the formants are usually relatively stable and moderate. Formant distribution: The distribution of formants is relatively stable and fluctuates little, reflecting the stable and relaxed emotions.
[0135] (5) Voice Quality. Characteristics: Voice quality involves the smoothness, roughness, etc. of the voice, and emotional changes can affect the vibration mode of the vocal cords. Implementation: Evaluate voice quality by analyzing the Harmonic-to-Noise Ratio (HNR) or Jitter and shimmer perturbation. The voice may be rougher when tense and smoother when relaxed.
[0136] (6) Pauses and Rhythm. Characteristics: The frequency and duration of pauses and the rhythm of speech can also reflect emotions. Implementation: Detect the silent segments in the speech and analyze the frequency and duration of pauses. Pauses may decrease when anxious and increase when sad.
[0137] Specific relationships between emotions and pauses, rhythm: 1) Anger: Short and frequent pauses, fast speech rate, high voice intensity. May be accompanied by sudden pauses or rhythm changes. 2) Sadness: Long and frequent pauses, slow speech rate, low voice intensity. The rhythm may seem drawn - out or incoherent. 3) Anxiety: Frequent and irregular pauses, speech rate may be fast or slow. Unstable rhythm, may be accompanied by repetition or stuttering. 4) Excitement: Few pauses, fast speech rate, smooth rhythm. High voice intensity, obvious intonation changes. 5) Relaxation: Natural pauses, moderate speech rate, stable rhythm. Moderate voice intensity, gentle intonation.
[0138] (7) Text information. Implementation: Combine the text information recognized by ASR (Automatic Speech Recognition), and then combine with LLM (Large Language Model) to determine emotions, so as to provide the emotion type determined solely from the text.
[0139] The training unit 111 is used to construct a training set corresponding to each predictor and train the emotion recognition model using the training set corresponding to each predictor respectively.
[0140] In this embodiment, the training unit 111 constructs a training set corresponding to each predictor and trains the emotion recognition model using the training set corresponding to each predictor respectively, including:
[0141] Collect user speech according to each predictor;
[0142] Mark the collected user speech and use the marked user speech to construct a training set corresponding to each predictor;
[0143] Use the training set corresponding to each predictor to train a model with prediction function respectively, and obtain an emotion recognition model corresponding to each predictor.
[0144] Among them, the predictive factors and corresponding emotion types in the user's speech can be marked.
[0145] Among them, the model with prediction function may include, but is not limited to, decision trees, large language models, etc.
[0146] Through the above embodiments, voice features such as pitch, intensity, speech rate, spectral features, voice quality, pauses, and rhythm can be analyzed, corresponding models can be trained respectively, and combined with text information and the determination results of the LLM large model, the emotional changes of the speaker can be effectively detected.
[0147] The learning unit 112 is used to perform ensemble learning on multiple trained emotion recognition models using a Boosting Tree to obtain a target emotion recognition model.
[0148] In this embodiment, the learning unit 112 performing ensemble learning on multiple trained emotion recognition models using a Boosting Tree to obtain a target emotion recognition model includes:
[0149] Construct an initial model;
[0150] Taking the initial model as the first-round model, in each round of training, an emotion recognition model is added one by one on the basis of the current-round model, and iterative training is performed;
[0151] When all emotion recognition models have completed iteration, stop training, and determine the currently obtained model as the target emotion recognition model;
[0152] Among them, in each round of iteration, the residual of the current-round model is fitted using the newly added emotion recognition model.
[0153] Among them, the initial model can be a simple model such as a constant model. For example: Suppose there are 5 samples in the training data, and the corresponding emotion intensity values are -0.2, 0.1, 0.3, -0.1, 0.2 respectively. First calculate the sum of these values: (-0.2 + 0.1 + 0.3 - 0.1 + 0.2) = 0.3, and then divide by the number of samples 5 to get an average value of 0.3 ÷ 5 = 0.06. Then the constant model will predict the emotion intensity of all new input data to be 0.06, that is, it is considered to be in a relatively neutral but slightly positive emotional state.
[0154] Of course, the initial prediction value of the initial model can also be the emotion recognition result of the majority type, the median, etc.
[0155] Among them, the fitting of the residual of the current-round model using the newly added emotion recognition model includes:
[0156] For each training sample, calculate the residual between the predicted value of the current-round model and the true value to obtain the residual corresponding to each training sample.
[0157] Use a decision tree to fit the residual corresponding to each training sample so that the value output by each leaf node on the decision tree is minimized.
[0158] Among them, the residual represents the difference between the true value and the predicted value of the current model. For each sample, calculating the residual between the predicted value of the current model and the true value represents the part that the current model fails to explain.
[0159] Among them, when using the boosting tree for ensemble learning of multiple emotion recognition models obtained by training, update the predicted value of the newly added emotion recognition model in each iteration to the predicted value of the current-round model to obtain the predicted value of the model obtained after each round of training.
[0160] Among them, calculate the product of the predicted value of the newly added emotion recognition model and the configured value in each iteration to obtain the update step size, and use the update step size to control the update amplitude to prevent overfitting.
[0161] Among them, a certain number of iterations can also be configured, and training stops when the number of iterations is reached. Or the overall performance of the model can also be detected, and training stops when the overall performance no longer improves.
[0162] Among them, the boosting tree can significantly improve the prediction accuracy by combining multiple weak models, is applicable to regression, classification, and ranking problems, and has a certain robustness to missing values and outliers.
[0163] In the above embodiment, ensemble learning is performed using a multiple-model combination boosting tree. The core idea of the boosting tree is to improve the performance of the overall model by continuously correcting the errors of the previous model, so that a high accuracy of the overall model can be achieved even when the accuracy of a single model is limited.
[0164] The acquisition unit 113 is configured to, in response to an emotion recognition instruction for a target user, collect the voice information of the target user as the voice to be recognized according to the voice determination anchor point.
[0165] In this embodiment, the emotion recognition instruction can be triggered automatically. For example, in the field of healthcare, for intelligent healthcare customer service, when it is detected that a user inputs voice, the real-time emotion recognition of the customer can be automatically triggered, so as to pay attention to the emotional changes of the customer and provide more considerate consultation services and care. For mental health counselors, when having a conversation with a patient, they can monitor the patient's emotional state in real time according to the patient's voice, so as to relieve the patient's emotions more scientifically and in a timely manner. For example, patients with depression or anxiety often have unique emotional characteristics in their voices. Through voice emotion recognition technology, analyzing the emotional state in the patient's voice can provide a reference basis for doctors to assist in the early screening and diagnosis of mental diseases.
[0166] Of course, in the financial field, emotion recognition can also be used to assist in better serving customers. For example, in the financial customer service scenario, with the help of voice emotion recognition technology, the emotional changes in the customer's voice can be captured in real time. When it is detected that the customer has negative emotions such as irritability or anger, the call can be transferred to an experienced customer service staff in a timely manner to give priority to handling the customer's problem and avoid customer loss. By analyzing the emotional information in a large amount of customer service voice data, financial institutions can understand the customer's satisfaction with products and services, find out the key links of customer dissatisfaction, and optimize product design and service processes. In the loan application or financial transaction scenario, abnormal emotions may be a manifestation of fraud. Voice emotion recognition technology can be used as an auxiliary means to analyze the voice emotions of customers when applying for loans or conducting important transactions. If a customer shows excessive nervousness, anxiety or unnatural emotions, it may mean that there are potential risks, prompting financial institutions to conduct more in-depth investigations and audits. For the sales process of high-risk investment products, voice emotion recognition can also be used to monitor the emotional state of investors, judge whether their understanding and tolerance of risks match the products they invest in, avoid investors making impulsive decisions to buy products that are not suitable for themselves, and reduce investment risks.
[0167] In this embodiment, the user voice can be collected according to the duration and number of words specified by the voice determination anchor point to improve the usability of the voice, thereby improving the accuracy of emotion recognition.
[0168] The input unit 114 is configured to input the voice to be recognized into the target emotion recognition model and obtain the recognition results of each tree in the target emotion recognition model.
[0169] Among them, the recognition result of each tree can be a specific value or an emotion type. For different types of recognition tasks, the model output will also be different.
[0170] The determination unit 110 is further configured to determine the recognition type according to the emotion recognition instruction.
[0171] In this embodiment, the recognition types may include classification types and regression types.
[0172] Among them, the classification type refers to the need to determine specific emotion types.
[0173] Among them, the regression type refers to outputting a specific numerical value, and the degree of emotion can be determined through this numerical value.
[0174] The processing unit 115 is configured to process the recognition results of each tree according to the recognition type to obtain the emotion recognition result of the target user.
[0175] In this embodiment, the processing unit 115 processes the recognition results of each tree according to the recognition type to obtain the emotion recognition result of the target user, including:
[0176] Obtain the weight of each tree;
[0177] Determine the prediction value of each tree according to the recognition result of each tree;
[0178] Calculate the weighted sum according to the weight of each tree and the prediction value of each tree to obtain the target prediction value;
[0179] When the recognition type is a classification type, obtain a pre-configured classification threshold, compare the target prediction value with the classification threshold to obtain a comparison result, and determine the emotion recognition result of the target user according to the comparison result; or
[0180] When the recognition type is a regression type, determine the target prediction value as the emotion recognition result of the target user.
[0181] Among them, the weight of each tree is determined by its performance during the training process.
[0182] For example: for the classification problem, emotions are divided into a finite number of discrete categories. For example, emotions are divided into fixed categories such as "happy", "sad", "angry", "calm", etc. When training the model, these categories are used as labels to enable the model to learn the associations between different features and various types of emotions. During prediction, the model will output the probabilities belonging to each category according to the input features. If it is higher than a certain threshold, it means happy, and if it is lower than a certain threshold, it means not happy.
[0183] For another example, for the regression problem, the emotional intensity is quantified as a continuous value. For example, the degree of pleasure of the emotion is continuously scored from -10 (extremely sad) to 10 (extremely happy). When training the model, sample data with such continuous emotional intensity values is used to enable the model to learn how to predict the corresponding value based on the input features. When predicting, the model will output a specific value representing the recognized emotional intensity. For example, if the predicted current emotional intensity value is 3, it means being in a relatively positive emotional state but not yet reaching the extremely happy level.
[0184] In the above embodiment, the boosting tree gradually corrects the residuals of the model by iteratively training multiple decision trees, and finally obtains the decision-making decision by weighted summing the prediction results of all the trees. This method can effectively improve the performance of the model, especially perform well in dealing with complex non-linear relationships, with strong robustness and high accuracy, truly meeting the needs of actual scenarios such as medical health and finance.
[0185] It can be seen from the above technical solutions that the present invention can use the training set corresponding to each prediction factor to train the emotion recognition model respectively, and use the boosting tree to perform ensemble learning on the multiple trained emotion recognition models to obtain the target emotion recognition model, avoiding problems such as poor robustness of a single model and inability to fully consider individual differences, and realizing a large voice emotion determination model that has multi-dimensional information, multiple model boosting trees for enhancement, and comprehensively uses voice changes and semantic information; collect the voice information of the target user as the voice to be recognized according to the voice determination anchor point, input the voice to be recognized into the target emotion recognition model, and process the recognition results of each tree according to the recognition type to obtain the emotion recognition result of the target user. By configuring the voice determination anchor point, it is possible to avoid the short voice judgment at the word level, thereby greatly reducing the judgment noise and improving the accuracy of emotion recognition.
[0186] As Figure 3 shown, it is a schematic structural diagram of a computer device according to a preferred embodiment of the method for realizing emotion recognition of the present invention.
[0187] The computer device 1 may include a memory 12, a processor 13, and a bus (the arrow in the figure is the bus), and may also include a computer program stored in the memory 12 and executable on the processor 13, such as an emotion recognition program.
[0188] Those skilled in the art can understand that the schematic diagram is only an example of the computer device 1, and does not constitute a limitation on the computer device 1. The computer device 1 can be either a bus structure or a star structure. The computer device 1 may also include more or fewer other hardware or software than shown in the figure, or different component arrangements. For example, the computer device 1 may also include input / output devices, network access devices, etc.
[0189] It should be noted that the computer device 1 is only an example. Other existing or future possible electronic products that can be adapted to the present invention should also be included within the protection scope of the present invention and are hereby incorporated by reference.
[0190] Among them, the memory 12 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 12 can be an internal storage unit of the computer device 1, such as the mobile hard disk of the computer device 1. In other embodiments, the memory 12 can also be an external storage device of the computer device 1, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device 1. Further, the memory 12 can also include both the internal storage unit and the external storage device of the computer device 1. The memory 12 can be used not only to store application software installed on the computer device 1 and various types of data, such as the code of the emotion recognition program, etc., but also to temporarily store data that has been output or will be output.
[0191] In some embodiments, the processor 13 can be composed of integrated circuits. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions, including the combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 13 is the control core (Control Unit) of the computer device 1, connecting various components of the entire computer device 1 through various interfaces and lines, and by running or executing programs or modules stored in the memory 12 (such as executing the emotion recognition program, etc.), and calling the data stored in the memory 12, to execute various functions of the computer device 1 and process data.
[0192] The processor 13 executes the operating system of the computer device 1 and various installed application programs. The processor 13 executes the application programs to implement the steps in the above-mentioned various embodiments of the emotion recognition method, such as Figure 1 the steps shown.
[0193] Exemplarily, the computer program may be divided into one or more modules / units, and the one or more modules / units are stored in the memory 12 and executed by the processor 13 to implement the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into a determination unit 110, a training unit 111, a learning unit 112, a collection unit 113, an input unit 114, and a processing unit 115.
[0194] The integrated units implemented in the form of software function modules may be stored in a computer-readable storage medium. The software function modules stored in a storage medium include several instructions for causing a computer device (which may be a personal computer, a computer device, or a network device, etc.) or a processor to execute a part of the emotion recognition method according to each embodiment of the present invention.
[0195] If the modules / units integrated in the computer device 1 are implemented in the form of software function units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned method embodiments of the present invention, it may also be completed by a computer program instructing relevant hardware devices. The computer program may be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above-mentioned method embodiments may be implemented.
[0196] Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory, etc.
[0197] Furthermore, the computer-readable storage medium mainly includes a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc.
[0198] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.
[0199] The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, in Figure 3 it is only represented by a single straight line, but it does not mean that there is only one bus or one type of bus. The bus is arranged to realize the connection and communication between the memory 12 and at least one processor 13, etc.
[0200] Although not shown, the computer device 1 may further include a power supply (such as a battery) for powering each component. Preferably, the power supply can be logically connected to the at least one processor 13 through a power management device, so as to realize functions such as charge management, discharge management, and power consumption management through the power management device. The power supply may further include any components such as one or more DC or AC power supplies, a recharge device, a power failure detection circuit, a power converter or an inverter, and a power status indicator. The computer device 1 may further include a variety of sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0201] Furthermore, the computer device 1 may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is usually used to establish a communication connection between the computer device 1 and other computer devices.
[0202] Optionally, the computer device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the computer device 1 and to display a visual user interface.
[0203] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.
[0204] Those skilled in the art can understand that Figure 3 the structure shown does not constitute a limitation on the computer device 1, and it may include fewer or more components than shown, or combine certain components, or have a different component arrangement.
[0205] In combination with Figure 1 , the memory 12 in the computer device 1 stores multiple instructions to implement an emotion recognition method, and the processor 13 can execute the multiple instructions to implement:
[0206] Determine a voice decision anchor point and determine multiple predictors;
[0207] Construct a training set corresponding to each predictor, and use the training set corresponding to each predictor to train an emotion recognition model respectively;
[0208] Use a boosting tree to perform ensemble learning on the multiple trained emotion recognition models to obtain a target emotion recognition model;
[0209] In response to an emotion recognition instruction for a target user, collect the voice information of the target user as the voice to be recognized according to the voice decision anchor point;
[0210] Input the voice to be recognized into the target emotion recognition model, and obtain the recognition results of each tree in the target emotion recognition model;
[0211] Determine the recognition type according to the emotion recognition instruction;
[0212] Process the recognition results of each tree according to the recognition type to obtain the emotion recognition result of the target user.
[0213] Specifically, the specific implementation method of the above instructions by the processor 13 can refer to Figure 1Descriptions of relevant steps in corresponding embodiments are not elaborated here.
[0214] It should be noted that all data involved in this case are legally obtained. The non-company software tools or components appearing in the embodiments of this application are only introduced by way of example and do not represent actual use.
[0215] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation.
[0216] The present invention can be used in many general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0217] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0218] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a hardware plus software functional module.
[0219] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.
[0220] Therefore, in any aspect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Thus, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed by the present invention. Any reference signs in the claims should not be construed as limiting the claims concerned.
[0221] In addition, it is obvious that the word "comprising" does not exclude other elements or steps, and the singular does not exclude the plural. A plurality of elements or devices stated in the present invention can also be implemented by one element or device through software or hardware. Words such as "first" and "second" are used to denote names and do not denote any particular order.
[0222] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. An emotion recognition method, characterized in that: The emotion recognition method comprises: Determining a speech determination anchor point and determining a plurality of prediction factors; Construct a training set corresponding to each predictor, and use the training set corresponding to each predictor to train the emotion recognition model respectively; The boosted tree is used to perform integrated learning on multiple emotion recognition models obtained through training to obtain the target emotion recognition model; In response to an emotion recognition instruction for a target user, collecting voice information of the target user as the voice to be recognized according to the voice determination anchor point; Inputting the speech to be recognized into the target emotion recognition model, and obtaining the recognition result of each tree in the target emotion recognition model; Determining a recognition type according to the emotion recognition instruction; The recognition result of each tree is processed according to the recognition type to obtain the emotion recognition result of the target user.
2. The emotion recognition method according to claim 1, characterized in that: The step of determining a speech determination anchor point and determining a plurality of prediction factors comprises: Get historical emotion recognition data; Performing validity analysis based on the historical emotion recognition data to obtain effective speech duration and effective word count thresholds as the speech determination anchor points; Performing feature extraction on the historical emotion recognition data to obtain a plurality of features associated with emotion recognition; Performing importance analysis on the multiple features to obtain analysis results; The plurality of predictors are selected from the plurality of features according to the analysis results.
3. The emotion recognition method according to claim 1, characterized in that: The step of constructing a training set corresponding to each prediction factor and using the training set corresponding to each prediction factor to train the emotion recognition model includes: Collecting user speech according to each predictor; Label the collected user voices, and use the labeled user voices to construct a training set corresponding to each prediction factor; The training set corresponding to each predictor is used to train a model with prediction function respectively, so as to obtain an emotion recognition model corresponding to each predictor.
4. The emotion recognition method according to claim 1, characterized in that: The method of using the boosted tree to perform integrated learning on multiple emotion recognition models obtained through training to obtain a target emotion recognition model includes: Build an initial model; Using the initial model as the first round model, adding emotion recognition models one by one on the basis of the current round model in each round of training, and performing iterative training; When all emotion recognition models have completed iteration, the training is stopped, and the currently obtained model is determined as the target emotion recognition model; In each round of iteration, the newly added emotion recognition model is used to fit the residual of the current round model.
5. The emotion recognition method according to claim 4, characterized in that: The residual error of the current round model fitted by the newly added emotion recognition model includes: For each training sample, the residual between the predicted value and the true value of the current round model is calculated to obtain the residual corresponding to each training sample; The decision tree is used to fit the residual corresponding to each training sample so as to minimize the value output by each leaf node on the decision tree.
6. The emotion recognition method according to claim 1, characterized in that: The method further comprises: When the boosted tree is used to perform integrated learning on the multiple emotion recognition models obtained through training, the prediction value of the newly added emotion recognition model in each round of iteration is updated to the prediction value of the model in the current round, and the prediction value of the model obtained after each round of training is obtained; The update step length is obtained by calculating the product of the predicted value of the newly added emotion recognition model and the configuration value in each iteration, and the update amplitude is controlled by using the update step length.
7. The emotion recognition method according to claim 1, characterized in that: The step of processing the recognition result of each tree according to the recognition type to obtain the emotion recognition result of the target user includes: Get the weight of each tree; Determine the predicted value of each tree based on the identification results of each tree; Calculate the weighted sum according to the weight of each tree and the predicted value of each tree to get the target predicted value; When the recognition type is a classification type, obtaining a pre-configured classification threshold, and comparing the target prediction value with the classification threshold to obtain a comparison result, and determining the emotion recognition result of the target user according to the comparison result; or When the recognition type is a regression type, the target prediction value is determined as the emotion recognition result of the target user.
8. An emotion recognition device, characterized in that: The emotion recognition device comprises: A determination unit, used to determine a speech determination anchor point and a plurality of prediction factors; A training unit, used for constructing a training set corresponding to each predictor, and using the training set corresponding to each predictor to train the emotion recognition model respectively; A learning unit, used for performing integrated learning on multiple emotion recognition models obtained through training using a boosted tree to obtain a target emotion recognition model; A collection unit, configured to respond to an emotion recognition instruction of a target user and collect voice information of the target user as a voice to be recognized according to the voice determination anchor point; An input unit, used to input the speech to be recognized into the target emotion recognition model, and obtain the recognition result of each tree in the target emotion recognition model; The determination unit is further configured to determine a recognition type according to the emotion recognition instruction; A processing unit is used to process the recognition result of each tree according to the recognition type to obtain the emotion recognition result of the target user.
9. A computer device, characterized in that: The computer device comprises: a memory storing at least one instruction; and A processor executes instructions stored in the memory to implement the emotion recognition method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, and the at least one instruction is executed by a processor in a computer device to implement the emotion recognition method as described in any one of claims 1 to 7.