Emotion recognition method and apparatus, and device and medium

WO2026194260A1PCT designated stage Publication Date: 2026-09-24PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/135634
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-20
Filing Date
2025-11-18
Publication Date
2026-09-24

Smart Images

  • Figure CN2025135634_24092026_PF_FP_ABST
    Figure CN2025135634_24092026_PF_FP_ABST
Patent Text Reader

Abstract

An emotion recognition method and apparatus, and a device and a medium. The method comprises: determining a voice determination anchor point and determining a plurality of prediction factors (S10); constructing a training set corresponding to each prediction factor, and using the training set corresponding to each prediction factor to train an emotion recognition model (S11); using a boosting tree to perform ensemble learning on a plurality of emotion recognition models obtained by means of training, so as to obtain a target emotion recognition model (S12); in response to an emotion recognition instruction for a target user, and according to the voice determination anchor point, collecting voice information of the target user as voice to be recognized (S13); inputting said voice into the target emotion recognition model, and acquiring a recognition result of each tree in the target emotion recognition model (S14); on the basis of the emotion recognition instruction, determining a recognition type (S15); and on the basis of the recognition type, processing the recognition result of each tree, so as to obtain an emotion recognition result for the target user (S16).
Need to check novelty before this filing date? Find Prior Art

Description

Emotion recognition methods, devices, equipment and media

[0001] This application claims priority to Chinese Patent Application No. 202510344420.X, filed on March 20, 2025, entitled “Emotion Recognition Method, Apparatus, Device and Medium”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the fields of artificial intelligence and medical health technology, and in particular to an emotion recognition method, device, equipment and medium. Background Technology

[0003] In recent years, with the continuous development of artificial intelligence, AI technology has been widely applied in various fields such as finance and healthcare. One important challenge is how to recognize user emotions based on their voice. For example, in the healthcare field, effectively recognizing customer emotions to better serve them is key to intelligent medical customer service.

[0004] In traditional technologies, the following methods are mainly used for emotion recognition based on user voice:

[0005] (1) Real-time emotion recognition of sound signals is performed using real-time signal processing technology and deep learning models. For example, features are extracted from sound signals and emotion classification is performed using convolutional neural networks (CNN) and recurrent neural networks (RNN).

[0006] This method uses a deep learning approach, which is an end-to-end learning approach. However, it does not employ effective feature engineering, resulting in a huge amount of data required for training, poor robustness, and difficulty in improving accuracy, leading to poor practical results.

[0007] (2) Combine sound, text and visual information for emotion recognition and use deep learning models for multimodal feature fusion.

[0008] While this method seems promising, the cost of acquiring and annotating the corpus is enormous, and many scenarios (such as call centers) lack visual information, thus limiting its application.

[0009] (3) Extract features such as pitch, intensity, and speech rate of the sound signal and classify them using a support vector machine (SVM).

[0010] This method belongs to traditional machine learning, which is usually binary classification. It does not fully reflect the emotions reflected in the process changes, so it is not only ineffective, but also has poor applicability to various scenarios.

[0011] Based on the above methods, the inventors realized that existing speech recognition emotion solutions do not fully consider the emotional changes represented by temporal variations. For example, some people naturally speak quickly and loudly, which cannot be identified as anger or impatience. Furthermore, relying on a single method leads to a bottleneck in accuracy, making further improvement difficult. The GAP (Global Average Processing Time) of accuracy makes it difficult to satisfy business users in practical applications, thus hindering its actual deployment and profitability. Summary of the Invention

[0012] In view of the above, it is necessary to provide an emotion recognition method, device, equipment, and medium to solve the problem of low accuracy in emotion recognition.

[0013] Firstly, this application provides an emotion recognition method, the emotion recognition method comprising:

[0014] Determine the speech judgment anchor point and identify multiple prediction factors;

[0015] Construct a training set corresponding to each predictor, and train the emotion recognition model separately using the training set corresponding to each predictor.

[0016] By using boosting trees to ensemble and learn multiple trained emotion recognition models, a target emotion recognition model is obtained.

[0017] In response to the emotion recognition instruction of the target user, the voice information of the target user is collected as the voice to be recognized according to the voice determination anchor point;

[0018] The speech to be recognized is input into the target emotion recognition model, and the recognition result of each tree in the target emotion recognition model is obtained;

[0019] The recognition type is determined according to the emotion recognition instruction;

[0020] The recognition results of each tree are processed according to the recognition type to obtain the emotion recognition result of the target user.

[0021] Secondly, this application also provides an emotion recognition device, the emotion recognition device comprising:

[0022] The determination unit is used to determine the speech determination anchor point and to determine multiple prediction factors;

[0023] The training unit is used to construct a training set corresponding to each predictor and to train the emotion recognition model using the training set corresponding to each predictor.

[0024] The learning unit is used to integrate and learn multiple emotion recognition models trained by boosting trees to obtain the target emotion recognition model.

[0025] The acquisition unit is used to collect the voice information of the target user as the voice to be recognized in response to the emotion recognition instruction of the target user;

[0026] The input unit is used to input the speech to be recognized into the target emotion recognition model and obtain the recognition result of each tree in the target emotion recognition model;

[0027] The determining unit is further configured to determine the recognition type based on the emotion recognition instruction;

[0028] The processing unit is used to process the recognition results of each tree according to the recognition type to obtain the emotion recognition result of the target user.

[0029] Thirdly, this application also provides a computer device, the computer device comprising:

[0030] Memory, storing at least one instruction; and

[0031] The processor executes instructions stored in the memory to perform the following steps:

[0032] Determine the speech judgment anchor point and identify multiple prediction factors;

[0033] Construct a training set corresponding to each predictor, and train the emotion recognition model separately using the training set corresponding to each predictor.

[0034] By using boosting trees to ensemble and learn multiple trained emotion recognition models, a target emotion recognition model is obtained.

[0035] In response to the emotion recognition instruction of the target user, the voice information of the target user is collected as the voice to be recognized according to the voice determination anchor point;

[0036] The speech to be recognized is input into the target emotion recognition model, and the recognition result of each tree in the target emotion recognition model is obtained;

[0037] The recognition type is determined according to the emotion recognition instruction;

[0038] The recognition results of each tree are processed according to the recognition type to obtain the emotion recognition result of the target user.

[0039] Fourthly, this application also provides a non-volatile computer-readable storage medium storing at least one instruction, which is executed by a processor in a computer device to perform the following steps:

[0040] Determine the speech judgment anchor point and identify multiple prediction factors;

[0041] Construct a training set corresponding to each predictor, and train the emotion recognition model separately using the training set corresponding to each predictor.

[0042] By using boosting trees to ensemble and learn multiple trained emotion recognition models, a target emotion recognition model is obtained.

[0043] In response to the emotion recognition instruction of the target user, the voice information of the target user is collected as the voice to be recognized according to the voice determination anchor point;

[0044] The speech to be recognized is input into the target emotion recognition model, and the recognition result of each tree in the target emotion recognition model is obtained;

[0045] The recognition type is determined according to the emotion recognition instruction;

[0046] The recognition results of each tree are processed according to the recognition type to obtain the emotion recognition result of the target user.

[0047] As can be seen from the above technical solutions, this application can train emotion recognition models separately using the training set corresponding to each predictor factor, and use boosting trees to integrate and learn multiple trained emotion recognition models to obtain the target emotion recognition model. This avoids the problems of poor robustness of a single model and inability to fully consider individual differences, and realizes a large-scale voice emotion judgment model with multi-dimensional information, boosting and enhancement of multiple models, and comprehensive use of voice changes and semantic information. By configuring speech judgment anchors, speech judgments that are too short at the word level can be avoided, thereby significantly reducing judgment noise and improving the accuracy of emotion recognition. Attached Figure Description

[0048] Figure 1 is a flowchart of a preferred embodiment of the emotion recognition method of this application.

[0049] Figure 2 is a functional block diagram of a preferred embodiment of the emotion recognition device of this application.

[0050] Figure 3 is a schematic diagram of the structure of a computer device implementing the emotion recognition method of this application. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of this application clearer, the application will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0052] Figure 1 shows a flowchart of a preferred embodiment of the emotion recognition method of this application. Depending on different needs, the order of the steps in this flowchart can be changed, and some steps can be omitted.

[0053] The emotion recognition method is applied to one or more computer devices. The computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0054] The computer device can be any electronic product that can interact with the user, such as a personal computer, tablet computer, smartphone, personal digital assistant (PDA), game console, interactive network television (IPTV), smart wearable device, etc.

[0055] The computer equipment may also include network equipment and / or user equipment. The network equipment includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.

[0056] The server can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0057] Artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0058] Foundational artificial intelligence technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0059] The network in which the computer device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, and virtual private network (VPN).

[0060] S10, determine the speech determination anchor point and determine multiple prediction factors.

[0061] The predictive factors may include factors that play a key role in emotion recognition.

[0062] In this embodiment, determining the speech determination anchor point and determining multiple prediction factors includes:

[0063] Obtain historical emotion recognition data;

[0064] Based on the historical emotion recognition data, an effectiveness analysis is performed to obtain the effective speech duration and effective word count thresholds as the speech judgment anchor points.

[0065] Feature extraction is performed on the historical emotion recognition data to obtain multiple features associated with emotion recognition;

[0066] An importance analysis was performed on the aforementioned features to obtain the analysis results;

[0067] Based on the analysis results, the multiple predictive factors are selected from the multiple features.

[0068] The effective speech duration and the effective word count threshold can be segmented using VAD (Voice Activity Detection) and ASR (Automatic Speech Recognition) technologies to obtain the optimal speech duration and word count threshold. For example, the effective speech duration can be configured to 8 seconds, and the word count threshold can be configured to 16 words. That is, only when the speech duration is greater than or equal to 8 seconds and the input word count is greater than or equal to 16 words can the input speech be used for emotion recognition. This fully considers the characteristics of emotion and the relationship between speech sentences, avoiding unnecessary consumption of judgment resources and judgment noise caused by excessively short speech. In principle, human emotions are unlikely to change abruptly at the word level, thus avoiding judgment of excessively short speech at the word level and significantly reducing judgment noise.

[0069] In this process, methods such as random forests can be used to perform importance analysis on the multiple features, and the features with higher importance can be obtained as the multiple prediction factors.

[0070] For example, the multiple predictive factors may include, but are not limited to, one or more of the following factors:

[0071] (1) Pitch. Characteristics: Pitch reflects the frequency of vocal cord vibration and changes with emotional fluctuations. Implementation: By extracting the fundamental frequency (F0), its range of variation and fluctuation patterns are analyzed. For example, when angry, the pitch usually rises and fluctuates greatly, while when sad, it falls and becomes stable.

[0072] Specific manifestations: 1) Increased pitch. Characteristics: Increased fundamental frequency (F0), higher sound frequency; Emotions: anger, excitement, joy, anxiety. 2) Decreased pitch. Characteristics: Decreased fundamental frequency (F0), lower sound frequency; Emotions: sadness, frustration, fatigue, calmness. 3) Fluctuating pitch. Characteristics: Unstable fundamental frequency (F0), sometimes high, sometimes low; Emotions: anxiety, tension, complex emotions.

[0073] (2) Intensity. Characteristics: Intensity is related to the amplitude of the sound; emotional changes affect the volume of speech. Implementation: Calculate the root mean square (RMS) energy of the sound signal and analyze its changes. For example, the volume usually increases when angry and decreases when sad.

[0074] Specific manifestations: 1) Increased sound intensity. Characteristics: Increased amplitude of sound, significantly higher volume; Emotions: anger, excitement, joy, anxiety. 2) Decreased sound intensity. Characteristics: Decreased amplitude of sound, significantly lower volume; Emotions: sadness, frustration, fatigue, calmness. 3) Fluctuating sound intensity. Characteristics: Unstable volume, sometimes loud, sometimes soft; Emotions: anxiety, tension, complex emotions.

[0075] (3) Speech Rate. Characteristics: Speech rate refers to the number of syllables pronounced per unit of time. Emotional changes affect speaking speed. Implementation: Calculates the pronunciation speed of syllables or words using speech recognition technology. For example, speech speed increases when anxious and decreases when relaxed.

[0076] Specific manifestations: 1) Increased speech rate. Characteristics: Shorter intervals between syllables and words, resulting in a significantly faster overall speech rate; Emotions: Anxiety, tension, excitement, anger. 2) Decreased speech rate. Characteristics: Longer intervals between syllables and words, resulting in a significantly slower overall speech rate; Emotions: Sadness, frustration, fatigue, contemplation. 3) Fluctuating speech rate. Characteristics: Unstable speech rate, sometimes fast, sometimes slow; Emotions: Hesitation, uncertainty, complex emotions.

[0077] (4) Spectral Features. Features: Spectral features, such as formant frequencies and bandwidth, reflect changes in the shape of the vocal tract. Implementation: Fourier transform or Linear Predictive Coding (LPC) is used to extract spectral features. The position and distribution of formants will differ under different emotions.

[0078] Specific manifestations: 1) Anger. Formant location: When angry, vocal cord tension increases and the shape of the vocal tract changes, usually leading to an increase in the frequency of the first formant (F1) and the second formant (F2). Formant distribution: The distribution of formants may be more dispersed and fluctuate significantly, reflecting emotional excitement and instability. 2) Excitement and joy. Formant location: When excited and joyful, vocal cord tension increases and the shape of the vocal tract changes, usually leading to an increase in the frequency of the first formant (F1) and the second formant (F2). Formant distribution: The distribution of formants is more concentrated and fluctuates more significantly, reflecting active and positive emotions. 3) Anxiety and tension. Formant location: When anxious and tense, vocal cord tension increases and the shape of the vocal tract changes, usually leading to an increase in the frequency of the first formant (F1) and the second formant (F2), but unstable fluctuations may also occur. Formant distribution: The distribution of formants may be unstable, sometimes high and sometimes low, reflecting emotional fluctuations and uncertainty. 4) Sadness and frustration. Formant Position: When sad and depressed, the vocal cords relax, and the shape of the vocal tract changes, usually resulting in a decrease in the frequency of the first formant (F1) and the second formant (F2). Formant Distribution: The formants are relatively concentrated and fluctuate little, reflecting a low and stable mood. 5) Calm and Relaxed. Formant Position: When calm and relaxed, the vocal cords and vocal tract are in a natural state, and the formant positions are usually relatively stable and moderate. Formant Distribution: The formants are relatively stable and fluctuate little, reflecting a stable and relaxed mood.

[0079] (5) Voice Quality. Characteristics: Voice quality involves the smoothness and roughness of the sound, and emotional changes can affect the vibration pattern of the vocal cords. Implementation: Voice quality is assessed by analyzing the harmonic-to-noise ratio (HNR) or jitter and shimmer perturbations. A tense voice may be rougher, while a relaxed voice is smoother.

[0080] (6) Pauses and Rhythm. Characteristics: The frequency and duration of pauses, as well as the rhythm of speech, can reflect emotions. Implementation: Detect silent segments in speech and analyze the frequency and duration of pauses. Pauses may decrease during anxiety and increase during sadness.

[0081] The specific relationship between emotions and pauses / rhythms: 1) Anger: Short and frequent pauses, fast speech, and high volume. May be accompanied by sudden pauses or rhythm changes. 2) Sadness: Longer and frequent pauses, slow speech, and lower volume. The rhythm may seem sluggish or disjointed. 3) Anxiety: Frequent and irregular pauses, speech speed may fluctuate. The rhythm is unstable and may be accompanied by repetition or stuttering. 4) Excitement: Fewer pauses, fast speech, and smooth rhythm. High volume and significant tone variation. 5) Relaxation: Natural pauses, moderate speech speed, and steady rhythm. Moderate volume and gentle tone.

[0082] (7) Text information. Implementation: Combine the text information recognized by ASR (Automatic Speech Recognition) with LLM (Large Language Model) for emotion determination, thereby providing the emotion type determined solely from the text.

[0083] S11, construct a training set corresponding to each predictor, and use the training set corresponding to each predictor to train the emotion recognition model respectively.

[0084] In this embodiment, constructing a training set corresponding to each predictor and training the emotion recognition model using the training set corresponding to each predictor includes:

[0085] Collect user voice data based on each predictor factor;

[0086] The collected user speech is labeled, and the labeled user speech is used to construct a training set corresponding to each prediction factor;

[0087] By training a model with predictive function using the training set corresponding to each predictor, an emotion recognition model corresponding to each predictor is obtained.

[0088] This allows for the labeling of predictor values ​​and corresponding emotion types in user speech.

[0089] The predictive model may include, but is not limited to, decision trees, large language models, etc.

[0090] Through the above embodiments, it is possible to analyze sound features such as pitch, intensity, speech rate, spectral characteristics, tone quality, pauses, and rhythm, and train corresponding models accordingly. By combining text information and the judgment results of the LLM large model, it is possible to effectively detect the speaker's emotional changes.

[0091] S12, using a boosting tree to integrate and learn multiple trained emotion recognition models to obtain the target emotion recognition model.

[0092] In this embodiment, the step of ensemble learning of multiple trained emotion recognition models using boosting trees to obtain the target emotion recognition model includes:

[0093] Build the initial model;

[0094] Using the initial model as the first-round model, in each round of training, an emotion recognition model is added one by one based on the current model, and iterative training is performed;

[0095] When all emotion recognition models have completed the iteration, training is stopped, and the currently obtained model is determined as the target emotion recognition model.

[0096] In each iteration, the residual of the current model is fitted using a newly added emotion recognition model.

[0097] The initial model can be a simple model such as a constant model. For example, suppose there are 5 samples in the training data, with corresponding emotion intensity values ​​of -0.2, 0.1, 0.3, -0.1, and 0.2. First, calculate the sum of these values: (-0.2 + 0.1 + 0.3 - 0.1 + 0.2) = 0.3, then divide by the sample size of 5 to get the mean: 0.3 ÷ 5 = 0.06. The constant model would then predict an emotion intensity of 0.06 for all new input data, meaning it considers the data to be in a relatively neutral but slightly positive emotional state.

[0098] Of course, the results of most types of emotion recognition, the median, etc., can also be used as the initial prediction values ​​of the initial model.

[0099] The residuals of fitting the current model using the newly added emotion recognition model include:

[0100] For each training sample, the residual between the predicted value and the true value of the model in that round is calculated to obtain the residual corresponding to each training sample;

[0101] The residual corresponding to each training sample is fitted using a decision tree to minimize the output value of each leaf node on the decision tree.

[0102] The residual represents the difference between the actual value and the current model's prediction. For each sample, the residual between the current model's prediction and the actual value is calculated, representing the portion that the current model fails to explain.

[0103] In the process of integrating multiple emotion recognition models trained by the boosting tree, the prediction value of the newly added emotion recognition model in each iteration is updated to the prediction value of the model in the current iteration, so as to obtain the prediction value of the model after each round of training.

[0104] Specifically, the update step size is obtained by multiplying the predicted value of the newly added emotion recognition model with the configuration value in each iteration, and the update step size is used to control the update amplitude to prevent overfitting.

[0105] This can also be configured with a certain number of iterations, at which point training stops. Alternatively, the overall performance of the model can be monitored, and training can stop when the overall performance no longer improves.

[0106] Among them, boosting trees can significantly improve prediction accuracy by combining multiple weak models. They are suitable for regression, classification and ranking problems, and have a certain degree of robustness to missing values ​​and outliers.

[0107] In the above embodiments, multiple models are combined to boost trees for ensemble learning. The core idea of ​​boost trees is to improve the performance of the overall model by continuously correcting the errors of the previous model, so as to achieve high accuracy of the overall model even when the accuracy of a single model is limited.

[0108] S13, in response to the emotion recognition instruction of the target user, collect the voice information of the target user as the voice to be recognized according to the voice determination anchor point.

[0109] In this embodiment, the emotion recognition command can be automatically triggered. For example, in the healthcare field, for smart medical customer service, when a user's voice input is detected, real-time emotion recognition of the customer can be automatically triggered, thereby paying attention to the customer's emotional changes and providing more considerate consultation services and care. For mental health counselors, when conversing with patients, they can monitor the patient's emotional state in real time based on the patient's voice, so as to provide more scientific and timely emotional relief. For example, patients with depression or anxiety often have unique emotional characteristics in their voices. By analyzing the emotional state in the patient's voice through voice emotion recognition technology, doctors can be provided with diagnostic references to assist in the early screening and diagnosis of mental illnesses.

[0110] Of course, within the financial sector, emotion recognition can also be used to better serve customers. For example, in financial customer service scenarios, voice emotion recognition technology can capture real-time changes in customer emotions during speech. When negative emotions such as frustration or anger are detected, the call can be promptly transferred to experienced customer service personnel to prioritize handling customer issues and prevent customer churn. By analyzing emotional information from large amounts of customer service voice data, financial institutions can understand customer satisfaction with products and services, identify key areas of customer dissatisfaction, and optimize product design and service processes. In loan applications or financial transactions, abnormal emotions may be a sign of fraud. Voice emotion recognition technology can serve as an auxiliary tool to analyze customer emotions during loan applications or important transactions. If a customer exhibits excessive tension, anxiety, or unnatural emotions, it may indicate potential risks, prompting financial institutions to conduct more in-depth investigations and reviews. In the sales process of high-risk investment products, voice emotion recognition can also be used to monitor investors' emotional states, assessing whether their risk perception and tolerance match the investment product, preventing impulsive decisions and reducing investment risk.

[0111] In this embodiment, user voice can be collected according to the duration and number of words specified by the voice determination anchor point to improve the usability of the voice and thus improve the accuracy of emotion recognition.

[0112] S14, input the speech to be recognized into the target emotion recognition model, and obtain the recognition result of each tree in the target emotion recognition model.

[0113] Each tree's recognition result can be either a specific numerical value or an emotion type. The model's output will also differ depending on the type of recognition task.

[0114] S15, determine the recognition type according to the emotion recognition instruction.

[0115] In this embodiment, the identification type may include classification type and regression type.

[0116] The classification type refers to the need to determine the specific emotion type.

[0117] The regression type refers to outputting a specific numerical value, which can be used to determine the degree of emotion.

[0118] S16, Process the recognition results of each tree according to the recognition type to obtain the emotion recognition result of the target user.

[0119] In this embodiment, processing the recognition results of each tree according to the recognition type to obtain the emotion recognition result of the target user includes:

[0120] Get the weight of each tree;

[0121] The predicted value for each tree is determined based on the identification results of each tree;

[0122] The target predicted value is obtained by calculating a weighted sum based on the weight of each tree and the predicted value of each tree.

[0123] When the recognition type is a classification type, a pre-configured classification threshold is obtained, and the target predicted value is compared with the classification threshold to obtain a comparison result. Based on the comparison result, the emotion recognition result of the target user is determined; or

[0124] When the recognition type is a regression type, the target predicted value is determined as the emotion recognition result of the target user.

[0125] The weight of each tree is determined by its performance during training.

[0126] For example, in the classification problem, emotions are divided into a finite number of discrete categories, such as fixed categories like "happy," "sad," "angry," and "calm." When training the model, these categories are used as labels, allowing the model to learn the association between different features and each emotion category. During prediction, the model outputs the probability of belonging to each category based on the input features; a probability higher than a certain threshold indicates happiness, and a probability lower than a certain threshold indicates unhappiness.

[0127] For example, in the regression problem, the intensity of emotion can be quantified as a continuous numerical value, such as continuously rating the level of pleasure from -10 (extreme sadness) to 10 (extreme happiness). When training the model, sample data with these continuous emotion intensity values ​​is used, allowing the model to learn how to predict the corresponding numerical value based on the input features. During prediction, the model outputs a specific numerical value representing the identified emotion intensity. For instance, if the predicted emotion intensity value is 3, it indicates a relatively positive emotional state, but not yet reaching the level of extreme happiness.

[0128] In the above embodiments, boosting trees iteratively train multiple decision trees to gradually correct the model's residuals, and finally obtain the decision by weighted summation of the predictions from all trees. This method can effectively improve the model's performance, especially when dealing with complex nonlinear relationships. It is robust, accurate, and truly meets the needs of practical scenarios such as healthcare and finance.

[0129] As can be seen from the above technical solutions, this application can train emotion recognition models separately using the training set corresponding to each predictor factor, and use boosting trees to integrate and learn multiple trained emotion recognition models to obtain the target emotion recognition model. This avoids the problems of poor robustness of a single model and inability to fully consider individual differences, and realizes a large-scale voice emotion judgment model with multi-dimensional information, boosting and enhancement of multiple models, and comprehensive use of voice changes and semantic information. The speech information of the target user is collected as the speech to be recognized according to the speech judgment anchor point. The speech to be recognized is input into the target emotion recognition model, and the recognition results of each tree are processed according to the recognition type to obtain the emotion recognition result of the target user. By configuring the speech judgment anchor point, speech judgment at the word level that is too short can be avoided, thereby greatly reducing judgment noise and improving the accuracy of emotion recognition.

[0130] Figure 2 shows a functional block diagram of a preferred embodiment of the emotion recognition device of this application. The emotion recognition device 11 includes a determining unit 110, a training unit 111, a learning unit 112, a data acquisition unit 113, an input unit 114, and a processing unit 115. The module / unit referred to in this application refers to a series of computer program segments that can be executed by a processor and perform a fixed function, and are stored in memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.

[0131] The determining unit 110 is used to determine the speech determination anchor point and to determine multiple prediction factors.

[0132] The predictive factors may include factors that play a key role in emotion recognition.

[0133] In this embodiment, the determining unit 110 determines the speech determination anchor point and determines multiple prediction factors, including:

[0134] Obtain historical emotion recognition data;

[0135] Based on the historical emotion recognition data, an effectiveness analysis is performed to obtain the effective speech duration and effective word count thresholds as the speech judgment anchor points.

[0136] Feature extraction is performed on the historical emotion recognition data to obtain multiple features associated with emotion recognition;

[0137] An importance analysis was performed on the aforementioned features to obtain the analysis results;

[0138] Based on the analysis results, the multiple predictive factors are selected from the multiple features.

[0139] The effective speech duration and the effective word count threshold can be segmented using VAD (Voice Activity Detection) and ASR (Automatic Speech Recognition) technologies to obtain the optimal speech duration and word count threshold. For example, the effective speech duration can be configured to 8 seconds, and the word count threshold can be configured to 16 words. That is, only when the speech duration is greater than or equal to 8 seconds and the input word count is greater than or equal to 16 words can the input speech be used for emotion recognition. This fully considers the characteristics of emotion and the relationship between speech sentences, avoiding unnecessary consumption of judgment resources and judgment noise caused by excessively short speech. In principle, human emotions are unlikely to change abruptly at the word level, thus avoiding judgment of excessively short speech at the word level and significantly reducing judgment noise.

[0140] In this process, methods such as random forests can be used to perform importance analysis on the multiple features, and the features with higher importance can be obtained as the multiple prediction factors.

[0141] For example, the multiple predictive factors may include, but are not limited to, one or more of the following factors:

[0142] (1) Pitch. Characteristics: Pitch reflects the frequency of vocal cord vibration and changes with emotional fluctuations. Implementation: By extracting the fundamental frequency (F0), its range of variation and fluctuation patterns are analyzed. For example, when angry, the pitch usually rises and fluctuates greatly, while when sad, it falls and becomes stable.

[0143] Specific manifestations: 1) Increased pitch. Characteristics: Increased fundamental frequency (F0), higher sound frequency; Emotions: anger, excitement, joy, anxiety. 2) Decreased pitch. Characteristics: Decreased fundamental frequency (F0), lower sound frequency; Emotions: sadness, frustration, fatigue, calmness. 3) Fluctuating pitch. Characteristics: Unstable fundamental frequency (F0), sometimes high, sometimes low; Emotions: anxiety, tension, complex emotions.

[0144] (2) Intensity. Characteristics: Intensity is related to the amplitude of the sound; emotional changes affect the volume of speech. Implementation: Calculate the root mean square (RMS) energy of the sound signal and analyze its changes. For example, the volume usually increases when angry and decreases when sad.

[0145] Specific manifestations: 1) Increased sound intensity. Characteristics: Increased amplitude of sound, significantly higher volume; Emotions: anger, excitement, joy, anxiety. 2) Decreased sound intensity. Characteristics: Decreased amplitude of sound, significantly lower volume; Emotions: sadness, frustration, fatigue, calmness. 3) Fluctuating sound intensity. Characteristics: Unstable volume, sometimes loud, sometimes soft; Emotions: anxiety, tension, complex emotions.

[0146] (3) Speech Rate. Characteristics: Speech rate refers to the number of syllables pronounced per unit of time. Emotional changes affect speaking speed. Implementation: Calculates the pronunciation speed of syllables or words using speech recognition technology. For example, speech speed increases when anxious and decreases when relaxed.

[0147] Specific manifestations: 1) Increased speech rate. Characteristics: Shorter intervals between syllables and words, resulting in a significantly faster overall speech rate; Emotions: Anxiety, tension, excitement, anger. 2) Decreased speech rate. Characteristics: Longer intervals between syllables and words, resulting in a significantly slower overall speech rate; Emotions: Sadness, frustration, fatigue, contemplation. 3) Fluctuating speech rate. Characteristics: Unstable speech rate, sometimes fast, sometimes slow; Emotions: Hesitation, uncertainty, complex emotions.

[0148] (4) Spectral Features. Features: Spectral features, such as formant frequencies and bandwidth, reflect changes in the shape of the vocal tract. Implementation: Fourier transform or Linear Predictive Coding (LPC) is used to extract spectral features. The position and distribution of formants will differ under different emotions.

[0149] Specific manifestations: 1) Anger. Formant location: When angry, vocal cord tension increases and the shape of the vocal tract changes, usually leading to an increase in the frequency of the first formant (F1) and the second formant (F2). Formant distribution: The distribution of formants may be more dispersed and fluctuate significantly, reflecting emotional excitement and instability. 2) Excitement and joy. Formant location: When excited and joyful, vocal cord tension increases and the shape of the vocal tract changes, usually leading to an increase in the frequency of the first formant (F1) and the second formant (F2). Formant distribution: The distribution of formants is more concentrated and fluctuates more significantly, reflecting active and positive emotions. 3) Anxiety and tension. Formant location: When anxious and tense, vocal cord tension increases and the shape of the vocal tract changes, usually leading to an increase in the frequency of the first formant (F1) and the second formant (F2), but unstable fluctuations may also occur. Formant distribution: The distribution of formants may be unstable, sometimes high and sometimes low, reflecting emotional fluctuations and uncertainty. 4) Sadness and frustration. Formant Position: When sad and depressed, the vocal cords relax, and the shape of the vocal tract changes, usually resulting in a decrease in the frequency of the first formant (F1) and the second formant (F2). Formant Distribution: The formants are relatively concentrated and fluctuate little, reflecting a low and stable mood. 5) Calm and Relaxed. Formant Position: When calm and relaxed, the vocal cords and vocal tract are in a natural state, and the formant positions are usually relatively stable and moderate. Formant Distribution: The formants are relatively stable and fluctuate little, reflecting a stable and relaxed mood.

[0150] (5) Voice Quality. Characteristics: Voice quality involves the smoothness and roughness of the sound, and emotional changes can affect the vibration pattern of the vocal cords. Implementation: Voice quality is assessed by analyzing the harmonic-to-noise ratio (HNR) or jitter and shimmer perturbations. A tense voice may be rougher, while a relaxed voice is smoother.

[0151] (6) Pauses and Rhythm. Characteristics: The frequency and duration of pauses, as well as the rhythm of speech, can reflect emotions. Implementation: Detect silent segments in speech and analyze the frequency and duration of pauses. Pauses may decrease during anxiety and increase during sadness.

[0152] The specific relationship between emotions and pauses / rhythms: 1) Anger: Short and frequent pauses, fast speech, and high volume. May be accompanied by sudden pauses or rhythm changes. 2) Sadness: Longer and frequent pauses, slow speech, and lower volume. The rhythm may seem sluggish or disjointed. 3) Anxiety: Frequent and irregular pauses, speech speed may fluctuate. The rhythm is unstable and may be accompanied by repetition or stuttering. 4) Excitement: Fewer pauses, fast speech, and smooth rhythm. High volume and significant tone variation. 5) Relaxation: Natural pauses, moderate speech speed, and steady rhythm. Moderate volume and gentle tone.

[0153] (7) Text information. Implementation: Combine the text information recognized by ASR (Automatic Speech Recognition) with LLM (Large Language Model) for emotion determination, thereby providing the emotion type determined solely from the text.

[0154] The training unit 111 is used to construct a training set corresponding to each predictor and to train the emotion recognition model using the training set corresponding to each predictor.

[0155] In this embodiment, the training unit 111 constructs a training set corresponding to each predictor factor, and trains the emotion recognition model using the training set corresponding to each predictor factor, including:

[0156] Collect user voice data based on each predictor factor;

[0157] The collected user speech is labeled, and the labeled user speech is used to construct a training set corresponding to each prediction factor;

[0158] By training a model with predictive function using the training set corresponding to each predictor, an emotion recognition model corresponding to each predictor is obtained.

[0159] This allows for the labeling of predictor values ​​and corresponding emotion types in user speech.

[0160] The predictive model may include, but is not limited to, decision trees, large language models, etc.

[0161] Through the above embodiments, it is possible to analyze sound features such as pitch, intensity, speech rate, spectral characteristics, tone quality, pauses, and rhythm, and train corresponding models accordingly. By combining text information and the judgment results of the LLM large model, it is possible to effectively detect the speaker's emotional changes.

[0162] The learning unit 112 is used to integrate and learn multiple trained emotion recognition models using a boosting tree to obtain a target emotion recognition model.

[0163] In this embodiment, the learning unit 112 uses boosting trees to perform ensemble learning on multiple trained emotion recognition models to obtain a target emotion recognition model, including:

[0164] Build the initial model;

[0165] Using the initial model as the first-round model, in each round of training, an emotion recognition model is added one by one based on the current model, and iterative training is performed;

[0166] When all emotion recognition models have completed the iteration, training is stopped, and the currently obtained model is determined as the target emotion recognition model.

[0167] In each iteration, the residual of the current model is fitted using a newly added emotion recognition model.

[0168] The initial model can be a simple model such as a constant model. For example, suppose there are 5 samples in the training data, with corresponding emotion intensity values ​​of -0.2, 0.1, 0.3, -0.1, and 0.2. First, calculate the sum of these values: (-0.2 + 0.1 + 0.3 - 0.1 + 0.2) = 0.3, then divide by the sample size of 5 to get the mean: 0.3 ÷ 5 = 0.06. The constant model would then predict an emotion intensity of 0.06 for all new input data, meaning it considers the data to be in a relatively neutral but slightly positive emotional state.

[0169] Of course, the results of most types of emotion recognition, the median, etc., can also be used as the initial prediction values ​​of the initial model.

[0170] The residuals of fitting the current model using the newly added emotion recognition model include:

[0171] For each training sample, the residual between the predicted value and the true value of the model in that round is calculated to obtain the residual corresponding to each training sample;

[0172] The residual corresponding to each training sample is fitted using a decision tree to minimize the output value of each leaf node on the decision tree.

[0173] The residual represents the difference between the actual value and the current model's prediction. For each sample, the residual between the current model's prediction and the actual value is calculated, representing the portion that the current model fails to explain.

[0174] In the process of integrating multiple emotion recognition models trained by the boosting tree, the prediction value of the newly added emotion recognition model in each iteration is updated to the prediction value of the model in the current iteration, so as to obtain the prediction value of the model after each round of training.

[0175] Specifically, the update step size is obtained by multiplying the predicted value of the newly added emotion recognition model with the configuration value in each iteration, and the update step size is used to control the update amplitude to prevent overfitting.

[0176] This can also be configured with a certain number of iterations, at which point training stops. Alternatively, the overall performance of the model can be monitored, and training can stop when the overall performance no longer improves.

[0177] Among them, boosting trees can significantly improve prediction accuracy by combining multiple weak models. They are suitable for regression, classification and ranking problems, and have a certain degree of robustness to missing values ​​and outliers.

[0178] In the above embodiments, multiple models are combined to boost trees for ensemble learning. The core idea of ​​boost trees is to improve the performance of the overall model by continuously correcting the errors of the previous model, so as to achieve high accuracy of the overall model even when the accuracy of a single model is limited.

[0179] The acquisition unit 113 is used to collect the voice information of the target user as the voice to be recognized in response to the emotion recognition instruction of the target user.

[0180] In this embodiment, the emotion recognition command can be automatically triggered. For example, in the healthcare field, for smart medical customer service, when a user's voice input is detected, real-time emotion recognition of the customer can be automatically triggered, thereby paying attention to the customer's emotional changes and providing more considerate consultation services and care. For mental health counselors, when conversing with patients, they can monitor the patient's emotional state in real time based on the patient's voice, so as to provide more scientific and timely emotional relief. For example, patients with depression or anxiety often have unique emotional characteristics in their voices. By analyzing the emotional state in the patient's voice through voice emotion recognition technology, doctors can be provided with diagnostic references to assist in the early screening and diagnosis of mental illnesses.

[0181] Of course, within the financial sector, emotion recognition can also be used to better serve customers. For example, in financial customer service scenarios, voice emotion recognition technology can capture real-time changes in customer emotions during speech. When negative emotions such as frustration or anger are detected, the call can be promptly transferred to experienced customer service personnel to prioritize handling customer issues and prevent customer churn. By analyzing emotional information from large amounts of customer service voice data, financial institutions can understand customer satisfaction with products and services, identify key areas of customer dissatisfaction, and optimize product design and service processes. In loan applications or financial transactions, abnormal emotions may be a sign of fraud. Voice emotion recognition technology can serve as an auxiliary tool to analyze customer emotions during loan applications or important transactions. If a customer exhibits excessive tension, anxiety, or unnatural emotions, it may indicate potential risks, prompting financial institutions to conduct more in-depth investigations and reviews. In the sales process of high-risk investment products, voice emotion recognition can also be used to monitor investors' emotional states, assessing whether their risk perception and tolerance match the investment product, preventing impulsive decisions and reducing investment risk.

[0182] In this embodiment, user voice can be collected according to the duration and number of words specified by the voice determination anchor point to improve the usability of the voice and thus improve the accuracy of emotion recognition.

[0183] The input unit 114 is used to input the speech to be recognized into the target emotion recognition model and obtain the recognition result of each tree in the target emotion recognition model.

[0184] Each tree's recognition result can be either a specific numerical value or an emotion type. The model's output will also differ depending on the type of recognition task.

[0185] The determining unit 110 is further configured to determine the recognition type based on the emotion recognition instruction.

[0186] In this embodiment, the identification type may include classification type and regression type.

[0187] The classification type refers to the need to determine the specific emotion type.

[0188] The regression type refers to outputting a specific numerical value, which can be used to determine the degree of emotion.

[0189] The processing unit 115 is used to process the recognition results of each tree according to the recognition type to obtain the emotion recognition result of the target user.

[0190] In this embodiment, the processing unit 115 processes the recognition results of each tree according to the recognition type to obtain the emotion recognition result of the target user, including:

[0191] Get the weight of each tree;

[0192] The predicted value for each tree is determined based on the identification results of each tree;

[0193] The target predicted value is obtained by calculating a weighted sum based on the weight of each tree and the predicted value of each tree.

[0194] When the recognition type is a classification type, a pre-configured classification threshold is obtained, and the target predicted value is compared with the classification threshold to obtain a comparison result. Based on the comparison result, the emotion recognition result of the target user is determined; or

[0195] When the recognition type is a regression type, the target predicted value is determined as the emotion recognition result of the target user.

[0196] The weight of each tree is determined by its performance during training.

[0197] For example, in the classification problem, emotions are divided into a finite number of discrete categories, such as fixed categories like "happy," "sad," "angry," and "calm." When training the model, these categories are used as labels, allowing the model to learn the association between different features and each emotion category. During prediction, the model outputs the probability of belonging to each category based on the input features; a probability higher than a certain threshold indicates happiness, and a probability lower than a certain threshold indicates unhappiness.

[0198] For example, in the regression problem, the intensity of emotion can be quantified as a continuous numerical value, such as continuously rating the level of pleasure from -10 (extreme sadness) to 10 (extreme happiness). When training the model, sample data with these continuous emotion intensity values ​​is used, allowing the model to learn how to predict the corresponding numerical value based on the input features. During prediction, the model outputs a specific numerical value representing the identified emotion intensity. For instance, if the predicted emotion intensity value is 3, it indicates a relatively positive emotional state, but not yet reaching the level of extreme happiness.

[0199] In the above embodiments, boosting trees iteratively train multiple decision trees to gradually correct the model's residuals, and finally obtain the decision by weighted summation of the predictions from all trees. This method can effectively improve the model's performance, especially when dealing with complex nonlinear relationships. It is robust, accurate, and truly meets the needs of practical scenarios such as healthcare and finance.

[0200] As can be seen from the above technical solutions, this application can train emotion recognition models separately using the training set corresponding to each predictor factor, and use boosting trees to integrate and learn multiple trained emotion recognition models to obtain the target emotion recognition model. This avoids the problems of poor robustness of a single model and inability to fully consider individual differences, and realizes a large-scale voice emotion judgment model with multi-dimensional information, boosting and enhancement of multiple models, and comprehensive use of voice changes and semantic information. The speech information of the target user is collected as the speech to be recognized according to the speech judgment anchor point. The speech to be recognized is input into the target emotion recognition model, and the recognition results of each tree are processed according to the recognition type to obtain the emotion recognition result of the target user. By configuring the speech judgment anchor point, speech judgment at the word level that is too short can be avoided, thereby greatly reducing judgment noise and improving the accuracy of emotion recognition.

[0201] Figure 3 shows a schematic diagram of the structure of a computer device that implements the emotion recognition method of this application.

[0202] The computer device 1 may include a memory 12, a processor 13, and a bus (the arrow in the figure represents the bus), and may also include a computer program, such as an emotion recognition program, stored in the memory 12 and executable on the processor 13.

[0203] Those skilled in the art will understand that the schematic diagram is merely an example of computer device 1 and does not constitute a limitation on computer device 1. Computer device 1 can be either a bus topology or a star topology. Computer device 1 may also include more or fewer other hardware or software than shown in the diagram, or different component arrangements. For example, computer device 1 may also include input / output devices, network access devices, etc.

[0204] It should be noted that the computer device 1 described is merely an example. Other existing or future electronic products that are suitable for this application should also be included within the scope of protection of this application and are incorporated herein by reference.

[0205] The memory 12 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 12 can be an internal storage unit of the computer device 1, such as a portable hard drive of the computer device 1. In other embodiments, the memory 12 can be an external storage device of the computer device 1, such as a plug-in portable hard drive, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 1. Furthermore, the memory 12 can include both internal and external storage units of the computer device 1. The memory 12 can be used not only to store application software and various types of data installed on the computer device 1, such as the code of an emotion recognition program, but also to temporarily store data that has been output or will be output.

[0206] In some embodiments, the processor 13 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control unit of the computer device 1, connecting various components of the computer device 1 via various interfaces and lines. It performs various functions of the computer device 1 and processes data by running or executing programs or modules stored in the memory 12 (e.g., executing emotion recognition programs) and accessing data stored in the memory 12.

[0207] The processor 13 executes the operating system of the computer device 1 and various installed applications. The processor 13 executes the applications to implement the steps in the various emotion recognition method embodiments described above, such as the steps shown in FIG1.

[0208] For example, the computer program may be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to complete this application. The one or more modules / units may be a series of computer-readable instruction segments capable of performing a specific function, which describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into a determining unit 110, a training unit 111, a learning unit 112, an acquisition unit 113, an input unit 114, and a processing unit 115.

[0209] The integrated unit implemented as a software functional module described above can be stored in a computer-readable storage medium. This software functional module, stored in a storage medium, includes several instructions to cause a computer device (which may be a personal computer, computer equipment, or network device, etc.) or processor to execute portions of the emotion recognition methods described in the various embodiments of this application.

[0210] If the modules / units integrated in the computer device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware devices. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above.

[0211] The computer program includes computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory, etc.

[0212] Furthermore, the non-volatile computer-readable storage medium may primarily include a stored program area and a stored data area, wherein the stored program area may store the operating system, an application program required for at least one function, etc.; and the stored data area may store data created based on the use of blockchain nodes, etc.

[0213] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.

[0214] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only a single straight line is used in Figure 3, but this does not mean that there is only one bus or one type of bus. The bus is configured to implement communication between the memory 12 and at least one processor 13, etc.

[0215] Although not shown, the computer device 1 may also include a power supply (such as a battery) to power various components. Preferably, the power supply can be logically connected to the at least one processor 13 through a power management device, thereby enabling functions such as charging management, discharging management, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The computer device 1 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.

[0216] Furthermore, the computer device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, a Bluetooth interface, etc.), which is typically used to establish a communication connection between the computer device 1 and other computer devices.

[0217] Optionally, the computer device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the computer device 1 and to display a visual user interface.

[0218] It should be understood that the embodiments described are for illustrative purposes only and are not limited to this structure in the scope of the patent application.

[0219] Those skilled in the art will understand that the structure shown in FIG3 does not constitute a limitation on the computer device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0220] Referring to Figure 1, the memory 12 in the computer device 1 stores multiple instructions to implement an emotion recognition method, and the processor 13 can execute the multiple instructions to achieve the following:

[0221] Determine the speech judgment anchor point and identify multiple prediction factors;

[0222] Construct a training set corresponding to each predictor, and train the emotion recognition model separately using the training set corresponding to each predictor.

[0223] By using boosting trees to ensemble and learn multiple trained emotion recognition models, a target emotion recognition model is obtained.

[0224] In response to the emotion recognition instruction of the target user, the voice information of the target user is collected as the voice to be recognized according to the voice determination anchor point;

[0225] The speech to be recognized is input into the target emotion recognition model, and the recognition result of each tree in the target emotion recognition model is obtained;

[0226] The recognition type is determined according to the emotion recognition instruction;

[0227] The recognition results of each tree are processed according to the recognition type to obtain the emotion recognition result of the target user.

[0228] Specifically, the specific implementation method of the processor 13 for the above instructions can be referred to the description of the relevant steps in the embodiment corresponding to Figure 1, which will not be repeated here.

[0229] It should be noted that all data involved in this case was legally obtained. Software tools or components not belonging to this company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.

[0230] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.

[0231] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0232] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0233] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.

[0234] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application.

[0235] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of this application is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within this application. No appended diagram markings in the claims should be construed as limiting the scope of the claims.

[0236] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices described in this application may also be implemented by a single unit or device through software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any specific order.

[0237] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit it. Although this application has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of this application without departing from the spirit and scope of the technical solutions of this application.

Claims

1. An emotion recognition method, wherein, The emotion recognition method includes: Determine the speech judgment anchor point and identify multiple prediction factors; Construct a training set corresponding to each predictor, and train the emotion recognition model separately using the training set corresponding to each predictor. By using boosting trees to ensemble and learn multiple trained emotion recognition models, a target emotion recognition model is obtained. In response to the emotion recognition instruction of the target user, the voice information of the target user is collected as the voice to be recognized according to the voice determination anchor point; The speech to be recognized is input into the target emotion recognition model, and the recognition result of each tree in the target emotion recognition model is obtained; The recognition type is determined according to the emotion recognition instruction; The recognition results of each tree are processed according to the recognition type to obtain the emotion recognition result of the target user.

2. The emotion recognition method as described in claim 1, wherein, The determination of the speech determination anchor point and the determination of multiple prediction factors include: Obtain historical emotion recognition data; Based on the historical emotion recognition data, an effectiveness analysis is performed to obtain the effective speech duration and effective word count thresholds as the speech judgment anchor points. Feature extraction is performed on the historical emotion recognition data to obtain multiple features associated with emotion recognition; An importance analysis was performed on the aforementioned features to obtain the analysis results; Based on the analysis results, the multiple predictive factors are selected from the multiple features.

3. The emotion recognition method as described in claim 1, wherein, The step of constructing a training set corresponding to each predictor and training an emotion recognition model using the training set corresponding to each predictor includes: Collect user voice data based on each predictor factor; The collected user speech is labeled, and the labeled user speech is used to construct a training set corresponding to each prediction factor; By training a model with predictive function using the training set corresponding to each predictor, an emotion recognition model corresponding to each predictor is obtained.

4. The emotion recognition method as described in claim 1, wherein, The method of integrating and learning multiple trained emotion recognition models using boosting trees to obtain the target emotion recognition model includes: Build the initial model; Using the initial model as the first-round model, in each round of training, an emotion recognition model is added one by one based on the current model, and iterative training is performed; When all emotion recognition models have completed the iteration, training is stopped, and the currently obtained model is determined as the target emotion recognition model. In each iteration, the residual of the current model is fitted using a newly added emotion recognition model.

5. The emotion recognition method as described in claim 4, wherein, The residuals of the current model fitted using the newly added emotion recognition model include: For each training sample, the residual between the predicted value and the true value of the model in that round is calculated to obtain the residual corresponding to each training sample; The residual corresponding to each training sample is fitted using a decision tree to minimize the output value of each leaf node on the decision tree.

6. The emotion recognition method as described in claim 1, wherein, The method further includes: When using the boosting tree to perform ensemble learning on multiple emotion recognition models trained, the predicted value of the newly added emotion recognition model in each iteration is updated to the predicted value of the model in the current iteration, so as to obtain the predicted value of the model after each round of training. Specifically, the update step size is obtained by multiplying the predicted value of the newly added emotion recognition model with the configuration value in each iteration, and the update step size is used to control the update magnitude.

7. The emotion recognition method as described in claim 1, wherein, The process of processing the recognition results of each tree according to the recognition type to obtain the emotion recognition result of the target user includes: Get the weight of each tree; The predicted value for each tree is determined based on the identification results of each tree; The target predicted value is obtained by calculating a weighted sum based on the weight of each tree and the predicted value of each tree. When the recognition type is a classification type, a pre-configured classification threshold is obtained, and the target predicted value is compared with the classification threshold to obtain a comparison result. Based on the comparison result, the emotion recognition result of the target user is determined; or When the recognition type is a regression type, the target predicted value is determined as the emotion recognition result of the target user.

8. An emotion recognition device, wherein, The emotion recognition device includes: The determination unit is used to determine the speech determination anchor point and to determine multiple prediction factors; The training unit is used to construct a training set corresponding to each predictor and to train the emotion recognition model using the training set corresponding to each predictor. The learning unit is used to integrate and learn multiple emotion recognition models trained by boosting trees to obtain the target emotion recognition model. The acquisition unit is used to collect the voice information of the target user as the voice to be recognized in response to the emotion recognition instruction of the target user; The input unit is used to input the speech to be recognized into the target emotion recognition model and obtain the recognition result of each tree in the target emotion recognition model; The determining unit is further configured to determine the recognition type based on the emotion recognition instruction; The processing unit is used to process the recognition results of each tree according to the recognition type to obtain the emotion recognition result of the target user.

9. A computer device, wherein, The computer device includes: Memory, storing at least one instruction; and The processor executes instructions stored in the memory to perform the following steps: Determine the speech judgment anchor point and identify multiple prediction factors; Construct a training set corresponding to each predictor, and train the emotion recognition model separately using the training set corresponding to each predictor. By using boosting trees to ensemble and learn multiple trained emotion recognition models, a target emotion recognition model is obtained. In response to the emotion recognition instruction of the target user, the voice information of the target user is collected as the voice to be recognized according to the voice determination anchor point; The speech to be recognized is input into the target emotion recognition model, and the recognition result of each tree in the target emotion recognition model is obtained; The recognition type is determined according to the emotion recognition instruction; The recognition results of each tree are processed according to the recognition type to obtain the emotion recognition result of the target user.

10. The computer device as claimed in claim 9, wherein, The determination of the speech determination anchor point and the determination of multiple prediction factors include: Obtain historical emotion recognition data; Based on the historical emotion recognition data, an effectiveness analysis is performed to obtain the effective speech duration and effective word count thresholds as the speech judgment anchor points. Feature extraction is performed on the historical emotion recognition data to obtain multiple features associated with emotion recognition; An importance analysis was performed on the aforementioned features to obtain the analysis results; Based on the analysis results, the multiple predictive factors are selected from the multiple features.

11. The computer device as claimed in claim 9, wherein, The step of constructing a training set corresponding to each predictor and training an emotion recognition model using the training set corresponding to each predictor includes: Collect user voice data based on each predictor factor; The collected user speech is labeled, and the labeled user speech is used to construct a training set corresponding to each prediction factor; By training a model with predictive function using the training set corresponding to each predictor, an emotion recognition model corresponding to each predictor is obtained.

12. The computer device as claimed in claim 9, wherein, The method of integrating and learning multiple trained emotion recognition models using boosting trees to obtain the target emotion recognition model includes: Build the initial model; Using the initial model as the first-round model, in each round of training, an emotion recognition model is added one by one based on the current model, and iterative training is performed; When all emotion recognition models have completed the iteration, training is stopped, and the currently obtained model is determined as the target emotion recognition model. In each iteration, the residual of the current model is fitted using a newly added emotion recognition model.

13. The computer device as claimed in claim 12, wherein, The residuals of the current model fitted using the newly added emotion recognition model include: For each training sample, the residual between the predicted value and the true value of the model in that round is calculated to obtain the residual corresponding to each training sample; The residual corresponding to each training sample is fitted using a decision tree to minimize the output value of each leaf node on the decision tree.

14. The computer device as claimed in claim 9, wherein, When the processor executes instructions stored in the memory, it further includes: When using the boosting tree to perform ensemble learning on multiple emotion recognition models trained, the predicted value of the newly added emotion recognition model in each iteration is updated to the predicted value of the model in the current iteration, so as to obtain the predicted value of the model after each round of training. Specifically, the update step size is obtained by multiplying the predicted value of the newly added emotion recognition model with the configuration value in each iteration, and the update step size is used to control the update magnitude.

15. The computer device as claimed in claim 9, wherein, The process of processing the recognition results of each tree according to the recognition type to obtain the emotion recognition result of the target user includes: Get the weight of each tree; The predicted value for each tree is determined based on the identification results of each tree; The target predicted value is obtained by calculating a weighted sum based on the weight of each tree and the predicted value of each tree. When the recognition type is a classification type, a pre-configured classification threshold is obtained, and the target predicted value is compared with the classification threshold to obtain a comparison result. Based on the comparison result, the emotion recognition result of the target user is determined; or When the recognition type is a regression type, the target predicted value is determined as the emotion recognition result of the target user.

16. A non-volatile computer-readable storage medium, wherein: The non-volatile computer-readable storage medium stores at least one instruction, which is executed by a processor in a computer device to perform the following steps: Determine the speech judgment anchor point and identify multiple prediction factors; Construct a training set corresponding to each predictor, and train the emotion recognition model separately using the training set corresponding to each predictor. By using boosting trees to ensemble and learn multiple trained emotion recognition models, a target emotion recognition model is obtained. In response to the emotion recognition instruction of the target user, the voice information of the target user is collected as the voice to be recognized according to the voice determination anchor point; The speech to be recognized is input into the target emotion recognition model, and the recognition result of each tree in the target emotion recognition model is obtained; The recognition type is determined according to the emotion recognition instruction; The recognition results of each tree are processed according to the recognition type to obtain the emotion recognition result of the target user.

17. The non-volatile computer-readable storage medium of claim 16, wherein, The determination of the speech determination anchor point and the determination of multiple prediction factors include: Obtain historical emotion recognition data; Based on the historical emotion recognition data, an effectiveness analysis is performed to obtain the effective speech duration and effective word count thresholds as the speech judgment anchor points. Feature extraction is performed on the historical emotion recognition data to obtain multiple features associated with emotion recognition; An importance analysis was performed on the aforementioned features to obtain the analysis results; Based on the analysis results, the multiple predictive factors are selected from the multiple features.

18. The non-volatile computer-readable storage medium of claim 16, wherein, The step of constructing a training set corresponding to each predictor and training an emotion recognition model using the training set corresponding to each predictor includes: Collect user voice data based on each predictor factor; The collected user speech is labeled, and the labeled user speech is used to construct a training set corresponding to each prediction factor; By training a model with predictive function using the training set corresponding to each predictor, an emotion recognition model corresponding to each predictor is obtained.

19. The non-volatile computer-readable storage medium of claim 16, wherein, The method of integrating and learning multiple trained emotion recognition models using boosting trees to obtain the target emotion recognition model includes: Build the initial model; Using the initial model as the first-round model, in each round of training, an emotion recognition model is added one by one based on the current model, and iterative training is performed; When all emotion recognition models have completed the iteration, training is stopped, and the currently obtained model is determined as the target emotion recognition model. In each iteration, the residual of the current model is fitted using a newly added emotion recognition model.

20. The non-volatile computer-readable storage medium of claim 19, wherein, The residuals of the current model fitted using the newly added emotion recognition model include: For each training sample, the residual between the predicted value and the true value of the model in that round is calculated to obtain the residual corresponding to each training sample; The residual corresponding to each training sample is fitted using a decision tree to minimize the output value of each leaf node on the decision tree.