User emotion recognition method and device, electronic equipment and storage medium
By analyzing the target person's historical video call records, generating an emotion change curve, and performing dynamic weight adjustment and standardization, the accuracy and stability issues of emotion recognition in existing video calls are solved. This enables personalized emotion recognition and intuitive presentation, improving the human-computer interaction experience in scenarios such as video calls and intelligent customer service.
Patent Information
- Application Number
- CN202511468412.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-10-15
AI Technical Summary
Existing video call emotion recognition technologies rely on speech signal processing methods, which are affected by environmental noise, individual differences, and the diversity of emotional expressions, resulting in insufficient accuracy and stability, and making it difficult to fully capture the richness and complexity of emotional expressions.
By mining the target person's historical video call records, an emotion change curve is generated and emotion change habits are extracted. Based on the emotion change habits, the real-time multimodal information is dynamically weighted and standardized, and then weighted and fused to generate a real-time emotion score and display a prompt interface.
It improves the accuracy and stability of emotion recognition, enabling the transformation of emotion recognition from general to personalized, enhances the flexibility and reliability of multimodal information fusion, realizes the intuitive presentation of emotion recognition results and optimization of communication strategies, and improves the human-computer interaction experience in scenarios such as video calls and intelligent customer service.
Smart Images

Figure CN120951293A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis technology, and more specifically, to a user emotion recognition method, device, electronic device, and storage medium. Background Technology
[0002] With the continuous development of mobile communication technology, users have placed higher demands on the quality and interactive experience of video calls. Among these demands, accurately acquiring and understanding the emotional states of both parties in a call is crucial for improving communication efficiency and user satisfaction. Through emotion recognition, the call system can better perceive the user's psychological dynamics, thereby promoting more natural and effective communication.
[0003] Currently, emotion recognition technology largely relies on speech signal processing methods, extracting various features such as tone, speech rate, pitch, and volume, and combining them with machine learning algorithms to judge the speaker's emotions. However, due to the influence of environmental noise, individual speaker differences, and the diversity of emotional expression on the speech signal itself, the accuracy of emotion judgment exhibits certain fluctuations. Furthermore, the richness and complexity of emotional expression are difficult to fully capture through speech signals, affecting the stability and reliability of the overall recognition results. Summary of the Invention
[0004] The problem addressed by this invention is how to improve the accuracy of emotion recognition in video calls.
[0005] To address the aforementioned problems, this invention provides a user emotion recognition method, apparatus, electronic device, and storage medium.
[0006] In a first aspect, the present invention provides a user emotion recognition method, comprising: Based on the historical video call records of the target person obtained in advance, a target emotion change curve is generated, and trend analysis is performed on the target emotion change curve to obtain emotion change habits; The real-time call information of the target person is obtained, and the real-time call information is dynamically weighted based on the emotional change habits to obtain the output weight. The real-time call information includes text information, facial expression information and voice information. The real-time call information is standardized, and the processed text information, facial expression information, and voice information are weighted and fused according to the output weights to obtain a real-time emotion score. The current emotion type of the target person is obtained based on the real-time emotion score, and a prompt interface is generated based on the current emotion type and displayed on the terminal interface.
[0007] Optionally, the step of generating a target emotion change curve based on the pre-acquired historical video call records of the target person, and performing trend analysis on the target emotion change curve to obtain emotion change habits, includes: Based on Mehrabian's rule, the historical multimodal features in the historical video call records are weighted and fused to obtain a historical emotion score. The historical multimodal features include historical call text information, historical facial expression information, and historical voice information. Generate a target emotion change curve with time as the horizontal axis and the historical emotion score as the vertical axis; Identify abrupt change nodes in the historical emotion scores of the target emotion change curve, associate them with the historical video call records corresponding to the abrupt change nodes, and obtain historical emotion slices; Based on the historical emotion scores, each historical emotion slice is labeled with an emotion type according to the emotion judgment criteria, and typical scenes and typical features under each emotion type are obtained based on the historical multimodal features corresponding to the same emotion type. Based on the typical scenarios, typical characteristics, and the emotion types, the emotion change habits are obtained.
[0008] Optionally, identifying abrupt change nodes in the historical emotion ratings within the target emotion change curve and associating them with historical video call records corresponding to those nodes includes: Set a mutation threshold, which is the maximum allowable fluctuation value of the historical emotion score within a preset continuous time window in the target emotion change curve; The target emotion change curve is traversed by sliding through the continuous time window to obtain the historical fluctuation value of the historical emotion score within each continuous time window. When the historical fluctuation value exceeds the mutation threshold, the start time of the continuous time window is marked as the mutation node of the historical emotion score. In the historical video call records, slices are obtained by extending a preset time before and after the mutation node, and the slices are associated with the corresponding historical video call records to obtain historical emotion slices.
[0009] Optionally, based on the historical emotion score, each historical emotion slice is labeled with an emotion type according to the emotion determination criteria, and typical scenes and typical features under each emotion type are obtained based on the historical multimodal features corresponding to the same emotion type, including: Obtain the historical emotion score corresponding to the historical emotion slice, and assign the emotion type label to each historical emotion slice according to the emotion judgment criteria, wherein the emotion type label includes happiness, anger, calmness and sadness; The K-means clustering algorithm is used to cluster the historical multimodal features under the same sentiment type label to obtain multiple feature clusters; Statistical analysis is performed on multiple feature clusters to obtain the typical scenarios and typical features; The typical scenarios and typical features are associated and stored with the corresponding emotion type labels to obtain emotion change habits.
[0010] Optionally, the step of performing statistical analysis on multiple feature clusters to obtain the typical scenario and the typical features includes: For each of the feature clusters, the corresponding historical call scenarios are extracted, the frequency of occurrence of the same historical call scenarios is counted, and the historical call scenarios with a frequency exceeding a preset frequency threshold are identified as the typical scenarios. The typical features are obtained by analyzing the historical multimodal features in each feature cluster, wherein the typical features include typical text features, typical facial expression features, and typical voice features.
[0011] Optionally, the step of dynamically weighting the real-time call information based on the emotional change habits to obtain output weights includes: Based on the text information, facial expression information and voice information of the call, the real-time call scene is extracted, and the first weight is obtained by pre-defined weight division and combination of the emotion type label associated with the typical scene that is the same as the real-time call scene; The text information of the call, the facial expression information, and the voice information are matched with the typical features to obtain the corresponding matching degree. The emotion type label corresponding to the matching degree being greater than the preset matching threshold is obtained. The weight adjustment value is obtained based on the matching degree, the preset matching threshold, and the preset adjustment coefficient. The second weight is obtained by dividing and combining the weight adjustment value and the preset weight corresponding to the emotion type label. Weight calibration values are obtained by performing weight calibration based on the acquisition quality of the call text information, facial expression information, and voice information. The first weight, the second weight, and the weight calibration value are input into the weight calibration model to obtain the output weight. The weight calibration model is trained on a preset model based on the historical first weight, the historical second weight, the historical weight calibration value, and the optimal weight. The historical first weight is a preset weight division combination corresponding to different emotion type labels. The historical second weight is adjusted after matching the historical multimodal features with the typical features. The historical weight calibration value is obtained based on the acquisition quality of the historical multimodal features. The optimal weight is the weight information manually labeled for each group of historical multimodal features.
[0012] Optionally, generating a prompt interface based on the current emotion type and displaying it on the terminal interface includes: Set a corresponding cue color and cue effect for each of the current emotion types; The prompt interface is generated based on the current emotion type and the corresponding prompt color and prompt effect, and then displayed on the terminal interface.
[0013] In a second aspect, the present invention provides a user emotion recognition device, comprising: The data processing module is used to generate a target emotion change curve based on the target person's historical video call records obtained in advance, and to perform trend analysis on the target emotion change curve to obtain emotion change habits. The weighting module is used to acquire the real-time call information of the target person, and to dynamically weight the real-time call information based on the emotional change habits to obtain the output weight. The real-time call information includes call text information, facial expression information and voice information. The scoring module is used to standardize the real-time call information, and to perform weighted fusion of the processed call text information, facial expression information and voice information according to the output weights to obtain a real-time emotion score. The display module is used to obtain the current emotion type of the target person based on the real-time emotion score, generate a prompt interface based on the current emotion type, and display it on the terminal interface.
[0014] Thirdly, the present invention provides an electronic device, including a memory and a processor; The memory is used to store computer programs; The processor is configured to implement the user emotion recognition method as described in the first aspect when executing the computer program.
[0015] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the user emotion recognition method as described in the first aspect.
[0016] The beneficial effects of user emotion recognition in this invention are as follows: By mining the target individual's historical video call records, generating the target's emotion change curve and extracting emotion change habits, the basis for emotion recognition shifts from "generalization" to "personalization," improving the adaptability of existing general models to individuals, providing accurate historical references for subsequent real-time recognition, and enhancing the targeting of emotion recognition. Based on emotion change habits, dynamic weight adjustments are made to real-time multimodal information. For example, when matching topics related to emotion fluctuations, the weight of text information is increased; when the feature matching degree is high, the weight of the corresponding modality is increased. This makes the weight allocation more suitable for complex real-time scenarios and user expression habits, enhancing the flexibility and reliability of multimodal information fusion. Standardization processing maps different modal features to a unified scale (such as the 0-1 range), eliminating interference caused by differences in units and magnitudes. Combined with dynamic output weights for weighted fusion, real-time emotion scoring can integrate effective information from various modalities, improving the accuracy and stability of the scoring. Based on real-time emotion scoring, specific emotion types are determined, and a personalized interface with prompts (such as colors and icons) is generated. This allows users or communication partners to perceive emotional states in real time, solving the problem of "vague application scenarios" in emotion recognition results. This helps optimize communication strategies (such as adjusting topics to address anxiety) and improves practicality. This invention, through a complete process of "personalized habit mining—dynamic weight adaptation—standardized fusion—intuitive presentation," deeply integrates users' historical behavioral patterns with real-time scene features, achieving an upgrade from "passive recognition" to "active adaptation." This significantly improves the accuracy, robustness, and practical value of emotion recognition and can be widely applied to scenarios such as video calls and intelligent customer service, improving the human-computer interaction experience. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the user emotion recognition method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the prompt interface according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the user emotion recognition device according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the electronic device structure according to an embodiment of the present invention. Detailed Implementation
[0018] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the accompanying drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0019] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0020] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to"; the term "based on" means "at least partially based on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; and the term "optionally" means "optional embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first," "second," etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0021] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0022] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0023] It is understood that any part of this application concerning data acquisition or collection has been authorized by the user.
[0024] To address the problems existing in the aforementioned related technologies, this embodiment provides a user emotion recognition method, device, electronic device, and storage medium.
[0025] like Figure 1 As shown in the figure, an embodiment of the present invention provides a user emotion recognition method, including: Step S1: Generate a target emotion change curve based on the target person's historical video call records obtained in advance, and perform trend analysis on the target emotion change curve to obtain emotion change habits.
[0026] Specifically, historical video call records include historical voice features, historical text features, and historical facial expression features used to generate target emotion change curves. Historical voice features are further subdivided into historical pitch change slope, historical speech rate fluctuation frequency, historical pitch peak value, historical sound intensity abrupt change duration, and historical noise levels. For example, Fourier transform is used to extract the historical pitch change slope (difference between the initial and final pitch values / 2 seconds) within each 2-second window; the historical speech rate fluctuation frequency for each sentence is calculated (the standard deviation of the duration of five adjacent syllables); the historical pitch peak value (the highest fundamental frequency of each call segment) and the historical sound intensity abrupt change duration (the duration of a sudden increase in sound pressure level from 60dB to above 80dB) are recorded; and a noise separation algorithm is used to label the historical background noise level (divided into no noise <20dB, low noise 20-40dB, and high noise >40dB). Historical text features are further subdivided into the frequency of occurrence of historical emotional sentiment words, historical semantic emotional sentiment values, and historical... For semantic communication scenarios, for example, based on sentiment dictionaries (such as CNKI sentiment dictionary), the frequency of occurrence of historical sentiment-biased words is counted (number of positive / negative words per 50 characters). Historical semantic sentiment bias values (range 0-1, 0 for extreme negative, 1 for extreme positive) are calculated using, for example, a bidirectional LSTM model. This is combined with topic models (such as LDA) to identify historical semantic communication scenarios (such as "academic discussion", "daily trivial matters", "business negotiation"). Historical facial expression features are further subdivided into historical changes in facial features and historical facial muscle movement rates. For example, the MTCNN algorithm is used to locate key points of facial features and calculate historical changes in facial features (such as the opening and closing of the corners of the eyes, the expansion rate of the nostrils). Historical facial muscle movement rates (average displacement velocity of key points within 20 consecutive frames, unit pixel / millisecond) are calculated using optical flow.
[0027] After obtaining the relevant features, principal component analysis (PCA) can be used to reduce the dimensionality of historical multimodal features. Principal components with a cumulative contribution rate of >85% are selected and mapped to historical sentiment scores (0-100) through a support vector regression (SVR) model. The target sentiment change curve is generated with a time granularity of 10 seconds.
[0028] Based on the obtained target emotion change curve, emotion change habits are analyzed. These habits can be identified as related topics and characteristic expression patterns that influence emotions. Wavelet transform can be used to decompose the curve into multiple scales to identify time nodes of abrupt changes in emotion scores (e.g., a change in score exceeding 30 points within 10 seconds is defined as a change in score). The semantic communication scenarios corresponding to these nodes are then associated to obtain the related topics of emotion fluctuations (e.g., "emotions are easily volatile when 'costs' are mentioned in business negotiations"). Hierarchical clustering algorithms are used to extract characteristic expression patterns from historical features of the same emotion type (e.g., "anger"). These patterns include (e.g., tone slope > 40 Hz / second, eye opening / closing < 30%, and frequency of negative words > 8 times / 50 words during anger).
[0029] Step S2: Obtain the real-time call information of the target person, and dynamically weight the real-time call information based on the emotional change habits to obtain the output weight. The real-time call information includes call text information, facial expression information and voice information.
[0030] Specifically, the text information of the call is further subdivided into the frequency of occurrence of sentiment-oriented words, semantic sentiment value, and semantic communication scenario. During real-time calls, the text is generated through a real-time speech-to-text engine (such as Google Speech-to-Text). The method for obtaining the specific features is the same as that for obtaining the historical text features mentioned above, and will not be repeated here. The facial expression information is further subdivided into changes in the shape of facial features and the rate of facial muscle movement. The method for obtaining the facial expression features is the same as that for obtaining the historical facial expression features mentioned above, and will not be repeated here. The voice information is further subdivided into the slope of tone changes, frequency of speech rate fluctuations, peak pitch, duration of sudden changes in sound intensity, and noise levels. The method for obtaining the voice information is the same as that for obtaining the historical voice features mentioned above, and will not be repeated here.
[0031] After obtaining real-time call information, initial weights are first set. For example, the initial weight for text information is 0.25, the initial weight for facial expression information is 0.35, and the initial weight for voice information is 0.4, reflecting the differences in information content of historical features. The obtained real-time call information is then matched with emotional change habits, such as related topics and feature expression patterns. The weights are adjusted based on the matching results to obtain the final output weights. For example, the cosine similarity between real-time call information and related topics, and the Euclidean distance between real-time call information and feature expression patterns are calculated. If the matching degree between the real-time semantic communication scenario and the related topics of emotional fluctuations is greater than the preset threshold of 0.8, the corresponding modality weight is increased by 0.15 (e.g., if the correlation of text information is high, the text weight is increased to 0.4, and other features are adjusted adaptively). When the Euclidean distance between real-time call information and feature expression patterns is less than the preset threshold of 0.3, the corresponding modality weight is increased by 0.1 (e.g., if the matching degree of facial features is high, the facial weight is increased to 0.45, and other features are adjusted adaptively). Based on this, the adjusted values are integrated through a linear programming model to ensure that the sum of the weights is 1, ultimately obtaining the output weights.
[0032] Step S3: Standardize the text information, facial expression information, and voice information of the call, and then perform weighted fusion of the processed text information, facial expression information, and voice information according to the output weights to obtain a real-time emotion score.
[0033] Specifically, since real-time call information has multi-source features, it is necessary to standardize the real-time call information before performing weighted full fusion to ensure more accurate data processing. For audio information, the pitch change slope can be standardized using the cumulative distribution function (CDF). A sample set of pitch change slopes from historical speech audio data is collected, and its cumulative distribution function is calculated. For the real-time pitch change slope value, the corresponding cumulative probability value (range 0-1) is determined by looking up the historical CDF; this value is the standardization result. If the real-time value exceeds the historical sample range (e.g., higher than the maximum value), it is treated as a boundary value (i.e., standardized value = 1). The speech rate fluctuation frequency can be standardized using logarithmic transformation and min-max normalization. A natural logarithmic transformation is performed on the real-time speech rate fluctuation frequency (formula: log(x+1), where x is the real-time value) to alleviate the right-skewed distribution of the data. The transformed values... The min-max normalization method is used: Standardized value = (Transformed value - Historical minimum transformed value) / (Historical maximum transformed value - Historical minimum transformed value), mapped to the 0-1 interval; Pitch peak value can be standardized using quantiles, with the 99th percentile value of the pitch peak value determined based on historical data as the upper limit threshold (e.g., 500Hz). If the real-time pitch peak value ≤ the upper limit threshold, the standardized value = real-time value / upper limit threshold; if the real-time value > the upper limit threshold, the standardized value = 1 (to avoid interference from extreme values on the overall scale); The duration of sound intensity change can be standardized using threshold truncation, with the standardized value = real-time sound intensity change duration / threshold. If the real-time value > the threshold, the standardized value = 1; if the real-time value < 0 (invalid value), the standardized value = 0. For textual information, the frequency of sentiment-related words can be standardized using Z-score and truncated. The mean (μ) and standard deviation (σ) of the frequency of sentiment-related words in historical data are calculated. The standardized value is calculated as (real-time frequency - μ) / σ. The result is then truncated: if the standardized value is > 3, it is treated as 3; if it is < -3, it is treated as -3. The final mapping to the 0-1 interval is: standardized value = (truncated value + 3) / 6. Semantic sentiment values can be standardized using linear offset. The original range is set to -1 to 1, where -1 represents extreme negativity and 1 represents extreme positivity. A linear transformation directly maps the original range to the 0-1 interval. The formula for standardization is: Standardized value = (Original semantic sentiment value + 1) / 2. For example, original value = -1 corresponds to 0, original value = 0 corresponds to 0.5, and original value = 1 corresponds to 1. The semantic communication scenario judgment adopts one-hot encoding for standardization. Based on historical data, all possible semantic communication scenarios (such as 5 types of scenarios) are defined to form a scenario set. For the semantic communication scenario identified in real time, a binary vector with a length equal to the number of scenarios in the scenario set is generated, where the position corresponding to the scenario is 1 and the rest are 0. For example, the "work" scenario corresponds to the vector [1,0,0,0,0], and "family" corresponds to [0,1,0,0,0].For facial expression information, changes in facial features can be standardized using relative deviation. The facial features of the target person in a calm state are used as the baseline (e.g., the angle of the corner of the mouth in a natural state = 0°). The deviation between the real-time form and the baseline value is calculated (Δ = real-time value - baseline value). Based on historical data, the maximum positive deviation (Δmax+) and the maximum negative deviation (Δmax-) are determined. If Δ ≥ 0: Standardized value = Δ / Δmax+ (mapped to 0-1, positive change); if Δ < 0: Standardized value = 1 - (Δ / Δmax-) (mapped to 0-1, negative change). Facial muscle movement rate can be standardized using an exponential function. Based on historical data, a saturation threshold for muscle movement rate is determined (e.g., 10 pixels / frame; exceeding this value no longer significantly increases the intensity of emotional expression). An exponential function mapping is used: Standardized value = 1 - e^(-k × real-time rate), where k is an adjustment parameter (fitted based on historical data, e.g., k = 0.3). This method is sensitive to low-rate changes (distinguishing subtle expressions) and tends to be less sensitive to high-rate changes (avoiding over-amplification).
[0034] The standardized real-time call information is weighted and summed separately. For example, text feature score = 0.6 × emotional vocabulary score + 0.4 × semantic tendency score. Then, the weighted sum is fused according to the output weight. For example, real-time emotion score = total text score × 0.35 + total facial score × 0.4 + total voice score × 0.25 to obtain the final real-time emotion score.
[0035] Step S4: Obtain the current emotion type of the target person based on the real-time emotion score, generate a prompt interface based on the current emotion type, and display it on the terminal interface.
[0036] Specifically, fuzzy logic reasoning is used to set the membership function between the scoring range and the emotion type. For example, 0-25 points belong to "anger" (0.8) and "sadness" (0.2); 26-45 points belong to "anxiety" (0.7) and "calm" (0.3); 46-75 points belong to "calm" (0.9); and 76-100 points belong to "pleasure" (0.8) and "excitement" (0.2). The final type is determined by combining the matching results of real-time features and typical patterns (e.g., "pleasure" requires the corners of the mouth to be raised by more than 15°).
[0037] Dynamic prompts are configured for different emotion types: "Anger" displays a red dynamic warning box; "Joy" displays a yellow gradient background; and "Calm" displays a blue static box. These prompts are displayed in real-time on the terminal interface (such as the sidebar of a video call software), including emotion type labels and characteristic prompts (e.g., "Abnormal pitch peak detected, current emotion: 'Anger'").
[0038] In this embodiment, by mining the target individual's historical video call records, a target emotion change curve is generated and emotion change habits are extracted. This shifts the basis of emotion recognition from "general" to "personalized," improving the adaptability of existing general models to individuals and providing accurate historical references for subsequent real-time recognition, thus enhancing the targeting of emotion recognition. Based on emotion change habits, dynamic weight adjustments are made to real-time multimodal information. For example, the weight of text information is increased when matching topics related to emotion fluctuations, and the weight of the corresponding modality is increased when the feature matching degree is high. This makes the weight allocation more suitable for complex real-time scenarios and user expression habits, enhancing the flexibility and reliability of multimodal information fusion. Standardization processes map different modal features to a unified scale (e.g., the 0-1 range), eliminating interference caused by differences in units and magnitudes. Combined with dynamic output weights for weighted fusion, real-time emotion scoring can integrate effective information from various modalities, improving the accuracy and stability of the scoring. Based on real-time emotion scoring, specific emotion types are determined, and a personalized interface with prompts (such as colors and icons) is generated. This allows users or communication partners to perceive emotional states in real time, solving the problem of "vague application scenarios" in emotion recognition results. This helps optimize communication strategies (such as adjusting topics to address anxiety) and improves practicality. This invention, through a complete process of "personalized habit mining—dynamic weight adaptation—standardized fusion—intuitive presentation," deeply integrates users' historical behavioral patterns with real-time scene features, achieving an upgrade from "passive recognition" to "active adaptation." This significantly improves the accuracy, robustness, and practical value of emotion recognition and can be widely applied to scenarios such as video calls and intelligent customer service, improving the human-computer interaction experience.
[0039] Optionally, the step of generating a target emotion change curve based on the pre-acquired historical video call records of the target person, and performing trend analysis on the target emotion change curve to obtain emotion change habits, includes: Based on Mehrabian's rule, the historical multimodal features in the historical video call records are weighted and fused to obtain a historical emotion score. The historical multimodal features include historical call text information, historical facial expression information, and historical voice information.
[0040] Specifically, the core of Mehrabian's rule is a fixed weighting ratio of "7% language content + 38% voice tone + 55% facial expression". When the emotional habits of the target person are unclear, the commonly used Mehrabian rule is used to weight and fuse historical multimodal features from historical video call records to generate a historical emotion score. The historical emotion score = (weighted sum of text features × 7%) + (weighted sum of voice and audio features × 38%) + (weighted sum of facial expression features × 55%), mapped to a score of 0-100. Among them, the weighting within each modality feature can refer to preset standards. In voice and audio features, the slope of tone change accounts for 30%, the frequency of speech rate fluctuation accounts for 20%, the peak pitch accounts for 20%, the duration of sound intensity change accounts for 20%, and the noise situation accounts for 10%. In text features, the frequency of emotional words accounts for 40%, the semantic emotional tendency value accounts for 40%, and the semantic communication scenario accounts for 20%. In facial expression features, the change of facial features accounts for 0%, and the muscle movement rate accounts for 40%.
[0041] By adapting Mehrabian's Law, the universality of the law in the laws governing human emotional expression is preserved, making the historical emotional scores obtained more consistent with the target person's true emotional state when the target person's emotional change habits are unknown.
[0042] Generate a target emotion change curve with time as the horizontal axis and the historical emotion score as the vertical axis.
[0043] Specifically, linear interpolation can be used to connect the rating values of each window to generate a continuous target emotion change curve. The horizontal axis represents the cumulative time since the start of the call (in seconds), and the vertical axis represents the historical emotion rating (0-100). Local extreme points (peaks and valleys) of the curve can also be marked to facilitate subsequent analysis.
[0044] Identify abrupt change nodes in the historical emotion score within the target emotion change curve, and associate them with the corresponding historical video call records to obtain historical emotion slices.
[0045] Specifically, judgment conditions are set, such as an absolute difference of more than 25 points in historical sentiment scores within two consecutive time windows (totaling 10 seconds), and a change rate of more than 50% in the score of the later window compared to the earlier window (e.g., a sudden increase from 40 points to 70 points, with a change rate of 75%). A sliding window is used to traverse the curve. When the above conditions are met, the start time of the later window can be marked as a mutation node. Centered on the mutation node, historical video call segments (including voice, text, and facial expression data) are extracted forward, for example, by 5 seconds, and backward, for example, by 2 seconds, to form a 7-second historical sentiment slice. At the same time, the slice can be labeled with the corresponding timestamp, historical sentiment score, and mutation direction (sudden increase / sudden decrease).
[0046] Based on the historical emotion scores, each historical emotion slice is labeled with an emotion type according to the emotion determination criteria. Typical scenes and typical features under each emotion type are obtained based on the historical multimodal features corresponding to the same emotion type.
[0047] Based on the typical scenarios, typical characteristics, and the emotion types, the emotion change habits are obtained.
[0048] Optionally, identifying abrupt change nodes in the historical emotion ratings within the target emotion change curve and associating them with historical video call records corresponding to those nodes includes: A mutation threshold is set, which is the maximum allowable fluctuation value of the historical emotion score within a preset continuous time window in the target emotion change curve.
[0049] Specifically, research shows that significant changes in human emotions typically occur within 5-10 seconds, with 8 seconds being the optimal value balancing sensitivity and stability. Therefore, based on the temporal characteristics of emotional expression in historical video calls, the continuous time window length is set to 8 seconds. Fluctuation values (i.e., the difference between the highest and lowest scores within a window) within all 8-second windows of the target person's historical emotional ratings are collected to form a fluctuation value sample set. The 90th percentile of the sample set is used as a baseline threshold. If 90% of the fluctuation values in the sample set are ≤22 points, the baseline threshold of 22 points is set as the mutation threshold. Alternatively, it can be adjusted based on emotion type; for example, for strong emotion types such as "anger" and "excitement," the threshold is increased by 10% (i.e., 24 points), and for weak emotion types such as "calm," the threshold is decreased by 10% (i.e., 20 points). The final mutation threshold is set at 22 points.
[0050] By sliding through the target emotion change curve in the continuous time window, the historical fluctuation value of the historical emotion score within each continuous time window is obtained. When the historical fluctuation value exceeds the mutation threshold, the start time of the continuous time window is marked as the mutation node of the historical emotion score.
[0051] Specifically, using continuous time windows (8 seconds) as units, the algorithm slides from the beginning of the target emotion change curve, with each slide step lasting 2 seconds (ensuring overlap between windows to avoid missing abrupt changes), until the end of the curve. For each window, the maximum and minimum values of all historical emotion scores are extracted, and the absolute difference between the two is calculated to obtain the historical fluctuation value. When the historical fluctuation value of a window is greater than the abrupt change threshold of the corresponding emotion type (e.g., for a window corresponding to "anger," a fluctuation value of 25 points is greater than 24 points), the starting time point of that window is marked as the abrupt change node of the historical emotion score. Alternatively, the trend point within the window can be combined with the fluctuation value determination result. For example, when the historical fluctuation value of a window is greater than the abrupt change threshold of the corresponding emotion type, and the window contains at least 3 non-continuous score increase / decrease points (excluding misjudgments caused by a single extreme value), the starting time point of that window is marked as the abrupt change node of the historical emotion score, and the fluctuation value and emotion type corresponding to the node are recorded (based on the average score within the window). By using a sliding window and overlapping step design, a full-coverage scan of the emotion curve is achieved. By using fluctuation values or a combination of fluctuation values and trend points, the system effectively distinguishes between real mutations and noise interference, thereby improving the recall and accuracy of mutation node identification.
[0052] In the historical video call records, slices are obtained by extending a preset time before and after the mutation node, and the slices are associated with the corresponding historical video call records to obtain historical emotion slices.
[0053] Specifically, in order to obtain contextual relevance analysis of emotional mutations, a video segment with a total duration of 45 seconds is formed by extending forward 20 seconds (covering the emotional incubation period before the mutation) and backward 25 seconds (covering the emotional duration after the mutation) from the mutation node. The resulting slice is separated from this slice and retains complete multimodal data. Metadata obtained from historical call video records is used to associate the slice with the corresponding mutation node timestamp, historical fluctuation value, and multimodal information to obtain historical emotional slices.
[0054] By varying the duration of the segment before and after the change, the "cause and effect" of the change are fully preserved (such as a slight increase in tone before the change and a sustained expression after the change). Combined with the complete extraction of multimodal information, the historical emotion slices not only include rating data, but also cover the speech, text and facial features that support the change in emotion. This solves the problem of missing context caused by the short slice duration and provides information-rich samples for subsequent typical scenarios and feature extraction.
[0055] Optionally, based on the historical emotion score, each historical emotion slice is labeled with an emotion type according to the emotion determination criteria, and typical scenes and typical features under each emotion type are obtained based on the historical multimodal features corresponding to the same emotion type, including: Obtain the historical emotion score corresponding to the historical emotion slice, and assign an emotion type label to each historical emotion slice according to the emotion judgment criteria, wherein the emotion type label includes happiness, anger, calmness and sadness.
[0056] Specifically, based on the distribution characteristics of historical emotion scores, four score ranges are defined to correspond to different emotion types: historical emotion scores of 75-100 correspond to "happiness" (characterized by predominantly positive emotions, with higher scores indicating stronger feelings of pleasure); historical emotion scores of 50-74 correspond to "calm" (stable emotions with no significant fluctuations); historical emotion scores of 25-49 correspond to "sadness" (predominantly negative emotions, accompanied by low-energy characteristics); and historical emotion scores of 0-24 correspond to "anger" (intense negative emotions, accompanied by high-energy characteristics). The average historical emotion score for each historical emotion slice is extracted, which can be determined by the arithmetic mean of scores at all time points within the slice. Emotion type labels are then directly matched according to the aforementioned ranges. For example, a slice with an average score of 82 is labeled "happiness," and an average score of 33 is labeled "sadness."
[0057] By directly mapping the rating range to the emotion type, standardized labeling of emotion tags is achieved, avoiding the subjectivity of manual labeling. The four emotion types cover the core categories of basic human emotions, and the correlation with the rating range conforms to the quantitative law of emotion intensity, solving the problem of the disconnect between emotion type and rating, and providing a unified classification benchmark for subsequent cluster analysis.
[0058] The K-means clustering algorithm was used to cluster the historical multimodal features under the same emotion type label, resulting in multiple feature clusters.
[0059] Specifically, historical multimodal features of the same emotion type label (e.g., "anger") are standardized, for example, by mapping them to the 0-1 interval using the min-max method. The preprocessed features are then integrated into a high-dimensional feature vector, such as each sample containing 10-dimensional features. The optimal number of clusters K is determined based on the silhouette coefficient method; for example, K=3 for the "anger" emotion, meaning 3 feature clusters. The cluster centers are iteratively updated using Euclidean distance as the similarity measure until the sum of squared errors within the cluster converges (e.g., the number of iterations ≤ 500). Each feature cluster is then output, and each cluster contains historical multimodal feature samples with high feature similarity.
[0060] By using K-means clustering, multimodal features under the same emotion type are divided into several sub-clusters, revealing different expression modes of the same emotion (such as "anger" may be expressed as "loud accusation" or "silent repression"). This solves the problem of difficulty in extracting typical patterns caused by feature mixing under a single emotion type, and lays the foundation for accurately mining typical scenarios and features.
[0061] Statistical analysis is performed on multiple feature clusters to obtain the typical scenarios and typical features.
[0062] Specifically, typical scenarios can be communication scenarios and noisy environments extracted based on historical multimodal features. For example, semantic communication scenarios can be extracted from text information, such as matching keywords to a pre-defined scenario library (e.g., "product complaint," "salary negotiation," "parent-child education," etc.). Noise scenarios can be determined based on noise levels, such as high-noise scenarios. For ambiguous scenarios (e.g., unclear keywords), contextual features of voice and audio information (e.g., "complaint" is often accompanied by high speech rate and high volume) are used to assist in the judgment, ensuring that the scenario classification accuracy is ≥90%. Typical features are the multimodal features of the target person under the corresponding emotional type. For example, when the target person is angry, typical voice features are: pitch change slope >30Hz / second, speech rate fluctuation frequency >0.6 (coefficient of variation), peak pitch >350Hz, and duration of sudden change in volume >1.2 seconds; typical text features are: frequency of negative words >8 times / 200 words, semantic sentiment tendency value <-0.6; and typical facial features are: eyebrow drooping amplitude >8mm, corner of mouth drooping angle >12°, and facial muscle movement rate >5 pixels / frame.
[0063] The typical scenarios and typical features are associated and stored with the corresponding emotion type labels to obtain emotion change habits.
[0064] Specifically, a three-dimensional data table is constructed: "Emotion Type - Typical Scene - Typical Feature". For example, the emotion type is "Anger" → the typical scene is "Contract Dispute" → the typical features are (tone change slope 35-45Hz / second, negative words account for 30%, eyebrow droop >15°). Each entry includes the feature statistics, scene proportion, and sample size (e.g., based on 50 historical emotion slices). In addition, when new historical video call records are added, the typical scene and feature statistics are re-clustered and updated (e.g., when the sample size increases to 60, the mean and quantiles are adjusted) to ensure that emotion change habits are continuously optimized with data accumulation.
[0065] By using structured storage, we can achieve a precise association between emotion type, scene, and features. This allows emotion change habits to include both the pattern of "what emotion is generated in what scene" (e.g., "contract disputes easily lead to anger") and the features of "how the emotion is expressed" (e.g., specific expressions such as tone of voice and facial expressions). This solves the problem of vague descriptions of emotion habits that are difficult to apply to real-time recognition, and provides a rule base that can be directly called upon for subsequent dynamic weighting.
[0066] Optionally, the step of performing statistical analysis on multiple feature clusters to obtain the typical scenario and the typical features includes: For each feature cluster, the corresponding historical call scenarios are extracted, the frequency of occurrence of the same historical call scenarios is counted, and the historical call scenarios with a frequency exceeding a preset frequency threshold are identified as the typical scenarios.
[0067] Specifically, the frequency of each historical call scenario within a cluster is counted (e.g., "product complaint" appears 28 times in cluster 1, and "salary negotiation" appears 8 times); a preset frequency threshold is set: based on the total number of samples within the cluster (e.g., cluster 1 contains 50 samples), the threshold is 30% of the total number of samples (i.e., 15 times); scenarios with a frequency exceeding the threshold are identified as typical scenarios of that cluster (e.g., "product complaint" becomes a typical scenario of cluster 1 because it appears 28 times > 15 times); if multiple scenarios exceed the threshold (e.g., "family dispute" appears 20 times and "neighborhood conflict" appears 18 times in cluster 2, totaling 50), then the top one in frequency is selected as the core typical scenario.
[0068] The typical features are obtained by analyzing the historical multimodal features in each feature cluster, wherein the typical features include typical text features, typical facial expression features, and typical voice features.
[0069] Specifically, for historical multimodal features of the same emotion type labels, box plot analysis was used to remove outliers (such as values exceeding 1.5 times the interquartile range). The mean and 95% confidence interval of the features of the remaining different emotion type labels were calculated to obtain typical features: such as typical speech features of "anger": pitch change slope > 30Hz / second, speech rate fluctuation frequency > 0.6 (coefficient of variation), peak pitch > 350Hz, duration of intensity change > 1.2 seconds; typical text features: frequency of negative words > 8 times / 200 words, semantic sentiment tendency value < -0.6; typical facial features: eyebrow drooping amplitude > 8mm, corner of mouth drooping angle > 12°, facial muscle movement rate > 5 pixels / frame.
[0070] Optionally, the step of dynamically weighting the real-time call information based on the emotional change habits to obtain output weights includes: Based on the text information, facial expression information, and voice information of the call, the real-time call scene is extracted, and the first weight is obtained by pre-defined weight division and combination of the emotion type tags associated with the typical scene that is the same as the real-time call scene.
[0071] Specifically, the real-time call scenario is first obtained, for example, by extracting keywords (such as "performance" and "assessment") from the text information of the real-time call, and combining them with the semantic sentiment tendency value (such as -0.3, which is neutral), and matching them with a preset scenario library (such as "work evaluation" and "casual chat") through the TF-IDF algorithm to determine the real-time call scenario (such as "work evaluation"). If the text information is ambiguous (such as less than 3 keywords), the slope of the tone change of the real-time voice information (such as 15Hz / second, which is smooth) and the muscle movement rate of facial expressions (such as 2 pixels / frame, which is stable) are combined to assist in the scenario judgment, ensuring that the scenario recognition accuracy is ≥85%.
[0072] Secondly, the preset weighting combination of emotion type labels associated with typical scenarios with the same real-time call scenario is called. It should be noted that since different features are expressed differently under different emotions, each emotion type label is assigned a weighting. For example, the "job evaluation" scenario is associated with the "anxiety" emotion type, and its preset weighting combination is: text information 0.3, facial expression information 0.3, and voice information 0.4 (based on historical data, voice features have the highest contribution to anxiety in this scenario). This combination is directly used as the first weight (text: 0.3, facial expression: 0.3, voice: 0.4).
[0073] Real-time scenes are extracted by multimodal information fusion to ensure the accuracy of scene matching; the first weight is quickly generated based on the preset weight combination, so that the weight division initially fits the scene characteristics, solving the problem of weight and scene disconnection, and providing a basic benchmark for subsequent dynamic adjustment.
[0074] The text information, facial expression information, and voice information of the call are matched with the typical features to obtain the corresponding matching degree. The emotion type label corresponding to the matching degree being greater than the preset matching threshold is obtained. The weight adjustment value is obtained based on the matching degree, the preset matching threshold, and the preset adjustment coefficient. The second weight is obtained by dividing and combining the weight adjustment value and the preset weight corresponding to the emotion type label.
[0075] Specifically, the multimodal features and typical features of real-time call information were matched item by item: Voice features: the overlap between the pitch change slope (real-time 18Hz / second) and the typical feature of "anxiety" (15-25Hz / second) was 80%; the overlap between the speech rate fluctuation frequency (real-time 0.5 times / second) and the typical value (0.4-0.6 times / second) was 90%; the overall voice matching degree = (80%+90%) / 2 = 85%; Text features: the overlap between the frequency of negative words (real-time 4 times / 150 words) and the typical value (3-5 times / 150 words) was 100%; semantic sentiment tendency value... The overlap between the real-time value (-0.4) and the typical value (-0.5 to -0.3) is 80%, and the overall text matching degree is (100% + 80%) / 2 = 90%; Facial expression features: the overlap between the eyebrow raising amplitude (real-time 4mm) and the typical value (3-6mm) is 100%, and the overlap between the muscle movement rate (real-time 3 pixels / frame) and the typical value (2-4 pixels / frame) is 100%, and the overall facial matching degree is 100%; The overall matching degree is calculated using a weighted average: voice 40% × 85% + text 30% × 90% + face 30% × 100% = 90.5%.
[0076] The preset matching threshold is set to 70%. The current matching degree is 90.5% > 70%. The preset weight combination associated with the "anxiety" emotion type is (text 0.3, face 0.3, voice 0.4). The adjustment coefficient α is calculated as (matching degree - threshold) / (1 - threshold) = (90.5% - 70%) / 30% ≈ 0.68. The maximum adjustment range for a single modality is ±0.1. The weight of the face feature (100%) with the highest matching degree is increased by 0.3 + 0.1 × α ≈ 0.37. The weight of the text feature (90%) with the second highest matching degree is fine-tuned by 0.3 + 0.1 × (α × 0.8) ≈ 0.35. The weight of the voice feature is decreased accordingly by 0.4 - (0.07 + 0.05) = 0.28. Finally, the second weight is obtained as: text 0.35, face 0.37, voice 0.28 (total of 1).
[0077] Measuring matching degree by feature overlap provides a basis for weight adjustment; dynamically allocating weights based on matching degree highlights the contribution of high-matching modalities (e.g., increasing weight when facial feature matching degree is highest), solving the problem that fixed weights cannot respond to real-time feature differences and improving the adaptability of weights to the current emotional expression. Weight calibration values are obtained by performing weight calibration based on the acquisition quality of the call text information, facial expression information, and voice information.
[0078] Specifically, the acquisition quality is first evaluated, including by assessing the signal-to-noise ratio (SNR=25dB) of the acquired sound information. If it is below 30dB, it is judged as "medium quality" and its weight needs to be reduced; if the speech-to-text accuracy is 95% (error rate <5%), it is judged as "excellent quality" and its weight is maintained; if the facial feature point detection success rate is 98% (occlusion <2%), it is judged as "excellent quality" and its weight is maintained.
[0079] Secondly, quality calibration rules are set. For example, the modal weight of "medium quality" is reduced by 0.05, while "excellent quality" remains unchanged. The weight of voice information (original second weight 0.28) is reduced by 0.05, resulting in a calibrated voice weight of 0.23. The weights of text and face are adjusted accordingly to keep the sum of 1: text 0.35 + 0.03 = 0.38, face 0.37 + 0.02 = 0.39. The final calibrated weight values are: text 0.38, face 0.39, and voice 0.23.
[0080] The first weight, the second weight, and the weight calibration value are input into the weight calibration model to obtain the output weight. The weight calibration model is trained on a preset model based on the historical first weight, the historical second weight, the historical weight calibration value, and the optimal weight. The historical first weight is a preset weight division combination corresponding to different emotion type labels. The historical second weight is adjusted after matching the historical multimodal features with the typical features. The historical weight calibration value is obtained based on the acquisition quality of the historical multimodal features. The optimal weight is the weight information manually labeled for each group of historical multimodal features.
[0081] Specifically, mainstream models such as random forest regression, deep neural networks (DNN), convolutional neural networks (CNN), or recurrent neural networks (RNN) are used as the preset models. The input features are the historical first weight, the historical second weight, and the historical weight calibration value. The output is the manually labeled optimal weight. The training data contains 1000 sets of historical samples (covering different emotion types and scenarios). Five-fold cross-validation is used to optimize the model parameters (e.g., 100 decision trees, maximum depth 8) so that the model prediction error (MSE) is ≤0.001. The first weight (0.3, 0.3, 0.4), the second weight (0.35, 0.37, 0.28), and the weight calibration value (0.38, 0.39, 0.23) are input into the trained random forest model. The model outputs the final weights through voting using ensemble decision trees: text 0.36, face 0.38, and voice 0.26 (summing up to 1), which is the output weight.
[0082] Optionally, generating a prompt interface based on the current emotion type and displaying it on the terminal interface includes: Set a corresponding cue color and cue effect for each current emotion type.
[0083] Specifically, such as Figure 2 As shown, when the current emotion type is "happy" or "joyful", the prompt color is a warm color such as orange or yellow, and the prompt effects include rapid flashing, jumping light effects, and lively and dynamic effects; when the current emotion type is "angry" or "warning", the prompt color is a warning color such as red or black, and the prompt effects include strong screen flashing, edge light bursting effects, and abrupt and strong dynamic effects; when the current emotion type is "depressed" or "sad", the prompt color is a soft and dark color such as gray or gray-blue, and the prompt effects include low-frequency flashing light effects and slowly fluctuating dynamic effects; when the current emotion type is "peaceful" or "calm", no prompt color or prompt effect is added.
[0084] The prompt interface is generated based on the current emotion type and the corresponding prompt color and prompt effect, and then displayed on the terminal interface.
[0085] Specifically, the Qt interface rendering engine can be used to build the prompt interface. The interface layout is a floating window in the upper right corner of the terminal interface (size 200×100 pixels, which does not obscure the main content of the call); the interface elements include: emotion type label (such as "current emotion: happy", font size 14pt, color consistent with the prompt color); dynamic effect layer (located below the label, occupying 60% of the interface area, and using OpenGL to implement light effect rendering); and brief feature prompt (such as "speech speed is fast: 3.2 words / second", font size 10pt, gray).
[0086] It should be noted that the prompt interface is set to "top-level display" through the terminal operating system's window management API (such as Windows' Win32 API) to ensure that it is not obscured by other windows; it is displayed when the real-time emotion score remains within a certain emotion type range (such as the excitement range of 85-100 points) for more than 10 seconds to avoid frequent switching caused by instantaneous fluctuations; and it is automatically hidden when the semantic communication scenario is "private conversation" (determined by the text keywords "confidential" or "privacy") to protect user privacy.
[0087] like Figure 3 As shown, an embodiment of the present invention provides a user emotion recognition device 300, comprising: Data processing module 310 is used to generate a target emotion change curve based on the historical video call records of the target person obtained in advance, and to perform trend analysis on the target emotion change curve to obtain emotion change habits; The weighting module 320 is used to acquire the real-time call information of the target person, and to dynamically weight the real-time call information based on the emotional change habits to obtain the output weight. The real-time call information includes call text information, facial expression information and voice information. The scoring module 330 is used to standardize the real-time call information and perform weighted fusion of the processed call text information, facial expression information and voice information according to the output weight to obtain a real-time emotion score. The display module 340 is used to obtain the current emotion type of the target person based on the real-time emotion score, generate a prompt interface based on the current emotion type, and display it on the terminal interface.
[0088] like Figure 4 As shown, an electronic device 400 provided in this embodiment of the invention includes a memory 410 and a processor 420; the memory 410 is used to store a computer program; the processor 420 is used to implement the user emotion recognition method as described above when the computer program is executed.
[0089] Alternatively, an electronic device 400 includes a memory 410 and a processor 420 coupled to the memory 410; the memory 410 is configured to store a computer program; and the processor 420 is configured to perform the following operations when the computer program is executed: Based on the historical video call records of the target person obtained in advance, a target emotion change curve is generated, and trend analysis is performed on the target emotion change curve to obtain emotion change habits; The real-time call information of the target person is obtained, and the real-time call information is dynamically weighted based on the emotional change habits to obtain the output weight. The real-time call information includes text information, facial expression information and voice information. The real-time call information is standardized, and the processed text information, facial expression information, and voice information are weighted and fused according to the output weights to obtain a real-time emotion score. The current emotion type of the target person is obtained based on the real-time emotion score, and a prompt interface is generated based on the current emotion type and displayed on the terminal interface.
[0090] This invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the user emotion recognition method described above.
[0091] Alternatively, a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the following operations: Based on the historical video call records of the target person obtained in advance, a target emotion change curve is generated, and trend analysis is performed on the target emotion change curve to obtain emotion change habits; The real-time call information of the target person is obtained, and the real-time call information is dynamically weighted based on the emotional change habits to obtain the output weight. The real-time call information includes text information, facial expression information and voice information. The real-time call information is standardized, and the processed text information, facial expression information, and voice information are weighted and fused according to the output weights to obtain a real-time emotion score. The current emotion type of the target person is obtained based on the real-time emotion score, and a prompt interface is generated based on the current emotion type and displayed on the terminal interface.
[0092] The present invention will now be described an electronic device 400 that can serve as a server or client of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. Electronic device 400 is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic device 400 can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0093] Electronic device 400 includes a computing unit that can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) or a computer program loaded from a storage unit into random access memory (RAM). The RAM may also store various programs and data required for device operation. The computing unit, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0094] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc. In this application, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of the present invention according to actual needs. Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units can be implemented in hardware or as software functional units.
[0095] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.
Claims
1. A user emotion recognition method, characterized in that, include: Based on the historical video call records of the target person obtained in advance, a target emotion change curve is generated, and trend analysis is performed on the target emotion change curve to obtain emotion change habits; The real-time call information of the target person is obtained, and the real-time call information is dynamically weighted based on the emotional change habits to obtain the output weight. The real-time call information includes text information, facial expression information and voice information. The real-time call information is standardized, and the processed text information, facial expression information, and voice information are weighted and fused according to the output weights to obtain a real-time emotion score. The current emotion type of the target person is obtained based on the real-time emotion score, and a prompt interface is generated based on the current emotion type and displayed on the terminal interface.
2. The user emotion recognition method according to claim 1, characterized in that, The process of generating a target emotional change curve based on pre-acquired historical video call records of the target individual, performing trend analysis on the target emotional change curve to obtain emotional change habits, includes: Based on Mehrabian's rule, the historical multimodal features in the historical video call records are weighted and fused to obtain a historical emotion score. The historical multimodal features include historical call text information, historical facial expression information, and historical voice information. Generate a target emotion change curve with time as the horizontal axis and the historical emotion score as the vertical axis; Identify abrupt change nodes in the historical emotion scores of the target emotion change curve, associate them with the historical video call records corresponding to the abrupt change nodes, and obtain historical emotion slices; Based on the historical emotion scores, each historical emotion slice is labeled with an emotion type according to the emotion judgment criteria, and typical scenes and typical features under each emotion type are obtained based on the historical multimodal features corresponding to the same emotion type. Based on the typical scenarios, typical characteristics, and the emotion types, the emotion change habits are obtained.
3. The user emotion recognition method according to claim 2, characterized in that, The step of identifying abrupt change nodes in the historical emotion score within the target emotion change curve and associating them with historical video call records corresponding to those nodes includes: Set a mutation threshold, which is the maximum allowable fluctuation value of the historical emotion score within a preset continuous time window in the target emotion change curve; The target emotion change curve is traversed by sliding through the continuous time window to obtain the historical fluctuation value of the historical emotion score within each continuous time window. When the historical fluctuation value exceeds the mutation threshold, the start time of the continuous time window is marked as the mutation node of the historical emotion score. In the historical video call records, slices are obtained by extending a preset time before and after the mutation node, and the slices are associated with the corresponding historical video call records to obtain historical emotion slices.
4. The user emotion recognition method according to claim 3, characterized in that, Based on the historical emotion scores, each historical emotion slice is labeled with an emotion type according to the emotion determination criteria. Typical scenarios and features for each emotion type are obtained based on the historical multimodal features corresponding to the same emotion type, including: Obtain the historical emotion score corresponding to the historical emotion slice, and assign the emotion type label to each historical emotion slice according to the emotion judgment criteria, wherein the emotion type label includes happiness, anger, calmness and sadness; The K-means clustering algorithm is used to cluster the historical multimodal features under the same sentiment type label to obtain multiple feature clusters; Statistical analysis is performed on multiple feature clusters to obtain the typical scenarios and typical features; The typical scenarios and typical features are associated and stored with the corresponding emotion type labels to obtain emotion change habits.
5. The user emotion recognition method according to claim 4, characterized in that, The statistical analysis of multiple feature clusters to obtain the typical scenario and the typical features includes: For each of the feature clusters, the corresponding historical call scenarios are extracted, the frequency of occurrence of the same historical call scenarios is counted, and the historical call scenarios with a frequency exceeding a preset frequency threshold are identified as the typical scenarios. The typical features are obtained by analyzing the historical multimodal features in each feature cluster, wherein the typical features include typical text features, typical facial expression features, and typical voice features.
6. The user emotion recognition method according to claim 2, characterized in that, The step of dynamically weighting the real-time call information based on the emotional change habits to obtain output weights includes: Based on the text information, facial expression information and voice information of the call, the real-time call scene is extracted, and the first weight is obtained by pre-defined weight division and combination of the emotion type label associated with the typical scene that is the same as the real-time call scene; The text information of the call, the facial expression information, and the voice information are matched with the typical features to obtain the corresponding matching degree. The emotion type label corresponding to the matching degree being greater than the preset matching threshold is obtained. The weight adjustment value is obtained based on the matching degree, the preset matching threshold, and the preset adjustment coefficient. The second weight is obtained by dividing and combining the weight adjustment value and the preset weight corresponding to the emotion type label. Weight calibration values are obtained by performing weight calibration based on the acquisition quality of the call text information, facial expression information, and voice information. The first weight, the second weight, and the weight calibration value are input into the weight calibration model to obtain the output weight. The weight calibration model is trained on a preset model based on the historical first weight, the historical second weight, the historical weight calibration value, and the optimal weight. The historical first weight is a preset weight division combination corresponding to different emotion type labels. The historical second weight is adjusted after matching the historical multimodal features with the typical features. The historical weight calibration value is obtained based on the acquisition quality of the historical multimodal features. The optimal weight is the weight information manually labeled for each group of historical multimodal features.
7. The user emotion recognition method according to claim 1, characterized in that, The step of generating a prompt interface based on the current emotion type and displaying it on the terminal interface includes: Set a corresponding cue color and cue effect for each of the current emotion types; The prompt interface is generated based on the current emotion type and the corresponding prompt color and prompt effect, and then displayed on the terminal interface.
8. A user emotion recognition device, characterized in that, include: The data processing module is used to generate a target emotion change curve based on the target person's historical video call records obtained in advance, and to perform trend analysis on the target emotion change curve to obtain emotion change habits. The weighting module is used to acquire the real-time call information of the target person, and to dynamically weight the real-time call information based on the emotional change habits to obtain the output weight. The real-time call information includes call text information, facial expression information and voice information. The scoring module is used to standardize the real-time call information, and to perform weighted fusion of the processed call text information, facial expression information and voice information according to the output weights to obtain a real-time emotion score. The display module is used to obtain the current emotion type of the target person based on the real-time emotion score, generate a prompt interface based on the current emotion type, and display it on the terminal interface.
9. An electronic device, characterized in that, Including memory and processor; The memory is used to store computer programs; The processor is configured to implement the user emotion recognition method as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the user emotion recognition method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Digital human posture adjusting method and device combined with image recognition
CN119473209A
Emotion recognition method based on large model and related device
CN119904901A
User emotion recognition method based on AI and voice data
CN120148561A
Intelligent customer life cycle management AiCRM method and system
CN120494834A
Sale partner training method and system based on GPT and multi-modal large model
CN120672533A
Cited By
Emotion recognition method and device in video scene, equipment and medium
CN121421538A