A user emotion recognition method and device, electronic equipment and storage medium

By analyzing the target person's historical video call records, generating an emotion change curve, and performing dynamic weight adjustment and standardization, the accuracy and stability issues of emotion recognition in video calls are solved. This enables personalized emotion recognition and intuitive presentation, improving the practicality of video calls and intelligent customer service.

CN120951293BActive Publication Date: 2025-12-16CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511468412.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2025-12-16
Estimated Expiration
2045-10-15

AI Technical Summary

Technical Problem

Existing video call emotion recognition technologies are affected by environmental noise, individual differences, and the diversity of emotional expressions, resulting in insufficient accuracy and stability, making it difficult to fully capture the richness and complexity of emotional expressions.

Method used

By mining the target person's historical video call records to generate an emotion change curve, extracting emotion change habits, dynamically adjusting and standardizing real-time call information based on emotion change habits, performing multimodal information weighted fusion, generating a real-time emotion score and displaying a prompt interface.

Benefits of technology

It improves the accuracy and stability of emotion recognition, makes the recognition results more closely match individual characteristics, enhances the flexibility and reliability of multimodal information fusion, provides intuitive communication strategy guidance, and improves the practicality of scenarios such as video calls and intelligent customer service.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120951293B_ABST
    Figure CN120951293B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data analysis, and provides a user emotion recognition method and device, electronic equipment and a storage medium, the method comprising the following steps: generating a target emotion change curve according to a historical video call record of a target person acquired in advance, performing trend analysis on the target emotion change curve to obtain an emotion change habit; obtaining real-time call information of the target person, performing dynamic weight division on the real-time call information based on the emotion change habit to obtain an output weight; performing standardization processing on the real-time call information, performing weighted fusion on the processed real-time call information according to the output weight to obtain a real-time emotion score; obtaining a current emotion type of the target person according to the real-time emotion score, and generating a prompt interface according to the current emotion type and displaying the prompt interface on a terminal interface. The application significantly improves the precision, robustness and practical value of emotion recognition, can be widely applied to video call, intelligent customer service and other scenes, and improves the human-computer interaction experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data analysis, in particular to a user emotion recognition method and device, electronic equipment and storage medium. BACKGROUND

[0002] With the continuous development of mobile communication technology, users have higher requirements for the quality and interactive experience of video calls. Among them, accurately obtaining and understanding the emotional state of both parties in the call is of great significance in improving communication efficiency and user satisfaction. Through emotion recognition, the call system can better perceive the psychological dynamics of users, thereby promoting more natural and effective communication.

[0003] Currently, emotion recognition technology relies on speech signal processing methods, which extract features such as intonation, speech rate, pitch, and sound intensity, and combine machine learning algorithms to judge the emotions of the speaker. However, due to the influence of environmental noise, individual differences of the speaker, and the diversity of emotional expression, the accuracy of emotion judgment shows certain volatility. In addition, the richness and complexity of emotional expression cannot be fully captured through speech signals, affecting the stability and reliability of the overall recognition effect. SUMMARY

[0004] The problem solved by the present application is how to improve the accuracy of character emotion recognition in video calls.

[0005] To solve the above problems, the present application provides a user emotion recognition method, device, electronic equipment and storage medium.

[0006] In a first aspect, the present application provides a user emotion recognition method, comprising:

[0007] Generating a target emotion change curve according to the pre-acquired historical video call records of the target character, performing trend analysis on the target emotion change curve, and obtaining an emotion change habit;

[0008] Obtaining real-time call information of the target character, and based on the emotion change habit, performing dynamic weight division on the real-time call information to obtain an output weight, wherein the real-time call information includes call text information, facial expression information and sound information;

[0009] Standardizing the real-time call information, and weighting and fusing the processed call text information, facial expression information and sound information according to the output weight to obtain a real-time emotion score;

[0010] According to the real-time emotion score, the current emotion type of the target character is obtained, and a prompt interface is generated according to the current emotion type and displayed on the terminal interface.

[0011] Optionally, the target emotion change curve is generated according to a historical video call record of a target person, the target emotion change curve is trend analyzed, and an emotion change habit is obtained, comprising:

[0012] The historical multi-modal features in the historical video call record are weighted and fused based on the Melara rule to obtain a historical emotion score, wherein the historical multi-modal features include historical call text information, historical facial expression information, and historical sound information;

[0013] The target emotion change curve is generated with time as the horizontal axis and the historical emotion score as the vertical axis;

[0014] A mutation node of the historical emotion score in the target emotion change curve is identified, a historical video call record corresponding to the mutation node is associated, and a historical emotion slice is obtained;

[0015] According to the historical emotion score, each historical emotion slice is labeled with an emotion type according to an emotion judgment standard, and a typical scene and a typical feature under each emotion type are obtained according to the historical multi-modal features corresponding to the same emotion type;

[0016] The emotion change habit is obtained according to the typical scene, the typical feature, and the emotion type.

[0017] Optionally, the mutation node of the historical emotion score in the target emotion change curve is identified, and the historical video call record corresponding to the mutation node is associated, comprising:

[0018] A mutation threshold is set, and the mutation threshold is a maximum allowed fluctuation value of the historical emotion score in a preset continuous time window in the target emotion change curve;

[0019] The target emotion change curve is traversed in the continuous time window, a historical fluctuation value of the historical emotion score in each continuous time window is obtained, and when the historical fluctuation value exceeds the mutation threshold, a starting time of the continuous time window is marked as a mutation node of the historical emotion score;

[0020] In the historical video call record, a slice is obtained by extending a preset time before and after the mutation node, the slice is associated with the corresponding historical video call record, and a historical emotion slice is obtained.

[0021] Optionally, according to the historical emotion score, each historical emotion slice is labeled with an emotion type according to an emotion judgment standard, and a typical scene and a typical feature under each emotion type are obtained according to the historical multi-modal features corresponding to the same emotion type, comprising:

[0022] obtaining the historical emotion scores corresponding to the historical emotion slices, and assigning the emotion type label to each of the historical emotion slices according to the emotion determination standard, wherein the emotion type label includes happiness, anger, calmness, and sadness;

[0023] adopting a K-means clustering algorithm to cluster the historical multi-modal features under the same emotion type label, to obtain a plurality of feature clustering clusters;

[0024] statistically analyzing the plurality of feature clustering clusters to obtain the typical scene and the typical feature;

[0025] storing the typical scene and the typical feature in association with the corresponding emotion type label to obtain an emotion change habit.

[0026] Optionally, the statistically analyzing the plurality of feature clustering clusters to obtain the typical scene and the typical feature includes:

[0027] for each of the feature clustering clusters, extracting a corresponding historical call scene, counting the occurrence frequency of the same historical call scene, and determining the historical call scene with an occurrence frequency exceeding a preset frequency threshold as the typical scene;

[0028] analyzing the historical multi-modal features in each of the feature clustering clusters to obtain the typical feature, wherein the typical feature includes a typical text feature, a typical facial expression feature, and a typical sound feature.

[0029] Optionally, the dynamically dividing the real-time call information based on the emotion change habit to obtain an output weight includes:

[0030] extracting a real-time call scene from the call text information, the facial expression information, and the sound information, obtaining a preset weight division combination corresponding to the emotion type label associated with the same typical scene as the real-time call scene to obtain a first weight;

[0031] matching the call text information, the facial expression information, and the sound information with the typical feature to obtain a corresponding matching degree, obtaining the emotion type label corresponding to the matching degree greater than a preset matching threshold, obtaining a weight adjustment value according to the matching degree, the preset matching threshold, and a preset adjustment coefficient, and obtaining a second weight based on the weight adjustment value and a preset weight division combination corresponding to the emotion type label;

[0032] weighting the call text information, the facial expression information, and the sound information according to the collection quality to obtain a weight calibration value;

[0033] inputting the first weight, the second weight and the weight calibration value into a weight calibration model to obtain the output weight, wherein the weight calibration model is obtained by training a preset model according to historical first weights, historical second weights, historical weight calibration values and optimal weights, the historical first weights are obtained by dividing and grouping preset weights corresponding to different emotion type labels, the historical second weights are obtained by adjusting the historical multi-modal features after matching the historical multi-modal features with the typical features, the historical weight calibration values are obtained according to the collection quality of the historical multi-modal features, and the optimal weights are weight information manually labeled for each group of the historical multi-modal features.

[0034] Optionally, the generating and displaying the prompt interface according to the current emotion type on the terminal interface comprises:

[0035] setting a corresponding prompt color and prompt effect for each current emotion type;

[0036] generating the prompt interface according to the current emotion type, the corresponding prompt color and the prompt effect, and displaying the prompt interface on the terminal interface.

[0037] In a second aspect, the present application provides a user emotion recognition device, comprising:

[0038] a data processing module configured to generate a target emotion change curve according to historical video call records of a target person, and perform trend analysis on the target emotion change curve to obtain an emotion change habit;

[0039] a weight division module configured to obtain real-time call information of the target person, and perform dynamic weight division on the real-time call information based on the emotion change habit to obtain an output weight, wherein the real-time call information comprises call text information, facial expression information and voice information;

[0040] a scoring module configured to perform standardization processing on the real-time call information, and perform weighted fusion on the processed call text information, facial expression information and voice information according to the output weight to obtain a real-time emotion score;

[0041] a display module configured to obtain a current emotion type of the target person according to the real-time emotion score, and generate a prompt interface according to the current emotion type and display the prompt interface on a terminal interface.

[0042] In a third aspect, the present application provides an electronic device comprising a memory and a processor.

[0043] The memory is configured to store a computer program.

[0044] The processor is configured to implement the user emotion recognition method according to the first aspect when executing the computer program.

[0045] In a fourth aspect, the present application provides a computer readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the user emotion recognition method according to the first aspect is implemented.

[0046] The user emotion recognition method of the present application has the following advantages: by mining the historical video call records of the target person, generating a target emotion change curve and extracting emotion change habits, the basis for emotion recognition is changed from "generalization" to "personalization", improving the individual adaptability of existing general models, providing accurate historical references for subsequent real-time recognition, and improving the pertinence of emotion recognition. Based on the emotion change habits, the dynamic weight adjustment of real-time multi-modal information is performed, for example, the weight of text information is increased when matching the emotion fluctuation related topics, and the weight of the corresponding mode is increased when the feature matching degree is high, so that the weight distribution is more suitable for real-time complex scenes and user expression habits, and the flexibility and reliability of multi-modal information fusion are enhanced. By standardizing the different modal features to a unified scale (such as the 0-1 interval), the interference caused by the unit and magnitude difference is eliminated, and then the dynamic output weight is combined for weighted fusion, so that the real-time emotion score can integrate the effective information of each mode, and the accuracy and stability of the score are improved. Based on the real-time emotion score, the specific emotion type is determined, and a dedicated prompt interface is generated, such as an interface containing dedicated prompt elements (such as color, icon), so that the user or the communication object can perceive the emotion state in real time, solve the problem of "application scene ambiguity" of emotion recognition results, and help to optimize the communication strategy (such as adjusting the topic for anxiety emotion), and improve the practicality. Through the complete process of "personalized habit mining-dynamic weight adaptation-standardization fusion-visualization presentation", the user's historical behavior rules are deeply combined with real-time scene features, realizing the upgrade from "passive recognition" to "active adaptation", significantly improving the precision, robustness and practical value of emotion recognition, and can be widely applied to video calls, intelligent customer service and other scenes, and improving the human-computer interaction experience. BRIEF DESCRIPTION OF DRAWINGS

[0047] Figure 1 FIG. 1 is a flowchart of a user emotion recognition method according to an embodiment of the present application;

[0048] Figure 2 FIG. 4 is a schematic diagram of a prompt interface according to an embodiment of the present application;

[0049] Figure 3 FIG. 5 is a structural schematic diagram of a user emotion recognition device according to an embodiment of the present application;

[0050] Figure 4 FIG. 6 is a structural schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0051] In order to make the above objectives, characteristics and advantages of the present application more apparent, specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, but rather, these embodiments are provided so as to more completely and thoroughly understand the present application. It should be understood that the drawings and embodiments of the present application are for illustrative purposes only, and are not intended to limit the scope of protection of the present application.

[0052] It should be understood that each of the steps described in the method embodiments of the present application can be performed in different orders, and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present application is not limited in this respect.

[0053] As used herein, the term "includes" and its variants are open-ended, meaning "includes but is not limited to"; the term "based on" is "based, at least in part, on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; the term "optionally" means "optional embodiments." Related definitions are given throughout the description. It should be noted that the concepts mentioned in the present application are merely used to distinguish different devices, modules or units, and are not intended to limit the functions performed by these devices, modules or units in order or interdependence.

[0054] It should be noted that the modification of "one" or "multiple" mentioned in the present application is illustrative rather than limiting, and those skilled in the art should understand that, unless otherwise explicitly indicated in the context, it should be understood as "one or more".

[0055] The names of the messages or information exchanged between the devices in the embodiments of the present application are only for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0056] It can be understood that any part related to data acquisition or collection involved in the present application has obtained the authorization of the user.

[0057] To solve the problems in the related art, the present embodiment provides a user emotion recognition method and device, electronic equipment and storage medium.

[0058] As shown in Figure 1 A user emotion recognition method provided by an embodiment of the present application includes:

[0059] Step S1, generating a target emotion change curve according to a pre-acquired historical video call record of a target person, performing trend analysis on the target emotion change curve to obtain an emotion change habit.

[0060] Specifically, the historical video call record includes historical sound features, historical text features and historical facial expression features used to generate the target emotion change curve. The historical sound features are specifically subdivided into historical intonation change slope, historical speech speed fluctuation frequency, historical pitch peak, historical sound intensity mutation duration and historical noise condition, etc. For example, the historical intonation change slope (difference between the initial and final pitch values / 2 seconds) in each 2-second window is extracted by Fourier transform, the historical speech speed fluctuation frequency (standard deviation of the duration of adjacent 5 syllables) of each sentence is counted, the historical pitch peak (the highest fundamental frequency of each call) is recorded, the historical sound intensity mutation duration (duration of sudden increase of sound pressure level from 60 dB to above 80 dB) is recorded, and the historical background noise level (divided into no noise <20 dB, low noise 20-40 dB, and high noise >40 dB) is labeled by a noise separation algorithm; the historical text features are specifically subdivided into historical emotion tendency word occurrence frequency, historical semantic emotion tendency value and historical semantic communication scene. For example, the historical emotion tendency word occurrence frequency (number of positive / negative words in every 50 characters) is counted based on an emotion dictionary (such as the HowNet emotion dictionary), the historical semantic emotion tendency value (range 0-1, 0 for extreme negative and 1 for extreme positive) is calculated by a bidirectional LSTM model, and the historical semantic communication scene (such as “academic discussion”, “daily trivia” and “business negotiation”) is identified by a topic model (such as LDA); the historical facial expression features are specifically subdivided into historical facial feature shape change and historical facial muscle movement rate. For example, the MTCNN algorithm is used to locate the key points of facial features, the historical facial feature shape change (such as eye corner opening degree and nose wing expansion rate) is calculated, and the historical facial muscle movement rate (average displacement speed of key points in consecutive 20 frames, unit: pixel / millisecond) is calculated by optical flow method.

[0061] After obtaining the relevant features, principal component analysis (PCA) can be used to reduce the dimensionality of the historical multi-modal features, and the principal components with cumulative contribution rate >85% are selected, which are mapped to historical emotion scores (0-100) by a support vector regression (SVR) model, and a target emotion change curve is generated with a time granularity of 10 seconds.

[0062] Based on the obtained target emotion change curve, the emotion change habit is analyzed, where the emotion change habit can be a related topic and a characteristic expression pattern affecting the emotion. The curve can be subjected to multi-scale decomposition by wavelet transform to identify time nodes of emotion score mutation (e.g., set the mutation criterion to be a score change > 30 points within 10 seconds), associate the nodes with semantic communication scenes, and obtain the associated topics of emotion fluctuation (e.g., "emotion is easy to fluctuate when 'cost' is mentioned in business negotiation"). The historical characteristics of the same emotion type (e.g., "anger") are subjected to hierarchical clustering algorithm to extract the characteristic expression pattern (e.g., the pitch change slope when angry > 40 Hz / s, the eye corner opening degree < 30%, and the negative vocabulary frequency > 8 times / 50 words).

[0063] In step S2, real-time communication information of the target person is obtained, and the real-time communication information is subjected to dynamic weight division based on the emotion change habit to obtain an output weight, where the real-time communication information includes communication text information, facial expression information, and sound information.

[0064] Specifically, the communication text information is specifically subdivided into emotional tendency vocabulary appearance frequency, semantic emotional tendency value, and semantic communication scene. The text is generated by a real-time speech-to-text engine (e.g., Google Speech-to-Text) during real-time communication, and the specific feature acquisition method is the same as the above-mentioned historical text feature acquisition method, which will not be repeated here. The facial expression information is specifically subdivided into facial feature change and facial muscle movement rate, and the acquisition method is the same as the above-mentioned historical facial expression feature acquisition method, which will not be repeated here. The sound information is specifically subdivided into pitch change slope, speech speed fluctuation frequency, sound height peak value, and sound intensity mutation length noise condition, and the acquisition method is the same as the above-mentioned historical sound feature acquisition method, which will not be repeated here.

[0065] After obtaining the real-time communication information, the initial weights are first set, for example, the initial weight of the text information is 0.25, the initial weight of the facial expression information is 0.35, and the initial weight of the sound information is 0.4, reflecting the difference in the amount of information of the historical characteristics. The obtained real-time communication information is matched with the emotional change habits such as the associated topics and the characteristic expression patterns, respectively, and the weights are adjusted according to the matching results to obtain the final output weights. For example, the cosine similarity of the real-time communication information and the associated topics is calculated, and the Euclidean distance of the real-time communication information and the characteristic expression pattern is calculated. If the matching degree of the real-time semantic communication scene and the associated topics of the emotional fluctuation is greater than a preset threshold 0.8, the corresponding modal weight is increased by 0.15 (for example, if the text information has a high correlation degree, the text weight is increased to 0.4, and other characteristics are adaptively adjusted). When the Euclidean distance of the real-time communication information and the characteristic expression pattern is less than a preset threshold 0.3, the corresponding modal weight is increased by 0.1 (for example, if the facial feature matching degree is high, the facial weight is increased to 0.45, and other characteristics are adaptively adjusted). On this basis, the linear programming model is used to integrate and adjust the values to ensure that the sum of the weights is 1, and finally the output weights are obtained.

[0066] Step S3, standardizing the communication text information, the facial expression information and the sound information, and weighting and fusing the processed communication text information, the facial expression information and the sound information according to the output weights to obtain a real-time emotion score.

[0067] Specifically, since the real-time call information is multi-source characteristics, in order to ensure more accurate data processing, the real-time call information needs to be standardized before weighted full fusion. For audio information, the pitch change slope can be standardized by using cumulative distribution function (CDF). The sample set of the pitch change slope in the historical voice audio data is collected, and the cumulative distribution function is calculated. For the real-time pitch change slope value, the corresponding cumulative probability value (range 0-1) is determined by looking up the historical CDF. The value is the standardization result. If the real-time value exceeds the historical sample range (such as higher than the maximum value), the boundary value is processed (i.e. the standardized value = 1); the speech rate fluctuation frequency can be standardized by using the logarithmic transformation and min-max normalization method. The natural logarithmic transformation is performed on the real-time speech rate fluctuation frequency (formula: log(x+1), x is the real-time value), which relieves the right-biased distribution of the data. The min-max normalization is used on the transformed value: standardized value = (transformed value-historical minimum transformed value) / (historical maximum transformed value-historical minimum transformed value), which is mapped to the 0-1 interval; the pitch peak value can be standardized by using quantile standardization. The 99% quantile value of the pitch peak value is determined based on the historical data as the upper threshold (such as 500Hz). If the real-time pitch peak value is less than or equal to the upper threshold, the standardized value = real-time value / upper threshold. If the real-time value is greater than the upper threshold, the standardized value = 1 (to avoid the interference of extreme values on the overall scale); the sound intensity mutation length can be standardized by using threshold truncation. The standardized value = real-time sound intensity mutation length / threshold. If the real-time value is greater than the threshold, the standardized value = 1. If the real-time value is less than 0 (invalid value), the standardized value = 0. For text information, the frequency of emotional tendency words can be standardized by using Z-score standardization and truncation processing. The mean (μ) and standard deviation (σ) of the frequency of emotional tendency words in the historical data are calculated. The standardized value = (real-time frequency-μ) / σ. The result is truncated, i.e. if the standardized value is greater than 3, it is processed as 3. If it is less than -3, it is processed as -3. Finally, it is mapped to the 0-1 interval: standardized value = (truncated value+3) / 6; the semantic sentiment value can be standardized by using linear offset standardization. The original range is set to -1 to 1. -1 is extreme negative and 1 is extreme positive. The original range is directly mapped to the 0-1 interval by linear transformation. The expression is: standardized value = (original semantic sentiment value+1) / 2. For example: original value =-1 corresponds to 0, original value = 0 corresponds to 0.5, and original value = 1 corresponds to 1; the semantic communication scene judgment is standardized by using one-hot encoding. Based on the historical data, all possible semantic communication scenes (such as 5 scenes) are defined to form a scene set. For the real-time recognized semantic communication scene, a binary vector with a length equal to the number of scene sets is generated. The position corresponding to the scene is 1 and the rest is 0. For example, the "work" scene corresponds to the vector [1,0,0,0,0] and the "family" corresponds to [0,1,0,0,0].For facial expression information, the facial feature morphological change can be standardized by relative deviation, taking the facial feature morphological value of the target person in a calm state as the reference value (e.g. the natural state angle of the mouth corner = 0°), calculating the deviation (Δ = real-time value - reference value) of the real-time morphological value from the reference value, determining the maximum positive deviation (Δmax+) and the maximum negative deviation (Δmax-) based on historical data, if Δ ≥ 0: standardized value = Δ / Δmax+ (mapped to 0-1, positive change), if Δ < 0: standardized value = 1-(Δ / Δmax-) (mapped to 0-1, negative change); the facial muscle movement rate can be standardized by an exponential function, and the saturation threshold of the muscle movement rate (e.g. 10 pixels / frame, after which the emotional expression intensity no longer increases significantly) is determined based on historical data, and the exponential function is used for mapping: standardized value = 1-e^(-k x real-time rate), where k is an adjustment parameter (based on historical data fitting, e.g. k = 0.3), which is sensitive to low-rate changes (distinguishing subtle expressions) and tends to be gentle for high-rate changes (avoiding over-amplification).

[0068] The weighted sum of each real-time call information after standardization is calculated, such as text feature score = 0.6 x sentiment vocabulary score + 0.4 x semantic tendency score, and then weighted fusion is performed according to the output weight, such as real-time emotional score = text total score x 0.35 + facial total score x 0.4 + sound total score x 0.25, to obtain the final real-time emotional score.

[0069] Step S4, obtaining the current emotional type of the target person according to the real-time emotional score, and generating a prompt interface according to the current emotional type and displaying it on the terminal interface.

[0070] Specifically, fuzzy logic reasoning is used, and membership functions of score intervals and emotional types are set, such as 0-25 points belong to "angry" (0.8), "sad" (0.2); 26-45 points belong to "anxious" (0.7), "calm" (0.3); 46-75 points belong to "calm" (0.9); 76-100 points belong to "happy" (0.8), "excited" (0.2), and the final type is determined in combination with the matching result of real-time features and typical patterns (e.g. "happy" requires a mouth corner up angle > 15°).

[0071] Dynamic prompt elements are configured for different emotional types: "angry" displays a red dynamic warning box; "happy" displays a yellow gradient background; "calm" displays a blue static box. Real-time display on the terminal interface (e.g. sidebar of a video call software), including emotional type label and feature prompt (e.g. "peak abnormality of sound height detected, current emotion: "angry").

[0072] In the embodiment, by mining the historical video call records of the target person, generating the target emotion change curve and extracting the emotion change habit, the basis for emotion recognition is changed from "generalization" to "personalization", the adaptability of the existing general model to individuals is improved, accurate historical reference is provided for subsequent real-time recognition, and the pertinence of emotion recognition is improved. Based on the emotion change habit, the dynamic weight adjustment of real-time multi-modal information is carried out, for example, the weight of text information is increased when matching the emotion fluctuation related topic, and the weight of the corresponding mode is increased when the feature matching degree is high, so that the weight distribution is more suitable for real-time complex scene and user expression habit, and the flexibility and reliability of multi-modal information fusion are enhanced. Through standardization processing, different modal features are mapped to a unified scale (such as 0-1 interval), and the interference caused by unit and magnitude difference is eliminated, and then combined with dynamic output weight for weighted fusion, so that the real-time emotion score can comprehensively consider the effective information of each mode, and the accuracy and stability of the score are improved. Based on the real-time emotion score, the specific emotion type is determined, and a dedicated prompt interface is generated, such as an interface containing a dedicated prompt element (such as color, icon), so that the user or communication object can realize real-time emotion state, solve the problem of "application scene ambiguity" of emotion recognition result, help to optimize the communication strategy (such as adjusting the topic for anxiety emotion), and improve the practicability. Through the complete process of "personalized habit mining-dynamic weight adaptation-standardization fusion-visualization presentation", the user's historical behavior rule is deeply combined with real-time scene features, from "passive recognition" to "active adaptation", the accuracy, robustness and practical value of emotion recognition are significantly improved, which can be widely applied to video calls, intelligent customer service and other scenes, and improve the human-computer interaction experience.

[0073] Optionally, the target emotion change curve is generated according to the pre-acquired historical video call records of the target person, the target emotion change curve is trend analyzed, and the emotion change habit is obtained, including:

[0074] The historical multi-modal features in the historical video call records are weighted and fused based on the Melabbin rule to obtain a historical emotion score, wherein the historical multi-modal features include historical call text information, historical facial expression information and historical sound information.

[0075] Specifically, the core of the Mehrabian rule is the fixed weight ratio of "7% language content + 38% voice tone + 55% facial expression". When the emotional change habit of the target person is not clear, the historical multimodal features in the historical video call record are weighted and fused by using the commonly used Mehrabian rule to generate a historical emotion score, historical emotion score = (weighted sum of text features x 7%) + (weighted sum of voice audio features x 38%) + (weighted sum of facial expression features x 55%), which is mapped to 0-100 points. Among them, the internal weighting of each modality feature can refer to the preset standard. In the voice audio feature, the tone change slope accounts for 30%, the speech speed fluctuation frequency accounts for 20%, the pitch peak accounts for 20%, the sound intensity mutation time accounts for 20%, and the noise accounts for 10%. In the text feature, the emotional vocabulary frequency accounts for 40%, the semantic sentiment tendency value accounts for 40%, and the semantic communication scene accounts for 20%. In the facial expression feature, the facial feature change accounts for 0%, and the muscle movement rate accounts for 40%.

[0076] By using the Mehrabian rule, the universality of the rule to human emotional expression rules is retained, so that when the emotional change habit of the target person is unknown, the historical emotion score obtained is more consistent with the real emotional state of the target person.

[0077] Generate the target emotion change curve with time as the horizontal axis and the historical emotion score as the vertical axis.

[0078] Specifically, the linear interpolation method can be used to connect the score values of each window to generate a continuous target emotion change curve, with the horizontal axis being the cumulative time (unit: seconds) after the call starts and the vertical axis being the historical emotion score (0-100). At the same time, the local extreme points (peak value, valley value) of the curve can be labeled, which is convenient for subsequent analysis.

[0079] Identify the mutation node of the historical emotion score in the target emotion change curve, and associate the historical video call record corresponding to the mutation node to obtain a historical emotion slice.

[0080] Specifically, a judgment condition is set, such as the absolute difference of the historical emotion score in two consecutive time windows (10 seconds in total) > 25 points, and the change rate of the score in the latter window is > 50% (such as from 40 points to 70 points, the change rate is 75%) compared with the former window. The sliding window is used to traverse the curve, and when the above condition is met, the start time of the latter window is marked as a mutation node. A historical video call segment (containing voice, text, and expression data) of, for example, 5 seconds forward and, for example, 2 seconds backward is intercepted with the mutation node as the center, forming a historical emotion slice with a length of 7 seconds. At the same time, the corresponding timestamp, historical emotion score, and mutation direction (rapid rise / rapid fall) of the slice can also be labeled.

[0081] According to the historical emotion scores, an emotion type is labeled for each of the historical emotion slices according to an emotion determination criterion, and a typical scene and a typical feature are obtained for each emotion type according to the historical multi-modal features corresponding to the same emotion type.

[0082] According to the typical scene, the typical feature, and the emotion type, the emotion change habit is obtained.

[0083] Optionally, the identifying the mutation node of the historical emotion score in the target emotion change curve comprises:

[0084] A mutation threshold is set, and the mutation threshold is a maximum allowed fluctuation value of the historical emotion score in a preset continuous time window in the target emotion change curve.

[0085] Specifically, research shows that a significant change in human emotion is usually completed within 5-10 seconds, and 8 seconds is an optimal value for balancing sensitivity and stability. Therefore, based on the time characteristics of emotion expression in historical video calls, the length of the continuous time window is set to 8 seconds. The fluctuation values (i.e., the difference between the highest score and the lowest score in the window) in all 8-second windows of the historical emotion scores of the target person are collected to form a fluctuation value sample set, and the 90th percentile value of the sample set is set as a reference threshold. For example, if 90% of the fluctuation values in the sample set are less than or equal to 22 points, the reference threshold of 22 points is set as the mutation threshold, or the mutation threshold can be modified in combination with the emotion type. For example, for strong emotion types such as "anger" and "excitement", the threshold is increased by 10% (i.e., 24 points), and for weak emotion types such as "calm", the threshold is decreased by 10% (i.e., 20 points). Finally, the mutation threshold is set to 22 points.

[0086] The target emotion change curve is traversed according to the continuous time window, and a historical fluctuation value of the historical emotion score in each continuous time window is obtained. When the historical fluctuation value exceeds the mutation threshold, the start time of the continuous time window is marked as a mutation node of the historical emotion score.

[0087] Specifically, starting from the beginning of the target emotion change curve, sliding in continuous time windows (8 seconds) with a sliding step of 2 seconds (ensuring window overlap to avoid missing mutations), until the end of the curve. For each window, extract the maximum and minimum values of all historical emotion scores, calculate the absolute difference between the two, and obtain the historical fluctuation value. When the historical fluctuation value of the window > the mutation threshold of the corresponding emotion type (e.g. for a window corresponding to "anger" emotion, fluctuation value 25 > 24), mark the starting time point of the window as the mutation node of the historical emotion score, or combine the trend point and fluctuation value determination results within the window, for example, when the historical fluctuation value of the window > the mutation threshold of the corresponding emotion type, and the window contains at least 3 non-consecutive score rising / falling points (excluding single extreme value caused false judgment), mark the starting time point of the window as the mutation node of the historical emotion score, and record the fluctuation value, emotion type (based on the average score within the window) corresponding to the node. Through the sliding window and overlapping step design, full coverage scanning of the emotion curve is realized, and through the fluctuation value or combined fluctuation value and trend point double determination, real mutations and noise interference are effectively distinguished, improving the recall rate and accuracy of mutation node identification.

[0088] In the historical video call record, the slice is obtained by extending a preset time before and after the mutation node, and the slice is associated with the corresponding historical video call record to obtain a historical emotion slice.

[0089] Specifically, to obtain the context association analysis of emotion mutation, the mutation node is taken as the center, extending 20 seconds forward (covering the emotion brewing period before the mutation) and 25 seconds backward (covering the emotion duration period after the mutation), forming a video segment with a total duration of 45 seconds, and the slice is separated from the segment and retains complete multi-modal data. According to the metadata obtained from the historical call video record, including the corresponding mutation node timestamp, historical fluctuation value, and multi-modal information, the historical emotion slice is obtained.

[0090] Through the differentiated forward and backward extension duration, the "cause and effect" of the mutation (such as the slight rise in tone before the mutation and the expression persistence after the mutation) is completely retained, and the complete extraction of multi-modal information enables the historical emotion slice to not only contain score data but also cover voice, text, and expression features supporting emotion change, solving the context missing problem caused by too short slice duration, and providing rich information samples for subsequent typical scene and feature extraction.

[0091] Optionally, based on the historical emotion score, the emotion type of each historical emotion slice is labeled according to the emotion determination standard, and the typical scenes and typical features under each emotion type are obtained according to the historical multi-modal features corresponding to the same emotion type, including:

[0092] The historical emotion score corresponding to the historical emotion slice is obtained, and the emotion type label is given to each historical emotion slice according to the emotion determination standard, wherein the emotion type label includes happy, angry, calm and sad.

[0093] Specifically, based on the distribution characteristics of the historical emotion score, the corresponding relationship between the four rating intervals and the emotion types is divided: the historical emotion score of 75-100 corresponds to “happy” (characterized by dominant positive emotion, the higher the score, the stronger the pleasure); the historical emotion score of 50-74 corresponds to “calm” (stable emotion, no significant fluctuation); the historical emotion score of 25-49 corresponds to “sad” (dominated by negative emotion, accompanied by low energy characteristics); the historical emotion score of 0-24 corresponds to “angry” (strong negative emotion, accompanied by high energy characteristics). The average historical emotion score of each historical emotion slice is extracted, which can obtain the arithmetic mean of the scores of all time points in the historical emotion slice, and the emotion type label is directly matched according to the above interval, for example, a certain slice has an average score of 82, which is labeled as “happy”; an average score of 33 is labeled as “sad”.

[0094] Through the direct mapping of the score interval and the emotion type, the standardization of the emotion label is realized, avoiding the subjectivity of manual labeling; the four emotion types cover the core category of human basic emotions, and the association with the score interval conforms to the quantization rule of emotion intensity, solving the problem of disconnection between emotion type and score, and providing a unified classification benchmark for subsequent clustering analysis.

[0095] The historical multi-modal features under the same emotion type label are clustered by using a K-means clustering algorithm to obtain a plurality of feature clustering clusters.

[0096] Specifically, the historical multi-modal features under the same emotion type label (such as “angry”) are standardized, for example, the min-max method is used to map to the 0-1 interval, and the preprocessed features are integrated into a high-dimensional feature vector, such as each sample containing 10-dimensional features. The optimal number of clusters K is determined based on the silhouette coefficient method, such as K=3 under the “angry” emotion, that is, 3 feature clustering clusters, and the Euclidean distance is used as the similarity measure, and the cluster center is iteratively updated until the within-cluster sum of squares converges (such as the number of iterations ≤500 times), and each feature clustering cluster is output, and each cluster contains historical multi-modal feature samples with high feature similarity.

[0097] By K-means clustering, the multi-modal features under the same emotion type are divided into several sub-clusters, revealing different expression modes of the same emotion (such as “angry” which may be expressed as “loud accusation” or “silent suppression”), solving the problem of difficulty in extracting typical patterns caused by mixed features under a single emotion type, and laying a foundation for accurately mining typical scenes and features.

[0098] statistically analyzing the plurality of feature clustering clusters to obtain the typical scene and the typical feature.

[0099] Specifically, the typical scene can be a communication scene, a noise environment, etc. extracted from historical multi-modal features, such as extracting a semantic communication scene from text information, such as matching keywords to a preset scene library “product complaint” “salary negotiation” “parent-child education” and the like, and determining a noise scene according to a noise level, such as a high noise scene. For ambiguous scenes (such as ambiguous keywords), the context features of the voice audio information (such as “complaints” often accompanied by high speech speed and high sound intensity) are combined to assist in judgment, ensuring that the scene classification accuracy is ≥ 90%. The typical feature is a multi-modal feature of the target person under the corresponding emotion type, such as the typical voice feature when the target person is in anger: pitch change slope > 30 Hz / s, speech speed fluctuation frequency > 0.6 (coefficient of variation), peak pitch > 350 Hz, and sound intensity mutation duration > 1.2 seconds; typical text features: negative word frequency > 8 times / 200 words, semantic sentiment tendency value < -0.6; typical facial features: eyebrow peak downward displacement > 8 mm, mouth corner downward angle > 12°, and facial muscle movement rate > 5 pixels / frame.

[0100] The typical scene and the typical feature are associated with the corresponding emotion type label and stored to obtain an emotion change habit.

[0101] Specifically, a “emotion type-typical scene-typical feature” three-dimensional data table is constructed, such as emotion type “angry” → typical scene “contract dispute” → typical feature (pitch change slope 35-45 Hz / s, negative word proportion 30%, eyebrow peak downward pressure > 15°), each entry contains the statistical quantity of the feature, the scene proportion and the sample size (such as based on 50 historical emotion slices). In addition, when a new historical video call record is added, the typical scene and the feature statistics are re-clustered and updated (such as when the sample size increases to 60, the mean and quantile value are adjusted), ensuring that the emotion change habit is continuously optimized with data accumulation.

[0102] Through structured storage, the accurate association of emotion type, scene and feature is realized, so that the emotion change habit not only contains the rule of “what emotion is produced in what scene” (such as “contract disputes are prone to anger”), but also contains the characteristics of “how the emotion is expressed” (such as the specific performance of tone and expression), solving the problem of ambiguous emotion habit description and difficult application to real-time recognition, and providing a directly callable rule library for subsequent dynamic weight division.

[0103] Optionally, the statistically analyzing the plurality of feature clustering clusters to obtain the typical scene and the typical feature comprises:

[0104] For each of the feature clustering cluster, the corresponding historical call scene is extracted, the occurrence frequency of the same historical call scene is counted, and the historical call scene with an occurrence frequency exceeding a preset frequency threshold is determined as the typical scene.

[0105] Specifically, the occurrence frequency of each historical call scene in the cluster is counted (e.g., "product complaint" occurs 28 times in cluster 1, and "salary negotiation" occurs 8 times); a preset frequency threshold is set: based on the total amount of samples in the cluster (e.g., cluster 1 contains 50 samples), the threshold is 30% of the total amount of samples (i.e., 15 times); the scene with an occurrence frequency exceeding the threshold is determined as the typical scene of the cluster (e.g., "product complaint" becomes the typical scene of cluster 1 because 28 times > 15 times); if multiple scenes exceed the threshold (e.g., "family dispute" 20 times and "neighborhood dispute" 18 times in cluster 2, total amount 50), the first one is taken as the core typical scene according to the frequency.

[0106] Based on the historical multi-modal features in each of the feature clustering clusters, the typical features are obtained, wherein the typical features include typical text features, typical facial expression features, and typical sound features.

[0107] Specifically, for historical multi-modal features of the same emotion type label, box plot analysis is used to eliminate outliers (e.g., values exceeding 1.5 times the interquartile range); the mean and 95% confidence interval of the features of the remaining different emotion type labels are calculated to obtain typical features: for example, the typical speech features of "angry": pitch change slope > 30 Hz / s, speech speed fluctuation frequency > 0.6 (coefficient of variation), pitch peak > 350 Hz, and sound intensity mutation duration > 1.2 seconds; typical text features: negative vocabulary frequency > 8 times / 200 words, semantic sentiment tendency value < -0.6; typical facial features: eyebrow peak downward displacement amplitude > 8 mm, mouth corner downward angle > 12°, and facial muscle movement rate > 5 pixels / frame.

[0108] Optionally, based on the emotion change habit, the real-time call information is dynamically weighted to obtain an output weight, including:

[0109] According to the call text information, the face expression information, and the sound information, a real-time call scene is extracted, a preset weight division combination corresponding to the emotion type label associated with the same typical scene of the real-time call scene is obtained to obtain a first weight.

[0110] Specifically, first, the real-time call scene is acquired, for example, keywords (such as "performance" and "assessment") are extracted from real-time call text information, combined with semantic sentiment tendency values (such as -0.3, slightly neutral), and a preset scene library (such as "work evaluation" and "daily casual conversation") is matched by a TF-IDF algorithm to determine the real-time call scene (such as "work evaluation"); if the text information is ambiguous (such as less than 3 keywords), the slope of the tone change (such as 15 Hz / s, gentle) of the real-time sound information and the muscle movement rate (such as 2 pixels / frame, stable) of the facial expression are combined to assist in scene judgment, ensuring that the scene recognition accuracy is greater than or equal to 85%.

[0111] Secondly, the preset weight division combination corresponding to the emotion type label associated with the typical scene of the real-time call scene is called, and it should be noted that since people in different emotions, different features reflect different authenticity, thus, each emotion type label is set to have a weight division, for example, the "work evaluation" scene is associated with the "anxiety" emotion type, and the preset weight combination thereof is: text information 0.3, facial expression information 0.3, and sound information 0.4 (based on the highest contribution of sound features to anxiety emotions in historical data under this scene), and the combination is directly used as the first weight (text: 0.3, face: 0.3, sound: 0.4).

[0112] The real-time scene is extracted by multi-modal information fusion to ensure the accuracy of scene matching; the first weight is quickly generated based on the preset weight combination, so that the weight division is initially matched with the scene characteristics, solving the problem of disconnection between weight and scene, and providing a basic benchmark for subsequent dynamic adjustment.

[0113] The call text information, the facial expression information, and the sound information are matched with the typical features to obtain a corresponding matching degree, the emotion type label corresponding to the matching degree greater than a preset matching threshold is acquired, a weight adjustment value is obtained according to the matching degree, the preset matching threshold, and a preset adjustment coefficient, and a second weight is obtained based on the weight adjustment value and the preset weight division combination corresponding to the emotion type label.

[0114] Specifically, the multi-modal features of real-time call information are matched with typical features item by item: sound features: the overlap degree of the pitch change slope (real-time 18 Hz / s) and the "anxiety" typical feature (15-25 Hz / s) is 80%, the overlap degree of the speech rate fluctuation frequency (real-time 0.5 times / s) and the typical value (0.4-0.6 times / s) is 90%, and the comprehensive sound matching degree is (80%+90%) / 2=85%; text features: the overlap degree of the negative vocabulary frequency (real-time 4 times / 150 words) and the typical value (3-5 times / 150 words) is 100%, the overlap degree of the semantic sentiment tendency value (real-time -0.4) and the typical value (-0.5 to -0.3) is 80%, and the comprehensive text matching degree is (100%+80%) / 2=90%; facial expression features: the overlap degree of the brow lifting amplitude (real-time 4 mm) and the typical value (3-6 mm) is 100%, the overlap degree of the muscle movement rate (real-time 3 pixels / frame) and the typical value (2-4 pixels / frame) is 100%, and the comprehensive facial matching degree is 100%; the overall matching degree is calculated by weighted average: sound 40% x 85% + text 30% x 90% + face 30% x 100% = 90.5%.

[0115] The preset matching threshold is set to 70%, the current matching degree is 90.5%> 70%, the preset weight combination (text 0.3, face 0.3, sound 0.4) associated with the "anxiety" emotion type is obtained; the adjustment coefficient a is calculated as (matching degree-threshold) / (1-threshold)=(90.5%-70%) / 30%≈0.68, the maximum adjustment range of a single mode is ±0.1; the weight of the facial feature (100%) with the highest matching degree is adjusted upward: 0.3+0.1 x a≈0.37; the text feature (90%) with the second highest matching degree is fine-tuned: 0.3+0.1 x (a x 0.8)≈0.35; the sound feature is adjusted downward accordingly: 0.4-(0.07+0.05)=0.28, and the final second weight is obtained: text 0.35, face 0.37, sound 0.28 (the sum is 1).

[0116] The matching degree is quantified by feature overlap, so that the weight adjustment has a basis; the weight is dynamically assigned based on the matching degree, highlighting the contribution of the high matching degree mode (such as the weight of the facial feature being the highest when the matching degree is the highest), solving the problem that the fixed weight cannot respond to real-time feature differences, and improving the adaptability of the weight to the current emotional expression

[0117] The weight is calibrated according to the collection quality of the call text information, the facial expression information and the sound information to obtain a weight calibration value.

[0118] Specifically, first, the quality of collection is evaluated, including evaluation by signal-to-noise ratio (SNR=25dB) of sound information collection, less than 30dB is determined as "medium quality", and the weight needs to be reduced; speech-to-text accuracy rate is 95% (error rate <5%), determined as "high quality", and the weight remains unchanged; and face feature point detection success rate is 98% (obstruction <2%), determined as "high quality", and the weight remains unchanged.

[0119] Secondly, quality calibration rules are set, for example, "medium quality" mode weight is reduced by 0.05, and "high quality" remains unchanged; the sound information (original second weight 0.28) is reduced by 0.05, and the calibrated sound weight is 0.23; the text and face weight are adjusted correspondingly to maintain the sum of 1: text 0.35+0.03=0.38, face 0.37+0.02=0.39, and finally the weight calibration value is obtained: text 0.38, face 0.39, and sound 0.23.

[0120] The first weight, the second weight and the weight calibration value are input into a weight calibration model to obtain the output weight, wherein the weight calibration model is obtained by training a preset model according to historical first weights, historical second weights, historical weight calibration values and optimal weights, the historical first weights are a combination of preset weight division corresponding to different emotion type labels, the historical second weights are obtained by adjusting after matching the historical multi-modal features with the typical features, the historical weight calibration values are obtained according to the collection quality of the historical multi-modal features, and the optimal weights are weight information manually labeled for each group of the historical multi-modal features.

[0121] Specifically, mainstream models such as random forest regression model, deep neural network (DNN), convolutional neural network (CNN) or recurrent neural network (RNN) are used as the preset model, the input features are historical first weights, historical second weights and historical weight calibration values, the output is the optimal weight manually labeled, the training data contains 1000 groups of historical samples (covering different emotion types and scenes), 5-fold cross-validation is used to optimize the model parameters (such as the number of decision trees is 100, and the maximum depth is 8), and the model prediction error (MSE) is ≤0.001. The first weight (0.3, 0.3, 0.4), the second weight (0.35, 0.37, 0.28) and the weight calibration value (0.38, 0.39, 0.23) are input into the trained random forest model, the model outputs the optimal weight by voting of the integrated decision trees: text 0.36, face 0.38, and sound 0.26 (the sum is 1), which is the output weight.

[0122] Optionally, the generating a prompt interface according to the current emotion type and displaying the prompt interface on the terminal interface comprises:

[0123] A corresponding prompt color and prompt effect are set for each current emotion type.

[0124] Specifically, as shown in the table, Figure 2 when the current emotion type is "happy" or "joyful", the prompt color is warm color such as orange or yellow, and the prompt effect includes fast flashing, light effect of jumping light points, and dynamic effect of lively and dynamic; when the current emotion type is "angry" or "warning", the prompt color is warning color such as red or black, and the prompt effect includes strong screen flashing, light effect of edge light flashing, and dynamic effect of abrupt and strong; when the current emotion type is "depressed" or "sad", the prompt color is soft and dim color such as gray or gray blue, and the prompt effect includes light effect of low flashing frequency and dynamic effect of slow fluctuation; when the current emotion type is "calm" or "peaceful", no prompt color and prompt effect are added.

[0125] The prompt interface is generated according to the current emotion type and the corresponding prompt color and prompt effect, and is displayed on the terminal interface.

[0126] Specifically, the Qt interface rendering engine can be used to construct the prompt interface, and the interface layout is a floating window (size: 200x100 pixels, without blocking the main content of the call) in the upper right corner of the terminal interface; the interface elements include: an emotion type label (such as "current emotion: happy", font size: 14pt, color consistent with the prompt color); a dynamic effect layer (located below the label, occupying 60% of the interface area, and the light effect is realized by OpenGL); and a feature brief prompt (such as "faster speech speed: 3.2 words per second", font size: 10pt, gray).

[0127] It should be noted that the prompt interface is set as "top layer display" through the window management API (such as Win32 API of Windows) of the terminal operating system, so as to ensure that it is not blocked by other windows; the prompt interface is displayed when the real-time emotion score is in a certain emotion type interval (such as the excitement interval of 85-100 points) for more than 10 seconds, so as to avoid frequent switching caused by transient fluctuations; the prompt interface is automatically hidden when the semantic communication scene is "private conversation" (determined by the text keywords "confidential" and "privacy"), so as to protect the privacy of the user.

[0128] As shown in the table, Figure 3 The user emotion recognition device 300 provided by the embodiment of the present application comprises:

[0129] The data processing module 310 is configured to generate a target emotion change curve according to a historical video call record of a target person, and perform trend analysis on the target emotion change curve to obtain an emotion change habit.

[0130] The weight division module 320 is configured to acquire real-time call information of the target person, perform dynamic weight division on the real-time call information based on the emotion change habit, and obtain an output weight, wherein the real-time call information includes call text information, facial expression information and voice information.

[0131] The scoring module 330 is configured to perform standardization processing on the real-time call information, perform weighted fusion on the processed call text information, facial expression information and voice information according to the output weight, and obtain a real-time emotion score.

[0132] The display module 340 is configured to obtain a current emotion type of the target person according to the real-time emotion score, and generate a prompt interface according to the current emotion type and display the prompt interface on a terminal interface.

[0133] As shown in Figure 4 The electronic device 400 provided by the embodiment of the present application includes a memory 410 and a processor 420; the memory 410 is configured to store a computer program; and the processor 420 is configured to implement the user emotion recognition method as described above when executing the computer program.

[0134] Alternatively, the electronic device 400 includes a memory 410 and a processor 420 coupled to the memory 410; the memory 410 is configured to store a computer program; and the processor 420 is configured to perform the following operations when executing the computer program:

[0135] generate a target emotion change curve according to the pre-acquired historical video call record of the target person, perform trend analysis on the target emotion change curve, and obtain an emotion change habit;

[0136] acquire real-time call information of the target person, perform dynamic weight division on the real-time call information based on the emotion change habit, and obtain an output weight, wherein the real-time call information includes call text information, facial expression information and voice information;

[0137] perform standardization processing on the real-time call information, perform weighted fusion on the processed call text information, facial expression information and voice information according to the output weight, and obtain a real-time emotion score;

[0138] obtain a current emotion type of the target person according to the real-time emotion score, and generate a prompt interface according to the current emotion type and display the prompt interface on a terminal interface.

[0139] The embodiment of the present application provides a computer readable storage medium, and the storage medium stores a computer program, and when the computer program is executed by a processor, the user emotion recognition method is realized.

[0140] Alternatively, a non-volatile computer readable storage medium, the storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the following operations:

[0141] According to the historical video call record of the target person obtained in advance, a target emotion change curve is generated, trend analysis is performed on the target emotion change curve, and an emotion change habit is obtained.

[0142] Real-time call information of the target person is acquired, the real-time call information is dynamically weighted based on the emotion change habit, and an output weight is obtained, wherein the real-time call information includes call text information, facial expression information and voice information.

[0143] The real-time call information is standardized, the call text information, the facial expression information and the voice information after processing are weighted and fused according to the output weight, and a real-time emotion score is obtained.

[0144] According to the real-time emotion score, a current emotion type of the target person is obtained, a prompt interface is generated according to the current emotion type, and the prompt interface is displayed on a terminal interface.

[0145] An electronic device 400 that can be a server or a client of the present application will now be described, which is an example of a hardware device that can be applied to various aspects of the present application. The electronic device 400 is intended to represent various forms of digital electronic computer devices, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device 400 can also represent various forms of mobile devices, such as personal digital processing, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections, and their functions, as described herein, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed herein.

[0146] The electronic device 400 includes a computing unit that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) or a computer program loaded into a random access memory (RAM) from a storage unit. In the RAM, various programs and data required for device operation can also be stored. The computing unit, the ROM, and the RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.

[0147] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing relevant hardware, and the program can be stored in a computer readable storage medium. When the program is executed, the program can include the processes of the above-mentioned embodiment methods. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM), a random access memory (RAM), or the like. In this application, the units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment of the present application. In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically independently, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0148] Although the present application is disclosed as above, the protection scope of the present application is not limited to this. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present application, and these changes and modifications will fall within the protection scope of the present application.

Claims

1. A user emotion recognition method, characterized in that, include: Based on pre-acquired historical video call records of the target individual, a target emotional change curve is generated. Trend analysis is then performed on the target emotional change curve to determine emotional change habits, including: Based on Mehrabian's rule, the historical multimodal features in the historical video call records are weighted and fused to obtain a historical emotion score. The historical multimodal features include historical call text information, historical facial expression information, and historical voice information. The historical video call records include historical voice features, historical text features, and historical facial expression features used to generate the target emotion change curve. The historical multimodal features are mapped to a historical emotion score through a support vector regression model, and the target emotion change curve is generated at the time granularity. Generate a target emotion change curve with time as the horizontal axis and the historical emotion score as the vertical axis; Identify abrupt change nodes in the historical emotion scores of the target emotion change curve, associate them with the historical video call records corresponding to the abrupt change nodes, and obtain historical emotion slices; Based on the historical emotion scores, each historical emotion slice is labeled with an emotion type according to the emotion judgment criteria, and typical scenes and typical features under each emotion type are obtained based on the historical multimodal features corresponding to the same emotion type. Based on the typical scenarios, typical characteristics, and the emotion types, the emotion change habits are obtained; The real-time call information of the target person is obtained, and the real-time call information is dynamically weighted based on the emotional change habits to obtain the output weight. The real-time call information includes text information, facial expression information and voice information. The real-time call information is standardized, and the processed text information, facial expression information, and voice information are weighted and fused according to the output weights to obtain a real-time emotion score. The current emotion type of the target person is obtained based on the real-time emotion score, and a prompt interface is generated based on the current emotion type and displayed on the terminal interface.

2. The user emotion recognition method according to claim 1, characterized in that, The step of identifying abrupt change nodes in the historical emotion score within the target emotion change curve and associating them with historical video call records corresponding to those nodes includes: Set a mutation threshold, which is the maximum allowable fluctuation value of the historical emotion score within a preset continuous time window in the target emotion change curve; The target emotion change curve is traversed by sliding through the continuous time window to obtain the historical fluctuation value of the historical emotion score within each continuous time window. When the historical fluctuation value exceeds the mutation threshold, the start time of the continuous time window is marked as the mutation node of the historical emotion score. In the historical video call records, slices are obtained by extending a preset time before and after the mutation node, and the slices are associated with the corresponding historical video call records to obtain historical emotion slices.

3. The user emotion recognition method according to claim 2, characterized in that, Based on the historical emotion scores, each historical emotion slice is labeled with an emotion type according to the emotion determination criteria. Typical scenarios and features for each emotion type are obtained based on the historical multimodal features corresponding to the same emotion type, including: Obtain the historical emotion score corresponding to the historical emotion slice, and assign the emotion type label to each historical emotion slice according to the emotion judgment criteria, wherein the emotion type label includes happiness, anger, calmness and sadness; The K-means clustering algorithm is used to cluster the historical multimodal features under the same sentiment type label to obtain multiple feature clusters; Statistical analysis is performed on multiple feature clusters to obtain the typical scenarios and typical features; The typical scenarios and typical features are associated and stored with the corresponding emotion type labels to obtain emotion change habits.

4. The user emotion recognition method according to claim 3, characterized in that, The statistical analysis of multiple feature clusters to obtain the typical scenario and the typical features includes: For each of the feature clusters, the corresponding historical call scenarios are extracted, the frequency of occurrence of the same historical call scenarios is counted, and the historical call scenarios with a frequency exceeding a preset frequency threshold are identified as the typical scenarios. The typical features are obtained by analyzing the historical multimodal features in each feature cluster, wherein the typical features include typical text features, typical facial expression features, and typical voice features.

5. The user emotion recognition method according to claim 1, characterized in that, The step of dynamically weighting the real-time call information based on the emotional change habits to obtain output weights includes: Based on the text information, facial expression information and voice information of the call, the real-time call scene is extracted, and the first weight is obtained by pre-defined weight division and combination of the emotion type label associated with the typical scene that is the same as the real-time call scene; The text information of the call, the facial expression information, and the voice information are matched with the typical features to obtain the corresponding matching degree. The emotion type label corresponding to the matching degree being greater than the preset matching threshold is obtained. The weight adjustment value is obtained based on the matching degree, the preset matching threshold, and the preset adjustment coefficient. The second weight is obtained by dividing and combining the weight adjustment value and the preset weight corresponding to the emotion type label. Weight calibration values ​​are obtained by performing weight calibration based on the acquisition quality of the call text information, facial expression information, and voice information. The first weight, the second weight, and the weight calibration value are input into the weight calibration model to obtain the output weight. The weight calibration model is trained on a preset model based on the historical first weight, the historical second weight, the historical weight calibration value, and the optimal weight. The historical first weight is a preset weight division combination corresponding to different emotion type labels. The historical second weight is adjusted after matching the historical multimodal features with the typical features. The historical weight calibration value is obtained based on the acquisition quality of the historical multimodal features. The optimal weight is the weight information manually labeled for each group of historical multimodal features.

6. The user emotion recognition method according to claim 1, characterized in that, The step of generating a prompt interface based on the current emotion type and displaying it on the terminal interface includes: Set a corresponding cue color and cue effect for each of the current emotion types; The prompt interface is generated based on the current emotion type and the corresponding prompt color and prompt effect, and then displayed on the terminal interface.

7. A user emotion recognition device, characterized in that, include: The data processing module is used to generate a target emotion change curve based on pre-acquired historical video call records of the target person, and to perform trend analysis on the target emotion change curve to obtain emotion change habits. This includes: weighted fusion of historical multimodal features in the historical video call records based on Mehrabian's rule to obtain a historical emotion score. The historical multimodal features include historical text information, historical facial expression information, and historical voice information. The historical video call records include historical voice features, historical text features, and historical facial expression features used to generate the target emotion change curve. The historical multimodal features are mapped to historical... Historical emotion scores are generated, and a target emotion change curve is generated with time granularity. The target emotion change curve is then generated with time as the horizontal axis and the historical emotion scores as the vertical axis. Abrupt change nodes in the historical emotion scores within the target emotion change curve are identified, and historical video call records corresponding to these nodes are associated to obtain historical emotion slices. Based on the historical emotion scores, each historical emotion slice is labeled with an emotion type according to emotion determination criteria. Typical scenarios and typical features under each emotion type are obtained based on the historical multimodal features corresponding to the same emotion type. Finally, the emotion change habits are obtained based on the typical scenarios, typical features, and the emotion type. The weighting module is used to acquire the real-time call information of the target person, and to dynamically weight the real-time call information based on the emotional change habits to obtain the output weight. The real-time call information includes call text information, facial expression information and voice information. The scoring module is used to standardize the real-time call information, and to perform weighted fusion of the processed call text information, facial expression information and voice information according to the output weights to obtain a real-time emotion score. The display module is used to obtain the current emotion type of the target person based on the real-time emotion score, generate a prompt interface based on the current emotion type, and display it on the terminal interface.

8. An electronic device, characterized in that, Including memory and processor; The memory is used to store computer programs; The processor is configured to implement the user emotion recognition method as described in any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the user emotion recognition method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Digital human posture adjusting method and device combined with image recognition

    CN119473209A

  • Emotion recognition method based on large model and related device

    CN119904901A