Emotion recognition methods and related devices based on large models

By acquiring multimodal data and combining it with a large model, this emotion recognition method solves the problem of low accuracy in traditional methods, achieving more accurate and reliable emotion recognition, and is applicable to emotion analysis in various scenarios.

CN119904901BActive Publication Date: 2025-12-02BEIJING LIXIN INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411984391.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-12-02
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

Traditional emotion recognition methods rely primarily on a single information source, resulting in low accuracy and an inability to fully capture the complexities of human emotional expression.

Method used

A large-model-based emotion recognition method is adopted to acquire multimodal data (text information, audio information, and image information), and emotion state analysis is performed through emotion recognition model. Combined with confidence assessment, user emotion recognition results are generated, including attention heatmaps and visual auxiliary analysis information.

Benefits of technology

It achieves cross-modal intelligent emotion recognition, improves the accuracy and efficiency of emotion recognition, expands the application scope of emotion recognition, and can adapt to the dominance of different modal data in various scenarios, providing reliable user emotion judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119904901B_ABST
    Figure CN119904901B_ABST
Patent Text Reader

Abstract

This application relates to the field of data processing, and in particular to a method and apparatus for emotion recognition based on a large model. The method includes: first, acquiring multimodal data to be processed; the multimodal data includes at least: text information, audio information, and image information input by the user; the image information includes at least: facial expression images and / or body movement images. Next, the multimodal data is input into an emotion recognition model for emotion state analysis to obtain the user's predicted emotion state. Finally, a confidence assessment is performed based on the predicted emotion state, and a user emotion recognition result is generated based on the predicted emotion state and the corresponding assessment result. This method captures richer emotional details from multimodal data through an emotion recognition model, achieving cross-modal intelligent emotion recognition, and further improves the accuracy and efficiency of emotion recognition by combining confidence assessment, thus expanding the application scope of emotion recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and in particular to a large-model-based emotion recognition method and related apparatus. Background Technology

[0002] Emotion recognition is an important research area in artificial intelligence and computer vision, with applications in many scenarios such as intelligent assistants, sentiment analysis, and mental health diagnosis. Traditional emotion recognition methods are mainly based on single information sources, such as speech, text, or video images, but these methods often struggle to accurately capture the complex emotions expressed by humans.

[0003] In related technologies, emotion recognition is typically performed using only voice or text information. This ignores the valuable information contained in other forms of information (such as images), resulting in lower accuracy in emotion recognition. Therefore, there is an urgent need to propose a novel technical solution to address at least one of the aforementioned issues. Summary of the Invention

[0004] This application addresses the technical problems existing in related technologies by providing a large-model-based emotion recognition method and related apparatus to solve at least one technical problem existing in related technologies.

[0005] In a first aspect, embodiments of this application provide an emotion recognition method based on a large model, the method comprising:

[0006] Acquire multimodal data to be processed; the multimodal data includes at least: text information, audio information, and image information input by the user; the image information includes at least: facial expression images and / or body movement images;

[0007] The multimodal data is input into the emotion recognition model to analyze the emotion state and obtain the user's predicted emotion state.

[0008] A confidence assessment is performed based on the predicted emotional state, and a user emotion recognition result is generated based on the predicted emotional state and the corresponding assessment result; the user emotion recognition result includes at least: an attention heatmap for displaying the user's emotional state from multiple dimensions, and / or visual auxiliary analysis information; the auxiliary analysis information includes at least: voice emotion analysis, text emotion analysis, and facial expression emotion analysis.

[0009] Secondly, embodiments of this application provide an emotion recognition device based on a large model, which includes at least the following units:

[0010] The acquisition unit is configured to acquire multimodal data to be processed; the multimodal data includes at least: text information, audio information, and image information input by the user; the image information includes at least: facial expression images and / or body movement images;

[0011] The analysis unit is configured to input the multimodal data into an emotion recognition model to perform emotion state analysis and obtain the user's predicted emotion state.

[0012] The output unit is configured to perform a confidence assessment based on the predicted emotional state, and generate a user emotion recognition result based on the predicted emotional state and the corresponding assessment result; the user emotion recognition result includes at least: an attention heatmap for displaying the user's emotional state from multiple dimensions, and / or visual auxiliary analysis information; the auxiliary analysis information includes at least: voice emotion analysis, text emotion analysis, and facial expression emotion analysis.

[0013] Thirdly, embodiments of this application provide an electronic device, the electronic device comprising:

[0014] At least one processor, memory, and input / output unit;

[0015] The memory is used to store computer programs, and the processor is used to call the computer programs stored in the memory to execute the first aspect of the large model-based emotion recognition method.

[0016] Fourthly, a computer-readable storage medium is provided, comprising instructions that, when executed on a computer, cause the computer to perform the large-model-based emotion recognition method of the first aspect.

[0017] The beneficial effects of this application embodiment are as follows: First, multimodal data to be processed is acquired; the multimodal data includes at least: text information, audio information, and image information input by the user; the image information includes at least: facial expression images and / or body movement images. Then, the multimodal data is input into an emotion recognition model for emotion state analysis to obtain the user's corresponding predicted emotion state. Finally, a confidence assessment is performed based on the predicted emotion state, and a user emotion recognition result is generated based on the predicted emotion state and the corresponding assessment result. The user emotion recognition result includes at least: an attention heatmap for displaying the user's emotion state from multiple dimensions, and / or visual auxiliary analysis information; the auxiliary analysis information includes at least: voice emotion analysis, text emotion analysis, and facial expression emotion analysis. In this application embodiment, the emotion recognition model captures richer emotional details from multimodal data, achieving cross-modal intelligent emotion recognition, and further improving the accuracy and efficiency of emotion recognition by combining confidence assessment, thus expanding the application scope of emotion recognition. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating an emotion recognition method based on a large model according to an embodiment of this application.

[0019] Figure 2 This is a schematic diagram of the structure of an emotion recognition device based on a large model according to an embodiment of this application;

[0020] Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application;

[0021] Figure 4 This is a schematic diagram of the structure of a media device according to an embodiment of this application. Detailed Implementation

[0022] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0023] To address at least one technical problem in related technologies, embodiments of this application provide a method and apparatus for emotion recognition based on a large model. In this technical solution, firstly, multimodal data to be processed is acquired; the multimodal data includes at least: user-input text information, audio information, and image information; the image information includes at least: facial expression images and / or body movement images. Then, the multimodal data is input into an emotion recognition model for emotion state analysis to obtain the user's predicted emotion state. Finally, a confidence assessment is performed based on the predicted emotion state, and a user emotion recognition result is generated based on the predicted emotion state and the corresponding assessment result. The user emotion recognition result includes at least: an attention heatmap for displaying the user's emotion state from multiple dimensions, and / or visual auxiliary analysis information; the auxiliary analysis information includes at least: voice emotion analysis, text emotion analysis, and facial expression emotion analysis.

[0024] The technical solution provided in this application, in its first aspect, acquires multimodal data including user-input text, audio, and image information (such as facial expression images and / or body movement images), moving beyond the limitations of single-modal analysis of user emotions. For example, relying solely on text may not accurately determine a user's true emotions. However, by combining audio information such as tone of voice with image information like facial expressions, the user's actual emotional state can be captured more comprehensively from multiple perspectives, reducing emotion recognition bias caused by insufficient or inaccurate single-modal information. In different application scenarios, users express emotions in different ways. In some online communication scenarios, text may be dominant, while in video call scenarios, audio and image information are more crucial. This method can adapt to these diverse scenarios, comprehensively analyzing data regardless of which modality dominates, thus enabling its widespread application in various situations such as psychological counseling, online customer service, social media interaction, and remote conferencing, improving the universality of emotion recognition.

[0025] Secondly, multimodal data is input into a specialized emotion recognition model for emotion state analysis. Leveraging the model's powerful learning and feature extraction capabilities, deep-seated emotional features and intermodal relationships can be uncovered from different modalities, leading to relatively accurate predictions of emotional states. For example, it can distinguish subtle emotional differences; within the broad category of "happiness," it can further differentiate between excited and calm happiness, thus improving the accuracy of emotion recognition. Considering that emotions are often not static but dynamically change over time, the model can incorporate time series analysis techniques (such as capturing the dynamic transition patterns between emotions mentioned in the text) to better understand the evolution of emotions in continuous scenarios. This allows for a better grasp of both short-term fluctuations and long-term trends in user emotions, which is highly helpful in understanding the emotional trajectory of users during extended interactions.

[0026] Thirdly, confidence assessment based on predicted emotional states provides a reliable basis for the final emotion recognition results. By assessing uncertainty using methods such as Bayesian deep learning and quantifying reliability using the framework of evidence theory, users can clearly understand the credibility of each emotion recognition result, avoiding blind reliance on inaccurate predictions. This is especially instructive in key decision-making scenarios (such as providing targeted services based on user emotions). The generated user emotion recognition results include attention heatmaps to display user emotional states from multiple dimensions, and visual auxiliary analysis information (such as voice sentiment analysis, text sentiment analysis, and facial expression sentiment analysis), making the emotion recognition results more intuitive and easy to understand. For example, attention heatmaps can clearly show which factors play a key role in emotion judgment at different modalities or times; various visualized sentiment analyses allow users to see at a glance the user's emotional state reflected in different aspects, facilitating further in-depth analysis and the development of targeted response measures.

[0027] In summary, this large-model-based emotion recognition method offers a more comprehensive, accurate, reliable, and intuitive way to identify and display user emotions. It plays a positive and important role in improving user experience, optimizing service quality, and supporting various business decisions that require understanding user emotions. This solution captures richer emotional details from multimodal data through an emotion recognition model, achieving cross-modal intelligent emotion recognition. Furthermore, it combines confidence assessment to further improve the accuracy and efficiency of emotion recognition, expanding its application scope.

[0028] The technical solution of this application, and the emotion recognition scheme based on a large model provided in the embodiments of this application, can also be executed by an electronic device, which can be a server, server cluster, or cloud server. The electronic device can also be a terminal device such as a mobile phone, computer, tablet computer, wearable device, or dedicated device (such as a dedicated terminal device with a large model-based emotion recognition method system). These electronic devices can also be equipped with the chips or other hardware processing units described in the above embodiments. Alternatively, these electronic devices can also install a service program for executing the large model-based emotion recognition scheme.

[0029] Figure 1 A flowchart illustrating a large-model-based emotion recognition method provided in this application embodiment is shown below. Figure 1 As shown, the method includes the following steps:

[0030] 101. Obtain the multimodal data to be processed;

[0031] 102. Input the multimodal data into the emotion recognition model to analyze the emotion state and obtain the predicted emotion state corresponding to the user.

[0032] 103. Confidence assessment is performed based on the predicted emotional state, and user emotion recognition results are generated based on the predicted emotional state and the corresponding assessment results.

[0033] In this embodiment, the image information includes at least facial expression images and / or body movement images. Specifically, the image information has a defined scope, encompassing at least facial expression images and body movement images. Facial expression images can intuitively present a person's emotional state, such as joy, anger, sorrow, and happiness; for example, smiles, frowns, and wide eyes can all reflect different moods. Body movement images are equally crucial; relaxed body movements may suggest a relaxed and happy mood, while curled-up or tense body postures may reflect anxiety or unease. By including these facial expression images and / or body movement images within the scope of image information, important visual evidence can be provided for subsequent work such as emotion recognition based on multimodal data, assisting in a more comprehensive and accurate understanding of the user's emotional state.

[0034] In this embodiment, the user emotion recognition result includes at least: an attention heatmap for displaying the user's emotional state from multiple dimensions, and / or visual auxiliary analysis information. Specifically, an attention heatmap is a way to display the user's emotional state from multiple dimensions. It uses visualization to present the key areas or crucial information that the model focuses on during emotion recognition in the form of a heatmap. For example, when comprehensively analyzing multimodal data, for text, the heatmap may highlight words that play a key role in judging emotion; for images, it will focus on the facial features (such as eyes, mouth, etc.) that best express emotion or the most representative parts of body movements. From different modalities and different feature dimensions, variations in color intensity and other factors intuitively demonstrate the importance of each part to the overall emotional state judgment, allowing users to quickly understand the model's focus during emotion analysis and the contribution of each part of the information.

[0035] Visualized auxiliary analysis information is content that displays and supplements the user's emotion recognition results from another perspective, and it at least covers aspects such as voice sentiment analysis, text sentiment analysis, and facial expression sentiment analysis. Further, optionally, the auxiliary analysis information includes at least: voice sentiment analysis, text sentiment analysis, and facial expression sentiment analysis.

[0036] For example, by professionally analyzing collected user audio information, acoustic features such as pitch, speech rate, and timbre can be extracted to determine the underlying emotional nuances. For instance, a higher pitch and faster speech rate might indicate excitement or agitation, while a low, slow voice might suggest sadness or depression. These emotional analysis results can then be presented visually, such as using bar charts to display the intensity of speech features corresponding to different emotional categories, or using line graphs to show the trend of emotion changes over time in a speech segment, making the user's emotional state immediately apparent.

[0037] For example, natural language processing (NLP) techniques can be used to analyze user-input text, including words, sentence structure, and semantics, to identify the emotional tone conveyed by the text. For instance, text containing more positive words is likely to reflect positive emotions, while text with a large number of negative words often reflects negative emotions. Visually, word clouds can be used to display frequently occurring emotional keywords in the text, or pie charts can be used to show the proportion of different emotional categories (such as positive, negative, and neutral) in the text, helping users clearly grasp the emotional state of the text.

[0038] For example, based on the facial expression images mentioned earlier, we can analyze the morphological changes of facial features (such as an upturned mouth indicating happiness, a furrowed brow indicating sadness) and the overall combination features of facial expressions to determine the corresponding emotional state. Through visualization methods, such as displaying different emotional categories corresponding to different expressions using image sets, or creating dynamic visualization charts of emotional changes in expressions (showing the emotional shifts of expressions over time), we can intuitively present the emotional trajectory reflected by the user's facial expressions, aiding in a more comprehensive understanding of the user's emotional expression.

[0039] By providing such user emotion recognition results, whether it be attention heatmaps or visual auxiliary analysis information, users can understand each other's emotional state from multiple dimensions in an intuitive and clear way, providing a strong basis for subsequent decision-making, interaction, or the provision of targeted services.

[0040] As an optional embodiment, in 101, a text input module is used to receive and format text data, providing semantic input for subsequent processing. A voice input module captures and digitizes audio signals, converting them into acoustic data that the system can process. An image input module acquires and standardizes visual data, providing facial expression and body language information. Exemplarily, multimodal data is received through three parallel input channels. In the preprocessing layer, text data undergoes word segmentation, noise reduction, and standardization to clean and normalize the text data; voice signals undergo noise reduction, segmentation, and normalization to eliminate interference, extract key acoustic features, and are segmented into speech segments using an adaptive segmentation algorithm; image data is processed using a deep learning framework for noise reduction, illumination normalization, pose correction, and size unification. Finally, the corresponding emotional feature information of the three modalities is labeled to achieve interoperability between the three modalities.

[0041] Specifically, in 101, the text input module plays a crucial role in receiving text data. It can interface with text information from various sources, such as text content directly entered by the user on the interactive interface, like statements entered in the chat window or comments posted in the comment section; it can also import text content from other documents and files, and is compatible with text input in different formats, ensuring that as comprehensive as possible the acquisition of text information relevant to the user is achieved.

[0042] In step 101, after receiving text data, the text input module performs formatting processing. For example, it standardizes the text encoding format to avoid garbled characters and other issues that may affect subsequent analysis due to inconsistent encoding; it standardizes and organizes the punctuation marks in the text to make the sentence structure clearer, facilitating semantic understanding by subsequent natural language processing algorithms; and it removes unnecessary spaces, line breaks, and other irrelevant formatting elements to organize the text into a neat and standardized format. This provides clear, accurate, and easily analyzable semantic input for subsequent processing, ensuring the smooth execution of subsequent text-based sentiment analysis and other operations.

[0043] In the 101, the voice input module has the ability to capture audio signals, and it can acquire audio information in a variety of ways. For example, it can use the device's built-in microphone, such as a mobile phone microphone or computer microphone, to collect the sound signal produced by the user's speech in real time; it can also receive audio streams input from external audio devices, such as audio files recorded by external professional recording equipment, to ensure that it can collect audio signals containing the user's voice generated in different scenarios.

[0044] In the 101 module, after capturing audio signals, the voice input module digitizes these analog audio signals. Through specific sampling and quantization techniques, the continuous audio waveform is converted into discrete digital signals, making them recognizable and processable by the computer system. Subsequently, these digitized audio signals are further converted into acoustic data that the system can process. This may involve extracting basic acoustic features of the audio, such as Mel-frequency cepstral coefficients (MFCC), pitch, and volume. The audio information is then transformed into a specific data format suitable for subsequent operations such as voice emotion analysis and fusion with other modal data, laying the foundation for comprehensive emotion recognition.

[0045] In the 101, the image input module is responsible for acquiring visual data, and the acquisition methods are diverse. On the one hand, it can capture images in real time through camera devices, such as taking pictures of users' facial expressions and body movements with the front camera of a mobile phone, or acquiring images of people in a video conference scene through an external computer camera; on the other hand, it can also receive stored image files, such as user photos selected from the album, keyframe images extracted from videos, etc., to ensure that sufficient visual data related to the user's emotional expression can be obtained.

[0046] In step 101, after acquiring visual data, the image input module performs standardization operations on this image data. For example, it standardizes the image size to meet the requirements of subsequent model processing, avoiding the impact of excessive image size differences on analysis efficiency and accuracy; it adjusts the image color mode to ensure the consistency and accuracy of color information; and it normalizes the image to ensure that the pixel values ​​are within an appropriate range. After these standardization processes, the image data can better provide key information such as facial expressions and body language, facilitating subsequent work such as sentiment analysis and multimodal fusion based on image features.

[0047] Each of these three modules performs its respective function, effectively processing multimodal data such as text, speech, and images. This provides an accurate, standardized, and suitable data foundation for the entire emotion recognition system based on multimodal data, ensuring the smooth operation of subsequent stages and ultimately achieving accurate user emotion recognition.

[0048] As an optional embodiment, in step 102, the multimodal data is input into an emotion recognition model for emotion state analysis to obtain the user's predicted emotion state, including:

[0049] The emotional features corresponding to the multimodal data are extracted; the emotional features are fused in a multimodal manner to obtain emotional fusion features; the emotional fusion features are input into an emotion recognition model to analyze the emotional state and obtain the predicted emotional state of the user.

[0050] In this embodiment of the application, the emotional feature information includes at least one of the following: textual emotional features, acoustic emotional features, and visual emotional features.

[0051] In this optional embodiment, processing multimodal data through a series of steps to obtain the user's predicted emotional state has multiple technical advantages. First, emotional feature information, including textual sentiment features, acoustic sentiment features, and visual sentiment features, is extracted from the multimodal data. This allows for a comprehensive mining of the emotional content contained in different modalities, avoiding the omission of important emotional cues due to reliance on a single modality. For example, textual sentiment features can reveal emotional tendencies through word semantics, acoustic sentiment features can reflect emotional states through tone and speech rate, and visual sentiment features can convey emotional information through facial expressions and body movements. Combining these features allows for a more complete capture of the user's true emotions.

[0052] Next, these emotional feature information are fused in a multimodal manner to obtain emotional fusion features. This breaks down the boundaries between different modalities, allowing the emotional features of each modality to complement and synergize with each other. The fused features contain richer and more comprehensive emotional information, better reflecting the expression of users' emotions in different dimensions and enhancing the accuracy of the overall grasp of users' emotions.

[0053] Finally, the fused emotion features are input into the emotion recognition model for emotion state analysis to obtain predicted emotion states. Leveraging the powerful learning and analysis capabilities of the emotion recognition model, based on the comprehensive and integrated features after fusion, it can accurately analyze the user's specific emotional state, distinguish the subtle differences between different emotions, and also take into account the dynamic changes of emotions over time and other factors. This provides a reliable and accurate foundation for subsequent confidence assessments and the generation of comprehensive user emotion recognition results, helping to achieve more scientific and accurate identification and presentation of user emotions, and has important value in many application scenarios that require understanding user emotions.

[0054] The following sections describe the implementation methods for extracting corresponding emotional feature information from the multimodal data in the above steps, from different modal perspectives.

[0055] For text information, the above steps of extracting the corresponding sentiment features from the multimodal data can be implemented as follows:

[0056] A multi-level text sentiment feature vector is extracted from the text information using a generalized linear model (GLM). The text sentiment feature vector includes at least: word-level sentiment feature vectors, sentence-level sentiment feature vectors, and context-level paragraph sentiment feature vectors. Multi-task learning is performed on the text sentiment feature vectors to obtain the implicit micro-emotional change features. The micro-emotional change features include at least: modal particle features and implicit expression features. The text sentiment feature vectors and the micro-emotional change features are used as the text sentiment features.

[0057] Specifically, a generalized linear model (GLM) is used to mine sentiment features in textual information, leveraging its powerful modeling capabilities to construct multi-level text sentiment feature vectors. These include word-level sentiment feature vectors, which capture the emotional tone inherent in individual words; for example, words like "joy" and "sadness," which inherently carry clear emotional tendencies, have vectors that reflect this emotional attribute; sentence-level sentiment feature vectors, which focus on the emotion conveyed by the entire sentence, considering word collocation, grammatical structure, and other factors to determine whether the overall emotional state of the sentence is positive, negative, or neutral; and context-level paragraph sentiment feature vectors, which take a more macro-level approach, considering logical relationships and narrative flow between paragraphs to grasp the overall emotional direction of the text.

[0058] Furthermore, based on the constructed text sentiment feature vector, multi-task learning is conducted. This process aims to discover the subtle emotional changes hidden within the text sentiment feature vector, which are crucial for accurately grasping the text's emotions. These include modal particle features, such as "ah," "ya," and "ne," which convey different emotions in different contexts. Adding them to a declarative sentence might soften the emotion or enhance the tone of an exclamatory sentence, highlighting a strong emotion. There are also implicit expressive features, such as the emotions implied by metaphors and symbols, for example, using "dark days" to allude to difficult and unpleasant times. Through multi-task learning, these easily overlooked but important subtle emotional changes are extracted and combined with the previously constructed text sentiment feature vector to form a complete text sentiment feature, thus reflecting the emotional information carried by the text more comprehensively and meticulously.

[0059] For audio information, pre-trained models such as Wav2Vec and HuBERT are used to obtain multi-dimensional acoustic features, including acoustic parameters such as fundamental frequency, formants, energy envelope, and spectral centroid, as well as prosodic features such as pitch curves, speech rate, and pauses. As an optional embodiment, the above steps of extracting corresponding emotional feature information from the multimodal data can be implemented as follows:

[0060] The audio information is converted into acoustic physical vectors using the Wav2Vec model. The acoustic physical vectors include at least: spectral features and / or prosodic features. The spectral features include at least: Mel-frequency cepstral coefficients (MFCC) and spectral centroid feature information. The prosodic features include at least: pitch curve features, speech rate variation features, pause features, and breathing features. The acoustic physical vectors and the audio information are input into the HuBERT model to identify the emotional cues hidden in the user's voiceprint, thereby obtaining the acoustic emotional features.

[0061] Specifically, for audio information, the Wav2Vec model is first used to transform it, resulting in an acoustic physical vector. This vector contains important acoustic information such as spectral features and prosodic features. Regarding spectral features, there are Mel-frequency cepstral coefficients (MFCCs), which simulate the human ear's perception of different frequencies, reflecting the frequency characteristics of audio, as well as spectral centroid features, indicating the location of the center of gravity in the sound spectrum. Prosodic features encompass pitch curve features, i.e., the curve formed by changes in pitch, which can intuitively reflect the speaker's emotional changes; for example, pitch often rises when excited. Speech rate variations reflect the speed and rhythm of speech. Pauses suggest thinking, emphasis, etc., and breathing also reflects the rhythm and emotional state of speech. These features together constitute the acoustic physical vector, presenting the basic characteristics of audio from an acoustic perspective.

[0062] Furthermore, the acoustic physical vectors obtained through the Wav2Vec model, along with the original audio information, are input into the HuBERT model. Leveraging its strengths in speech processing, the HuBERT model can deeply analyze this information and identify the emotional cues hidden in the user's voiceprint. For example, when angry, the voiceprint may exhibit specific frequency and rhythmic changes, while when joyful, it will display different characteristics, ultimately yielding acoustic emotional features and accurately capturing the emotional content contained in the audio information.

[0063] It is further important to understand that related technologies, when processing audio information to extract emotional features, often focus only on a portion of acoustic features, such as only focusing on one or two parameters in the spectral features, or simply considering a few prosodic features such as pitch and speech rate. In contrast, the embodiments of this application first comprehensively acquire acoustic physical vectors, including spectral features such as Mel-frequency cepstral coefficients (MFCC) and spectral centroid features, as well as prosodic features such as pitch curve features, speech rate variation features, pause features, and breathing features, through the Wav2Vec model. This describes the basic characteristics of audio information from multiple dimensions, leaving no acoustic detail that may contain emotional cues unexplored. In comparison, this achieves more comprehensive and detailed audio feature extraction, laying a more solid foundation for subsequent accurate sentiment analysis.

[0064] When further utilizing the HuberT model to identify emotional cues, it not only relies on the acoustic physical vectors extracted by the Wav2Vec model but also inputs the original audio information. Related techniques may only perform subsequent processing based on the extracted features, easily losing some implicit information in the original audio that is difficult to fully represent through features alone. The combined approach of this application's embodiment allows the HuberT model to fully mine information from both the original audio and the extracted features, maximizing the use of audio data to discover hidden emotional cues in the user's voiceprint, further improving the completeness and accuracy of emotional feature extraction.

[0065] From the perspective of model synergy, many related technologies employ a single model to complete the entire audio emotion feature extraction process, and the limitations of the model itself can affect the final extraction results. However, this application's embodiment first uses the Wav2Vec model, which excels at converting audio into rich feature vectors, for initial feature extraction. Then, it leverages the Hubert model, which has unique advantages in speech processing and can deeply mine emotional cues, for further analysis. This achieves synergistic work and complementary advantages between the two advanced models. Wav2Vec provides high-quality feature input to Hubert, enabling Hubert to better recognize emotions, making the entire audio emotion feature extraction process more efficient and accurate, representing a significant improvement over related technologies that rely solely on a single model.

[0066] From the perspective of emotional cue mining depth, in the crucial step of identifying the emotional cues hidden in a user's voiceprint, related technologies may struggle to accurately capture the subtle frequency and rhythmic changes in voiceprints that occur with emotional shifts due to insufficient feature comprehensiveness or limited model capabilities, leading to inaccurate emotion judgments. This application's embodiments, through comprehensive feature extraction and dual-model collaboration, can keenly identify specific frequency and rhythmic changes in voiceprints under different emotions such as anger and joy, more deeply mining the emotional cues contained in the voiceprint, and thus accurately acquiring acoustic emotional features. This significantly improves the quality of emotional feature extraction from audio information, enabling subsequent audio-based emotion recognition and other related work to be carried out more reliably.

[0067] In summary, in terms of extracting emotional features from audio information, this technology has achieved improvements in several key aspects compared to related technologies through more comprehensive feature extraction, the synergistic advantages of dual models, and deeper emotional clue mining. This helps to capture the emotional content contained in audio information more accurately.

[0068] For image information, visual feature extraction utilizes large models such as Vision Transformer and Swing Transformer to extract multi-scale visual features, including facial key points, micro-expressions, eye-tracking features, and pose estimation. As an optional embodiment, the above steps of extracting corresponding emotional feature information from the multimodal data can be implemented as follows:

[0069] Using the Vision Transformer model, initial visual features at multiple scales are extracted from the image information. The initial visual features include at least: facial key point features, facial contour features, limb contour features, posture features, motion change features, and color distribution features. The initial visual features are then subjected to emotion state recognition and change trend analysis to obtain the visual emotion features.

[0070] Specifically, leveraging the powerful visual feature extraction capabilities of the Vision Transformer model, initial visual features at multiple scales are extracted from image information. These features are abundant, including facial keypoint features, such as the position and shape changes of facial features like the eyes and mouth, which play a crucial role in judging facial expressions and inferring emotional states; facial contour features can help identify changes in the overall facial shape; limb contour features can reflect the posture and range of limb movements; posture features relate to the overall posture of a person, such as an upright posture which may be associated with confidence, while a hunched back may suggest a depressed mood; movement change features can capture the emotional information contained in dynamic movements such as waving and nodding; color distribution features are also indispensable, as different color combinations create different emotional atmospheres visually, such as warm colors conveying positive and enthusiastic emotions, while cool colors may be associated with calmness and sadness. Combining these multi-scale initial visual features provides rich visual material for subsequent sentiment analysis.

[0071] Then, the extracted initial visual features are further analyzed for emotional state recognition and trend analysis. By analyzing the dynamic changes of facial key points, the expression is determined to be different states such as smiling or frowning. Combined with features such as posture and body movements, the current emotional state presented in the image is identified, and the changing trends of these features over time or between different images are observed, such as the process of changing from a smiling expression to a calm expression. In this way, visual emotional features are obtained, and the emotional information contained in the image information is comprehensively and accurately grasped, laying the foundation for multimodal emotion fusion and subsequent emotion recognition work.

[0072] By performing the above steps for extracting sentiment features from text, audio, and image data, we can fully explore the emotional information contained in each modality, providing a comprehensive and high-quality data foundation for subsequent multimodal fusion and more accurate emotion recognition.

[0073] For example, in the feature extraction process of the above three modalities, the feature extractor in this application embodiment can adopt domain adaptation and progressive transfer learning strategies to improve the discriminative and generalization capabilities of features through multi-task learning and contrastive learning.

[0074] Optionally, before extracting the corresponding sentiment feature information from the multimodal data, the multimodal data can be further split according to data type to obtain text preprocessed data, audio preprocessed data, and image preprocessed data. Then, the text preprocessed data is preprocessed to obtain preprocessed text information. The preprocessing operations according to the text data type include: word segmentation, noise reduction, and standardization. The audio preprocessed data is preprocessed to obtain preprocessed audio information. The preprocessing operations according to the audio data type include: noise reduction, segmentation, and normalization. The image preprocessed data is preprocessed to obtain preprocessed image information. The preprocessing operations according to the image data type include: noise reduction, resizing, and color standardization.

[0075] As an optional embodiment, step 103 involves multimodal fusion of the emotional feature information to obtain emotional fusion features, including:

[0076] The emotional features are mapped to a high-dimensional semantic space according to different modalities using a nonlinear projection matrix, resulting in a multimodal mapping feature matrix. Through adversarial training, the multimodal mapping feature matrix undergoes modality dissimilarity elimination processing to obtain a first mapping feature matrix. A multi-layer, multi-head cross-attention mechanism is employed to establish fine-grained correlations between different modalities in the first mapping feature matrix, resulting in a second mapping feature matrix. Based on a Transformer encoder improved using a meta-learning strategy, the second mapping feature matrix is ​​adaptively evaluated according to the importance of different modalities, and dynamic weights are assigned according to the evaluation results to obtain the emotional fusion features.

[0077] Specifically, the first step is to collect sentiment feature information extracted from different modalities (such as text, speech, and images). These features have their own different forms of representation; for example, text features may be multi-level sentiment feature vectors, speech features may be related data such as spectrum and prosody, and image features may be multi-scale visual features.

[0078] Next, nonlinear projection matrices are used to operate on the emotional features of these different modalities. The aim is to map the features of each modality, originally existing in different dimensions and representations, into a unified high-dimensional semantic space. Through this mapping process, each modality's features have a corresponding representation in this new semantic space, and the combination of all modality feature representations forms a multimodal mapping feature matrix. For example, text features, originally existing as word vectors, are transformed into a new vector form that reflects their semantic meaning in a high-dimensional semantic space after being mapped by the projection matrix; speech features are also transformed from spectral data representations into corresponding vector forms in this semantic space; the same applies to image features, which are transformed from visual features into vector representations in a high-dimensional semantic space. This prepares the ground for subsequent modal difference elimination and further fusion.

[0079] After obtaining the multimodal mapping feature matrix, adversarial training is used for processing. This process involves setting up a structure similar to a discriminator.

[0080] Specifically, under this adversarial training mechanism, features from different modalities begin to "compete" with each other. The basic principle is to teach the model to distinguish between different modal features while simultaneously prompting it to reduce the inherent differences between modalities. For example, the discriminator tries to identify whether a feature vector comes from a text modality or a speech modality, but the model adjusts the representations of each modal feature to make them increasingly similar at the semantic level, making it difficult for the discriminator to accurately distinguish them. Through this continuous adversarial learning process, the differences between modalities gradually decrease, and the final feature matrix after modality difference elimination is the first mapping feature matrix. This makes the features of each modality more semantically unified, which is more conducive to subsequent fusion operations.

[0081] Based on the first mapping feature matrix, a multi-layer, multi-head cross-attention mechanism is employed to establish fine-grained correlations between different modalities. Under this mechanism, features from different modalities interact and cross-attention. For example, text features can focus on related information in speech and image features; similarly, speech and image features can focus on corresponding information in other modalities. Through this mutual attention and interaction, not only are more detailed and in-depth correlations established between modalities, but the model can also automatically learn the relative importance of each modal feature within the overall fusion context. For instance, when analyzing the emotional content of a video containing audio narration, text subtitles, and related displayed images, this mechanism clearly identifies the key roles played by text, speech, and image modalities in expressing a specific emotion. After this processing, the second mapping feature matrix is ​​obtained, further enriching and refining the fusion relationships between modalities.

[0082] Finally, the second mapping feature matrix is ​​processed using a Transformer encoder improved based on a meta-learning strategy. This Transformer encoder adaptively evaluates the importance of each modality's features based on the specific application scenario, the data involved, and the comprehensive factors demonstrated in the preceding steps. For example, in certain emotional expression scenarios, the emotional information contained in the image modality may be more crucial for the overall judgment, and this encoder can identify this.

[0083] Then, based on the evaluation results of the importance of each modality's features, the weights corresponding to each modality are dynamically adjusted. If a certain modality is determined to be more important for emotional expression, its weight is increased accordingly; conversely, its weight is decreased. This flexible weight adjustment method fully considers the different contributions of each modality in real-world situations, achieving more realistic multimodal fusion. The final result is a fully fused emotional feature, which integrates the advantages of each modality and can be better used for subsequent tasks such as emotion recognition.

[0084] Moreover, when encountering situations where image modalities are missing, such as when the image display function is turned off during a video call, the system upon which this Transformer encoder, based on a meta-learning strategy, relies can also activate a compensation mechanism. It can utilize existing text, speech, and other modal information, as well as previously learned patterns and rules, to fill the data gaps caused by the absence of image modalities, ensuring that the entire multimodal fusion process can still proceed stably and guaranteeing the accuracy and reliability of the final emotional fusion features.

[0085] For example, see Figure 2 As shown in Figure 103, the feature alignment module first uses a nonlinear projection matrix to map features from different modalities to a unified semantic space, and then eliminates modal differences through adversarial training. The attention fusion module employs a multi-layer, multi-head cross-attention mechanism to establish fine-grained associations between modalities, while introducing a gating mechanism to control information flow. The dynamic weight allocation module, based on an improved transformer encoder structure and combined with a meta-learning strategy, achieves adaptive evaluation and dynamic adjustment of feature importance. Furthermore, the system introduces a modality missing data compensation mechanism, which can maintain stable performance even when certain modal data is missing.

[0086] As an optional embodiment, in step 103, before performing modality difference elimination processing on the multimodal mapping feature matrix through adversarial training to obtain the first mapping feature matrix, corresponding gated unit models can be established according to different modalities. Then, using the gated unit models, invalid matrix elements corresponding to different modalities are filtered and deleted from the multimodal mapping feature matrix; wherein, invalid matrix elements are those that satisfy the preset filtering conditions corresponding to their respective modalities.

[0087] Specifically, corresponding gating unit models are established for different modalities (such as text, speech, and image). Each gating unit model is designed based on the characteristics and data features of the corresponding modality. It can understand and process the information representation forms and potential semantic relationships unique to that modality. For example, for the text modality, the gating unit model will consider factors such as the grammatical structure and lexical semantics of the text; for the speech modality, it will combine acoustic features and prosodic characteristics; and for the image modality, it will be constructed with reference to visual features and spatial layout, so as to adapt to the data processing requirements of the corresponding modality.

[0088] Each modality is assigned its own preset filtering conditions, which are derived from the analysis and experience of the data in the process of emotional feature expression and fusion for that modality. For example, in the text modality, the preset filtering conditions may involve feature vector elements corresponding to certain low-frequency words that have no obvious effect on emotion judgment, or feature elements corresponding to some text fragments with messy grammatical structure, ambiguous semantics, and difficulty in conveying effective emotional information; in the speech modality, it may be acoustic feature elements that are too low in volume, have too messy spectral features, and do not conform to the normal emotional expression rules of speech; for the image modality, it may be visual feature elements corresponding to background areas in the image that are unrelated to the emotional expression of the person, or feature elements corresponding to parts of the image where emotional information cannot be accurately extracted due to image quality problems (such as excessive blurring or severe occlusion).

[0089] After establishing the gating unit models corresponding to each modality and setting preset filtering conditions, the multimodal mapping feature matrix is ​​input into the corresponding gating unit models. These gating unit models act like intelligent "sieves," checking and judging the modality-related portions of the multimodal mapping feature matrix element by element according to their preset filtering conditions. Matrix elements that meet the preset filtering conditions, i.e., those deemed invalid, are removed from the multimodal mapping feature matrix and deleted. Thus, after processing by the gating unit models corresponding to each modality, elements in the multimodal mapping feature matrix that are considered unhelpful or even potentially interfering with subsequent sentiment fusion and analysis are removed, resulting in a relatively purer feature matrix that is more conducive to modal difference elimination and subsequent fusion—an intermediate matrix state preparing for the generation of the first mapping feature matrix.

[0090] Through the steps described above, in the process of multimodal data fusion, the original feature information extracted from different modalities often contains elements that have no practical value for sentiment analysis and fusion, or even interfere with it. By establishing a gating unit model for filtering and deletion, these invalid matrix elements can be accurately removed, preventing them from mixing into the effective information during subsequent modal difference elimination and further fusion processes, thus avoiding information confusion and difficulties in model learning. This allows the entire multimodal fusion process to focus more on valuable sentiment feature information, improving the efficiency and quality of fusion.

[0091] After screening, each modality retains its most representative features that are more strongly correlated with emotional expression. This makes the features of each modality more refined and accurate. For the text modality, the remaining features are those that more clearly convey emotional semantics; for the speech modality, the remaining features are effective acoustic features that better reflect emotional rhythm and acoustic characteristics; and for the image modality, the remaining features are key visual features closely related to the emotional expression of the characters. In this way, during subsequent modality elimination and fusion, each modality can better cooperate and work together based on high-quality features, which helps to more accurately extract the emotional information contained in multimodal data and improve the reliability and effectiveness of the final emotional fusion features.

[0092] Since invalid elements are removed in advance before the modal difference elimination process, it provides a better data foundation for subsequent modal difference elimination operations such as adversarial training. During the process of making different modal features closer at the semantic level through adversarial training, the model can focus more on learning and adjusting the differences between truly meaningful features, reducing the additional computational burden caused by the existence of invalid information and possible incorrect learning situations. Ultimately, it optimizes the effect of multimodal fusion as a whole, enabling the obtained first mapping feature matrix and the subsequent generated sentiment fusion features to more accurately reflect the sentiment states contained in the multimodal data, providing stronger support for related applications such as emotion recognition.

[0093] In summary, the operation of establishing a gating unit model according to different modalities and screening and deleting invalid matrix elements, based on reasonable principles, can bring various positive effects in the process of multimodal sentiment feature fusion, contributing to improving the performance and accuracy of the entire sentiment analysis system.

[0094] In the above steps, determining the preset screening conditions under different modalities is a complex but crucial process that requires considering multiple factors. Specifically as follows:

[0095] For the text modality, first, the sentiment dictionary can be used to determine the sentiment tendency and intensity of words. For words with unclear sentiment tendency (such as neutral words) or extremely low sentiment intensity, in certain specific scenarios, if these words do not play a key role in text sentiment expression, the corresponding feature elements can be considered as screening conditions. For example, some very common conjunctions, auxiliary words, etc., like "de", "di", "he", etc., generally do not provide substantial help for sentiment judgment in most cases. At the same time, analyze the semantic importance of words in a specific field or topic. If the text is about technical documents, some general words unrelated to technology may be relatively unimportant; if it is an emotional literary work, those descriptive words that cannot reflect the emotions of characters (such as simple environmental description words in some cases) may be regarded as screening conditions for invalid elements.

[0096] Next, observe the grammar and syntactic structure of sentences. For sentence parts with chaotic structures, not conforming to normal grammar rules and unable to accurately understand their semantics, they can be set as screening conditions. For example, some text fragments with ambiguous semantics due to typos or grammar errors, their corresponding feature elements may interfere with sentiment analysis.

[0097] Consider the contribution of sentence components to emotional expression. In a sentence, the subject and predicate usually carry the main emotional information, while lengthy modifiers or parenthetical phrases, if they do not substantially affect the emotional judgment, can also be used as filtering criteria. For example, in the sentence "That decoration, which actually seemed a bit superfluous, made him happy," the modifier "actually seemed a bit superfluous" might not be crucial to the emotional judgment of "happy" in some situations, and could be considered as a filtering option.

[0098] Furthermore, by statistically analyzing the frequency of words in the text, words that appear infrequently and are not strongly associated with sentiment expression can be used as filtering criteria. For example, in a sentiment analysis task, if a rare word appears only once and its sentiment tendency cannot be clearly identified in the sentiment dictionary or the contextual semantics, then its corresponding feature element may be filtered out.

[0099] Analyze the distribution of words across different sentiment categories in texts. If a word appears evenly across all sentiment categories (e.g., positive, negative, neutral) without a clear sentiment bias, it can be considered as a selection criterion.

[0100] For speech modalities, the screening criteria for spectral features are determined based on the range of human auditory perception and common spectral patterns in speech emotion expression. For example, a reasonable range for Mel-frequency cepstral coefficients (MFCCs) is set. If the MFCC values ​​in certain frequency bands exceed the normal speech range (e.g., abnormal frequencies that may be caused by noise or equipment malfunction), these corresponding acoustic physical vector elements can be used as screening criteria.

[0101] For prosodic features, the normal range of pitch, speech rate, and volume should be considered. For example, segments with excessively low volume (which may be due to environmental noise interference or a weak speech signal) or excessively high volume (which may be due to abnormal screaming or other non-emotionally expressive sounds), segments with pitches exceeding the normal range of human speech expression (e.g., excessively high or low pitches may indicate signal errors), and segments with excessively fast or slow speech rates (which do not conform to the normal rhythm of emotional expression, such as words per second exceeding the normal range) can be used as filtering criteria.

[0102] For example, we can study the typical patterns of speech feature changes under different emotional states. For instance, in expressions of anger, speech typically has a higher pitch, greater volume, and faster pace, while in expressions of sadness, the pace may be slower and the pitch lower. Feature elements that do not conform to common emotional speech patterns can be considered as screening criteria. We can also observe the relationship between pauses, breaths, and other features in speech signals and emotional expression. In normal emotional expression, pauses and breaths follow certain patterns; for example, pauses occur during thinking or emphasis. If abnormally frequent or excessively long pauses occur (which may be due to signal interruption or other interference), or if a breathing pattern does not conform to the logic of emotional expression, the corresponding feature elements can be used as screening criteria.

[0103] Consider the background noise characteristics of the voice recording environment. By analyzing the environmental noise, identify the acoustic features primarily caused by noise as screening criteria. For example, if there is continuous machine noise in the background, its corresponding frequency range and acoustic pattern can be marked as invalid elements and removed during the screening process.

[0104] For parts of the speech signal that may be affected by environmental reflections, echoes, etc., analyze the degree of interference with emotional expression. If factors such as echoes cause distortion of speech features and affect emotional judgment, the feature elements corresponding to these distorted parts can be used as screening criteria.

[0105] For image modalities, the focus is primarily on visual features related to facial expressions, with an emphasis on features of key areas such as the eyes and mouth. If parts of the face in an image are occluded, the facial keypoint features or expression features corresponding to the occluded portion can be used as filtering criteria. For example, if most of the eyes and mouth are obscured by hair or hands, these occluded parts may not be helpful for expression sentiment analysis, and their corresponding feature elements can be filtered out.

[0106] Analyze the correlation between body movements and posture features and emotional expression. Feature elements corresponding to minor body movements unrelated to emotional expression (such as unconscious shaking) or postures that do not have emotional connotations in the current scene (such as a simple standing posture, in some cases) can be used as screening criteria.

[0107] Consider the relationship between color distribution features in an image and emotional atmosphere. If a certain color has a high or low proportion in an image and is not closely associated with common emotional colors (for example, in an image mainly used for emotional analysis of people, a large area of ​​pure black background has no obvious emotional connotation), its corresponding color distribution feature elements can be used as a screening criterion.

[0108] For blurry image areas, especially those where the blur affects the recognition of facial expressions and body movements, their corresponding visual feature elements are used as filtering criteria. For example, features corresponding to parts of the image where facial expressions cannot be clearly recognized due to low pixel count of the shooting device or motion blur.

[0109] This study analyzes the impact of factors such as occlusion and shadows in images on emotional feature extraction. If an object obscures most of a person's body, and the emotional state of the obscured portion cannot be inferred from other information, the visual feature elements corresponding to these obscured areas can be used as filtering criteria.

[0110] Consider the image resolution and level of detail. If the image resolution is too low, making it impossible to accurately extract key facial landmarks, poses, and other emotion-related visual features, then the feature elements corresponding to these parts that are difficult to provide effective information due to insufficient resolution can be used as filtering criteria.

[0111] As an optional embodiment, before performing multimodal fusion on the emotional feature information to obtain emotional fusion features in step 103, if the number of image information in the emotional feature information does not meet the preset image mode threshold, image data compensation can also be performed on the emotional feature information.

[0112] Specifically, when the amount of image information is insufficient (e.g., it does not reach the preset image mode threshold), especially when image modalities are missing (such as in extreme cases like a video call being closed), the compensation mechanism relies on existing text, speech, and other modal information to find potential correlations with image features. For example, in everyday communication scenarios, the scenes described in text, the actions of people, or the emotional states expressed in speech often have a certain logical connection with the facial expressions and postures of people in the corresponding images. By learning and analyzing a large amount of labeled multimodal data, the correspondence patterns between these modalities can be summarized. For example, when the text mentions "laughing happily" and the tone of voice has a cheerful rhythm, the person in the image may usually have facial expressions such as upturned corners of the mouth and squinting eyes, as well as relatively relaxed body movements. Based on this pattern, text and speech information can be used to infer the content of potentially missing image information.

[0113] The initial training phase accumulated rich prior knowledge and, through model learning, demonstrated the distribution patterns of different modal features under various emotional states. For example, it was learned that expressing anger not only involves a rise in pitch and increased speech rate in speech, but also the use of intense vocabulary in text, and the typical facial expressions and body language of people in images, such as frowning and clenching fists. When image information is missing, this prior knowledge and model learning results are used to reasonably infer the approximate appearance of corresponding image features based on the emotional characteristics exhibited by the current text and speech, thus filling the data gaps caused by the lack of image modalities.

[0114] In the early stages of multimodal fusion, features from different modalities are mapped to a unified high-dimensional semantic space through operations such as nonlinear projection matrices, establishing feature mapping relationships between modalities. During image data compensation, these semantic space mapping relationships are utilized to further fill in image data gaps. For example, within this unified semantic space, there are relatively stable geometric or semantic relationships between text feature vectors, speech feature vectors, and image feature vectors. When some image feature vectors are missing, the reasonable range or approximate shape of the image feature vectors can be inferred from the positions of other modalities in the semantic space and their mapping relationships. This allows for the generation of corresponding compensation image data, ensuring that relatively complete image features participate in the subsequent fusion process and maintaining its stability.

[0115] In multimodal fusion, information from different modalities complements and works synergistically. If too much modal information is missing from the image and no compensation is made, this balance between modalities will be disrupted, making subsequent multimodal fusion-based operations (such as using multi-layer multi-head cross-attention mechanisms or Transformer encoders improved based on meta-learning strategies) difficult to carry out smoothly because of the lack of data support from the crucial dimension of image features. Image data compensation mechanisms can promptly fill these gaps, ensuring that each stage of the fusion process has relatively complete multimodal data available, maintaining the stable operation of the entire fusion process, and preventing interruptions or errors due to insufficient image modal information.

[0116] Image information itself contains rich emotional cues; facial expressions and body language are key factors in determining emotional states. When image data is insufficient, compensation is performed to ensure that the final fused image features can reconstruct a complete emotional picture as much as possible. This allows for better integration with text and speech features, leveraging the strengths of each modality to extract emotional information and thus improving the accuracy of the final fused emotional features. For example, in an emotion analysis scenario, missing image data might have caused the loss of cues of happiness conveyed through facial expressions. However, compensation can incorporate this emotional information, allowing the fused emotional features to more comprehensively and accurately reflect the true emotional state, thereby providing a more reliable data foundation for subsequent emotion recognition and other tasks.

[0117] In real-world applications, image modalities are frequently missing. By employing an image data compensation mechanism, the model can better handle such incomplete modal data, improving its generalization ability under different data conditions. Regardless of whether image modalities are missing, the model can strive to complete multimodal fusion and output relatively accurate sentiment fusion features through reasonable compensation. This allows the model to be more widely applicable to various complex and changing real-world application scenarios. For example, even when network instability leads to lost video call images or partial image data corruption, it can still effectively perform sentiment analysis, ensuring the normal operation and functionality of the system.

[0118] In summary, the aforementioned image data compensation mechanism can bring about many positive effects in the multimodal fusion process, which helps to improve the performance and reliability of the entire multimodal sentiment analysis system.

[0119] As an optional embodiment, in step 103, the emotion fusion features are input into an emotion recognition model for emotion state analysis to obtain the user's predicted emotion state, including:

[0120] The emotion fusion features are used for emotion classification through a deep multilayer perceptron to obtain the corresponding user emotion categories. An emotion state transition matrix is ​​constructed from the emotion fusion features based on a bidirectional long short-term memory network. The emotion state transition matrix represents the dynamic evolution process between user emotions. A time attention mechanism is used to extract the user's short-term emotion fluctuation trend and long-term emotion change trend from the emotion fusion features. Based on the user emotion category, the emotion state transition matrix, the short-term emotion fluctuation trend, and the long-term emotion change trend, the user's future emotion change tendency and corresponding predicted emotion category are predicted.

[0121] Specifically, firstly, the sentiment fusion features obtained from the preceding multimodal fusion steps are used as input data and fed into a deep multilayer perceptron (MLP). An MLP is a powerful neural network architecture that processes the input sentiment fusion features through multiple fully connected layers. During this process, neurons in each layer collaborate to perform nonlinear transformations and combinations on the features, attempting to uncover hidden pattern information that can distinguish different emotion categories. For example, the sentiment fusion features may contain various emotion-related cues extracted and fused from multimodal data such as text, speech, and images. The MLP integrates these cues to determine the corresponding emotion category, ultimately outputting specific user emotion categories such as anger, satisfaction, and anxiety, achieving preliminary emotion classification.

[0122] When using MLP for emotion classification, two optimization techniques—batch normalization and residual connections—were employed to improve model performance and stability. Batch normalization normalizes the input data for each layer of the network, keeping the mean and variance within a relatively stable range. This helps accelerate model convergence, avoids training instability caused by large differences in data distribution, and also reduces overfitting to some extent, allowing the model to perform stably and well when faced with different emotion fusion feature inputs. Residual connections allow a layer in the network to directly pass input information to subsequent layers, solving problems such as vanishing or exploding gradients that may occur as network depth increases. This enables the model to better learn complex emotion classification patterns, further enhancing model stability and ensuring more reliable output of user emotion categories.

[0123] Next, based on the emotion fusion features, a bidirectional Long Short-Term Memory (LSTM) network is used to construct an emotion state transition matrix. Bidirectional LSTM has a unique structural advantage, capable of simultaneously processing both forward and reverse information in sequential data. For emotion fusion features, which contain multimodal emotional information and may be correlated over time (such as emotional information in a continuous video or dialogue scene), bidirectional LSTM can capture the underlying emotional evolution patterns from two directions. By analyzing and learning the emotion fusion features at different times, it summarizes information such as the probability of emotions transitioning from one state to another, thus constructing the emotion state transition matrix. This matrix clearly represents the dynamic evolution of user emotions; for example, it can reflect the probability of transitioning from a happy emotional state to a calm state, or the probability of transitioning from a sad emotional state to an angry emotional state, providing important evidence for a deeper understanding of emotional changes in continuous scenarios.

[0124] Building upon the characteristics of emotional fusion, this study employs a time-attention mechanism to uncover trends in emotional change. This mechanism dynamically focuses on emotional information at key moments in the sequence, based on the relative importance of emotional fusion characteristics at different times. In the short term, it captures fluctuations in emotions over short time intervals. For example, in a brief conversation, emotions might instantly shift from curiosity to surprise; these instantaneous fluctuations can be keenly observed, forming short-term emotional trends. In the long term, it analyzes the overall development of emotions over longer periods. For instance, while watching a movie, how does emotion gradually transition from initial relaxation to tension, and finally to relief? This extracts long-term emotional trends, providing a more comprehensive view of how user emotions manifest over time.

[0125] Finally, by synthesizing the information obtained from the preceding data, including user emotion categories, emotion state transition matrices, short-term emotion fluctuation trends, and long-term emotion change trends, we can predict the user's emotional tendencies and corresponding predicted emotion categories for future periods. For example, if the current user's emotion category is known to be satisfaction, by combining the probability of transitioning from satisfaction to other emotions in the emotion state transition matrix, and by referring to short-term emotion fluctuation trends (whether there are signs of rapid changes in emotion recently) and long-term emotion change trends (whether the overall emotion is developing in a positive or negative direction), we can use certain algorithms and models to calculate and predict the user's possible emotional tendencies in the coming period. For example, the user's emotion may continue to remain in a satisfied state, or there is a certain probability that it will change to a happy state. This allows us to determine the corresponding predicted emotion category, providing more accurate feedback for subsequent actions, such as recommending suitable content or adjusting interaction methods based on the prediction results, to better adapt to the user's emotional state.

[0126] Through the above series of steps, we can fully utilize the characteristics of emotion fusion and leverage various advanced neural network technologies and mechanisms to comprehensively and deeply analyze users' emotional states and make reasonable predictions about future emotional situations. This has significant value in many application scenarios involving user emotion analysis.

[0127] As an optional embodiment, the confidence assessment based on the predicted sentiment state in step 104 can be implemented as follows:

[0128] A Bayesian deep learning method is used to convert the predicted emotional state into an emotional state probability distribution variable for uncertainty assessment. In the inference stage of the emotion recognition model, different emotion recognition sub-networks are constructed through multiple forward propagations, and the emotional state probability distribution variable is processed by the different emotion recognition sub-networks to obtain multiple corresponding predicted emotional states. The confidence level of multiple predicted emotional states is evaluated based on Monte Carlo Dropout. The reliability of the emotion category is quantitatively evaluated based on the confidence level, and the final target predicted emotional state is determined based on the quantitative evaluation results.

[0129] Traditional deep learning typically treats model parameters as fixed values ​​for training and prediction. However, Bayesian deep learning incorporates Bayesian inference, treating model parameters as random variables with probability distributions. The advantage of this approach is that it can more comprehensively describe the uncertainty of the model when faced with data uncertainty and the complexity of the model itself, rather than providing a deterministic prediction.

[0130] In this context, the predicted emotional state is transformed into a probability distribution variable for uncertainty assessment using Bayesian deep learning. For example, the user's predicted emotional category might initially be a definite "happy" emotion. However, through Bayesian deep learning, this "happy" prediction is represented as a probability distribution. This distribution includes a certain probability of "happy," while also having a smaller probability of other related emotional categories such as "calm" or "slightly excited." This probability distribution reflects the degree of uncertainty in the prediction, as different emotional categories are assigned corresponding probability values.

[0131] In the inference phase of the emotion recognition model (that is, the phase where the model is trained and used for prediction), multiple forward propagation methods are employed. During each forward propagation, the model calculates and outputs a result based on the input data (in this case, the transformed emotion state probability distribution variable). Performing this operation multiple times is equivalent to processing the same input from different perspectives or based on different "states" of the model.

[0132] During each forward propagation, the model can employ randomization mechanisms (such as randomly deactivating neurons) to make each computation appear as if a different subnetwork is at work. These different emotion recognition subnetworks, each with slightly different structures or parameter states (differences generated through randomization within the overall model framework), all process the same probability distribution variable for emotional states, ultimately yielding multiple predicted emotional states. These multiple predicted emotional states reflect the model's judgment of the input from different perspectives, and because of the differences between the subnetworks, these predictions will also differ, further highlighting the uncertainty of prediction.

[0133] Monte Carlo Dropout is a method that extends the Dropout technique (randomly deactivating some neurons during training to prevent overfitting) to the inference stage. During inference, Dropout is applied randomly multiple times, resulting in different sub-networks for prediction each time, similar to the process of building different emotion recognition sub-networks. Confidence is assessed by observing the distribution of multiple predicted emotion states obtained from these different emotion recognition sub-networks. If multiple predictions are relatively concentrated, for example, most pointing to the emotion category of "happiness" with only a few deviating, it indicates that the model has low uncertainty regarding the prediction of "happiness," and correspondingly, high confidence. Conversely, if the predictions are very scattered, with significant proportions from various emotion categories, it means high uncertainty and low confidence. In this way, the dispersion of prediction results is quantified into a confidence index, intuitively reflecting the reliability of each predicted emotion state.

[0134] Next, based on the calculated confidence scores, the reliability of the emotion categories is further quantitatively evaluated. For example, a threshold can be set; emotion categories with confidence scores above this threshold are considered highly reliable, while those below the threshold are considered less reliable. For different predicted emotional states, corresponding quantitative reliability scores are assigned according to their confidence scores, thus enabling a clear comparison of the reliability of different emotion categories as the final prediction results.

[0135] Finally, the target predicted sentiment state is determined based on the results of the quantitative assessment. Typically, the sentiment category with the highest reliability quantitative assessment score (i.e., the highest confidence level) is selected as the final prediction. This way, when outputting the predicted sentiment state, not only is the sentiment category provided, but also information about its reliability is included, making the entire prediction result more scientific, rigorous, and valuable.

[0136] Understandably, by transforming predicted emotional states into probability distribution variables and utilizing multiple forward propagation and different emotion recognition subnetworks, the uncertainty inherent in the prediction results can be fully explored and revealed. Unlike traditional single deterministic predictions that ignore the various possibilities that the data and model themselves may bring, this method comprehensively presents the probability of different emotion categories occurring from a probabilistic perspective. This better reflects the ambiguity and uncertainty of emotional judgment in reality, allowing users to have a more accurate grasp of the prediction results.

[0137] Based on Monte Carlo Dropout assessment of confidence and subsequent reliability quantification, a detailed reliability judgment was made for each predicted sentiment state. This ensures that the most reliable result is selected when finally determining the target predicted sentiment state, avoiding blind reliance on potentially inaccurate single predictions and reducing the risk of misjudging user sentiment due to model errors or data noise. This significantly improves the accuracy and reliability of the output predicted sentiment states, providing a more solid data foundation for subsequent applications based on sentiment analysis (such as personalized recommendations and user interaction optimization).

[0138] The entire confidence assessment process is transparent and quantifiable. Users can see not only the final predicted emotional state but also the corresponding confidence level and how it was determined through a series of rigorous evaluation steps. Compared to traditional "black box" emotion recognition models, this approach increases the model's interpretability, allowing users to better understand the model's decision-making basis and reliability. It is more readily accepted and trusted in practical applications, especially in scenarios requiring high decision-making accuracy (such as applications involving emotion judgment in fields like medicine and psychological counseling).

[0139] In summary, the confidence assessment method in this optional embodiment is based on reasonable principles and can bring positive effects in improving prediction accuracy, reliability, and enhancing model interpretability, thus helping to improve the performance and practical value of the entire emotion recognition system.

[0140] As an optional embodiment, generating user emotion recognition results based on the predicted emotional state and the corresponding evaluation results in step 104 can be implemented as follows: generating a corresponding attention heatmap based on the predicted emotional state; and / or generating the visual auxiliary analysis information based on the predicted emotional state to improve the interpretability of the recognition results.

[0141] In this embodiment, the attention heatmap is a visualization tool designed to show the distribution of key areas or crucial information that the model focuses on during emotion recognition. Variations in color intensity represent the importance of different parts to the final predicted emotional state; darker colors indicate greater importance in emotion judgment, while lighter colors indicate less importance. Its main function is to help users intuitively understand how the model makes emotion judgments based on input multimodal data (such as text, speech, and images), thereby improving the interpretability of the overall emotion recognition results.

[0142] When generating attention heatmaps based on predicted sentiment states, different modalities of data are processed differently. Taking text as an example, when analyzing text to determine sentiment categories, the model assigns different attention weights to words in the text. Words that play a key role in judging the predicted sentiment state (such as adjectives with strong emotional connotations, verbs expressing key emotional tendencies, etc.) are given higher weights, and the areas corresponding to these words are displayed in darker colors when generating the heatmap. For example, when judging the sentiment of the sentence "Today was a really happy day" as "happy," the area containing the word "happy" might appear as a darker color in the heatmap, indicating that the model has focused on this keyword that embodies emotion.

[0143] For speech modalities, the model may focus on certain key parts of the acoustic features. For example, when judging positive emotions, the time periods corresponding to features such as higher pitch and faster speech rate will be represented by darker colors in the heatmap, showing the importance of these acoustic features in emotion judgment.

[0144] In terms of image modality, key facial expressions (such as the eyes and mouth) and parts of body movements closely related to emotional expression are displayed in heatmaps with corresponding color intensities based on their contribution to the predicted emotional state. For example, the area where the corners of the mouth turn up when smiling, or the area where the eyes squint when laughing, will be darker in color on the heatmap if these features are crucial for determining the current emotion as "happy," visually presenting the emotion-related parts that the model focuses on in the image information.

[0145] By integrating the attention information corresponding to each part under different modalities, a comprehensive, cross-modal attention heatmap is finally generated, which shows the focus of the model when making a prediction of the emotional state of the target, allowing users to clearly see the role of each modality and its different elements in emotion recognition.

[0146] Visualized auxiliary analysis information includes multiple aspects, such as voice sentiment analysis, text sentiment analysis, and facial expression sentiment analysis. It further analyzes and displays the predicted emotional state from different angles, also in order to improve the interpretability of the recognition results and allow users to have a more comprehensive and in-depth understanding of the basis and specific circumstances of emotion recognition.

[0147] By professionally analyzing the collected speech data, acoustic features such as pitch, speech rate, and volume are extracted. These features are then correlated with corresponding emotion categories and visualized. For example, a line graph can be used to show how the pitch of a speech changes over time, while different colored lines mark the predicted emotional states corresponding to different stages of pitch change (e.g., rising pitch corresponds to excitement, represented by a red line; stable pitch corresponds to calmness, represented by a green line, etc.), allowing users to intuitively see the connection between changes in the acoustic features of speech and emotions. Bar charts can also be used to compare differences in average speech rate, volume, and other features under different emotional states, clearly demonstrating the supporting role of speech features in emotion judgment.

[0148] For the input text information, natural language processing techniques are used to analyze the words, sentence structure, and semantics to uncover the sentiment tendency. In visualization, word clouds can be used to display frequently occurring sentiment keywords in the text. For example, in text judged to be positive, positive words such as "happy," "wonderful," and "like" will be displayed in larger font in the word cloud, intuitively reflecting the emotional tone of the text. Pie charts can also be used to show the proportion of different sentiment categories (such as positive, negative, and neutral) in the entire text, helping users to clearly grasp the emotional state of the text and the contribution of each part to the overall predicted sentiment.

[0149] By analyzing facial expressions and body movements in images, different facial expression states and their corresponding emotional meanings can be identified and visualized. For example, a dynamic visualization chart of facial expression changes can be created, showing the process of a person's facial expression changing from calm to smiling, and then to laughing over time or in different situations. The predicted emotional state corresponding to each expression state is also labeled, clearly demonstrating the relationship between facial expression changes and emotional evolution. Alternatively, a collection of images can be used to display the typical emotions corresponding to different expressions (frowning, glaring, upturned corners of the mouth, etc.), making it easier for users to understand how facial expression features in images support the prediction of emotional states.

[0150] By generating such rich and diverse visual auxiliary analysis information, combined with attention heatmaps, the entire emotion recognition process and basis can be presented to users from multiple dimensions in an intuitive and easy-to-understand way. This makes the emotion recognition results more transparent and interpretable, helping users to better utilize these emotion recognition results for subsequent decision-making, service optimization, and other related work.

[0151] In summary, by generating attention heatmaps based on predicted emotional states and visual auxiliary analysis information based on predicted emotional states, the interpretability of emotion recognition results can be significantly improved. This allows the emotion recognition system to not only output a simple emotion category, but also provide a comprehensive, intuitive, and easy-to-understand emotion analysis display, thereby enhancing the system's practicality and credibility.

[0152] As an optional implementation, the data preparation phase first requires collecting multimodal data related to emotion recognition, including text, speech, and images. This data forms the foundation for generating attention heatmaps, each containing different dimensions of emotional cues. For example, text data might be users' chat logs or comments; speech data might be recorded audio of users speaking; and image data could be video screenshots or photos containing users' facial expressions and body movements.

[0153] Feature extraction is performed on data of different modalities. For text data, features such as word vectors, parts of speech, and syntactic structure are extracted. Simultaneously, natural language processing techniques are used to analyze sentiment keywords and semantic relationships within the text. These features help determine the contribution of each part of the text to the predicted sentiment state. For example, word embedding techniques are used to convert each word in the text into a vector representation, facilitating computer understanding and processing of its semantic information.

[0154] In terms of speech data, audio processing algorithms are used to extract acoustic features such as Mel-frequency cepstral coefficients (MFCC), pitch, speech rate, and pauses. These features reflect the prosody and timbre of speech, which are closely related to emotional expression and are important criteria for judging emotions. They are also key factors in determining the attention of speech segments when generating heatmaps.

[0155] For image data, computer vision technology is used to extract facial key point coordinates, facial expression features (such as the degree of eye opening and closing, the upturn or downturn of the corners of the mouth, etc.), body movement features (such as the extension or bending of the arms, body posture, etc.), as well as visual features such as color and texture of the image, in order to prepare for analyzing the importance of different regions in the image for emotion judgment.

[0156] Based on the target emotion prediction algorithm and related natural language processing models, this study analyzes the importance of each element in the text to the emotion judgment. For example, if the target predicted emotion is "happy," then words in the text that directly express happiness, such as "joy," "pleasant," and "laughter," will be assigned higher weights. At the same time, descriptive statements that are semantically related to happiness, such as "sunny days make people feel good," will also have their related words and phrases assigned corresponding weights based on their contribution to the overall expression of happiness.

[0157] Furthermore, sentence structure and grammatical relationships also influence weighting. For example, in a sentence emphasizing tone like "It's really so happy!", words like "really" and "so" that strengthen the emotional tone will receive relatively higher weights because they reinforce the expression of "happy". By comprehensively considering these factors, a corresponding attention weight is assigned to each word, phrase, and other element in the text for emotion judgment. The higher the weight value, the more critical it is to obtaining the target predicted emotional state.

[0158] To predict the emotional state based on the target, the weights are determined by analyzing the correlation between speech acoustic features and the emotion. Taking the emotion of "happiness" as an example, happy speech is typically characterized by relatively high pitch, relatively fast speech rate, fewer pauses, and a rising intonation. Therefore, when analyzing the corresponding speech data, time periods with higher pitch, faster speech rates, and speech segments exhibiting rising intonation are considered to play an important role in judging the emotion of "happiness" and are thus assigned higher weights.

[0159] Simultaneously, the synergistic effect between different acoustic features is also taken into account. For example, a combination of rising pitch and increased speech rate may be more representative of expressing "happiness" and will be assigned a higher weight than a single prominent feature. By analyzing the time series of various acoustic features in the speech data and their fit with the target predicted emotional state, the attention weight of the speech segment corresponding to each time segment in emotion judgment is determined.

[0160] In image-based modalities, when predicting the emotional state of a target, the focus is on regions associated with facial expressions and body movements. For expressing "happiness," typical facial features include squinting eyes and upturned corners of the mouth; therefore, the areas containing the eyes and mouth are crucial for judging happiness and are given high weight. Similarly, regarding body movements, relaxed postures are often associated with positive emotions, and the corresponding body regions are assigned weights based on their importance in emotional expression.

[0161] Furthermore, the combination of different features in an image and the overall visual effect also affect the weighting. For example, a smiling expression paired with a relaxed body posture presents a stronger atmosphere of happiness. In this case, the weight of the facial and body-related areas involved in judging this emotion will be reasonably increased by taking these factors into account, thereby determining the attention weight of each part of the image in emotion judgment.

[0162] During the heatmap generation stage, professional visualization libraries (such as Matplotlib and Seaborn in Python) can be selected to generate heatmaps. The coordinate system of the heatmap is determined, generally using word order in text, time series of speech, or pixel positions in images as coordinate axes. Appropriate settings are made according to the characteristics of different data modalities to accurately display the attention given to each part.

[0163] The attention weights of different parts of each modality, as determined earlier, are mapped to a color space. Generally, light and dark colors are used to visually represent the weights; for example, higher weight values ​​correspond to darker colors (warm colors such as red and orange are often used to represent areas of high attention), while lower weight values ​​correspond to lighter colors (cool colors such as light blue and light gray represent areas of low attention). This color mapping method allows the weights to be presented in a visual form.

[0164] Alternatively, if it is necessary to display the overall attention situation across modalities, the attention heatmaps corresponding to different modalities such as text, speech, and images can be integrated. For example, in the same visualization interface, a text area, a speech timeline area, and an image display area can be divided separately. Then, the heatmaps of each modality can be arranged and combined according to their corresponding positions and proportions to generate a comprehensive attention heatmap. This comprehensively displays the importance distribution of different modalities and different elements within each modality when obtaining the predicted emotional state of the target, allowing users to clearly see the focus of the model throughout the entire emotion recognition process.

[0165] The generated heatmap underwent detailed adjustments, including setting the range of color bars, annotating the meaning of the axes, and adding titles and descriptions to make it clearer, more intuitive, and easier to understand. For example, the weight range corresponding to different shades of color was labeled next to the color bars, allowing users to accurately grasp the level of attention represented by different colors. Corresponding text content, audio time points, or specific image location information were noted on the axes, making it convenient for users to compare the data with the original data to view the attention levels.

[0166] By following the detailed steps above, an attention heatmap can be generated based on the predicted emotional state of the target, clearly presenting the key points that the model focuses on during the emotion recognition process in a visual way, thereby improving the interpretability of the emotion recognition results.

[0167] As an optional embodiment, after generating the user emotion recognition result based on the predicted emotional state and the corresponding evaluation result in step 104, user feedback information on the user emotion recognition result can also be used, and online learning can be performed based on the feedback information to obtain model optimization parameters. Furthermore, the model optimization parameters are used in the model adjustment process of the emotion recognition model to improve the emotion recognition accuracy of the emotion recognition model.

[0168] Specifically, after generating and presenting user emotion recognition results based on predicted emotional states and corresponding evaluation results, users are encouraged to provide feedback on these results. This feedback can take various forms, such as users directly confirming the accuracy of the identified emotion category (e.g., marking it as correct or incorrect), or providing more detailed suggestions for improvement, explaining the differences between their actual feelings and the recognition results. For example, if the emotion recognition model outputs a predicted emotional state of "happiness," but the user feels their actual emotion at the time was "calm," and they report this difference to the system, this becomes an important basis for subsequent model optimization.

[0169] Collected user feedback is considered as additional labeled data and linked to the multimodal data (text, speech, images, etc.) used to train the model and their corresponding true emotion labels (in the labeled training dataset). Because there is a discrepancy between the model's output (predicted emotional state) and the true emotion in user feedback, these discrepancies can be analyzed to infer potential problems with the current model parameter settings and identify areas for adjustment. For example, if a user's "calm" emotion is repeatedly misclassified as "happy," it may mean that the model is not allocating appropriate weights to modal features when processing certain combinations of features (such as a calm tone of voice with a slight smile), or that some parameters in the neural network layers used for emotion classification and feature extraction are not accurately capturing the key information distinguishing between these two emotions.

[0170] By employing appropriate optimization algorithms (such as stochastic gradient descent and its variants), and based on the model prediction error revealed by feedback information, the algorithm calculates how to adjust the model parameters to reduce this error. Specifically, guided by the loss function (a function that measures the difference between the model's prediction and the actual result), the calculation of the loss function is adjusted based on feedback information to more accurately reflect the model's current performance issues. Then, using optimization algorithms, based on the gradient information of the loss function with respect to the model parameters, the model parameters are gradually updated, enabling the model to move towards reducing the difference between predictions and the user's actual emotions in subsequent predictions. This process is an online learning process based on feedback information. By continuously incorporating new feedback data for such learning and parameter adjustment, the model achieves continuous optimization.

[0171] After obtaining optimized model parameters through online learning, these new parameters are applied to the corresponding structure of the emotion recognition model, replacing the old parameters. For example, the weight parameters of each layer in the deep multilayer perceptron (used for emotion classification) and the relevant weights and bias parameters in the bidirectional long short-term memory network (used to construct the emotion state transition matrix) are updated. This allows the overall structure and function of the model to be adjusted based on the new parameters, thereby changing the model's processing methods for feature extraction, fusion, and emotion judgment of multimodal data. The goal is to more accurately output predicted emotional states that match the user's true emotions in subsequent emotion recognition tasks, thus improving the accuracy of emotion recognition.

[0172] It's worth noting that by incorporating user feedback for online learning and model parameter adjustments, the model can correct its errors and shortcomings based on users' real-world experiences. For example, initially, due to limited training data or uneven data distribution, there might have been more misjudgments of emotions in certain scenarios. However, by continuously absorbing user feedback and optimizing parameters, the model can better adapt to various complex and changing real-world situations, more accurately identify users' true emotions, reduce the deviation between predicted emotional states and actual emotions, and significantly improve the accuracy of emotion recognition.

[0173] Different users may express and experience emotions in significantly different ways, and real-world application scenarios are also diverse. Utilizing user feedback for online learning allows the model to access a wider range of emotion samples that more closely resemble real-world usage scenarios. This continuous learning and adaptation to new situations broadens the model's applicability and enhances its generalization ability across different user groups and application scenarios. Regardless of the style of text expression, speech tone characteristics, or facial expressions and body language in images, the model can more reasonably identify emotions based on optimized parameters, improving overall performance.

[0174] Accurate emotion recognition is crucial for many applications involving human-computer interaction and user services. As the accuracy of emotion recognition models improves, systems can provide more appropriate responses and services based on recognition results that more closely match the user's true emotions. For example, in intelligent customer service scenarios, accurately identifying a user's dissatisfaction allows for timely and effective reassurance measures and more targeted solutions, thereby improving user satisfaction with the overall service, enhancing the positive interaction between the user and the system, and making the system more practical in real-world applications.

[0175] In related technologies, model training is often performed offline using fixed datasets, making it difficult to adjust to new situations in real-world use once deployed. However, the online learning mechanism based on user feedback in this embodiment allows the emotion recognition model to dynamically optimize its parameters during operation, enabling continuous iterative updates. This ensures the model stays abreast of changes in user needs and application scenarios, maintains good performance, extends its effective lifespan, and reduces the cost of frequently redeveloping and training new models.

[0176] In conclusion, utilizing user feedback on emotion recognition results for online learning and model parameter adjustment, based on sound principles, can bring positive and significant effects in improving emotion recognition accuracy, enhancing model adaptability, and improving user experience. This helps emotion recognition systems better serve various practical application scenarios.

[0177] In this application's technical solution, firstly, multimodal data to be processed is acquired; the multimodal data includes at least: user-input text information, audio information, and image information; the image information includes at least: facial expression images and / or body movement images. Then, the multimodal data is input into an emotion recognition model for emotion state analysis to obtain the user's predicted emotion state. Finally, a confidence assessment is performed based on the predicted emotion state, and a user emotion recognition result is generated based on the predicted emotion state and the corresponding assessment result. The user emotion recognition result includes at least: an attention heatmap for displaying the user's emotion state from multiple dimensions, and / or visual auxiliary analysis information; the auxiliary analysis information includes at least: voice emotion analysis, text emotion analysis, and facial expression emotion analysis. This application's technical solution captures richer emotional details from multimodal data through an emotion recognition model, achieving cross-modal intelligent emotion recognition, and further improves the accuracy and efficiency of emotion recognition by combining confidence assessment, thus expanding the application scope of emotion recognition.

[0178] In one alternative embodiment, an example of user emotion recognition and service strategy optimization based on multimodal data in an intelligent customer service scenario.

[0179] In intelligent customer service scenarios, users interact with the customer service system through various methods such as text chat, voice conversations, and video calls. The system aims to accurately identify users' emotional states in order to dynamically optimize service strategies and improve customer experience and satisfaction.

[0180] In the input and preprocessing stages of multimodal data, for text input, after receiving the text sent by the user in the chat window, regular expressions and stop word lists are used to remove noise characters and meaningless words. A pre-trained large language model (such as Qwen or GLM) is used to complete text segmentation and semantic vectorization, and syntactic, semantic, and sentiment polarity features are extracted. For voice input, audio signals are obtained from call recordings, noise reduction algorithms (such as Spectral Subtraction) are used to eliminate environmental noise, effective speech segments are separated using speech endpoint detection technology, and acoustic features (such as pitch, formants, and rhythm features) are extracted using the Wav2Vec model. For image input, user facial images are obtained from video calls, OpenCV is used to solve uneven lighting problems and adjust image brightness, facial landmark detection and pose correction are performed using Dlib, and after standardizing the image size, visual features are extracted using the Swing Transformer.

[0181] In the feature extraction and representation learning stage, text feature extraction uses a GLM model to generate multi-level sentiment feature vectors covering word, sentence, and contextual discourse levels, and employs a multi-task learning strategy to capture subtle emotional changes in the text, such as interjections and metaphorical expressions. Speech feature extraction uses a Wav2Vec model to extract spectral features (such as MFCC and spectral centroid) and prosodic features (such as pitch curves and speech rate), combined with a HuBERT model to identify emotional cues such as anger and joy implied in the voiceprint. Image feature extraction uses a Vision Transformer to extract multi-scale visual features, identifies facial expression changes (such as smiling and frowning) through facial keypoint analysis, and analyzes emotional trends by combining posture features.

[0182] Next comes the multimodal feature fusion stage. First, projection matrices are used to map text, speech, and image features to a unified high-dimensional semantic space, and adversarial training is used to reduce intermodal differences. Then, a multi-layer, multi-head cross-attention mechanism is employed to establish intermodal associations and learn the importance of each modality's features. Simultaneously, a gating unit is introduced to dynamically control information flow and prevent interference from invalid information. Based on a meta-learning strategy, an improved Transformer is used to evaluate the importance of each modal input and dynamically adjust the weights. When image modalities are missing (e.g., video calls are closed), a compensation mechanism is used to fill data gaps, ensuring fusion stability.

[0183] In the emotion recognition and state tracking stage, a deep multilayer perceptron (MLP) is used to classify the fused features and output the user's emotion category (such as anger, satisfaction, and anxiety). Batch normalization and residual connections are used to optimize training and enhance model stability. An emotion state transition matrix based on bidirectional LSTM is constructed to capture the dynamic evolution between emotions. A time attention mechanism is used to model the short-term fluctuations and long-term trends of user emotions. Combined with sequential emotion changes, the possible future emotional tendencies of users are predicted, providing accurate feedback for service strategy adjustments.

[0184] Finally, in the results output and optimization stage, Bayesian deep learning is used to calculate the uncertainty of the classification results, Monte Carlo Dropout is used to evaluate confidence, and the reliability of each emotion category is quantified using the evidence theory framework. Attention heatmaps are used to visually display the emotional cues of multimodal features, and auxiliary information such as speech pitch waveforms, text sentiment analysis, and facial expression feature images are provided to enhance the interpretability of the results. Simultaneously, user feedback during interaction is used for online learning to continuously optimize model parameters and improve emotion recognition accuracy.

[0185] This application's embodiments enable accurate identification of users' emotional changes in intelligent customer service scenarios. For example, it accurately captures the user's emotional shift from anxiety to satisfaction and dynamically adjusts service strategies accordingly, such as shifting from urgent explanation to gentle reassurance or proactively providing solutions, thereby significantly improving the user experience. Furthermore, the multimodal fusion design of this method gives it a strong ability to capture complex emotional expressions, maintaining high performance even in the event of data loss. This effectively meets the application needs of intelligent customer service in multiple scenarios, providing strong support for the efficient operation and high-quality service of intelligent customer service systems.

[0186] In another embodiment of this application, a large-model-based emotion recognition device is also provided. See [link to related document]. Figure 2 The device comprises the following units:

[0187] The acquisition unit is configured to acquire multimodal data to be processed; the multimodal data includes at least: text information, audio information, and image information input by the user; the image information includes at least: facial expression images and / or body movement images;

[0188] The analysis unit is configured to input the multimodal data into an emotion recognition model to perform emotion state analysis and obtain the user's predicted emotion state.

[0189] The output unit is configured to perform a confidence assessment based on the predicted emotional state, and generate a user emotion recognition result based on the predicted emotional state and the corresponding assessment result; the user emotion recognition result includes at least: an attention heatmap for displaying the user's emotional state from multiple dimensions, and / or visual auxiliary analysis information; the auxiliary analysis information includes at least: voice emotion analysis, text emotion analysis, and facial expression emotion analysis.

[0190] Optionally, the analysis unit, which inputs the multimodal data into an emotion recognition model for emotion state analysis to obtain the user's predicted emotion state, is configured to: extract corresponding emotional feature information from the multimodal data; the emotional feature information includes at least one of the following: textual emotion features, acoustic emotion features, and visual emotion features; perform multimodal fusion on the emotional feature information to obtain emotion fusion features; and input the emotion fusion features into the emotion recognition model for emotion state analysis to obtain the user's predicted emotion state.

[0191] Further optionally, the analysis unit extracts the corresponding sentiment feature information from the multimodal data, which is configured to: extract multi-level text sentiment feature vectors from the text information using a generalized linear model; the text sentiment feature vectors include at least: word-level sentiment feature vectors, sentence-level sentiment feature vectors, and context-level paragraph sentiment feature vectors; perform multi-task learning on the text sentiment feature vectors to obtain the micro-emotional change features implicit in the text sentiment feature vectors; the micro-emotional change features include at least: modal particle features and implicit expression features; and use the text sentiment feature vectors and the micro-emotional change features as the text sentiment features.

[0192] Optionally, the analysis unit extracts corresponding emotional feature information from the multimodal data, configured to: convert the audio information into acoustic physical vectors using a Wav2Vec model; the acoustic physical vectors include at least: spectral features and / or prosodic features; the spectral features include at least: Mel-frequency cepstral coefficients (MFCC) and spectral centroid features; the prosodic features include at least: pitch curve features, speech rate variation features, pause features, and breathing features; the acoustic physical vectors and the audio information are input into a HuBERT model to identify the emotional cues hidden in the user's voiceprint, thereby obtaining the acoustic emotional features.

[0193] Further optionally, the analysis unit extracts corresponding emotional feature information from the multimodal data, which is configured to: extract initial visual features at multiple scales from the image information using a Vision Transformer model; the initial visual features include at least: facial key point features, facial contour features, limb contour features, posture features, action change features, and color distribution features; and perform emotional state recognition and change trend analysis on the initial visual features to obtain the visual emotional features.

[0194] Optionally, before extracting the corresponding sentiment feature information from the multimodal data, the analysis unit is further configured to: split the multimodal data according to data type to obtain text preprocessed data, audio preprocessed data, and image preprocessed data; perform preprocessing operations on the text preprocessed data to obtain preprocessed text information; wherein the preprocessing operations according to the text data type include: word segmentation, noise reduction, and standardization; perform preprocessing operations on the audio preprocessed data to obtain preprocessed audio information; wherein the preprocessing operations according to the audio data type include: noise reduction, segmentation, and normalization; perform preprocessing operations on the image preprocessed data to obtain preprocessed image information; wherein the preprocessing operations according to the image data type include: noise reduction, size adjustment, and color standardization.

[0195] Optionally, the analysis unit performs multimodal fusion on the emotional feature information to obtain emotional fusion features, configured as follows: The emotional feature information is mapped to a high-dimensional semantic space according to different modalities using a nonlinear projection matrix to obtain a multimodal mapping feature matrix; the multimodal mapping feature matrix is ​​subjected to modality difference elimination processing through adversarial training to obtain a first mapping feature matrix; a multi-layer multi-head cross-attention mechanism is used to establish fine-grained correlations between different modalities in the first mapping feature matrix to obtain a second mapping feature matrix; and a Transformer encoder improved based on a meta-learning strategy is used to adaptively evaluate the second mapping feature matrix according to the importance of different modalities, and dynamically assign weights according to the evaluation results to obtain the emotional fusion features.

[0196] Further optionally, the analysis unit, before performing modality difference elimination processing on the multimodal mapping feature matrix through adversarial training to obtain the first mapping feature matrix, is further configured to: establish corresponding gating unit models according to different modalities; use the gating unit models to filter and delete invalid matrix elements corresponding to different modalities from the multimodal mapping feature matrix; wherein, invalid matrix elements are those that meet the preset filtering conditions corresponding to their respective modalities.

[0197] Further optionally, before the analysis unit performs multimodal fusion on the emotional feature information to obtain the emotional fusion features, it is further configured to: if the number of image information in the emotional feature information does not meet the preset image mode threshold, then perform image data compensation on the emotional feature information.

[0198] Optionally, the analysis unit inputs the emotion fusion features into an emotion recognition model for emotion state analysis to obtain the user's predicted emotion state. This is configured to: classify the emotion fusion features using a deep multilayer perceptron to obtain the corresponding user emotion category; construct a corresponding emotion state transition matrix based on a bidirectional long short-term memory network; the emotion state transition matrix represents the dynamic evolution process between user emotions; extract the user's short-term emotion fluctuation trend and long-term emotion change trend from the emotion fusion features using a time attention mechanism; and predict the user's future emotion change tendency and corresponding predicted emotion category based on the user's emotion category, the emotion state transition matrix, the short-term emotion fluctuation trend, and the long-term emotion change trend.

[0199] Further optionally, the output unit, based on the predicted emotional state, performs confidence assessment and is configured to use Bayesian deep learning to convert the predicted emotional state into an emotional state probability distribution variable for uncertainty assessment; in the inference phase of the emotion recognition model, different emotion recognition sub-networks are constructed through multiple forward propagations, and the emotional state probability distribution variable is processed by the different emotion recognition sub-networks to obtain multiple corresponding predicted emotional states; the confidence levels corresponding to the multiple predicted emotional states are assessed based on Monte Carlo Dropout; the reliability of the emotion category is quantitatively assessed based on the confidence levels, and the target predicted emotional state of the final output is determined based on the quantitative assessment results.

[0200] Further optionally, the output unit, which generates user emotion recognition results based on the predicted emotional state and the corresponding evaluation results, is configured to: generate a corresponding attention heatmap based on the target predicted emotional state; and / or generate the visual auxiliary analysis information based on the predicted emotional state to improve the interpretability of the recognition results.

[0201] Further optionally, after generating the user emotion recognition result based on the predicted emotional state and the corresponding evaluation result, the output unit is further configured to: use the user's feedback information on the user emotion recognition result, and perform online learning based on the feedback information to obtain model optimization parameters; use the model optimization parameters in the model adjustment process of the emotion recognition model to improve the emotion recognition accuracy of the emotion recognition model.

[0202] The system can implement various steps in the above method embodiments, which will not be elaborated here.

[0203] Please see Figure 3 , Figure 3 A schematic diagram illustrating an embodiment of the electronic device provided in this application. For example... Figure 3 As shown, this application provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor 320. When the processor 320 executes the computer program 311, it implements an emotion recognition method based on a large model.

[0204] Please see Figure 4 , Figure 4 This is a schematic diagram illustrating an embodiment of a computer-readable storage medium provided in this application. For example... Figure 4 As shown, this embodiment provides a computer-readable storage medium 400 storing a computer program 411, which, when executed by a processor, implements a large-model-based emotion recognition method. It should be noted that the descriptions of each embodiment in the above embodiments have different focuses; parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments. Those skilled in the art should understand that the embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application also intends to include these modifications and variations.

Claims

1. A large-model-based emotion recognition method, characterized in that, The method includes: Acquire multimodal data to be processed; the multimodal data includes at least: text information, audio information, and image information input by the user; the image information includes at least: facial expression images and / or body movement images; The multimodal data is input into an emotion recognition model for emotion state analysis to obtain the user's predicted emotion state. This includes: classifying the emotion fusion features using a deep multilayer perceptron to obtain the corresponding user emotion category; constructing the emotion fusion features into a corresponding emotion state transition matrix based on a bidirectional long short-term memory network; the emotion state transition matrix is ​​used to represent the dynamic evolution process between user emotions; extracting the user's short-term emotion fluctuation trend and long-term emotion change trend from the emotion fusion features using a time attention mechanism; and predicting the user's emotion change tendency and corresponding predicted emotion category in future time periods based on the user's emotion category, the emotion state transition matrix, the short-term emotion fluctuation trend, and the long-term emotion change trend. The process involves assessing the confidence level of the predicted emotional state and generating a user emotion recognition result based on the predicted emotional state and the corresponding assessment result. This includes: using Bayesian deep learning to convert the predicted emotional state into a probability distribution variable for uncertainty assessment; constructing different emotion recognition sub-networks through multiple forward propagations during the inference phase of the emotion recognition model, and processing the probability distribution variable using these sub-networks to obtain multiple predicted emotional states; assessing the confidence level of the multiple predicted emotional states using Monte Carlo Dropout; quantifying the reliability of the emotion category based on the confidence level, and determining the final target predicted emotional state based on the quantification assessment result; generating a corresponding attention heatmap based on the target predicted emotional state; and the user emotion recognition result includes at least: an attention heatmap for displaying the user's emotional state from multiple dimensions, and / or visual auxiliary analysis information; the auxiliary analysis information includes at least: voice emotion analysis, text emotion analysis, and facial expression emotion analysis.

2. The emotion recognition method based on a large model according to claim 1, characterized in that, The step of inputting the multimodal data into an emotion recognition model for emotion state analysis to obtain the user's predicted emotion state includes: Extract the corresponding sentiment features from the multimodal data; the sentiment features include at least one of the following: text sentiment features, acoustic sentiment features, and visual sentiment features; The emotional feature information is fused using multimodal methods to obtain emotional fusion features; The emotion fusion features are input into the emotion recognition model to analyze the emotion state and obtain the user's predicted emotion state.

3. The emotion recognition method based on a large model according to claim 2, characterized in that, The step of extracting the corresponding sentiment feature information from the multimodal data includes: Using a generalized linear model, multi-level text sentiment feature vectors are extracted from the text information; the text sentiment feature vectors include at least: word-level sentiment feature vectors, sentence-level sentiment feature vectors, and context-level paragraph sentiment feature vectors. Multi-task learning is performed on the text sentiment feature vector to obtain the implicit micro-emotional change features in the text sentiment feature vector; the micro-emotional change features include at least: modal particle features and implicit expression features; The text sentiment feature vector and the micro-emotional change features are used as the text sentiment features.

4. The emotion recognition method based on a large model according to claim 2, characterized in that, The step of extracting the corresponding sentiment feature information from the multimodal data includes: The audio information is converted into acoustic physical vectors using the Wav2Vec model; the acoustic physical vectors include at least: spectral features and / or prosodic features; the spectral features include at least: Mel-frequency cepstral coefficients (MFCC) and spectral centroid feature information; the prosodic features include at least: pitch curve features, speech rate variation features, pause features, and breathing features; The acoustic physical vectors and the audio information are input into the HuBERT model to identify the emotional cues hidden in the user's voiceprint and obtain the acoustic emotional features.

5. The emotion recognition method based on a large model according to claim 2, characterized in that, The step of extracting the corresponding sentiment feature information from the multimodal data includes: The Vision Transformer model is used to extract initial visual features at multiple scales from the image information; the initial visual features include at least: facial key point features, facial contour features, limb contour features, posture features, motion change features, and color distribution features. The initial visual features are subjected to emotional state recognition and change trend analysis to obtain the visual emotional features.

6. The emotion recognition method based on a large model according to claim 2, characterized in that, Before extracting the corresponding sentiment feature information from the multimodal data, the method further includes: The multimodal data is split according to data type to obtain text preprocessing data, audio preprocessing data, and image preprocessing data; The preprocessed text data is subjected to preprocessing operations to obtain preprocessed text information; wherein, the preprocessing operations according to the text data type include: word segmentation, noise reduction, and standardization; The audio preprocessing data is subjected to preprocessing operations to obtain preprocessed audio information; wherein, the preprocessing operations according to the audio data type include: noise reduction, segmentation, and normalization. The image preprocessing data is subjected to preprocessing operations to obtain preprocessed image information; wherein, the preprocessing operations according to the image data type include: noise reduction, size adjustment, and color standardization.

7. The emotion recognition method based on a large model according to claim 2, characterized in that, The process of multimodal fusion of the emotional feature information to obtain emotional fusion features includes: By using a nonlinear projection matrix, the emotional feature information is mapped to a high-dimensional semantic space according to different modalities, resulting in a multimodal mapping feature matrix; By performing adversarial training, the multimodal mapping feature matrix is ​​subjected to modal difference elimination processing to obtain the first mapping feature matrix; A multi-layer, multi-head cross-attention mechanism is used to establish fine-grained correlations between different modalities in the first mapping feature matrix, thereby obtaining the second mapping feature matrix; The Transformer encoder, improved based on the meta-learning strategy, adaptively evaluates the second mapping feature matrix according to the importance of different modalities and dynamically assigns weights according to the evaluation results to obtain the emotion fusion features.

8. The emotion recognition method based on a large model according to claim 7, characterized in that, Before performing modality difference elimination processing on the multimodal mapping feature matrix through adversarial training to obtain the first mapping feature matrix, the method further includes: Establish corresponding gated unit models according to different modes; Using the gated unit model, invalid matrix elements corresponding to different modes are filtered and deleted from the multimodal mapping feature matrix; wherein, invalid matrix elements are those that satisfy the preset filtering conditions corresponding to their respective modes; Before performing multimodal fusion on the emotional feature information to obtain the emotional fusion features, the method further includes: If the amount of image information in the emotional feature information does not meet the preset image mode threshold, then image data compensation is performed on the emotional feature information.

9. An emotion recognition device based on a large model, characterized in that, The device includes the following units, wherein, The acquisition unit is configured to acquire multimodal data to be processed; the multimodal data includes at least: text information, audio information, and image information input by the user; the image information includes at least: facial expression images and / or body movement images; The analysis unit is configured to input the multimodal data into an emotion recognition model for emotion state analysis to obtain the user's predicted emotion state. This includes: classifying the emotion fusion features using a deep multilayer perceptron to obtain the corresponding user emotion category; constructing a corresponding emotion state transition matrix based on a bidirectional long short-term memory network; the emotion state transition matrix representing the dynamic evolution of user emotions; extracting the user's short-term emotion fluctuation trend and long-term emotion change trend from the emotion fusion features using a time attention mechanism; and predicting the user's future emotion change tendency and corresponding predicted emotion category based on the user's emotion category, the emotion state transition matrix, the short-term emotion fluctuation trend, and the long-term emotion change trend. The output unit is configured to perform confidence assessment based on the predicted emotional state and generate a user emotion recognition result based on the predicted emotional state and the corresponding assessment result, including: using Bayesian deep learning to convert the predicted emotional state into an emotional state probability distribution variable for uncertainty assessment; during the inference phase of the emotion recognition model, constructing different emotion recognition sub-networks through multiple forward propagations, and using the different emotion recognition sub-networks to process the emotional state probability distribution variable to obtain multiple corresponding predicted emotional states; evaluating the confidence of multiple predicted emotional states based on Monte Carlo Dropout; performing a reliability quantification assessment of the emotion category based on the confidence; and determining the final output target predicted emotional state based on the quantification assessment result; generating a corresponding attention heatmap based on the target predicted emotional state; the user emotion recognition result includes at least: an attention heatmap for displaying the user's emotional state from multiple dimensions, and / or visual auxiliary analysis information; the auxiliary analysis information includes at least: voice emotion analysis, text emotion analysis, and facial expression emotion analysis.

Citation Information

Patent Citations

  • Multi-modal emotion recognition method and system based on context awareness

    CN113947702A

  • Emotion analysis method and system based on multi-modal prefix and cross-modal attention

    CN117609882A