Multi-modal emotion recognition method and device, electronic equipment and storage medium
By using a multimodal emotion recognition method, which integrates text, audio, and video features through multiple fusion and classification modules, the problem of insufficient semantic association between modalities in existing technologies is solved, and more accurate emotion recognition is achieved.
Patent Information
- Application Number
- CN202511066415.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-18
AI Technical Summary
Existing multimodal emotion recognition methods cannot effectively mine the deep semantic and emotional connections between modalities, resulting in the fused features failing to effectively integrate the advantageous information of each modality, thus affecting the accuracy of emotion recognition.
By acquiring users' text, video, and audio data, and extracting features, the features are input into a pre-trained multimodal emotion recognition model. Multiple fusion modules are used for feature fusion and enhancement, including a first fusion module, a second fusion module, a third fusion module, and a fourth fusion module. Semantic associations between modalities are established, cross-modal consensus information is integrated, strong emotion-related signals are reinforced, and an emotion classification module is used for emotion classification.
It improves the accuracy and precision of sentiment classification results, enhances the model's generalization ability, and improves the accuracy and robustness of sentiment recognition.
Smart Images

Figure CN120974407A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of emotion recognition technology, and in particular to a multimodal emotion recognition method, apparatus, electronic device and storage medium. Background Technology
[0002] Emotion recognition technology serves as a core support for human-computer interaction and intelligent services. By analyzing users' text, voice, facial expressions, and other behavioral data, it can determine their emotional states (positive, negative, neutral, etc.) and specific emotional intensity, thereby enabling personalized responses such as emotional soothing in customer service scenarios and interactive adjustments in educational scenarios. However, single-modal emotion recognition (such as using only text or voice) is insufficient to meet accuracy requirements. Therefore, multimodal emotion fusion recognition, which integrates data from multiple sources such as text, audio, and vision, has become the mainstream approach. Its core lies in compensating for the information deficiencies of single-modal recognition by integrating multimodal information, thereby improving the robustness of emotion recognition.
[0003] However, most existing multimodal emotion recognition methods use "direct splicing" or "simple weighting" to fuse multimodal features. This method cannot uncover the deep semantic and emotional connections between modalities, resulting in the fused features failing to effectively integrate the advantageous information of each modality, thus affecting the accuracy of emotion recognition.
[0004] Therefore, there is an urgent need to propose a new method to solve the above problems. Summary of the Invention
[0005] This invention provides a multimodal emotion recognition method, apparatus, electronic device, and storage medium to improve the accuracy of emotion recognition.
[0006] In a first aspect, embodiments of the present invention provide a multimodal emotion recognition method, the method comprising:
[0007] Acquire user's text data, video data, and audio data; extract features from the text data to obtain text features, extract features from the video data to obtain facial features, and extract features from the audio data to obtain audio features;
[0008] The text features, audio features, and facial features are input into a pre-trained multimodal emotion recognition model to obtain the user's emotion classification result; the multimodal emotion recognition model includes a first fusion module, a second fusion module, a third fusion module, a fourth fusion module, and a classification module;
[0009] The first fusion module is used to fuse the audio features and the text features to obtain audio-text features; the second fusion module is used to fuse the facial features and the text features to obtain facial-text features; and the third fusion module is used to enhance the text features to obtain enhanced text features.
[0010] The fourth fusion module is used to fuse the audio text features, the facial text features, and the enhanced text features to obtain multimodal features;
[0011] The multimodal features are classified using the classification module to obtain the user's sentiment classification result.
[0012] The technical solution of this invention first acquires the user's text data, video data, and audio data; then, it extracts features from the text data to obtain text features, from the video data to obtain facial features, and from the audio data to obtain audio features, providing a data foundation for obtaining more accurate user sentiment classification results. Next, the text features, audio features, and facial features are input into a pre-trained multimodal sentiment recognition model, providing a data foundation for determining the user's sentiment classification results. Subsequently, the first fusion module fuses audio and text features to obtain audio-text features; the second fusion module fuses facial and text features to obtain facial-text features; and the third fusion module enhances the text features to obtain enhanced text features. This enhances the text features, preserving unique information from each modality (such as the acoustic properties of audio and the visual dynamics of the face), and establishes semantic connections between modalities using text features as an "intermediate hub." This allows the subsequent fourth fusion module to more efficiently mine deep intermodal connections when integrating audio-text features, facial text features, and enhanced text features, thereby improving the emotional representation ability of multimodal features and increasing the accuracy of subsequent emotional classification results. Then, the fourth fusion module fuses the audio-text features, facial text features, and enhanced text features to obtain multimodal features, integrating cross-modal consensus information and strengthening strong emotionally relevant signals. Simultaneously, it further weakens noise in the features, significantly improving the multimodal features' resistance to interference in complex environments. Building upon this foundation, the complementarity and synergy of multimodal information are enhanced, thereby improving the discriminative power and robustness of the features. This provides a clearer basis for the classification module, thus improving the accuracy and precision of sentiment classification results. Furthermore, the fused multimodal features can adapt to diverse scenario requirements, enhancing the model's generalization ability. Finally, the classification module is used to classify the multimodal features to obtain the user's sentiment classification results, improving the accuracy of the determined sentiment classification results and thus enhancing the accuracy of user sentiment recognition. Therefore, the technical solution of this invention solves the problem in existing technologies where the inability to mine deep semantic and sentiment connections between modalities leads to the inability of fused features to effectively integrate the advantageous information of each modality, thereby affecting the accuracy of sentiment recognition.
[0013] Secondly, embodiments of the present invention also provide a multimodal emotion recognition device, the device comprising:
[0014] The acquisition module is used to acquire the user's text data, video data, and audio data; to extract features from the text data to obtain text features, to extract features from the video data to obtain facial features, and to extract features from the audio data to obtain audio features;
[0015] The input module is used to input the text features, the audio features, and the facial features into a pre-trained multimodal emotion recognition model to obtain the user's emotion classification result; the multimodal emotion recognition model includes a first fusion module, a second fusion module, a third fusion module, a fourth fusion module, and a classification module;
[0016] An initial fusion module is used to perform feature fusion on the audio features and the text features using the first fusion module to obtain audio-text features; to perform feature fusion on the facial features and the text features using the second fusion module to obtain facial-text features; and to perform feature enhancement on the text features using the third fusion module to obtain enhanced text features.
[0017] A multimodal fusion module is used to fuse the audio text features, the facial text features, and the enhanced text features using the fourth fusion module to obtain multimodal features;
[0018] The sentiment classification module is used to classify the multimodal features using the classification module to obtain the sentiment classification result of the user.
[0019] Thirdly, embodiments of the present invention also provide an electronic device, the electronic device comprising:
[0020] At least one processor; and a memory communicatively connected to said at least one processor;
[0021] The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the multimodal emotion recognition method according to any embodiment of the present invention.
[0022] Fourthly, embodiments of the present invention also provide a storage medium containing computer-executable instructions, characterized in that the computer-executable instructions, when executed by a computer processor, implement the multimodal emotion recognition method described in any embodiment of the present invention.
[0023] It should be noted that the aforementioned computer instructions may be stored, in whole or in part, on a computer-readable storage medium. This computer-readable storage medium may be packaged together with the processor of the multimodal emotion recognition device, or it may be packaged separately from the processor of the multimodal emotion recognition device; this application does not impose any limitations on this.
[0024] The descriptions of the second, third, and fourth aspects in this application can be referenced to the detailed description of the first aspect; and the beneficial effects described in the second, third, and fourth aspects can be referenced to the analysis of the beneficial effects of the first aspect, which will not be repeated here.
[0025] In this application, the names of the aforementioned multimodal emotion recognition devices do not limit the devices or functional modules themselves. In actual implementation, these devices or functional modules may appear under other names. As long as the functions of each device or functional module are similar to those in this application, they fall within the scope of the claims of this application and their equivalents.
[0026] These or other aspects of this application will become more readily apparent in the following description. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 A flowchart illustrating a multimodal emotion recognition method provided in an embodiment of the present invention;
[0029] Figure 2 A flowchart of another multimodal emotion recognition method provided in an embodiment of the present invention;
[0030] Figure 3 This is a schematic diagram of the structure of a multimodal emotion recognition device provided in an embodiment of the present invention;
[0031] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0032] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.
[0033] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0034] The terms "first" and "second," etc., used in the specification and drawings of this application are used to distinguish different objects or to distinguish different treatments of the same object, rather than to describe a specific order of objects.
[0035] Furthermore, the terms "comprising" and "having," and any variations thereof, used in the description of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus.
[0036] Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but may also have additional steps not included in the figures. The process can correspond to a method, function, procedure, subroutine, subroutine, etc. Moreover, embodiments and features in the embodiments of the present invention can be combined with each other without conflict.
[0037] It should be noted that in the embodiments of this application, the words "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the words "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0038] In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0039] Figure 1 This is a flowchart illustrating a multimodal emotion recognition method provided in an embodiment of the present invention. This embodiment is applicable to situations requiring analysis of user emotions. The method can be executed by a multimodal emotion recognition device, which can be implemented in software and / or hardware. For example, the device can be an electronic device. (See reference...) Figure 1 The multimodal emotion recognition method in this embodiment specifically includes the following steps:
[0040] Step 110: Obtain the user's text data, video data, and audio data; extract features from the text data to obtain text features, extract features from the video data to obtain facial features, and extract features from the audio data to obtain audio features.
[0041] Specifically, user text data refers to a collection of information recorded and expressed in written form, such as chat logs, comments, and entered text content. User video data refers to dynamic video information containing user images. User audio data refers to data existing in sound form, including user voice information (such as speaking voice, tone of voice, etc.) and ambient sound. Text features refer to representative information extracted from text data, used to quantify the sentiment tendency of the text; for example, text features can be lexical sentiment tendency, semantic vectors, syntactic structure features, etc. Facial features refer to user facial features extracted from video data, used to reflect the user's visual emotions; for example, facial features can be facial key point coordinates (such as eye and mouth positions), expression categories (such as smiling, frowning, etc.), facial muscle movement amplitude, etc. Audio features refer to acoustic features extracted from audio data, used to quantify the emotional characteristics of speech; for example, audio features can be Mel-frequency cepstral coefficients, fundamental frequency, energy, spectral entropy, pitch, intensity, timbre, speech rate, pauses, etc.
[0042] In practice, the system first monitors the user's input status in input boxes (such as chat windows and form filling areas) on terminal devices (such as mobile phones, computers, and smart kiosks). Once a new text input is detected, the system immediately retrieves the latest text content in the input box and uses it as text data. Then, using the timestamp information attached to the text data as a reference point, it retrieves audio and video data for a pre-defined time period (e.g., 1 second). Specifically, it acquires a continuous video data stream through the terminal device's camera, then extracts video segments that meet the time requirements (the pre-defined time period before the timestamp) as video data; simultaneously, it acquires an audio data stream within the pre-defined time period through the terminal device's microphone to obtain audio data.
[0043] Next, for the text data, preprocessing is performed (such as removing special symbols, correcting typos, and word segmentation) to improve data quality. After preprocessing, text features can be extracted in two ways: one is to use text feature extraction algorithms (such as TF-IDF algorithm, BM25 algorithm, etc.) to extract features from the preprocessed text data; the other is to use pre-trained language models (such as BERT, Word2Vec, etc.) to extract features from the preprocessed text data. For video data, two extraction methods can be used: First, the video is processed frame by frame, and keyframes containing the user's face are extracted from continuous frames. Then, facial detection algorithms (such as MTCNN algorithm, single-stage detection algorithm based on deep learning, etc.) are used to locate the facial region. Subsequently, feature point detection models (such as the 68-point feature point detection model in the Dlib library, deep learning model based on heatmap regression, etc.) are used to extract the coordinate information of key points such as eyes, eyebrows, nose, and mouth. Further, expression recognition models (such as FER+ model, DenseNet model, etc.) are used to determine the expression categories such as smiling, frowning, and surprise. At the same time, the amplitude of facial muscle movement (such as the angle of mouth corners and the degree of eyebrow tilt) is calculated as dynamic features. Finally, the above data is summarized to obtain facial features. Second, pre-trained facial feature extraction models (such as VGG model, ShuffleNetV2 model, etc.) are used to extract features from the video data to obtain facial features. For audio data, preprocessing is performed first (such as framing, windowing, cutting continuous audio into short segments, and reducing spectral leakage) to optimize data quality. After preprocessing, audio features can be extracted in two ways: First, audio features can be obtained by calculating acoustic features, such as calculating Mel-frequency cepstral coefficients to capture spectral envelope features, statistically analyzing the trend of audio energy changes, and using spectral entropy to measure sound complexity. Second, pre-trained audio feature extraction models (such as deep neural network models based on MFCC) can be used directly to extract features from the audio data.
[0044] It should be noted that explicit authorization from users for the collection of such data must be obtained beforehand.
[0045] Furthermore, in practical applications, before acquiring user text, video, and audio data, the user's current interaction scenario should be determined. Common interaction scenarios include terminal application interaction, smart terminal interaction, voice call interaction, and video communication interaction. Different data acquisition strategies should be adopted for different interaction scenarios. For example, if the interaction scenario is determined to be a voice call interaction, since there is usually no direct text input and video data is generally not involved, the initial data acquired only includes the user's audio data. In this case, the audio data needs to be converted and processed using speech-to-text technology to generate the corresponding text data, while the default video features are used, before proceeding to the subsequent feature extraction stage.
[0046] In this embodiment, the above steps provide a data foundation for obtaining more accurate user sentiment classification results in the future.
[0047] Step 120: Input the text features, audio features, and facial features into the pre-trained multimodal emotion recognition model.
[0048] Specifically, a multimodal emotion recognition model refers to a model trained using historical feature data (including historical text features, historical audio features, and historical facial features) and corresponding actual emotion tags from various interaction scenarios. A multimodal emotion recognition model includes a first fusion module, a second fusion module, a third fusion module, a fourth fusion module, and a classification module. For example, the actual emotion tag could be happiness, anger, etc.
[0049] In practice, after obtaining text features, audio features, and facial features, these features can be input into a pre-trained multimodal emotion recognition model. Since the multimodal emotion recognition model is trained based on historical feature data and corresponding actual emotion labels, it can determine the user's emotion classification result based on the input user's text features, audio features, and facial features.
[0050] In this embodiment, the above steps provide a data foundation for subsequently determining the user's sentiment classification results.
[0051] Step 130: Use the first fusion module to fuse audio features and text features to obtain audio-text features; use the second fusion module to fuse facial features and text features to obtain facial-text features; use the third fusion module to enhance text features to obtain enhanced text features.
[0052] Specifically, the first fusion module refers to the module that fuses audio features and text features. The second fusion module refers to the module that fuses facial features and text features. The third fusion module refers to the module that enhances text features. Audio-text features refer to the comprehensive features obtained after fusing audio features and text features, combining information from speech emotion and text semantics. Facial-text features refer to the comprehensive features obtained after fusing facial features and text features, combining information from visual expression and text semantics. Enhanced text features refer to the features obtained after strengthening text features.
[0053] In practical implementation, a suitable feature fusion attention mechanism can be selected first based on the actual situation or requirements. For example, the feature fusion attention mechanism could be a multi-head attention mechanism, a cross-modal attention mechanism, or a self-attention mechanism. Then, the selected feature fusion attention mechanism is applied to fuse audio and text features to obtain audio-text features. Similarly, the fusion of facial and text features can follow the same logic: a suitable attention mechanism (such as a cross-attention mechanism) is selected based on the scenario requirements, and then the selected feature fusion attention mechanism is applied to fuse facial and text features to obtain facial-text features. Afterward, a suitable feature enhancement attention mechanism can be selected based on the actual situation or requirements. For example, the feature enhancement attention mechanism could be a multi-head attention mechanism, a hierarchical attention mechanism, or a local attention mechanism. Then, the selected feature enhancement attention mechanism is applied to assign differentiated weights to words within the text features, highlighting core semantic information, thereby obtaining enhanced text features.
[0054] In this embodiment, fusing audio features with text features overcomes the limitations of a single modality; fusing facial features with text features enhances the consistency between visual and semantic information. Both of these local fusion methods capture the correlation patterns between modalities, making the resulting audio-text features and facial-text features more emotionally discriminative than single-modal features. Simultaneously, enhancing text features effectively filters noise information in the text and strengthens core semantics. The enhanced text features more accurately reflect the user's semantic and emotional tendencies, providing a high-quality semantic benchmark for subsequent multimodal fusion. These steps, while preserving the unique information of each modality (such as the acoustic characteristics of audio and the visual dynamics of the face), establish semantic relationships between modalities using text features as an "intermediate hub." This allows the subsequent fourth fusion module to more efficiently mine deep relationships between modalities when integrating audio-text features, facial-text features, and enhanced text features, thereby improving the emotional representation capabilities of multimodal features and ultimately increasing the accuracy of subsequent emotional classification results.
[0055] Step 140: Use the fourth fusion module to fuse audio text features, facial text features, and enhanced text features to obtain multimodal features.
[0056] Specifically, the fourth fusion module refers to the module that fuses audio text features, facial text features, and enhanced text features. Multimodal features refer to the features formed by fusing audio text features, facial text features, and enhanced text features. They integrate information from multiple modalities and can more comprehensively and accurately represent the user's emotional state.
[0057] In practical implementation, after obtaining audio text features, facial text features, and enhanced text features, fully connected layers or convolutional layers can be used to align their dimensions, mapping them to the same feature space. Then, a suitable attention mechanism (such as multimodal cross-attention, modal attention, gated attention, cross-modal self-attention, etc.) is selected based on the actual situation or needs. The selected attention mechanism is then applied to fuse the audio text features, facial text features, and enhanced text features to obtain multimodal features. For example, if the selected attention mechanism is a multimodal cross-attention mechanism, the enhanced text features can be used as the query vector, the audio text features as the key vector, and the facial text features as the value vector. The weights of the multimodal cross-attention mechanism are calculated based on these three vectors, and then multiplied point-by-point by the facial text features to obtain the initial multimodal features. Finally, the initial multimodal features are residually connected with the enhanced text features to obtain the final multimodal features.
[0058] In addition, to reduce implementation complexity, a simpler fusion method can be adopted: directly concatenating or weighting the audio text features, facial text features, and enhanced text features to obtain multimodal features.
[0059] In this embodiment, the above steps integrate cross-modal consensus information (e.g., all three features point to "anger"), strengthening strong emotion-related signals. Simultaneously, noise in the features is further weakened (e.g., background noise in audio text features and illumination interference in facial text features), significantly improving the multimodal features' resistance to interference in complex environments. Based on this, the complementarity and synergy of multimodal information are enhanced, thereby improving the discriminative power and robustness of the features, providing clearer criteria for the classification module, and thus improving the accuracy and precision of emotion classification results. Furthermore, the fused multimodal features can adapt to diverse scenario requirements, enhancing the model's generalization ability.
[0060] Step 150: Use the classification module to classify the multimodal features to obtain the user's sentiment classification results.
[0061] Specifically, the classification module refers to the module that determines the sentiment category of multimodal features. The user's sentiment classification result refers to the final judgment made by the multimodal sentiment recognition model on the user's sentiment. For example, the sentiment classification result can be a sentiment polarity classification such as positive, negative, or neutral, or a specific emotion category such as joy, anger, or surprise.
[0062] In practice, the multimodal features can be linearly transformed first to obtain preliminary transformed features; then the preliminary transformed features can be nonlinearly compressed to obtain compressed features; then the compressed features can be linearly transformed using the second mapping unit to obtain the sentiment intensity mapping value of the multimodal features; and the sentiment classification result of the user can be determined based on the sentiment intensity mapping value.
[0063] Furthermore, after obtaining the user's sentiment classification results, a large language model for intelligent bank responses can be constructed based on a knowledge base of the banking business domain, combining retrieval enhancement generation technology and fine-tuning technology, and then trained. Subsequently, the acquired user text data, sentiment classification results, user information, and current business scenario (such as financial consultation) are input into the trained large language model. This allows for the generation of "warm" response language tailored to the user's text data, while ensuring professionalism is not compromised. This enables users to feel cared for and understood during their interaction with the system, thereby improving the user experience.
[0064] In this embodiment, the accuracy of the determined emotion classification result is improved through the above steps, thereby improving the accuracy of the user's emotion recognition.
[0065] The multimodal emotion recognition method provided in this invention first acquires the user's text data, video data, and audio data. It then extracts features from the text data to obtain text features, from the video data to obtain facial features, and from the audio data to obtain audio features, providing a data foundation for obtaining more accurate emotion classification results for the user. Next, the text features, audio features, and facial features are input into a pre-trained multimodal emotion recognition model, providing a data foundation for determining the user's emotion classification results. Subsequently, the first fusion module fuses audio and text features to obtain audio-text features; the second fusion module fuses facial and text features to obtain facial-text features; and the third fusion module enhances the text features to obtain enhanced text features. This enhances the text features, preserving unique information from each modality (such as the acoustic properties of audio and the visual dynamics of the face) while using text features as an "intermediate hub" to establish semantic relationships between modalities. This allows the subsequent fourth fusion module to more efficiently mine deep intermodal relationships when integrating audio-text features, facial text features, and enhanced text features, thereby improving the emotional representation ability of multimodal features and increasing the accuracy of subsequent emotional classification results. Then, the fourth fusion module fuses the audio-text features, facial text features, and enhanced text features to obtain multimodal features, integrating cross-modal consensus information and strengthening strong emotionally relevant signals. Simultaneously, it further weakens noise in the features, significantly improving the multimodal features' resistance to interference in complex environments. Building upon this foundation, the complementarity and synergy of multimodal information are enhanced, thereby improving the discriminative power and robustness of the features. This provides a clearer basis for the classification module, thus improving the accuracy and precision of sentiment classification results. Furthermore, the fused multimodal features can adapt to diverse scenario requirements, enhancing the model's generalization ability. Finally, the classification module is used to classify the multimodal features to obtain the user's sentiment classification results, improving the accuracy of the determined sentiment classification results and thus enhancing the accuracy of user sentiment recognition. Therefore, the technical solution of this invention solves the problem in existing technologies where the inability to mine deep semantic and sentiment connections between modalities leads to the inability of fused features to effectively integrate the advantageous information of each modality, thereby affecting the accuracy of sentiment recognition.
[0066] Figure 2 A flowchart of another multimodal emotion recognition method provided by an embodiment of the present invention is shown. This embodiment is a specific implementation based on the above embodiments. In this embodiment, the method may further include:
[0067] Step 210: Obtain the user's text data, video data, and audio data; extract features from the text data to obtain text features, extract features from the video data to obtain facial features, and extract features from the audio data to obtain audio features.
[0068] Furthermore, acquiring the user's text data, video data, and audio data includes: acquiring the user's text data at the current moment as text data; and acquiring continuous video data and audio data within a preset time period before and after the timestamp based on the timestamp of the text data to obtain video data and audio data.
[0069] Specifically, a preset time period refers to a duration set in advance based on actual conditions or needs. For example, a preset time period can be 1 second, 2 seconds, etc.
[0070] In practice, when a user interacts (such as entering a chat message, clicking a preset reply, or making a phone call), the text content they directly output is recorded. If the interaction is a phone call, the audio content is converted to text, and the data is automatically timestamped. This timestamped text content is directly used as text data and its metadata (such as text length and whether it contains special characters) is stored. Then, based on the timestamp, all consecutive frames within a preset time period before and after the timestamp are selected from the continuously acquired video stream and stitched together to form a complete video segment, thus obtaining the video data. Similarly, audio segments within a preset time period before and after the timestamp are selected from the continuously acquired audio stream, thus obtaining the audio data.
[0071] It should be noted that both the video and audio streams are continuously sampled at a fixed frequency (e.g., 30 frames per second for video and 16 kHz for audio), and each video frame and each audio sample carries a corresponding timestamp to ensure time synchronization with the text data.
[0072] In this embodiment, the above steps ensure the time synchronization of multimodal data, improve the reliability of feature association, and avoid mis-association of sentiment features due to time misalignment, providing a reliable data foundation for subsequent multimodal feature fusion. Furthermore, the preset time period can be dynamically adjusted according to the actual scenario, effectively balancing data volume and processing efficiency.
[0073] Step 211: Perform a nonlinear transformation on the text features to obtain intermediate text features; perform a nonlinear transformation on the audio features to obtain intermediate audio features; perform a nonlinear transformation on the facial features to obtain intermediate facial features.
[0074] Specifically, intermediate text features refer to the features obtained after nonlinear transformation of text features. Intermediate audio features refer to the features obtained after nonlinear transformation of audio features. Intermediate facial features refer to the features obtained after nonlinear transformation of facial features.
[0075] In practical implementation, independent nonlinear transformation units can be configured for each of the three features. Then, the corresponding units are used to perform nonlinear transformations on each feature. Specifically, fully connected layers or convolutional layers containing nonlinear activation functions can be used to perform nonlinear transformations on text features to obtain intermediate text features; recurrent neural network layers or fully connected layers with activation functions can be used to perform nonlinear transformations on audio features to obtain intermediate audio features; and convolutional layers or attention mechanisms combined with nonlinear activation functions can be used to perform nonlinear transformations on facial features to obtain intermediate facial features.
[0076] In this embodiment, the above steps achieve independent nonlinear transformation processing of the three features. This processing method ensures that each transformation can adapt to the characteristics of the corresponding modality, so that the output intermediate text features, intermediate audio features, and intermediate facial features can not only retain the core information of the original features, but also focus more on the complex patterns related to emotions within the modality (such as metaphorical emotions in text, non-stationary acoustic features in audio, transient facial expressions, etc.). This lays the foundation for subsequent multimodal feature dimension alignment and fusion, and effectively avoids the problem of affecting the accuracy of cross-modal association due to linear redundancy or noise interference of the original features.
[0077] Step 212: Map the intermediate text features to the target dimension to obtain the updated text features; map the intermediate audio features to the target dimension to obtain the updated audio features; map the intermediate facial features to the target dimension to obtain the updated facial features.
[0078] Specifically, the target dimension refers to the feature dimension that is pre-defined based on the actual situation or needs.
[0079] In practical implementation, independent dimension mapping units (such as fully connected layers) can be configured for the three intermediate features (i.e., intermediate text features, intermediate audio features, and intermediate facial features). Each unit is only responsible for the dimension transformation of a single modality to adapt to the feature distribution of the corresponding modality. Then, the corresponding unit is used to perform mapping processing on each intermediate feature. For example, the dimension mapping unit of the intermediate text feature takes the original dimension of its intermediate text feature (e.g., 768 dimensions) as input and outputs the target dimension (e.g., 512 dimensions) as output; the dimension mapping unit of the intermediate audio feature takes the original dimension of its intermediate audio feature (e.g., 1024 dimensions) as input and outputs the target dimension (e.g., 512 dimensions) as output; and the dimension mapping unit of the intermediate facial feature takes the original dimension of its intermediate facial feature (e.g., 2048 dimensions) as input and outputs the target dimension (e.g., 512 dimensions) as output.
[0080] In addition, to avoid overfitting, regularization constraints (such as L2 regularization) can be added during the mapping process to ensure that the mapped features are both concise and retain key information.
[0081] In this embodiment, through the above steps, the updated text features, audio features, and facial features are all unified into the target dimension while retaining the core emotional features of the corresponding modality. That is, they are embedded into the same dimensional space, which solves the problem of fusion difficulties caused by the inconsistency of the original feature dimensions (such as the inability to directly calculate the similarity between high-dimensional audio features and low-dimensional text features). This makes it easier for the mapped features to achieve cross-modal association, thereby improving the accuracy and processing efficiency of the subsequent multimodal emotion recognition model.
[0082] Step 213: Input the text features, audio features, and facial features into the pre-trained multimodal emotion recognition model.
[0083] Furthermore, the training process of the multimodal emotion recognition model includes: first, acquiring historical feature data (including historical text features, historical audio features, and historical facial features) and corresponding actual emotion tags for each interaction scenario. This data is then cleaned and preprocessed (e.g., unified in dimensionality) to ensure it meets the model's input requirements. Next, the acquired historical feature data and actual emotion tags are used as training data to train the deep learning model. This process typically includes two steps: forward propagation and backpropagation. Forward propagation involves passing the input data through the model to obtain a prediction result, and then calculating the loss between the prediction result and the true target. Backpropagation updates the model's parameters based on the loss function to reduce the gap between the prediction result and the true target. The deep learning model can be a combination of Transformer, recurrent neural networks, and convolutional neural networks, etc., and this embodiment of the invention does not impose any limitations on this. Finally, the model can be optimized based on the backpropagation algorithm (e.g., choosing gradient descent or Adam to determine hyperparameters such as the learning rate) until the loss function converges, resulting in the multimodal emotion recognition model. Therefore, by training a multimodal emotion recognition model with historical feature data (including historical text features, historical audio features, and historical facial features) and corresponding actual emotion tags in various interaction scenarios, the multimodal emotion recognition model can be used to determine the user's situation classification result based on the user's text features, audio features, and facial features.
[0084] Step 214: Use the first fusion module to perform feature fusion on audio features and text features to obtain audio-text features; use the second fusion module to perform feature fusion on facial features and text features to obtain facial-text features; use the third fusion module to enhance the text features to obtain enhanced text features.
[0085] Optionally, the first fusion module includes a first attention mechanism, where the query vector of the first attention mechanism is text features, the key vector of the first attention mechanism is audio features, and the value vector of the first attention mechanism is audio features; the second fusion module includes a second attention mechanism, where the query vector of the second attention mechanism is text features, the key vector of the second attention mechanism is facial features, and the value vector of the second attention mechanism is facial features; the third fusion module includes a third attention mechanism, where the query vector of the third attention mechanism is text features, the key vector of the third attention mechanism is text features, and the value vector of the third attention mechanism is text features. Here, the first attention mechanism refers to the attention mechanism applied to the first fusion module. The second attention mechanism refers to the attention mechanism applied to the second fusion module. The third attention mechanism refers to the attention mechanism applied to the third fusion module. For example, the first, second, and third attention mechanisms can all be multi-head self-attention mechanisms, with specific calculation formulas as follows:
[0086]
[0087] Where MultiHead(Q,K,V) represents the output features of the multi-head self-attention mechanism, Q is the query vector of the multi-head self-attention mechanism, K is the key vector of the multi-head self-attention mechanism, V is the value vector of the multi-head self-attention mechanism, and h is the number of "heads". Where i is the scaling factor, and W is the index of the "header" index. o The output weight matrix is T, where T is the matrix transpose, Concat means concatenation along the feature dimension, and softmax is the activation function.
[0088] For example, a first attention mechanism can be used (the query vector of the first attention mechanism is the text feature, the key vector of the first attention mechanism is the audio feature, and the value vector of the first attention mechanism is the audio feature) to enable the text feature to actively associate with the audio feature, achieving preliminary cross-modal fusion and obtaining the first initial fused feature; then, the text feature and the first initial fused feature are added element-wise to obtain the first additive fused feature, and then the first additive fused feature is normalized to obtain the first intermediate fused feature; then, a feedforward neural network is used to perform a position-wise nonlinear transformation on the first intermediate fused feature to obtain the first enhanced fused feature; finally, the first enhanced fused feature and the first intermediate fused feature are fused using a residual connection, and the fused feature is normalized to obtain the audio-text feature. Similarly, a second attention mechanism (where the query vector is text features, the key vector is facial features, and the value vector is facial features) can be used to actively associate text features with facial features, achieving initial cross-modal fusion and obtaining the second initial fused features. Then, the text features are added element-wise to the second initial fused features to obtain the second additive fused features. These features are then normalized to obtain the second intermediate fused features. Next, a feedforward neural network is used to perform a position-wise nonlinear transformation on the second intermediate fused features to obtain the second enhanced fused features. Finally, a residual connection is used to fuse the second enhanced fused features and the second intermediate fused features, and the fused features are normalized to obtain the facial text features. Therefore, by using the first and second attention mechanisms to actively associate text features with audio features and facial features respectively, the fusion process can focus on cross-modal information related to text semantics. This allows for targeted extraction of cross-modal information, filtering out irrelevant information, and retaining cross-modal features that are helpful in understanding text semantics. This results in audio text features and facial text features with richer semantic expression, thereby improving the accuracy of subsequent emotion recognition results.
[0089] Alternatively, a third attention mechanism can be used (where the query vector, key vector, and value vector are all text features) to enable self-attention interaction among text features, achieving initial feature enhancement. This results in initial enhanced features. The text features are then element-wise added to the initial enhanced features to obtain additive enhanced features, which are then normalized to obtain normalized enhanced features. A feedforward neural network is then used to perform a nonlinear transformation on the normalized enhanced features to obtain deep enhanced features. Finally, residual connections are used to fuse the deep enhanced features and normalized enhanced features, and the fused features are normalized to obtain enhanced text features. Therefore, utilizing the third attention mechanism enables self-attention interaction among text features, allowing for closer connections between different parts of the text, strengthening semantic logic, better expressing the inherent semantic structure of text features, and better integrating with features from other modalities. This provides more accurate and richer semantic cues, achieves more effective cross-modal information interaction, and ultimately improves the accuracy of sentiment recognition results.
[0090] Step 215: Use the fourth fusion module to fuse audio text features, facial text features, and enhanced text features to obtain multimodal features.
[0091] Step 216: Use the classification module to classify the multimodal features and obtain the user's sentiment classification results.
[0092] Optionally, the classification module includes a first mapping unit, an activation unit, and a second mapping unit.
[0093] Further, step 216 may specifically include: using the first mapping unit to perform a linear transformation on the multimodal features to obtain preliminary transformed features; using the activation unit to perform nonlinear compression on the preliminary transformed features to obtain compressed features; using the second mapping unit to perform a linear transformation on the compressed features to obtain the sentiment intensity mapping value of the multimodal features; and determining the user's sentiment classification result based on the sentiment intensity mapping value.
[0094] Specifically, the first mapping unit value refers to the unit that performs a linear transformation on the multimodal features. The activation unit refers to the unit that performs nonlinear compression on the initially transformed features. The second mapping unit refers to the unit that performs a linear transformation on the compressed features. The sentiment intensity mapping value refers to a numerical value used to represent the degree of association between the multimodal features and different sentiment categories; for example, the sentiment intensity mapping value can be 2, and its range can be [-3, 3]. The initially transformed features refer to the features obtained after performing a linear transformation on the multimodal features. The compressed features refer to the features obtained after performing nonlinear compression on the initially transformed features.
[0095] In practice, a fully connected layer can be used to perform a linear transformation on the multimodal features to obtain preliminary transformed features. The specific calculation formula is as follows:
[0096] F1 = W1F + b1
[0097] Where F represents the multimodal features, F1 represents the initial transformation features, W1 represents the learnable weight matrix of the fully connected layer (i.e., the first mapping unit), and b1 represents the learnable bias matrix of the fully connected layer (i.e., the first mapping unit).
[0098] Then, the initial transformed features are nonlinearly compressed using an activation function to obtain compressed features. The specific calculation formula is as follows:
[0099] F2 = σ(F1)
[0100] Where F2 is the compression feature and σ(x) is the activation function, for example: σ(x) = tanh(x).
[0101] Next, another fully connected layer is used to perform a linear transformation on the compressed features to obtain the sentiment intensity mapping value of the multimodal features. The specific calculation formula is as follows:
[0102] y = W²F² + b²
[0103] Where y is the sentiment intensity mapping value, W2 is the learnable weight matrix of the fully connected layer (i.e., the second mapping unit), and b2 is the learnable bias matrix of the fully connected layer (i.e., the second mapping unit).
[0104] Finally, the obtained sentiment intensity mapping values are mapped to specific sentiment categories to obtain the user's sentiment classification result. This process relies on the "sentiment intensity mapping value - sentiment category" association established during the model training phase.
[0105] In this embodiment, a first mapping unit performs a linear transformation on the multimodal features to obtain preliminary transformed features. This process removes noise from the multimodal features and strengthens the emotion discriminative features. Next, the preliminary transformed features are nonlinearly compressed to obtain compressed features. This enhances the features' ability to distinguish subtle emotional differences, avoids gradient vanishing or exploding, and implicitly normalizes the features, improving the stability of the subsequent linear transformation performed by the second mapping unit. Then, the second mapping unit performs a linear transformation on the compressed features, projecting the compressed general features onto a dimension matching the number of emotion categories. This yields the emotion intensity mapping value for the multimodal features, ensuring consistency between emotion intensity and category determination. Finally, the emotion classification result for the user is determined based on the emotion intensity mapping value, thus achieving the classification of the user's emotions.
[0106] Furthermore, after obtaining the user's sentiment classification result, the process also includes: determining the user's text sentiment classification result based on text features; determining the user's audio sentiment classification result based on audio features; determining the user's facial sentiment classification result based on facial features; and updating the sentiment classification result using the text sentiment classification result, audio sentiment classification result, and facial sentiment classification result to obtain the updated sentiment classification result.
[0107] Specifically, text sentiment classification results refer to user sentiment classification results determined solely based on text features. Audio sentiment classification results refer to user sentiment classification results determined solely based on audio features. Facial sentiment classification results refer to user sentiment classification results determined solely based on facial features. Updated sentiment classification results refer to more accurate and comprehensive user sentiment classification results obtained by updating the initial sentiment classification results by combining text sentiment classification results, audio sentiment classification results, and facial sentiment classification results.
[0108] In the specific implementation, after obtaining the user's sentiment classification result, text features can be input into a pre-trained text sentiment recognition model to obtain the user's text sentiment classification result; audio features can be input into a pre-trained audio sentiment recognition model to obtain the user's audio sentiment classification result; facial features can be input into a pre-trained facial sentiment recognition model to obtain the user's facial sentiment classification result. Finally, the text sentiment classification result, audio sentiment classification result, and facial sentiment classification result are weighted and summed with the original sentiment classification result to obtain the updated sentiment classification result. The pre-trained text sentiment recognition model refers to the model trained based on historical text features and corresponding actual sentiment labels for each interaction scenario. The pre-trained audio sentiment recognition model refers to the model trained based on historical audio features and corresponding actual sentiment labels for each interaction scenario. The pre-trained facial sentiment recognition model refers to the model trained based on historical facial features and corresponding actual sentiment labels for each interaction scenario.
[0109] It should be noted that the text sentiment classification results, audio sentiment classification results, and facial sentiment classification results in this process, as well as the original sentiment classification results (the results obtained by the multimodal sentiment recognition model), all include the user's specific sentiment classification category and its corresponding sentiment intensity mapping value. Furthermore, the value range of each sentiment intensity mapping value is kept consistent, so as to perform weighted processing on the above four types of results to obtain the updated sentiment classification results.
[0110] In this embodiment, the above steps can be used to perform "cross-validation" between the results of the multimodal model and the results of the single-modal model, correcting potential errors in the output results of the multimodal emotion recognition model, and thus obtaining more accurate emotion classification results.
[0111] Optionally, to improve processing efficiency and reduce time consumption, after obtaining the user's text features, audio features, and facial features, these three features can be input into a pre-trained multimodal emotion recognition model to obtain the user's emotion classification result. At the same time, the text emotion classification result can be determined based on the text features, the audio emotion classification result can be determined based on the audio features, and the facial emotion classification result can be determined based on the facial features. Finally, the final emotion classification result is obtained based on the above text, audio, and facial emotion classification results and the user emotion classification result output by the multimodal emotion recognition model.
[0112] Step 217: Input the user's sentiment classification result, user information, and current business scenario into the pre-trained response model, so that the response model can generate the response statement corresponding to the user's text data based on the preset prompt words corresponding to the current business scenario, the user's sentiment classification result, and the user's information.
[0113] Specifically, the response model refers to a pre-trained language model used to generate a response statement corresponding to the user's text data based on the user's sentiment classification, user information, and the current business scenario, combined with preset prompts. Preset prompts refer to fixed text templates pre-defined for different business scenarios, used to constrain the generation style and content scope of the response model. User information refers to basic user-related data (such as user account, age, gender, preference tags, etc.) and historical interaction data (such as text data from the past three interactions and corresponding response content, feedback ratings, etc.). The current business scenario refers to the specific scenario in which the user interacts with the system, for example: the current business scenario could be e-commerce customer service consultation (such as after-sales issues), banking customer service scenario (such as account inquiry and management, money transfer, credit and loan business, credit card business, investment and wealth management business, etc.). The response statement refers to the natural language sentence ultimately generated by the response model to respond to the user's current text data.
[0114] In practice, user information and the current business scenario can be obtained first through the user system interface. Then, the user's sentiment classification result, user information, and current business scenario are input into a pre-trained response model. After receiving the above information, the response model will extract the user's text data from the context (such as the latest dialogue history) and determine the corresponding preset prompt words based on the current business scenario. Then, based on the preset prompt words, user information, sentiment classification result, and extracted user text data, candidate response statements are generated. Subsequently, the response model itself scores the candidate response statements and selects the response statement with the highest score as the final output, thus obtaining the response statement corresponding to the user's text data.
[0115] Optionally, the acquired user text data, user sentiment classification results, user information, and current business scenario can be input into a pre-trained response model. In this way, the response model can directly generate a response statement based on the preset prompts corresponding to the current business scenario, combined with the user sentiment classification results and user information. This eliminates the need for the model to extract user text data from the context (such as the latest dialogue history), thereby reducing the model's data processing time, speeding up response generation, and reducing the information bias that may occur when the response model extracts text from the dialogue history, making the response statement more in line with the user's input intent and business scenario requirements.
[0116] Furthermore, the response model training process is as follows: First, historical text data, corresponding sentiment tags, and actual responses from different customers in various application scenarios are collected and used as the training sample set. Next, prompt words for each scenario are formulated based on actual conditions or the needs of specific application scenarios. Then, a suitable pre-trained large language model (such as ChatGLM, PaLM, LLaMa, etc.) is selected according to requirements, and a knowledge base containing knowledge of the intelligent response domain and banking business domain is imported into the selected large language model to enable it to learn the domain knowledge. The response model is then constructed by combining retrieval enhancement techniques and fine-tuning techniques. The aforementioned training sample set and prompt words for each scenario are then input into the constructed response model. The response model calculates the loss function value by comparing its output with the expected results in the training sample set. The model parameters are then iteratively adjusted based on the loss function value until a preset stopping condition is reached (such as reaching the maximum number of iterations). After completing the aforementioned training process, the model can be further fine-tuned based on the final loss function value and model performance (such as full parameter fine-tuning, adapter fine-tuning, incremental learning, etc.) to obtain the final trained response model.
[0117] In practical applications, AutoGPT technology and prompt word engineering can also be used to optimize the preset prompt words in the response model, thereby generating higher quality response statements.
[0118] In existing intelligent customer service response models, the reliance on fixed script templates leads to static and formulaic interaction strategies. This makes them ill-suited to adapting to varying user emotions and complex business scenarios. Furthermore, their inability to accurately identify user emotions often results in stiff and rigid responses. In contrast, this embodiment enhances the emotional relevance of generated responses through the aforementioned steps, making them more aligned with the user's current emotional state and improving the communication experience. It also increases the personalization of responses, enhancing the user's sense of exclusivity. Secondly, it improves the accuracy of responses within business scenarios, effectively reducing information discrepancies and ensuring close alignment between response content and specific business contexts. Moreover, this method can quickly output compliant responses without human intervention, significantly reducing labor costs. Finally, it generates more "warm" responses, allowing users to feel cared for and understood during their interaction with the system.
[0119] The multimodal emotion recognition method provided in this invention first acquires user text data, video data, and audio data. It then extracts features from the text data to obtain text features, from the video data to obtain facial features, and from the audio data to obtain audio features, providing a data foundation for obtaining more accurate emotion classification results. Next, it performs a nonlinear transformation on the text features to obtain intermediate text features; a nonlinear transformation on the audio features to obtain intermediate audio features; and a nonlinear transformation on the facial features to obtain intermediate facial features. This ensures that each transformation adapts to the characteristics of the corresponding modality, so that the output intermediate text features, intermediate audio features, and intermediate facial features retain the core information of the original features while focusing more on complex emotion-related patterns within the modality. This lays the foundation for subsequent multimodal feature dimension alignment and fusion, effectively avoiding the problem of cross-modal association accuracy being affected by linear redundancy or noise interference in the original features. Then, the intermediate text features are mapped to the target dimension to obtain updated text features; the intermediate audio features are mapped to the target dimension to obtain updated audio features; and the intermediate facial features are mapped to the target dimension to obtain updated facial features. This ensures that the updated text, audio, and facial features, while retaining the core emotional features of their respective modalities, are all unified to the target dimension, i.e., embedded into the same dimensional space. This solves the fusion difficulty caused by the inconsistency of the original feature dimensions, making the mapped features easier to achieve cross-modal association, thereby improving the accuracy and processing efficiency of the subsequent multimodal emotion recognition model. The text, audio, and facial features are then input into the pre-trained multimodal emotion recognition model, providing a data foundation for subsequently determining the user's emotion classification results. Subsequently, the first fusion module fuses audio and text features to obtain audio-text features; the second fusion module fuses facial and text features to obtain facial-text features; and the third fusion module enhances the text features to obtain enhanced text features. While preserving unique information of each modality (such as the acoustic characteristics of audio and the visual dynamics of the face), text features serve as an "intermediate hub" to establish semantic relationships between modalities. This allows the subsequent fourth fusion module to more efficiently mine deep intermodal relationships when integrating audio-text features, facial text features, and enhanced text features, thereby improving the emotional representation ability of multimodal features and thus increasing the accuracy of subsequent emotional classification results. Then, the fourth fusion module fuses audio-text features, facial text features, and enhanced text features to obtain multimodal features, integrating cross-modal consensus information and strengthening strong emotionally relevant signals; simultaneously, it further weakens noise in the features, significantly improving the multimodal features' resistance to interference in complex environments.Building upon this foundation, the complementarity and synergy of multimodal information are enhanced, thereby improving the discriminative power and robustness of features. This provides clearer criteria for the classification module, improving the accuracy and precision of sentiment classification results. Furthermore, the fused multimodal features can adapt to diverse scenario requirements, enhancing the model's generalization ability. Subsequently, the classification module is used to classify the multimodal features, obtaining the user's sentiment classification results, improving the accuracy of the determined sentiment classification results, and thus enhancing the accuracy of user sentiment recognition. Finally, the user's sentiment classification results, user information, and the current business scenario are input into a pre-trained response model. This allows the response model to generate response statements corresponding to the user's text data based on preset prompts for the current business scenario, the user's sentiment classification results, and the user's information. This enhances the sentiment adaptability of the generated response statements, making them more aligned with the user's current emotional state, thereby improving the user's communication experience. Simultaneously, it increases the personalization of responses, enhancing the user's sense of exclusivity. Secondly, it improves the accuracy of responses in business scenarios, effectively reducing information bias and ensuring that the response content closely matches the specific business scenario. Furthermore, this method can quickly generate compliant responses without human intervention, significantly reducing labor costs. Ultimately, it can generate more "warm" responses, allowing users to feel cared for and understood during their interaction with the system.
[0120] Figure 3 This is a schematic diagram of the structure of a multimodal emotion recognition device provided in an embodiment of the present invention. This device belongs to the same inventive concept as the multimodal emotion recognition methods in the above embodiments. For details not described in detail in the embodiments of the multimodal emotion recognition device, please refer to the embodiments of the above multimodal emotion recognition methods.
[0121] like Figure 3 As shown, the device includes:
[0122] The acquisition module 310 is used to acquire the user's text data, video data, and audio data; perform feature extraction on the text data to obtain text features, perform feature extraction on the video data to obtain facial features, and perform feature extraction on the audio data to obtain audio features;
[0123] The input module 320 is used to input the text features, the audio features and the facial features into a pre-trained multimodal emotion recognition model, the multimodal emotion recognition model including a first fusion module, a second fusion module, a third fusion module, a fourth fusion module and a classification module;
[0124] The initial fusion module 330 is used to perform feature fusion on the audio features and the text features using the first fusion module to obtain audio-text features; to perform feature fusion on the facial features and the text features using the second fusion module to obtain facial-text features; and to perform feature enhancement on the text features using the third fusion module to obtain enhanced text features.
[0125] The multimodal fusion module 340 is used to fuse the audio text features, the facial text features, and the enhanced text features using the fourth fusion module to obtain multimodal features;
[0126] The sentiment classification module 350 is used to classify the multimodal features using the classification module to obtain the sentiment classification result of the user.
[0127] Based on the above embodiments, the first fusion module includes a first attention mechanism, wherein the query vector of the first attention mechanism is the text feature, the key vector of the first attention mechanism is the audio feature, and the value vector of the first attention mechanism is the audio feature; the second fusion module includes a second attention mechanism, wherein the query vector of the second attention mechanism is the text feature, the key vector of the second attention mechanism is the facial feature, and the value vector of the second attention mechanism is the facial feature; the third fusion module includes a third attention mechanism, wherein the query vector of the third attention mechanism is the text feature, the key vector of the third attention mechanism is the text feature, and the value vector of the third attention mechanism is the text feature.
[0128] Based on the above embodiments, the classification module includes a first mapping unit, an activation unit, and a second mapping unit. The emotion classification module 350 is specifically used for:
[0129] The first mapping unit is used to perform a linear transformation on the multimodal features to obtain preliminary transformed features;
[0130] The activation unit is used to perform nonlinear compression on the preliminary transformation features to obtain compressed features;
[0131] The compressed features are linearly transformed using the second mapping unit to obtain the emotional intensity mapping value of the multimodal features;
[0132] The user's emotional classification result is determined based on the emotional intensity mapping value.
[0133] Based on the above embodiments, the device further includes:
[0134] The result update module is used to determine the user's text sentiment classification result based on the text features after obtaining the user's sentiment classification result; determine the user's audio sentiment classification result based on the audio features; determine the user's facial sentiment classification result based on the facial features; and update the sentiment classification result using the text sentiment classification result, the audio sentiment classification result, and the facial sentiment classification result to obtain the updated sentiment classification result.
[0135] Based on the above embodiments, the device further includes:
[0136] The response module, after obtaining the user's sentiment classification result, inputs the user's sentiment classification result, the user's information, and the current business scenario into a pre-trained response model, so that the response model generates a response statement corresponding to the user's text data based on the preset prompt words corresponding to the current business scenario, the user's sentiment classification result, and the user's information.
[0137] Based on the above embodiments, the acquisition module 310 acquires the user's text data, video data, and audio data, including:
[0138] Obtain the text data corresponding to the user at the current moment, and use it as the text data;
[0139] Based on the timestamp of the text data, continuous video data and audio data within a preset time period before and after the timestamp are obtained to obtain the video data and the audio data.
[0140] Based on the above embodiments, the device further includes:
[0141] The feature update module is used to perform a nonlinear transformation on the text features to obtain intermediate text features before inputting the text features, audio features, and facial features into a pre-trained multimodal emotion recognition model to obtain the user's emotion classification result. This transformation involves: performing a nonlinear transformation on the audio features to obtain intermediate audio features; performing a nonlinear transformation on the facial features to obtain intermediate facial features; mapping the intermediate text features to a target dimension to obtain updated text features; mapping the intermediate audio features to a target dimension to obtain updated audio features; and mapping the intermediate facial features to a target dimension to obtain updated facial features.
[0142] The multimodal emotion recognition device provided in the embodiments of the present invention can execute the multimodal emotion recognition method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.
[0143] It is worth noting that in the embodiments of the multimodal emotion recognition device described above, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.
[0144] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Figure 4 A block diagram of an exemplary electronic device 4 suitable for implementing embodiments of the present invention is shown. Figure 4 The electronic device 4 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0145] like Figure 4 As shown, electronic device 4 is represented in the form of a general-purpose computing electronic device. The components of electronic device 4 may include, but are not limited to: one or more processors or processing units 16, system memory 28, and bus 18 connecting different system components (including system memory 28 and processing unit 16).
[0146] Bus 18 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0147] Electronic device 4 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by electronic device 4, including volatile and non-volatile media, removable and non-removable media.
[0148] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Electronic device 4 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (… Figure 4 Not shown; usually referred to as a "hard drive"). Although Figure 4Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. System memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0149] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in system memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of the present invention.
[0150] Electronic device 4 can also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), and with one or more devices that enable a user to interact with electronic device 4, and / or with any device that enables electronic device 4 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed through input / output (I / O) interface 22. Furthermore, electronic device 4 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) through network adapter 20. Figure 4 As shown, network adapter 20 communicates with other modules of electronic device 4 via bus 18. It should be understood that, although... Figure 4 Not shown, it can be combined with electronic device 4 to use other hardware and / or software modules, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0151] Processing unit 16 executes various functional applications and page displays by running programs stored in system memory 28, such as implementing the multimodal emotion recognition method provided in this embodiment of the invention, which includes:
[0152] Acquire user's text data, video data, and audio data; extract features from the text data to obtain text features, extract features from the video data to obtain facial features, and extract features from the audio data to obtain audio features;
[0153] The text features, audio features, and facial features are input into a pre-trained multimodal emotion recognition model, which includes a first fusion module, a second fusion module, a third fusion module, a fourth fusion module, and a classification module.
[0154] The first fusion module is used to fuse the audio features and the text features to obtain audio-text features; the second fusion module is used to fuse the facial features and the text features to obtain facial-text features; and the third fusion module is used to enhance the text features to obtain enhanced text features.
[0155] The fourth fusion module is used to fuse the audio text features, the facial text features, and the enhanced text features to obtain multimodal features;
[0156] The multimodal features are classified using the classification module to obtain the user's sentiment classification result.
[0157] Of course, those skilled in the art will understand that the processor can also implement the technical solutions of the multimodal emotion recognition method provided in any embodiment of the present invention.
[0158] This invention provides a computer-readable storage medium storing a computer program thereon. When executed by a processor, the program implements, for example, the multimodal emotion recognition method provided in this invention, which includes:
[0159] Acquire user's text data, video data, and audio data; extract features from the text data to obtain text features, extract features from the video data to obtain facial features, and extract features from the audio data to obtain audio features;
[0160] The text features, audio features, and facial features are input into a pre-trained multimodal emotion recognition model, which includes a first fusion module, a second fusion module, a third fusion module, a fourth fusion module, and a classification module.
[0161] The first fusion module is used to fuse the audio features and the text features to obtain audio-text features; the second fusion module is used to fuse the facial features and the text features to obtain facial-text features; and the third fusion module is used to enhance the text features to obtain enhanced text features.
[0162] The fourth fusion module is used to fuse the audio text features, the facial text features, and the enhanced text features to obtain multimodal features;
[0163] The multimodal features are classified using the classification module to obtain the user's sentiment classification result.
[0164] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0165] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0166] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0167] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0168] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computing device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0169] Furthermore, the acquisition, storage, use, and processing of data in the technical solution of this invention all comply with relevant laws and regulations.
[0170] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. A multimodal emotion recognition method, characterized in that, The method includes: Acquire user's text data, video data, and audio data; extract features from the text data to obtain text features, extract features from the video data to obtain facial features, and extract features from the audio data to obtain audio features; The text features, audio features, and facial features are input into a pre-trained multimodal emotion recognition model, which includes a first fusion module, a second fusion module, a third fusion module, a fourth fusion module, and a classification module. The first fusion module is used to fuse the audio features and the text features to obtain audio-text features; the second fusion module is used to fuse the facial features and the text features to obtain facial-text features; and the third fusion module is used to enhance the text features to obtain enhanced text features. The fourth fusion module is used to fuse the audio text features, the facial text features, and the enhanced text features to obtain multimodal features; The multimodal features are classified using the classification module to obtain the user's sentiment classification result.
2. The multimodal emotion recognition method according to claim 1, characterized in that, The first fusion module includes a first attention mechanism, wherein the query vector of the first attention mechanism is the text feature, the key vector of the first attention mechanism is the audio feature, and the value vector of the first attention mechanism is the audio feature; the second fusion module includes a second attention mechanism, wherein the query vector of the second attention mechanism is the text feature, the key vector of the second attention mechanism is the facial feature, and the value vector of the second attention mechanism is the facial feature; the third fusion module includes a third attention mechanism, wherein the query vector of the third attention mechanism is the text feature, the key vector of the third attention mechanism is the text feature, and the value vector of the third attention mechanism is the text feature.
3. The multimodal emotion recognition method according to claim 1, characterized in that, The classification module includes a first mapping unit, an activation unit, and a second mapping unit. The classification module is used to classify the multimodal features to obtain the user's sentiment classification result, including: The first mapping unit is used to perform a linear transformation on the multimodal features to obtain preliminary transformed features; The activation unit is used to perform nonlinear compression on the preliminary transformation features to obtain compressed features; The compressed features are linearly transformed using the second mapping unit to obtain the emotional intensity mapping value of the multimodal features; The user's emotional classification result is determined based on the emotional intensity mapping value.
4. The multimodal emotion recognition method according to claim 1, characterized in that, After obtaining the user's sentiment classification result, the following is also included: The text emotion classification result of the user is determined based on the text features; the audio emotion classification result of the user is determined based on the audio features; the facial emotion classification result of the user is determined based on the facial features. The emotional classification result is updated by using the text emotional classification result, the audio emotional classification result, and the facial emotional classification result to obtain the updated emotional classification result.
5. The multimodal emotion recognition method according to claim 1, characterized in that, After obtaining the user's sentiment classification result, the following is also included: The user's sentiment classification result, the user's information, and the current business scenario are input into a pre-trained response model, so that the response model generates a response statement corresponding to the user's text data based on preset prompt words corresponding to the current business scenario, the user's sentiment classification result, and the user's information.
6. The multimodal emotion recognition method according to claim 1, characterized in that, Obtain user text data, video data, and audio data, including: Obtain the text data corresponding to the user at the current moment, and use it as the text data; Based on the timestamp of the text data, continuous video data and audio data within a preset time period before and after the timestamp are obtained to obtain the video data and the audio data.
7. The multimodal emotion recognition method according to claim 1, characterized in that, Before inputting the text features, audio features, and facial features into a pre-trained multimodal emotion recognition model to obtain the user's emotion classification result, the method further includes: The text features are subjected to a nonlinear transformation to obtain intermediate text features; the audio features are subjected to a nonlinear transformation to obtain intermediate audio features; and the facial features are subjected to a nonlinear transformation to obtain intermediate facial features. The intermediate text features are mapped to the target dimension to obtain updated text features; the intermediate audio features are mapped to the target dimension to obtain updated audio features; the intermediate facial features are mapped to the target dimension to obtain updated facial features.
8. A multimodal emotion recognition device, characterized in that, The device includes: The acquisition module is used to acquire the user's text data, video data, and audio data; to extract features from the text data to obtain text features, to extract features from the video data to obtain facial features, and to extract features from the audio data to obtain audio features; The input module is used to input the text features, the audio features, and the facial features into a pre-trained multimodal emotion recognition model, wherein the multimodal emotion recognition model includes a first fusion module, a second fusion module, a third fusion module, a fourth fusion module, and a classification module; An initial fusion module is used to perform feature fusion on the audio features and the text features using the first fusion module to obtain audio-text features; to perform feature fusion on the facial features and the text features using the second fusion module to obtain facial-text features; and to perform feature enhancement on the text features using the third fusion module to obtain enhanced text features. A multimodal fusion module is used to fuse the audio text features, the facial text features, and the enhanced text features using the fourth fusion module to obtain multimodal features; The sentiment classification module is used to classify the multimodal features using the classification module to obtain the sentiment classification result of the user.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to said at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the multimodal emotion recognition method according to any one of claims 1-7.
10. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the multimodal emotion recognition method according to any one of claims 1-7.
Citation Information
Cited By
Data processing method and device and related equipment
CN121278508A
Emotion recognition method and system based on multi-modal feature retrieval, terminal and storage medium
CN121479706A
An emotion recognition method and system based on multi-modal feature retrieval, a terminal and a storage medium
CN121479706B
Large language model question answering method and question answering system based on audio-text bimodality
CN121615793A