Multi-modal emotion recognition method and system for service-oriented robot

By fusing visual and speech dual-modal data and using dynamic time warping technology, the robustness and temporal offset problems of single-modal emotion recognition in complex environments are solved, achieving high accuracy and strong robustness in multimodal emotion recognition, which is suitable for natural interaction of service robots in home and public scenarios.

CN120995416AActive Publication Date: 2025-11-21SUZHOU CITY UNIV

Patent Information

Application Number
CN202511520732.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2025-11-21
Estimated Expiration
2045-10-23

AI Technical Summary

Technical Problem

In existing technologies, single-modal emotion recognition is not robust enough in complex environments, multimodal features have temporal shifts and modal heterogeneity, and static fusion strategies cannot adapt to dynamic scenarios, resulting in insufficient accuracy and robustness of emotion recognition.

Method used

By fusing visual and speech data, dynamic time warping technology is used to solve the temporal offset of audio and video streams, a dual-modal confidence quantification model is constructed, and a segmented dynamic weight allocation strategy is combined with a cross-modal temporal collaboration module to optimize feature mapping and fusion, thereby achieving end-to-end multimodal emotion recognition.

Benefits of technology

It improves the accuracy and robustness of emotion recognition, adapts to complex environments, is applicable to various service scenarios, and enables natural and intelligent human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120995416A_ABST
    Figure CN120995416A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and particularly relates to a service-oriented robot-oriented multi-modal emotion recognition method and system, and the method comprises the steps: collecting audio and video stream data of emotion changes of a user, and separating visual and voice data; extracting visual and voice emotion features through a pre-training model, and calculating prediction probability distribution of each mode; constructing a bimodal confidence quantitative model based on the distribution to obtain each modal confidence; and fusing the features by adopting a sectional type dynamic weight distribution strategy so as to identify the emotional state of the user. Visual and voice modes are fused, feature alignment is realized in combination with dynamic time warping, spatial optimization performance is shared and expressed through a confidence model, a dynamic weight strategy and a cross-modal time sequence cooperation module, and the method has high recognition accuracy, high robustness and real-time processing capacity in a complex environment and is suitable for various service scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology, and in particular refers to a multimodal emotion recognition method and system for service robots. Background Technology

[0002] With the large-scale application of service robots in scenarios such as home care and medical care, the ability of natural human-machine interaction has become a core indicator for measuring service quality. Emotion, as the underlying logic of human communication, is a key basis for service robots to understand needs and optimize service strategies—for example, in a home setting, it is necessary to recognize a child's crying to perform soothing actions, and in a medical setting, it is necessary to perceive the anxiety state of the elderly to dynamically adjust the interaction paradigm. However, single-modal emotion recognition technology is no longer sufficient to meet the practical application needs of service robots.

[0003] The operating environment of service robots is dynamic and complex: home environments are subject to interference such as television noise and backlighting, while public service environments face challenges such as multiple people obstructing the view and regional accents. In such environments, recognition mechanisms that rely solely on visual modalities (facial expressions) or voice modalities (tone features) are prone to failure—for example, users turning their backs to the robot may result in a lack of visual information, and environmental noise may cause distortion of voice features, both of which can lead to misjudgment of emotions and thus reduce the service experience.

[0004] Human emotional expression naturally possesses multimodal collaborative characteristics: "joy" is usually accompanied by a smile (visual feature), a light tone of voice (vocal feature), and positive semantics (textual feature). The complementarity and corroboration of multimodal signals provide redundant information for emotion recognition. Therefore, multimodal emotion recognition combining vision and speech has become an inevitable path for service robots to break through interaction bottlenecks. Its core value lies in improving the robustness of emotion judgment in complex environments through cross-modal information complementarity.

[0005] However, current emotion recognition technology for service robots faces the following core challenges: First, dynamic scene adaptation is difficult: the lighting and noise levels in homes, offices, and public spaces differ significantly, and fixed-weight fusion strategies cannot cope with dynamic changes in modal reliability. Secondly, there is a lack of temporal continuity: human emotions are a continuous process of change (such as a gradual change from calm to anger). The emotion judgment of a single frame is difficult to reflect the real emotional trend, while service robots need to provide continuous services based on temporal emotional changes.

[0006] In existing technologies, most multimodal emotion recognition solutions use static weight fusion or simple feature splicing, lacking specific optimizations for service robot scenarios: either they do not consider the hardware constraints of real-time processing, and the model complexity exceeds the carrying capacity of embedded devices; or they ignore temporal emotion analysis and cannot capture the emotional fluctuation patterns of users. Summary of the Invention

[0007] Therefore, the technical problems to be solved by this invention are the insufficient robustness of single-modal emotion recognition in complex environments, the temporal shift and modal heterogeneity of multimodal features, and the inability of static fusion strategies to adapt to dynamic scenarios.

[0008] To address the aforementioned technical problems, this invention provides a multimodal emotion recognition method and system for service robots, achieving high accuracy and robust end-to-end multimodal emotion recognition.

[0009] Specifically, the multimodal emotion recognition method for service robots includes the following steps: S1: Collect audio and video stream data showing the dynamic changes in user emotions, perform modal separation to obtain visual and speech data; use a pre-trained model to extract features from the two types of data respectively, generate visual emotion features and speech emotion features, and simultaneously calculate the predicted probability distribution value of each modality data. S2: Based on the predicted probability distribution value of each modality data, construct a bimodal confidence quantification model, and calculate the visual confidence of each modality data respectively; S3: Based on the visual confidence of each modality data, a segmented dynamic weight allocation strategy is adopted to fuse the visual emotion features and the voice emotion features to obtain fused features; S4: Identify the user's emotional state based on the fusion features.

[0010] In one embodiment of the present invention, S1, calculating the predicted probability distribution value for each modality of data includes: Based on the CNN architecture, an image classification model is constructed. The visual data is input into the pre-trained image classification model. The output layer of the image classification model is a fully connected layer containing multiple neurons. The softmax activation function is used. By processing the extracted image emotion features, the probability values ​​corresponding to multiple visual emotions are output. These probability values ​​together constitute the predicted probability distribution of the visual modality. Based on the CNN architecture, a speech classification model is constructed. The speech data is input into the pre-trained speech classification model. The output layer of the speech classification model is a fully connected layer with a softmax activation function. By processing the MFCC features of the speech, the probability values ​​corresponding to multiple speech emotions are output. These probability values ​​together constitute the predicted probability distribution of the speech modality.

[0011] In one embodiment of the present invention, S2, calculating the visual confidence level of each modality data includes: S21: Based on the predicted probability distribution of each modality feature The bimodal confidence quantification model is constructed as follows: , in, For modal confidence, Modal type, For visual modality, For speech modality, To predict the probability distribution Information entropy This represents the maximum predicted probability distribution value for mode m. Represents the predicted probability distribution of the visual modality. The predicted probability distribution of the speech modality is represented by α, β, and γ, which are weighting coefficients. S22: Based on the bimodal confidence quantification model, calculate the confidence level of each modal feature.

[0012] In one embodiment of the present invention, the segmented dynamic weight allocation strategy includes: Set confidence threshold Determine whether the maximum value of either the visual confidence score of the visual modality data or the speech confidence score of the speech modality data is greater than or equal to the confidence threshold. : If so, use the following formula to assign weights: ,in, The weights of mode m, Modal type, For visual modality, For speech modality, Let m be the confidence level of mode m. For visual confidence, Indicates speech confidence level. Indicates the magnification factor. This is the lower limit of the protection weight value; Otherwise, automatically switch to the high-confidence modality priority mode and use the following formula for weight allocation: , where I represents an indicator function that outputs 1 when the condition is met, and 0 otherwise.

[0013] In one embodiment of the present invention, before obtaining the fused features, the method further includes a step of temporal optimization and feature alignment of the visual emotion features and the speech emotion features, including: A cross-modal temporal collaboration module is constructed. By constructing an emotion state transition matrix, introducing inter-frame smoothing constraint loss, and adopting dynamic time warping technology, the temporal offset of visual emotion features and the speech emotion features, as well as the coherence of emotion transition, are adjusted. Furthermore, a cross-modal alignment loss function is constructed to optimize the mapping relationship between visual emotion features and speech emotion features in the shared representation space.

[0014] In one embodiment of the present invention, constructing an emotional state transition matrix includes: Define a set of emotional states S = { , ,…, }, Representing the total number of emotional state types, each element in the emotional state transition matrix... Indicates from state Transition to state The probability of is calculated using the following formula: , in, To start from emotions Shift to emotions Number of times, For emotions Total number of occurrences; All elements Arranged into The emotional state transition matrix, namely: .

[0015] In one embodiment of the present invention, the method for constructing the inter-frame smoothing constraint loss includes: The inter-frame smoothing constraint loss is calculated by real-time computation of the differences in sentiment distribution and model weights between adjacent frames, as follows: , in, Let KL divergence be the KL divergence. This represents the probability distribution of emotions at the current time t. for The probability distribution of emotions at the previous time t-1. These are the emotion-related weight parameters in the model at the current time t. express The weight parameters corresponding to the previous time t-1.

[0016] In one embodiment of the present invention, dynamic time warping technology is used to adjust the temporal offset between visual emotional features and the voice emotional features, including: Calculate the matching distance between each frame in the visual feature sequence and the audio feature sequence at each time point, and construct a distance matrix accordingly. Each element in the distance matrix corresponds to a set of distances between visual features and audio features. Through formula Find the optimal alignment path in the distance matrix; Based on the optimal alignment path, the visual emotion features and the speech emotion features are matched in time to ensure that emotion features with the same semantic meaning are synchronized. in, The first of the visual emotion features 1 eigenvector The first of the speech emotion features 1 eigenvector Represents Euclidean distance. To align the path, Representing a path An alignment point on top, It is an index of the visual feature sequence. It is an index of the audio feature sequence.

[0017] In one embodiment of the present invention, a cross-modal alignment loss function is constructed as follows: , Where N represents the number of samples, Indicates the sample index. For the first The original visual features of each sample For the first Original speech features of each sample; The feature mapping function representing the visual modality is used to map the original visual features. Projected into the shared representation space; The feature mapping function representing the speech modality is used to map the original speech features. Projected into the shared representation space; It is an L2 norm; For a three-element loss, , The distance between positive sample pairs is the distance between cross-modal features of the same emotion category. The distance between negative sample pairs, i.e., the distance between cross-modal features of different emotion categories. This represents the marginal value.

[0018] Based on the same inventive concept, this invention also provides a multimodal emotion recognition system for service robots, including a feature extraction and probability distribution calculation module, a dual-modal confidence assessment module, a dynamic feature fusion module, and an emotion state recognition module. The feature extraction and probability distribution calculation module is configured to: collect audio and video stream data showing dynamic changes in user emotions, perform modal separation to obtain visual data and speech data; extract features from the two types of data using a pre-trained model to generate visual emotion features and speech emotion features, and simultaneously calculate the predicted probability distribution value for each modality of data. The bimodal confidence assessment module is configured to: construct a bimodal confidence quantification model based on the predicted probability distribution value of each modality data, and calculate the visual confidence of each modality data respectively; The dynamic feature fusion module is configured to: based on the visual confidence of each modality data, use a segmented dynamic weight allocation strategy to fuse the visual emotion features and the voice emotion features to obtain fused features; The emotion state recognition module is configured to: identify the user's emotion state based on the fused features.

[0019] The present invention also provides an electronic device, which includes a processor, a memory, and a bus system. The processor and the memory are connected through the bus system. The memory is used to store instructions, and the processor is used to execute the instructions stored in the memory to implement the multimodal emotion recognition method described above.

[0020] The present invention also provides a computer storage medium storing a computer software product, the computer software product including a plurality of instructions for causing a computer device to execute the multimodal emotion recognition method described above.

[0021] The technical solution of the present invention has the following advantages compared with the prior art: This invention achieves precise feature alignment by fusing visual and speech dual-modal data and utilizing dynamic time warping technology to address temporal offsets in audio and video streams. It constructs a dual-modal confidence quantification model to evaluate the reliability of each modality, and combines a segmented dynamic weighting strategy. When both modalities are reliable, a proportional weighting with amplification factors and lower limit protection is used; when neither is reliable, a high-confidence modality priority mode is switched to improve fusion effectiveness. A cross-modal temporal collaboration module (emotional state transition matrix, inter-frame smoothing constraints) is introduced to enhance temporal coherence. A cross-modal shared representation space optimized feature mapping is constructed to solve the modal heterogeneity problem. This results in high recognition accuracy, strong robustness, and adaptability in complex environments, making it suitable for various service scenarios. Attached Figure Description

[0022] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings, wherein... Figure 1This is a flowchart illustrating a multimodal emotion recognition method for service robots provided in an embodiment of the present invention. Figure 2 This is a flowchart illustrating the process of assigning weight coefficients to each modality based on the visual confidence level of each modality data using a segmented dynamic weight allocation strategy. Figure 3 This is a schematic diagram of the structure of a multimodal emotion recognition system for service robots provided in an embodiment of the present invention; Explanation of the reference numerals in the instruction manual: 100, Feature extraction and probability distribution calculation module; 200, Bimodal confidence assessment module; 300, Dynamic feature fusion module; 400, Emotional state recognition module. Detailed Implementation

[0023] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand and implement the present invention. However, the embodiments described are not intended to limit the present invention.

[0024] Example 1: like Figure 1 As shown, this invention provides a multimodal emotion recognition method for service robots, which includes the following steps: S1: Collect audio and video stream data showing dynamic changes in user emotions and perform modal separation to obtain visual data and speech data; extract features from the visual data and speech data respectively using a pre-built pre-trained model to obtain visual emotional features and speech emotional features, and calculate the predicted probability distribution value of each modality data; S2: Based on the predicted probability distribution value of each modality data, construct a bimodal confidence quantification model, and calculate the visual confidence of each modality data respectively; S3: Based on the visual confidence of each modality data, a segmented dynamic weight allocation strategy is adopted to fuse the visual emotion features and the voice emotion features to obtain fused features; S4: Identify the user's emotional state based on the fusion features.

[0025] As can be seen from the above technical solution, this technical solution collects visual and speech dual-modal data, extracts features using a pre-trained model and calculates the predicted probability distribution, combines a dual-modal confidence quantification model to accurately evaluate the reliability of each modality, and then uses a segmented dynamic weight allocation strategy to fuse features, ultimately achieving emotion recognition. This solution has several significant advantages: First, by leveraging the complementarity of dual-modal data, the problem of single-modal data easily failing under complex environments (such as noise and changes in lighting) can be effectively overcome, thereby improving the robustness of emotion recognition. Secondly, through the dynamic weight allocation strategy, the fusion weights can be flexibly adjusted according to the real-time confidence of each modality, avoiding the limitations of static weights that cannot adapt to dynamic scenarios and enhancing adaptability to different environments. Third, feature extraction and probability calculation are performed based on pre-trained models, providing a reliable foundation for subsequent confidence assessment and fusion. This helps improve the accuracy and efficiency of emotion recognition. The overall solution is more in line with the needs of service robots in actual interaction scenarios, and can achieve more natural and intelligent human-machine emotional interaction.

[0026] Furthermore, in S1, the predicted probability distribution of the visual modality is calculated. ,include: An image classification model is constructed based on a Convolutional Neural Network (CNN) architecture. Its network topology consists of a sequentially connected input layer, a first convolutional layer, a second convolutional layer, a pooling layer, and a fully connected layer. Specifically, the first convolutional layer uses 3×3 convolutional kernels with a total of 32 kernels; the second convolutional layer uses 3×3 convolutional kernels with a total of 64 kernels. The model was trained based on the FER2013 facial expression public dataset. During preprocessing, pixel strings were converted into 48×48 grayscale images. The training configuration used the Adam optimizer (learning rate 0.0001), with classification cross-entropy as the loss function, combined with horizontal flipping data augmentation. The training parameters were set to batch size 64, training epochs 5, and validation ratio 20%, resulting in a pre-trained image classification model. The visual data is input into the pre-trained image classification model. After image emotion feature extraction and processing, the output layer (a fully connected layer containing multiple neurons, using the softmax activation function) outputs probability values ​​for multiple visual emotions. These probability values ​​collectively constitute the predicted probability distribution of the visual modality. .

[0027] Furthermore, in S1, the predicted probability distribution of the speech modality is calculated. ,include: Based on the Transformer model architecture, a speech classification model is constructed, including an encoder and a decoder. The encoder includes 6 encoder layers, and each encoder layer integrates a multi-head self-attention mechanism and a feedforward neural network. The model was trained based on the CASIA emotional speech dataset. During preprocessing, MFCC features were extracted (sampling rate 16000Hz, coefficient dimension 40, maximum frame length 400), and a data augmentation strategy involving random noise was employed. The training configuration used the Adam optimizer (learning rate 0.0001) with classification cross-entropy as the loss function. Training parameters were set to batch size 32 and training epochs 60. An early stopping mechanism (patience value of 8 epochs) and a plateau learning rate scheduling strategy were also introduced to obtain the pre-trained speech classification model. The speech data is input into the pre-trained speech classification model. After processing the Mel-frequency cepstral coefficient (MFCC) features, the output layer (a fully connected layer using the softmax activation function) outputs probability values ​​for multiple speech emotions. These probability values ​​collectively constitute the predicted probability distribution of the speech modality. .

[0028] Specifically, in S2, based on the predicted probability distribution value of each modality data, a bimodal confidence quantification model is constructed to calculate the visual confidence of each modality data, including: S21: Based on the predicted probability distribution of each modality feature The bimodal confidence quantification model is constructed as follows: , in, For modal confidence, Modal type, For visual modality, For speech modality, To predict the probability distribution Information entropy The maximum predicted probability distribution value for mode m; This represents the cosine similarity term, used to measure intermodal consistency. Represents the predicted probability distribution of the visual modality. This represents the predicted probability distribution of the speech modality, where α, β, and γ are weighting coefficients; preferably, ; S22: Based on the bimodal confidence quantification model, calculate the confidence level of each modal feature.

[0029] like Figure 2 As shown, in step S3, based on the visual confidence of each modality data, a segmented dynamic weight allocation strategy is used to assign weight coefficients to each modality data, including: S31: Set confidence threshold ; S32: Determine whether the maximum value of either the visual confidence score of the visual modality data or the speech confidence score of the speech modality data is greater than or equal to the confidence threshold. : If so, the weight allocation is performed using the introduced amplification factor and the set lower limit protection weight value, as follows: ,in, The weights of mode m, Modal type, For visual modality, For speech modality, Let m be the confidence level of mode m. For visual confidence, Indicates speech confidence level. Indicates the magnification factor. This is the lower limit of the protection weight value; Otherwise, automatically switch to the high-confidence modality priority mode and use the following formula for weight allocation: Where I represents the indicator function, when the following condition is met: The output is 1 if the condition is met, otherwise the output is 0.

[0030] The aforementioned piecewise function distinguishes between two scenarios: "at least one modality is reliable" and "both modalities are unreliable". It employs weighted fusion and single-modality priority strategies respectively, thereby improving recognition accuracy through dual-modal collaboration and ensuring system robustness in complex environments through rule constraints, thus avoiding performance loss caused by ineffective fusion.

[0031] Furthermore, before obtaining the fused features, the process includes a step of temporal optimization and feature alignment of the visual emotion features and the spoken emotion features, including: A cross-modal temporal collaboration module is constructed. This module adjusts the temporal offset of visual and vocal emotional features, as well as the coherence of emotional transitions, by building an emotion state transition matrix, introducing inter-frame smoothing constraint loss, and employing dynamic time warping techniques. This includes: Define a set of emotional states S = { , ,…, }, This represents the total number of emotional state types, such as 7 emotion categories: anger, disgust, fear, happiness, sadness, surprise, and neutral. Based on a Hidden Markov Model, this is achieved by traversing the temporal emotion tags of historical video stream data to statistically analyze emotional states. Transition to state Number of times And in combination with emotional state Total number of occurrences The probability of emotional state transition is calculated. : ,form Emotional state transition matrix: elements on the diagonal This matrix represents the probability of maintaining an emotional state, such as the probability of "happy → happy". It provides prior constraints for current emotion prediction: during inter-frame emotion prediction, if the currently predicted emotion transition probability (e.g., ...) → ) and in the matrix In the event of low-probability value conflicts (i.e., abrupt changes that do not conform to historical statistical patterns), the system will correct the prediction results based on the matrix probability distribution, reduce the weight of unreasonable shifts, thereby suppressing emotional abrupt changes, ensuring that the transition of emotional states conforms to the laws of natural evolution, and enhancing the coherence of temporal emotions.

[0032] To suppress sudden jumps in emotion prediction and enhance the coherence of temporal emotions, in addition to dynamic constraints through the aforementioned emotion state transition matrix, the model also calculates the inter-frame smoothing constraint loss by real-time calculation of the differences in emotion distribution and model weights between adjacent frames. This incorporates "suppressing mutations" into the optimization objective of model training, thereby forcing the stability of emotion changes from a real-time computational perspective.

[0033] Specifically, the specific calculation formula for the inter-frame smoothing constraint loss is as follows: , in, Let KL divergence be the KL divergence. Adjacent time points t and t The KL divergence of the probability distribution of emotions at -1 is used to measure the difference in emotion distribution. This represents the probability distribution of emotions at the current time t. for The probability distribution of emotions at the previous time t-1; The L2 norm of the model weight parameters at adjacent time points is used to quantify weight changes. These are the emotion-related weight parameters in the model at the current time t. express The weight parameters corresponding to the previous time t-1.

[0034] In summary, the emotion state transition matrix provides prior constraints based on historical data, regulating the possibility of emotion transitions from a macroscopic perspective; while the inter-frame smoothing constraint loss corrects model parameters in real time through backpropagation, ensuring the smoothness of emotion changes from a microscopic inter-frame difference perspective. The synergistic effect of these two mechanisms ensures that emotion prediction conforms to historical statistical patterns while adapting to subtle changes in real-time data, jointly improving the stability of temporal emotion recognition.

[0035] In the temporal sentiment analysis process, the sentiment state transition matrix provides a reasonable reference threshold for the sentiment changes between frames, while the inter-frame smoothing constraint loss further quantifies and corrects the deviation in the actual prediction based on this reference, forming a dual guarantee mechanism of "prior constraint + real-time optimization", ultimately achieving the naturalness and coherence of the sentiment state transition.

[0036] To address the temporal shifts in audio and video stream data caused by acquisition rhythm and transmission delays, dynamic time warping technology is employed to adjust the temporal shifts in visual and vocal emotional features, including: Computational visual feature sequence With audio feature sequences The matching distance at each time point in each frame is used to construct a distance matrix, where each element in the distance matrix corresponds to the squared Euclidean distance between a set of visual features and audio features. ; Solving for the optimal alignment path using dynamic programming algorithm That is, minimize the cumulative distance The optimal alignment path is found in the distance matrix to achieve “bent” matching of visual and speech features in the time dimension (such as multiple frames of visual features corresponding to multiple frames of audio features), correct the temporal offset, and ensure the temporal synchronization of the same semantic emotion feature. in, The first of the visual emotion features 1 eigenvector The first of the speech emotion features 1 eigenvector Represents Euclidean distance. To align the path, Representing a path An alignment point on top, It is an index of the visual feature sequence. It is an index of the audio feature sequence; And construct the cross-modal alignment loss function as follows: Where N represents the sample size. Indicates the sample index. For the first The original visual features of each sample For the first Original speech features of each sample; The feature mapping function representing the visual modality is used to map the original visual features. Projected into the shared representation space; The feature mapping function representing the speech modality is used to map the original speech features. Projected into the shared representation space; It is an L2 norm; The ternary loss is used to enhance feature discriminative power. , The distance between positive sample pairs is the distance between cross-modal features of the same emotion category, such as the video features of the i-th sample. With audio features The distance; The distance between negative sample pairs is the distance between cross-modal features of different emotion categories, such as the distance between the video features of the i-th sample and the audio features of the j-th sample, where i and j are of different categories. Represents the marginal value, forcing positive samples to be at a distance of... Distance between negative sample pairs It must be at least 0.5 smaller; otherwise, the loss value will not be zero, thus driving model optimization to widen the gap between outliers. The cross-modal alignment loss function optimizes the mapping relationship between visual and speech emotion features in the shared representation space, ensuring that the distance between dissimilar features is greater than the distance between similar features. This optimized mapping relationship ensures that visual and speech features have a consistent feature distribution in the shared space, addressing the modal heterogeneity problem and providing more homogeneous feature inputs for subsequent fusion.

[0037] The above steps, through offset correction and coherence enhancement in the temporal dimension, and cross-modal alignment in the feature space dimension, lay the foundation for the effective fusion of visual and speech emotion features, and improve the accuracy and robustness of multimodal emotion recognition.

[0038] Furthermore, this technical solution employs multiple optimization strategies at the engineering implementation level: Hardware acceleration is achieved through GPU parallel computing and TensorRT model optimization; memory management utilizes streaming feature extraction and frame buffer pool mechanisms to improve efficiency; real-time processing relies on multi-threaded pipelines and keyframe sampling algorithms to ensure response speed; adaptive processing integrates adaptive illumination enhancement and noise-robust feature extraction technologies to adapt to complex environments. The deployment architecture adopts a cloud-edge collaborative model, providing RESTful API interfaces to support video analysis requests and returning time-series data, dominant sentiment, and structured reports; containerized deployment based on Docker and Kubernetes ensures system scalability.

[0039] In terms of performance, the image model achieves an accuracy of 72.3% on the FER2013 dataset (45fps under RTX3090 graphics card environment), and the speech model achieves an accuracy of 68.5% on the CASIA dataset and supports real-time streaming processing. Ultimately, it realizes end-to-end multimodal emotion recognition, which can be applied to scenarios such as video content analysis, psychological counseling assistance, and security monitoring.

[0040] Example 2: Based on the same inventive concept as Embodiment 1, this invention also provides a multimodal emotion recognition system for service robots, used to implement the steps of the multimodal emotion recognition method for service robots described in Embodiment 1. Figure 3As shown, the system includes a feature extraction and probability distribution calculation module 100, a bimodal confidence assessment module 200, a dynamic feature fusion module 300, and an emotion state recognition module 400. The feature extraction and probability distribution calculation module 100 is configured to: collect audio and video stream data of dynamic changes in user emotions, perform modal separation to obtain visual data and speech data; perform feature extraction on the two types of data respectively through a pre-trained model to generate visual emotion features and speech emotion features, and simultaneously calculate the predicted probability distribution value of each modal data. The bimodal confidence assessment module 200 is configured to: construct a bimodal confidence quantification model based on the predicted probability distribution value of each modality data, and calculate the visual confidence of each modality data respectively; The dynamic feature fusion module 300 is configured to: based on the visual confidence of each modality data, use a segmented dynamic weight allocation strategy to fuse the visual emotion features and the voice emotion features to obtain fused features; The emotion state recognition module 400 is configured to: recognize the user's emotion state based on the fused features.

[0041] This embodiment proposes a multimodal emotion recognition system for service robots, which is used to implement the aforementioned multimodal emotion recognition method for service robots. Therefore, the specific implementation of the multimodal emotion recognition system for service robots can be found in the embodiment section of the aforementioned multimodal emotion recognition method for service robots. For example, the feature extraction and probability distribution calculation module 100, the bimodal confidence evaluation module 200, the dynamic feature fusion module 300, and the emotion state recognition module 400 are respectively used to implement steps S1, S2, S3, and S4 in the method described in Embodiment 1. Therefore, its specific implementation can be referred to the description of the corresponding embodiments. To avoid redundancy, it will not be repeated here.

[0042] Example 3: The present invention also provides an electronic device, which includes a processor, a memory, and a bus system. The processor and the memory are connected through the bus system. The memory is used to store instructions, and the processor is used to execute the instructions stored in the memory to implement the multimodal emotion recognition method described in Embodiment 1.

[0043] Example 4:

[0044] The present invention also provides a computer storage medium storing a computer software product, the computer software product including a plurality of instructions for causing a computer device to execute the multimodal emotion recognition method described in Embodiment 1.

[0045] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0046] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0047] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0048] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0049] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. A multimodal emotion recognition method for service robots, characterized in that, Includes the following steps: S1: Collect audio and video stream data showing the dynamic changes in user emotions, perform modal separation to obtain visual and speech data; use a pre-trained model to extract features from the two types of data respectively, generate visual emotion features and speech emotion features, and simultaneously calculate the predicted probability distribution value of each modality data. S2: Based on the predicted probability distribution value of each modality data, construct a bimodal confidence quantification model, and calculate the visual confidence of each modality data respectively; S3: Based on the visual confidence of each modality data, a segmented dynamic weight allocation strategy is adopted to fuse the visual emotion features and the voice emotion features to obtain fused features; S4: Identify the user's emotional state based on the fusion features.

2. The multimodal emotion recognition method for service robots according to claim 1, characterized in that, In S1, the predicted probability distribution value for each modality of data is calculated, including: Based on the CNN architecture, an image classification model is constructed. The visual data is input into the pre-trained image classification model. The output layer of the image classification model is a fully connected layer containing multiple neurons. The softmax activation function is used. By processing the extracted image emotion features, the probability values ​​corresponding to multiple visual emotions are output. These probability values ​​together constitute the predicted probability distribution of the visual modality. Based on the CNN architecture, a speech classification model is constructed. The speech data is input into the pre-trained speech classification model. The output layer of the speech classification model is a fully connected layer with a softmax activation function. By processing the MFCC features of the speech, the probability values ​​corresponding to multiple speech emotions are output. These probability values ​​together constitute the predicted probability distribution of the speech modality.

3. The multimodal emotion recognition method for service robots according to claim 1, characterized in that, In S2, the visual confidence score for each modality of data is calculated, including: S21: Based on the predicted probability distribution of each modality feature The bimodal confidence quantification model is constructed as follows: , in, For modal confidence, Modal type, For visual modality, For speech modality, To predict the probability distribution Information entropy This represents the maximum predicted probability distribution value for mode m. Represents the predicted probability distribution of the visual modality. The predicted probability distribution of the speech modality is represented by α, β, and γ, which are weighting coefficients. S22: Based on the bimodal confidence quantification model, calculate the confidence level of each modal feature.

4. The multimodal emotion recognition method for service robots according to claim 1, characterized in that, The segmented dynamic weight allocation strategy includes: Set confidence threshold Determine whether the maximum value of either the visual confidence score of the visual modality data or the speech confidence score of the speech modality data is greater than or equal to the confidence threshold. : If so, use the following formula to assign weights: ,in, The weights of mode m, Modal type, For visual modality, For speech modality, Let m be the confidence level of mode m. For visual confidence, Indicates speech confidence level. Indicates the magnification factor. This is the lower limit of the protection weight value; Otherwise, automatically switch to the high-confidence modality priority mode and use the following formula for weight allocation: Where I represents the indicator function, when the following condition is met: The output is 1 if the condition is met, otherwise the output is 0.

5. The multimodal emotion recognition method for service robots according to claim 1, characterized in that, Before obtaining the fused features, the method further includes a step of temporal optimization and feature alignment of the visual emotion features and the speech emotion features, including: A cross-modal temporal collaboration module is constructed. By constructing an emotion state transition matrix, introducing inter-frame smoothing constraint loss, and adopting dynamic time warping technology, the temporal offset of visual emotion features and the speech emotion features, as well as the coherence of emotion transition, are adjusted. Furthermore, a cross-modal alignment loss function is constructed to optimize the mapping relationship between visual emotion features and speech emotion features in the shared representation space.

6. The multimodal emotion recognition method for service robots according to claim 5, characterized in that, Constructing an emotional state transition matrix includes: Define a set of emotional states S = { , ,…, }, Representing the total number of emotional state types, each element in the emotional state transition matrix... Indicates from state Transition to state The probability of is calculated using the following formula: , in, To be from emotions Shift to emotions Number of times, For emotions Total number of occurrences; All elements Arranged into The emotional state transition matrix, namely: 。 7. The multimodal emotion recognition method for service robots according to claim 5, characterized in that, The method for constructing the inter-frame smoothing constraint loss includes: The inter-frame smoothing constraint loss is calculated by real-time computation of the differences in sentiment distribution and model weights between adjacent frames, as follows: , in, Let KL divergence be the KL divergence. This represents the probability distribution of emotions at the current time t. for The probability distribution of emotions at the previous time t-1. These are the emotion-related weight parameters in the model at the current time t. express The weight parameters corresponding to the previous time t-1.

8. The multimodal emotion recognition method for service robots according to claim 5, characterized in that, The temporal offset of visual emotional features and the spoken emotional features is adjusted using dynamic time warping techniques, including: Calculate the matching distance between each frame in the visual feature sequence and the audio feature sequence at each time point, and construct a distance matrix accordingly. Each element in the distance matrix corresponds to a set of distances between visual features and audio features. Through formula Find the optimal alignment path in the distance matrix; Based on the optimal alignment path, the visual emotion features and the speech emotion features are matched in time to ensure that emotion features with the same semantic meaning are synchronized. in, The first of the visual emotion features 1 eigenvector The first of the speech emotion features 1 eigenvector Represents Euclidean distance. To align the path, Representing a path An alignment point on top, It is an index of the visual feature sequence. It is an index of the audio feature sequence.

9. The multimodal emotion recognition method for service robots according to claim 5, characterized in that, Construct the cross-modal alignment loss function as follows: , Where N represents the number of samples, Indicates the sample index. For the first The original visual features of each sample For the first Original speech features of each sample; The feature mapping function representing the visual modality is used to map the original visual features. Projected into the shared representation space; The feature mapping function representing the speech modality is used to map the original speech features. Projected into the shared representation space; It is an L2 norm; For a three-element loss, , The distance between positive sample pairs is the distance between cross-modal features of the same emotion category. The distance between negative sample pairs, i.e., the distance between cross-modal features of different emotion categories. This represents the marginal value.

10. A multimodal emotion recognition system for service robots, characterized in that, include: The feature extraction and probability distribution calculation module is configured to: collect audio and video stream data of dynamic changes in user emotions, perform modal separation to obtain visual data and speech data; perform feature extraction on the two types of data respectively through a pre-trained model to generate visual emotion features and speech emotion features, and simultaneously calculate the predicted probability distribution value of each modality data; The bimodal confidence assessment module is configured to: construct a bimodal confidence quantification model based on the predicted probability distribution value of each modality data, and calculate the visual confidence of each modality data respectively; The dynamic feature fusion module is configured to: based on the visual confidence of each modality data, use a segmented dynamic weight allocation strategy to fuse the visual emotion features and the voice emotion features to obtain fused features; And an emotion state recognition module is configured to: recognize the user's emotion state based on the fused features.

11. An electronic device, characterized in that, The electronic device includes a processor, a memory, and a bus system, wherein the processor and the memory are connected through the bus system, the memory is used to store instructions, and the processor is used to execute the instructions stored in the memory to implement the multimodal emotion recognition method according to any one of claims 1 to 9.

12. A computer storage medium, characterized in that, The computer storage medium stores a computer software product, the computer software product including a plurality of instructions for causing a computer device to execute the multimodal emotion recognition method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Emotion recognition method based on multilevel decision fusion of fuzzy rule

    CN115731595A

  • Audio and video bimodal emotion recognition method based on decision level fusion

    CN119028377A

  • Robot-oriented serial multi-mode emotion recognition method

    CN120030498A

  • Dynamic self-adaptive multi-modal sentiment analysis fusion method and system

    CN120429691A

  • Sign language synthesis service method based on multiple modes

    CN120823294A

Cited By

  • Virtual image model construction method and system based on image cloning

    CN121349311A

  • Depression degree detection method and system based on bimodal emotion purification and alignment

    CN121502703A

  • Interactive data processing method and device, computer equipment and readable storage medium

    CN121561814A