Robot-oriented serial multi-mode emotion recognition method
By processing audio and image data in serial and using a cross-modal bidirectional feedback mechanism collaborative training method, the problem of low efficiency of parallel processing and joint training in the existing technology is solved, more efficient emotion recognition is achieved, and recognition accuracy and system processing efficiency are improved.
Patent Information
- Application Number
- CN202510452440.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-11
AI Technical Summary
In the prior art, parallel processing and joint training are inefficient, emotional timing information and feedback mechanism are not considered, and the fusion method does not consider the interaction relationship of multimodal data, resulting in low emotional recognition accuracy.
The serial multimodal emotion recognition method is adopted to analyze the audio data through the audio multimodal model and generate feedback text. The image multimodal model is inputted after the image data is time stamped and the cross-modal bidirectional feedback mechanism is collaboratively trained, and the weight matrix is dynamically corrected to optimize the model output.
It effectively explores the potential connection between audio and image data, captures the timing changes of emotional information, improves the accuracy, nature and fluency of emotion recognition, optimizes the use of computing resources and system complexity, and improves the adaptability and flexibility in multiple practical application scenarios.
Smart Images

Figure CN120030498A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a serial multimodal emotion recognition method and system for robots. Background Art
[0002] With the rapid development of artificial intelligence technology, emotion recognition has become an important research direction in the field of multimodal artificial intelligence. The goal of emotion recognition is to extract emotional information from multiple modal inputs (such as audio, images, text, etc.) and classify emotions. The main challenge of multimodal emotion recognition is how to extract effective emotional information from multiple input modalities and how to effectively fuse multimodal data, especially in terms of the complementarity and temporal nature of the two information sources, audio and image.
[0003] At present, multimodal emotion recognition methods often use parallel processing and joint training methods to process multimodal data. The parallel processing method processes audio and images as independent inputs, uses convolutional neural networks (CNN) to process image data, and uses recurrent neural networks (RNN) or long short-term memory networks (LSTM) to process audio data, and then fuses the features of these two data. Although this method can process multimodal data at the same time, it cannot fully utilize the potential dynamic interaction between audio and images. In the task of emotion analysis, the tone, intonation and speech speed of audio data have a great impact on the emotion recognition of image data (such as facial expressions, emotional states, etc.). The guiding role of audio data on image data has not been effectively explored and utilized, resulting in inaccurate fusion of multimodal data and reducing the overall accuracy of emotion recognition. The joint training method optimizes the fusion of audio and image features through multiple models or multiple tasks. The existing joint training methods often share the features of audio and image in the same representation space through multi-task learning (MTL) or shared representation learning, and then jointly optimize between tasks of multiple modalities. Although this method can improve the integration ability of the model, it requires a large amount of labeled data and computing resources, and the optimization process is relatively complex and easily affected by data imbalance or bias. At the same time, the lack of in-depth mining of timing information and feedback mechanisms leads to low accuracy in emotion recognition.
[0004] In addition, the existing technology often uses feature splicing and weighted fusion to fuse multimodal data, lacks deep modeling of the complex interactive relationship between multimodal data, resulting in the transmission and sharing of emotional information being inefficient and inaccurate. At the same time, the existing technology does not consider the temporal changes of emotional information in real application scenarios that need to process user emotional fluctuations, resulting in low efficiency in information feedback and fusion processing of audio and image data, and thus low accuracy in emotion recognition. Summary of the invention
[0005] To this end, the technical problem to be solved by the present invention is to overcome the low efficiency of parallel processing and joint training in the prior art, the failure to consider emotional timing information and feedback mechanism, and the failure of the fusion method to consider the interactive relationship of multimodal data, resulting in low accuracy of emotion recognition.
[0006] In order to solve the above technical problems, the present invention provides a serial multimodal emotion recognition method for robots, comprising:
[0007] Obtain the audio data sequence and image data sequence of the current emotional activity;
[0008] Input the preset audio instruction text and audio data sequence into the trained audio multimodal model, and output the audio feedback text of the current emotional activity; wherein the audio feedback text includes the volume dynamic range, fundamental frequency contour information and tone classification result of the audio data sequence;
[0009] After aligning the timestamps of the audio feedback text and the image data sequence, the trained image multimodal model is input to output the emotion recognition results of the current emotional activity and the image feedback text; the image feedback text includes voice response and action response information.
[0010] Preferably, an audio-text training set is constructed based on each audio data sequence and its corresponding audio emotion label in an open source data set, combined with the audio instruction text; wherein the open source data set contains audio data sequences and their corresponding image data sequences;
[0011] Using the audio-text training set, the initial audio multimodal model is trained by minimizing the first audio loss function to obtain a pre-trained audio multimodal model;
[0012] Input the audio-text training set into the pre-trained audio multimodal model, extract the audio data feature vector, and output each first audio feedback text; after aligning the timestamps of each first audio feedback text with each image data sequence, combine the image emotion labels corresponding to each image data sequence to construct the first image-text training set;
[0013] The first image-text training set is used to train the initial image multimodal model by minimizing the first image loss function, thereby obtaining a pre-trained image multimodal model.
[0014] Preferably, based on a cross-modal bidirectional feedback mechanism, the pre-trained audio multimodal model and image multimodal model are trained, including:
[0015] Input the first audio feedback text and image data sequence after timestamp alignment into the pre-trained image multimodal model, encode the first audio feedback text and image data at the same timestamp respectively through the CLIP architecture dual-tower encoder, and extract the first audio feedback text feature vector and the image data feature vector;
[0016] Generate a first modified weight matrix based on the image data feature vector extracted by the pre-trained image multimodal model and the audio data feature vector extracted by the pre-trained audio multimodal model;
[0017] According to the first corrected weight matrix, the emotion probability distribution output by the classification layer of the pre-trained audio multimodal model, and the emotion probability distribution output by the classification layer of the pre-trained image multimodal model, the audio emotion probability distribution output by the classification layer of the pre-trained audio multimodal model is corrected to obtain a first corrected audio emotion probability distribution;
[0018] Constructing a first reverse feedback loss function based on the first corrected audio emotion probability distribution and the image emotion probability distribution output by the classification layer of the pre-trained image multimodal model;
[0019] Based on the first reverse feedback loss function and its corresponding weight and the first audio loss function and its corresponding weight, a second audio loss function is constructed; using the audio-text training set, the pre-trained audio multimodal model is trained by minimizing the second audio loss function to obtain an audio multimodal model that has completed collaborative training;
[0020] Input the audio-text training set into the audio multimodal model after collaborative training, extract the audio data feature vector, and output each second audio feedback text; after aligning the timestamps of each second audio feedback text with each image data sequence, combine the image emotion labels corresponding to each image data sequence to construct the second image-text training set;
[0021] Based on the audio emotion probability distribution output by the classification layer of the co-trained audio multimodal model, a first positive feedback loss function is constructed;
[0022] Based on the first forward feedback loss function and its corresponding weights and the first image loss function and its corresponding weights, a second image loss function is constructed; using the second image-text training set, the pre-trained image multimodal model is trained by minimizing the second image loss function to obtain a collaboratively trained image multimodal model.
[0023] Preferably, the expression of the first modified weight matrix is:
[0024] ;
[0025] in, represents the first modified weight matrix; Represents the audio data feature vector extracted by the pre-trained audio multimodal model; Represents the feature vector of image data extracted by the pre-trained image multimodal model; Represents the Sigmoid activation function; represents the learnable weight matrix;
[0026] The expression of the first modified audio emotion probability distribution is:
[0027] ;
[0028] in, represents the first corrected audio emotion probability distribution; Represents an element-by-element multiplication operation; Represents the audio emotion probability distribution output by the classification layer of the pre-trained audio multimodal model; Represents the image emotion probability distribution output by the classification layer of the pre-trained image multimodal model.
[0029] Preferably, the first reverse feedback loss function The expression is:
[0030] ;
[0031] in, represents KL divergence; represents the first corrected audio emotion probability distribution; represents the image emotion probability distribution output by the classification layer of the pre-trained image multimodal model;
[0032] The first positive feedback loss function The expression is:
[0033] ;
[0034] in, represents KL divergence; represents the audio emotion probability distribution output by the classification layer of the co-trained audio multimodal model; Represents the image emotion probability distribution output by the classification layer when training the pre-trained image multimodal model.
[0035] Preferably, it also includes:
[0036] Input the second audio feedback text and image data sequence with aligned timestamps into the co-trained image multimodal model, encode the second audio feedback text and image data at the same timestamp respectively through the CLIP architecture dual-tower encoder, and extract the second audio feedback text feature vector and image data feature vector;
[0037] Generate a second modified weight matrix based on the image data feature vector extracted by the collaboratively trained image multimodal model and the audio data feature vector extracted by the collaboratively trained audio multimodal model;
[0038] According to the second corrected weight matrix, the emotion probability distribution output by the classification layer of the collaboratively trained audio multimodal model, and the emotion probability distribution output by the classification layer of the collaboratively trained image multimodal model, the audio emotion probability distribution output by the classification layer of the collaboratively trained audio multimodal model is corrected to obtain a second corrected audio emotion probability distribution;
[0039] Based on the second corrected audio emotion probability distribution and the image emotion probability distribution output by the classification layer of the image multimodal model after collaborative training, a second reverse feedback loss function is constructed; according to the modal confidence range, the corresponding weight of the second reverse feedback loss function is dynamically adjusted to construct a third audio loss function;
[0040] Using the audio-text training set, the audio multimodal model after the collaborative training is trained by minimizing the third audio loss function to obtain an audio multimodal model that completes the target training;
[0041] Input the audio-text training set into the target trained audio multimodal model, and output each third audio feedback text; after aligning the timestamps of each third audio feedback text with each image data sequence, combine the image emotion labels corresponding to each image data sequence to construct a third image-text training set;
[0042] Based on the audio emotion probability distribution output by the classification layer of the target trained audio multimodal model, a second positive feedback loss function is constructed; according to the modal confidence range, the corresponding weight of the second positive feedback loss function is dynamically adjusted to construct a third image loss function; wherein the modal confidence range is 0.1-0.3;
[0043] The collaboratively trained image multimodal model is trained by minimizing the third image loss function using the third image-text training set to obtain an image multimodal model that completes the target training.
[0044] Preferably, in the pre-training stage, the proportion of audio instruction text replaced by empty characters in the audio-text training set and the proportion of audio feedback text replaced by empty characters in the first image-text training set are both set to the first data ratio; in the collaborative training stage, the proportion of audio instruction text replaced by empty characters in the audio-text training set and the proportion of audio feedback text replaced by empty characters in the second image-text training set are both set to the second data ratio; in the target training stage, the proportion of audio instruction text replaced by empty characters in the audio-text training set and the proportion of audio feedback text replaced by empty characters in the third image-text training set are both set to the third data ratio.
[0045] Preferably, the step of training the initial audio multimodal model by minimizing the first audio loss function using the audio-text training set to obtain the pre-trained audio multimodal model comprises:
[0046] Based on the deep neural network, a teacher audio multimodal model is constructed; the audio data sequence and the audio instruction text are used as the input of the teacher audio multimodal model, and the audio emotion label corresponding to the audio data sequence is used as the output of the teacher audio multimodal model. The teacher audio multimodal model is trained using the audio-text training set to obtain a trained teacher audio multimodal model;
[0047] Based on the trained teacher audio multimodal model, a student audio multimodal model is constructed as the initial audio multimodal model; a loss function between the teacher audio multimodal model and the student audio multimodal model is constructed as the first audio loss function; using the audio-text training set, the student audio multimodal model is trained by minimizing the first audio loss function to obtain the trained student multimodal model as the pre-trained audio multimodal model.
[0048] The present invention also provides a serial multimodal emotion recognition system for robots, comprising:
[0049] A data acquisition module, used to obtain the audio data sequence and image data sequence of the current emotional activity;
[0050] The perception module is connected to the data acquisition module and the instruction text acquisition module, and includes:
[0051] An audio processing unit, used to input a preset audio instruction text and an audio data sequence into a trained audio multimodal model, and output an audio feedback text of the current emotional activity; wherein the audio feedback text includes the volume dynamic range, fundamental frequency contour information and tone classification result of the audio data sequence;
[0052] The image processing unit is communicatively connected with the audio processing unit; it is used to align the timestamps of the audio feedback text and the image data sequence, input the trained image multimodal model, and output the emotion recognition result of the current emotional activity and the image feedback text; wherein the image feedback text includes voice response and action response information.
[0053] Preferably, it also includes:
[0054] The interaction module is connected to the perception module in communication; it is used to convert the voice response information in the image feedback text into an audio stream, and obtain the stiffness data, motion angle data and motion time data of each joint of the robot according to the action response information in the image feedback text;
[0055] The voice playback module is located on the robot head and is connected to the interaction module for communication; it is used to play the audio stream at a preset volume;
[0056] The control module is located on the top of the robot and is connected to the interaction module for communication. It is used to control the robot to perform corresponding limb movements according to the stiffness data, motion angle data and motion time data of each joint of the robot.
[0057] The above technical solution of the present invention has the following beneficial effects compared with the prior art:
[0058] (1) The serial multimodal emotion recognition method for robots described in the present invention processes audio and image data in series, first uses an audio multimodal model to analyze the audio data, and then inputs the result into an image multimodal model as a condition, thereby fully exploring the potential connection and interaction between audio and image data, and effectively capturing the temporal changes of emotional information. The emotional change trend in audio analysis can dynamically affect the image analysis process, so that the model has a better understanding of the fluctuation of emotions in the time dimension, and further improves the accuracy, naturalness and fluency of emotion recognition. At the same time, this information fusion method that combines image static features with audio dynamic information makes the emotion analysis results more comprehensive and accurate, and improves the communication between the robot and the user. The naturalness and fluency of emotional interaction; the serial processing method of first processing the audio and then using the audio feedback to optimize the image processing can optimize the use of computing resources, reduce system complexity, improve the overall processing efficiency of the system, and show strong adaptability and flexibility in multiple practical application scenarios; in addition, through the audio instruction text, the audio multimodal model is guided to process audio data in different emotional scenarios, so that the model can better adapt to various emotional activity scenarios and improve the ability to recognize emotions in different scenarios; the audio and image multimodal models are trained using open source data sets and real-time collected data, so that the model can learn a rich variety of emotional expression patterns, enhancing the model's adaptability and generalization capabilities to different scenarios and data.
[0059] (2) The present invention discloses a serial multimodal emotion recognition method for robots, which adopts a progressive training strategy to train the audio multimodal model and the image multimodal model. In the pre-training stage, the audio multimodal model and the image multimodal model are trained separately. By adjusting the proportion of empty characters replacing text in the audio-text training set and the first image-text training set, the basic representation capabilities of independent modes such as audio and vision are quickly established to avoid premature introduction of multimodal noise leading to feature confusion. In the collaborative training stage, the audio multimodal model and the image multimodal model are trained collaboratively to dynamically correct conflicts, so that the audio multimodal model and the image multimodal model treat each modal information equally, and by reducing the proportion of empty characters replacing text in the audio-text training set and the second image-text training set, the proportion of multimodal data in the training set is increased, which promotes the audio multimodal model to be more sensitive to the image and text. The audio and image modalities learn and collaborate with each other to capture the complex relationship between audio and image, strengthen the complementarity between multimodal modes, and avoid single dominance. In order to further improve the accuracy and efficiency of model training, in the target training phase, while the audio multimodal model and the image multimodal model are collaboratively trained, the loss weight is adjusted according to the modal confidence threshold, and the proportion of multimodal data is further improved by further reducing the proportion of empty character replacement text in the audio-text training set and the third image-text training set. More refined learning and optimization of complex multimodal information enhances the model's ability to process multimodal data, suppresses interference from highly uncertain modalities, improves adaptability to different emotional activity scenarios, and improves the accuracy and stability of emotion recognition. This progressive training design improves training efficiency and achieves low-error, highly robust multimodal collaborative reasoning.
[0060] (3) The present invention discloses a serial multimodal emotion recognition method for robots. According to the reverse feedback strategy in the cross-modal bidirectional feedback mechanism, the image emotion probability distribution output by the classification layer of the image multimodal model is fed back to the audio multimodal model to generate a modified weight matrix to adjust the output of the classification layer of the audio multimodal model, which helps to enhance the information interaction and collaborative optimization between different modal models, so that the audio multimodal model can adjust its own output according to the information of the image multimodal modality, thereby improving the accuracy and robustness of the model in audio emotion classification. At the same time, the cross-modal bidirectional feedback mechanism enables a bidirectional information flow between the audio multimodal model and the image multimodal model, which helps to coordinate the optimization of the two models. In the feedback process, the emotion probability distribution output by the classification layer of the audio multimodal model will affect the image multimodal model; in the reverse feedback process, the emotion probability distribution output by the classification layer of the image multimodal model can be fed back to the audio multimodal model, promoting the deep fusion of audio and image modal information, so that the audio multimodal model can use the information of the image modality to better understand the audio data; in addition, the reverse feedback loss function and the forward feedback loss function are constructed to narrow the difference in the emotion probability distribution output by the classification layer of the audio multimodal model and the image multimodal model, so as to make the emotion category results of the two models consistent, further improve the stability and reliability of the model, and enable the multimodal model to give more accurate and consistent results in the emotion classification and recognition task. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below according to specific embodiments of the present invention in conjunction with the accompanying drawings, wherein:
[0062] Figure 1 It is a flow chart of a serial multimodal emotion recognition method for robots provided by the present invention;
[0063] Figure 2 is a flowchart of the training process of the audio multimodal model and the image multimodal model;
[0064] Figure 3 It is a schematic diagram of a serial multimodal emotion recognition system for robots provided by the present invention. DETAILED DESCRIPTION
[0065] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it, but the embodiments are not intended to limit the present invention.
[0066] Reference Figure 1 As shown, Figure 1 The present invention provides a serial multimodal emotion recognition method for robots; specifically comprising:
[0067] S11: Acquire the audio data sequence and image data sequence of the current emotional activity;
[0068] S12: Input the preset audio instruction text and audio data sequence into the trained audio multimodal model, and output the audio feedback text of the current emotional activity; wherein the audio feedback text includes the volume dynamic range, fundamental frequency contour information and tone classification result of the audio data sequence; wherein the preset audio instruction text is a text containing volume, fundamental frequency contour and tone information;
[0069] S13: After aligning the timestamps of the audio feedback text and the image data sequence, the trained image multimodal model is input to output the emotion recognition result of the current emotional activity and the image feedback text; wherein the image feedback text includes voice response and action response information.
[0070] Reference Figure 2 As shown, Figure 2 The flowchart of the training process of the audio multimodal model and the image multimodal model includes:
[0071] S21: Pre-train the initial audio multimodal model and image multimodal model, including:
[0072] S211: constructing an audio-text training set based on each audio data sequence and its corresponding audio emotion label in the open source data set and in combination with the audio instruction text; wherein the open source data set includes the audio data sequence and its corresponding image data sequence;
[0073] Using the audio-text training set, the initial audio multimodal model is trained by minimizing the first audio loss function to obtain a pre-trained audio multimodal model, including:
[0074] Based on deep neural network, a teacher audio multimodal model is constructed; the audio data sequence and audio instruction text are used as the input of the teacher audio multimodal model, and the audio emotion label corresponding to the audio data sequence is used as the output of the teacher audio multimodal model. The teacher audio multimodal model is trained using the audio-text training set to obtain a trained teacher audio multimodal model; in the teacher audio multimodal model constructed based on deep neural network, Mel spectrum analysis is used to extract the fundamental frequency contour, resonance peak shift and other acoustic features of the input audio data sequence, and at the same time, the input audio instruction text is feature extracted through the BERT encoder to establish an audio-text cross-modal spatiotemporal association matrix;
[0075] Using the knowledge distillation framework, based on the trained teacher audio multimodal model, a student audio multimodal model is constructed as the initial audio multimodal model; a loss function between the teacher audio multimodal model and the student audio multimodal model is constructed as the first audio loss function; using the audio-text training set, the student audio multimodal model is trained by minimizing the first audio loss function to obtain the trained student multimodal model as the pre-trained audio multimodal model;
[0076] Among them, the first audio loss function includes: feature alignment strategy loss, emotion distribution transfer mechanism loss and attention transfer loss, and its expression is:
[0077] ;
[0078] in, represents the feature alignment strategy loss; Represents the weight corresponding to the feature alignment strategy loss; Indicates the loss of the emotion distribution transfer mechanism; Represents the weight corresponding to the loss of the sentiment distribution transfer mechanism; Indicates loss of attention shifting mechanism; represents the weight corresponding to the attention transfer loss; in a specific embodiment of the present invention, , , ;
[0079] Among them, the loss function of the feature alignment strategy is the mean square error loss function, and its expression is:
[0080] ;
[0081] in, represents the first The feature output of the layer; Represents the first The feature output of the layer; represents the number of intermediate layers of the trained teacher audio multimodal model and the student audio multimodal model that are aligned; wherein, the mean square error loss function is used as the loss function of the feature alignment strategy. By optimizing the mean square error loss function, the difference between the intermediate layer features of the student audio multimodal model and the intermediate layer features of the teacher audio multimodal model can be reduced, and the difference in intermediate layer features between the student model and the teacher model can be constrained; at the same time, the teacher model has been trained with a large amount of data, and its intermediate layer features often contain richer and more representative audio information. Optimizing the feature alignment strategy with the mean square error loss function can enable the student audio multimodal model to learn a better feature representation of the intermediate layer of the teacher model, so as to better extract audio features and more accurately capture key features, thereby improving the accuracy of the task;
[0082] The loss function of the emotion distribution transfer mechanism is the KL divergence loss function, which is expressed as:
[0083] ;
[0084] in, represents the logits output of the trained teacher audio multimodal model; represents the logits output of the student audio multimodal model; represents the softmax function; represents the temperature hyperparameter of the KL divergence loss; represents KL divergence; wherein, the KL divergence loss function is used as the loss function of the emotion distribution transfer mechanism. By optimizing the KL divergence loss function, the difference between the audio emotion probability distribution output by the classification layer of the student audio multimodal model and the audio emotion probability distribution output by the classification layer of the teacher audio multimodal model can be reduced; at the same time, the audio emotion probability distribution output by the teacher model is based on its learning and understanding of a large amount of audio data. By optimizing the emotion distribution transfer mechanism through the KL divergence loss function, the student model can learn from the teacher model's experience in audio emotion classification, so as to better grasp the emotion categories corresponding to different audios, improve the performance in the audio emotion recognition task, and make the emotion classification results output by it more accurate;
[0085] The loss function of the attention transfer mechanism is a matrix norm loss function, which is expressed as:
[0086] ;
[0087] in, Indicates the number of attention heads; represents the first The attention matrix of the attention heads; represents the first The attention matrix of the attention heads; represents a dimension mapping function; wherein, the matrix norm loss function is used as the loss function of the attention mechanism. By optimizing the matrix norm loss function, the difference between the attention weight distribution of the student audio multimodal model to the audio data sequence and the attention weight distribution of the learning teacher audio multimodal model to the audio data sequence can be reduced; at the same time, the attention mechanism is optimized with the help of the matrix norm loss function, so that the student audio multimodal model can learn the attention weight distribution mode of the teacher audio multimodal model to the audio data sequence, so that the student model can more effectively focus on the key parts of the audio, ignore irrelevant information, and enhance the adaptability of the student model to different audio data;
[0088] S212: inputting the audio-text training set into the pre-trained audio multimodal model, extracting audio data feature vectors, and outputting each first audio feedback text; after aligning the timestamps of each first audio feedback text with each image data sequence, combining the image emotion labels corresponding to each image data sequence, constructing the first image-text training set; wherein, aligning the timestamps of the audio feedback text output by the audio multimodal model with the image data sequence can ensure that in the subsequent processing process, the audio feedback information and the corresponding image information at the same time point can be analyzed simultaneously, providing consistency guarantee in the time dimension for accurate emotion recognition;
[0089] Using the first image-text training set, the initial image multimodal model is trained by minimizing the first image loss function to obtain a pre-trained image multimodal model;
[0090] The image loss function includes: cross-modal contrast loss and cross-modal attention alignment loss, and its expression is:
[0091] ;
[0092] in, represents the cross-modal contrast loss; Represents the weight corresponding to the cross-modal contrast loss; represents the cross-modal attention alignment loss; Represents the weight corresponding to the cross-modal attention alignment loss;
[0093] Among them, the expression of the cross-modal contrast loss is:
[0094] ;
[0095] in, represents the number of samples; Indicates The image feature vector of samples; Indicates The text feature vector of samples; Indicates The text feature vector of samples; represents the temperature hyperparameter of the cross-modal contrastive loss; represents a natural constant; wherein, by optimizing the cross-modal contrast loss function constructed based on the cross-modal contrast learning mechanism, the distance between positive samples (audio feedback text and image data sequence pairs at the same timestamp) in the feature space can be shortened, and the distance between negative samples can be extended, thereby mining semantic associations and enhancing the effectiveness of semantic associations, and obtaining the semantic association space between text and image, so as to be able to mine the potential connection between audio feedback text and image data at the semantic level, and improve the model's ability to understand multimodal data;
[0096] The expression of the cross-modal attention alignment loss is:
[0097] ;
[0098] in, Indicates the number of sub-features in each sample; represents the cross entropy loss function; Indicates The image feature vector of the sample image sub-feature matrix; Indicates The text feature vector of the sample text sub-feature matrices; among them, by optimizing the cross-modal attention alignment loss function constructed based on the differentiable cross-modal attention alignment mechanism, the image multimodal model can focus on the key areas of text and image, achieve more refined local associations, and improve the multimodal data fusion effect, accuracy and flexibility; at the same time, based on the cross-modal contrast learning mechanism and the differentiable cross-modal attention alignment mechanism, the semantic association space between text and image and the local dynamic association space are learned, which helps the model capture more detailed semantic relationships and dynamic changes between different modal data, and improve the semantic understanding and association reasoning capabilities of multimodal data.
[0099] S22: Based on the cross-modal bidirectional feedback mechanism, the pre-trained audio multimodal model and image multimodal model are collaboratively trained, including:
[0100] S221: Input the first audio feedback text and image data sequence after timestamp alignment into the pre-trained image multimodal model, encode the first audio feedback text and image data at the same timestamp respectively through the dual-tower encoder of the CLIP architecture, and extract the first audio feedback text feature vector and the image data feature vector; wherein, through the visual branch channel in the dual-tower encoder of the CLIP architecture, the ViT-B / 32 visual Transformer is used to encode the JPG image in blocks (patch size 32×32), and through the text branch channel in the dual-tower encoder of the CLIP architecture, the deep semantics of the instruction text is extracted based on the RoBERTa-base model; wherein, through the dual-tower encoder of the CLIP architecture, the features of the image and text are extracted respectively, which can effectively fuse the features of the two modalities of audio feedback text and image to obtain the feature space of text and image. This cross-modal feature fusion can make full use of the complementary information of different modal data and enrich the model's understanding and representation capabilities of data;
[0101] Based on the cross-modal bidirectional feedback mechanism, a reverse correction path is constructed to feed back the image emotion probability distribution output by the image multimodal model to the audio multimodal model through the gated linear unit to generate a correction weight matrix to adjust the output of the classification layer of the audio multimodal model, which helps to enhance the information interaction and collaborative optimization between different modal models, so that the audio multimodal model can adjust its own output according to the information of the image multimodal modality, thereby improving the accuracy and robustness of the model in audio emotion classification. Specifically:
[0102] Based on the image data feature vector extracted by the pre-trained image multimodal model and the audio data feature vector extracted by the pre-trained audio multimodal model, a first modified weight matrix is generated, and its expression is:
[0103] ;
[0104] in, represents the first modified weight matrix; Represents the audio data feature vector extracted by the pre-trained audio multimodal model; Represents the feature vector of image data extracted by the pre-trained image multimodal model; Represents the Sigmoid activation function; represents the learnable weight matrix;
[0105] According to the first corrected weight matrix, the emotion probability distribution output by the classification layer of the pre-trained audio multimodal model, and the emotion probability distribution output by the classification layer of the pre-trained image multimodal model, the audio emotion probability distribution output by the classification layer of the pre-trained audio multimodal model is corrected to obtain a first corrected audio emotion probability distribution, which is expressed as follows:
[0106] ;
[0107] in, represents the first corrected audio emotion probability distribution; Represents an element-by-element multiplication operation; Represents the audio emotion probability distribution output by the classification layer of the pre-trained audio multimodal model; represents the image emotion probability distribution output by the classification layer of the pre-trained image multimodal model;
[0108] S222: Based on the first corrected audio emotion probability distribution and the image emotion probability distribution output by the classification layer of the pre-trained image multimodal model, a first reverse feedback loss function is constructed, and its expression is:
[0109] ;
[0110] in, represents KL divergence; represents the first corrected audio emotion probability distribution; represents the image emotion probability distribution output by the classification layer of the pre-trained image multimodal model;
[0111] Based on the first reverse feedback loss function and its corresponding weight and the first audio loss function and its corresponding weight, a second audio loss function is constructed, and its expression is:
[0112] ;
[0113] in, represents the updated weight corresponding to the feature alignment strategy loss; represents the updated weight corresponding to the loss of the sentiment distribution transfer mechanism; Represents the updated weight corresponding to the attention transfer loss; represents the reverse feedback loss function; Represents the weight corresponding to the inverse feedback loss function; , and Characterize the weight corresponding to the first audio loss function;
[0114] Using the audio-text training set, the pre-trained audio multimodal model is trained by minimizing the second audio loss function to obtain a collaboratively trained audio multimodal model;
[0115] Input the audio-text training set into the audio multimodal model after collaborative training, extract the audio data feature vector, and output each second audio feedback text; after aligning the timestamps of each second audio feedback text with each image data sequence, combine the image emotion labels corresponding to each image data sequence to construct the second image-text training set;
[0116] Based on the audio emotion probability distribution output by the classification layer of the audio multimodal model after collaborative training, the first positive feedback loss function is constructed, and its expression is:
[0117] ;
[0118] in, represents KL divergence; represents the audio emotion probability distribution output by the classification layer of the co-trained audio multimodal model; It represents the probability distribution of image emotions output by the classification layer when training the pre-trained image multimodal model;
[0119] S223: Based on the first forward feedback loss function and its corresponding weight and the first image loss function and its corresponding weight, construct a second image loss function, which is expressed as:
[0120] ;
[0121] in, Represents the updated weights corresponding to the cross-modal contrast loss; represents the updated weights corresponding to the cross-modal attention alignment loss; represents the positive feedback loss function; Represents the weight corresponding to the positive feedback loss function; and Both represent the weights corresponding to the first image loss function;
[0122] The pre-trained image multimodal model is trained by utilizing the second image-text training set by minimizing the second image loss function to obtain a co-trained image multimodal model.
[0123] S23: According to the modal confidence range, dynamically balance training is performed on the co-trained audio multimodal model and image multimodal model, including:
[0124] S231: Input the second audio feedback text and the image data sequence after the timestamps are aligned into the co-trained image multimodal model, and respectively encode the second audio feedback text and the image data at the same timestamp through the dual-tower encoder of the CLIP architecture, and extract the second audio feedback text feature vector and the image data feature vector;
[0125] Generate a second modified weight matrix based on the image data feature vector extracted by the collaboratively trained image multimodal model and the audio data feature vector extracted by the collaboratively trained audio multimodal model;
[0126] According to the second corrected weight matrix, the emotion probability distribution output by the classification layer of the collaboratively trained audio multimodal model, and the emotion probability distribution output by the classification layer of the collaboratively trained image multimodal model, the audio emotion probability distribution output by the classification layer of the collaboratively trained audio multimodal model is corrected to obtain a second corrected audio emotion probability distribution;
[0127] S232: constructing a second reverse feedback loss function based on the second corrected audio emotion probability distribution and the image emotion probability distribution output by the classification layer of the collaboratively trained image multimodal model; dynamically adjusting the corresponding weight of the second reverse feedback loss function according to the modal confidence range to construct a third audio loss function;
[0128] Using the audio-text training set, the audio multimodal model after the collaborative training is trained by minimizing the third audio loss function to obtain an audio multimodal model that completes the target training;
[0129] S233: input the audio-text training set into the target trained audio multimodal model, and output each third audio feedback text; after aligning the timestamps of each third audio feedback text with each image data sequence, combine the image emotion labels corresponding to each image data sequence to construct a third image-text training set;
[0130] Based on the audio emotion probability distribution output by the classification layer of the target trained audio multimodal model, a second positive feedback loss function is constructed; according to the modal confidence range, the corresponding weight of the second positive feedback loss function is dynamically adjusted to construct a third image loss function; wherein the modal confidence range is 0.1-0.3;
[0131] The collaboratively trained image multimodal model is trained by minimizing the third image loss function using the third image-text training set to obtain an image multimodal model that completes the target training.
[0132] In summary, in the pre-training stage, the proportion of audio instruction text replaced by empty characters in the audio-text training set and the proportion of audio feedback text replaced by empty characters in the first image-text training set are both set to the first data ratio; in the collaborative training stage, the proportion of audio instruction text replaced by empty characters in the audio-text training set and the proportion of audio feedback text replaced by empty characters in the second image-text training set are both set to the second data ratio; in the target training stage, the proportion of audio instruction text replaced by empty characters in the audio-text training set and the proportion of audio feedback text replaced by empty characters in the third image-text training set are both set to the third data ratio; the first data ratio is 30%. In the pre-training process, using less multimodal data allows the model to first learn the audio and image unimodal data separately. Feature representation, lay a solid foundation, avoid introducing too much complex multimodal information at the beginning, which makes it difficult for the model to learn and converge, just like letting students master basic knowledge first and then learn comprehensive knowledge; the second data ratio is 70%. In the process of two-way collaborative training, using more multimodal data can promote mutual learning and collaboration between audio and image modalities, optimize the model through a two-way feedback mechanism, improve the ability to fuse multimodal data, capture the complex relationship between audio and image, and better perform emotion recognition; the third data ratio is 90%. In the process of dynamic balance training, a higher proportion of multimodal data allows the model to make full use of previously learned knowledge, conduct more refined learning and optimization of complex multimodal information, adapt to different emotional expressions and scenes, and improve the accuracy and stability of emotion recognition;
[0133] In a specific embodiment of the present invention, the number of iterations of pre-training accounts for 10% of the entire model training process, which can allow the model to initially build a single-modal feature extraction capability within a reasonable time to prepare for multi-modal collaboration; the number of iterations of bidirectional collaborative training accounts for 60% of the entire model training process, which can ensure that the model fully learns in bidirectional feedback and collaborative optimization, reduce the audio-image modality conflict rate, and improve the emotion recognition ability in complex scenes, such as effectively compressing the recognition error in the face occlusion scene; the number of iterations of dynamic balance training accounts for 30% of the entire model training process, which can further improve the model's ability to process complex multi-modal information without causing the model to overfit due to too long a time.
[0134] Ultimately, the final audio and image multimodal serial model obtained through training reduced the conflict rate by 58%, and compressed the emotion recognition error from 22% to 7% in the face occlusion scenario through reverse feedback. In the inference stage, zero-delay reverse calibration was achieved through the pre-compiled correction matrix, and the memory usage only increased by 2.1MB, while still maintaining a real-time processing capability of 14ms / frame in the robot embedded system.
[0135] The advantages of the present invention are:
[0136] (1) First, the audio data is processed to generate audio text results, and then the image data is processed to ensure the efficient integration of multimodal data; the results of audio analysis are fed back to the image processing module as conditional information, which enhances the accuracy of emotion recognition, especially in the case of emotional fluctuations; the final emotional response is fed back through the robot's voice and action generation to achieve natural emotional interaction; by serializing audio processing and image processing, the audio emotional feedback directly affects the image emotion analysis results, enhancing the interaction between multimodal data; the large model not only analyzes emotions, but also generates targeted responses and provides feedback through robot actions and voice, improving the naturalness and accuracy of human-computer interaction; through this serial, multimodal data processing method, the accuracy of emotion recognition can be significantly improved, and robots can be provided with richer and more natural emotional interaction capabilities.
[0137] (2) In the audio multimodal model, a teacher audio multimodal model with a large parameter scale is first trained. This model uses a deep neural network architecture to perform spectral analysis on the audio waveform, and combines the semantic understanding of the instructional text to establish a cross-modal association between audio features and text descriptions. After the teacher model reaches stable performance, a more compact student model is constructed to achieve knowledge transfer through feature alignment and distribution transfer, that is, using the knowledge distillation framework, based on the trained teacher audio multimodal model, and through feature alignment strategy and transfer learning mechanism, a student audio multimodal model is constructed. While retaining the 95% recognition accuracy of the teacher audio multimodal model, the student audio multimodal model reduces the number of parameters to 30%, reducing the amount of model calculation, and successfully reduces the analysis delay of key acoustic features such as fundamental frequency contour and resonance peak shift to less than 15ms, so that the complex multimodal emotion analysis algorithm can run in real time on embedded devices such as robots, improving the overall processing efficiency of the model, thereby being more efficient in processing audio data, improving the robustness and accuracy of the model, improving the emotion recognition performance in various scenarios, and improving the emotion recognition ability in complex scenarios.
[0138] (3) The audio feedback text output by the audio multimodal model is a natural language description text, which includes processed information such as volume, pitch, tone, and emotion change trend (such as "two sudden pitch increases in the middle of the conversation, corresponding to a 62% increase in the probability of anger"); the structured text output by the audio multimodal model is constructed by analyzing multi-dimensional acoustic parameters such as volume dynamic range (unit: dB, sampling interval: 100ms), fundamental frequency contour (F0, accuracy: ±2Hz), tone classification label (such as "rapid / smooth / fluctuating"), and emotion intensity change gradient (ΔE), to construct a directional control system for robot interaction. That is, the volume and emotion intensity gradient are encoded into conditional vectors through a learnable projection matrix, and the cross-modal attention query weight of the visual Transformer is dynamically adjusted to make the image model focus in high-volume scenes. The robot uses the fundamental frequency contour mean and standard deviation to control the rhythmic parameters of speech synthesis through parameterized mapping, and realizes the adaptive adjustment of pitch offset and speech speed (such as compressing the speech response delay to 800ms when the fundamental frequency is greater than 220Hz). The tone label and the emotion gradient jointly drive the action intensity controller, and the joint stiffness and swing frequency are mapped by nonlinear functions (the finger tremor frequency is increased to 2Hz under a rapid tone), so that the amplitude of limb movements when the emotion intensity gradient is greater than 0.3 is enhanced to 1.2 times the baseline value. This design enables the robot to maintain an emotion recognition accuracy of 82% in a noisy environment with facial occlusion, the speech emotion matching MOS score is increased to 4.6, the action fluency and empathy satisfaction are increased by 63% and 29% respectively, forming a closed-loop optimization of the multimodal response chain driven by acoustic features.
[0139] (4) In the image multimodal model, the visual branch divides the image into small blocks for feature extraction, the text branch parses the instruction semantics, and associates the key areas of the image (such as facial expressions) with the text description through the dynamic attention mechanism; an innovative two-way feedback mechanism is introduced to use the image analysis results to reversely correct the emotional judgment of the audio model, and dynamically balance the consistency of the audio-visual data through weight fusion. At the same time, a phased training strategy is adopted to gradually increase the proportion of multimodal data and transition from single-modal learning to deep collaboration; finally, the audio multimodal model achieves millisecond-level response (14ms / frame) on the robot embedded device, and the recognition error of the face occlusion scene is reduced to 7%, and the memory usage is only increased by 2.1MB, which significantly improves the robot's ability to analyze complex emotions (such as forced smiles and hidden anger), providing reliable technical support for natural emotional interaction; and based on the analysis results, the language response and body response to the current emotional activity are output; a reverse attention pathway is constructed, so that a dual-mode model is formed between the audio multimodal model and the image multimodal model. The reverse information flow contributes to the coordinated optimization of the two models. In the forward process, the information of the audio multimodal model affects the image multimodal model. After the reverse attention pathway is turned on, the information of the image multimodal model can be fed back to the audio multimodal model, promoting the deep fusion of audio and image modal information, so that the audio multimodal model can use the information of the image modality to better understand the audio data. The feedback information of the reverse attention pathway can help the audio multimodal model adjust its feature representation by generating a modified weight matrix to adjust the output of the model classification layer, and provide additional supervision information for the optimization of the audio multimodal model, so that the audio multimodal model can not only consider its own characteristics and tasks when generating the audio emotion probability distribution, but also refer to the emotion information of the image modality, thereby achieving more accurate emotion classification, helping the model to discover more essential features, so that when facing different types of audio data and various complex practical application scenarios, it can more accurately perform emotion classification and feature extraction, and show better generalization performance.
[0140] (5) Based on the gradual nature of multimodal learning and the need for conflict resolution, the proportion of multimodal data in the three stages of the progressive training strategy is set to 30%, 70%, and 90%, respectively. In the early stages of training (the first 10% of the stage), the model quickly establishes the basic representation capabilities of independent modalities such as audio and vision through unimodal pre-training (multimodal data only accounts for 30%) to avoid feature confusion caused by the premature introduction of multimodal noise. For example, in the emotion recognition task, this stage focuses on learning the time-frequency characteristics of the Mel spectrum from pure audio data, or extracting facial key points from pure visual data, laying the foundation for subsequent fusion; then enters the mid-term (the middle 60% stage), and the proportion of multimodal data increases to 70%. Through two-way collaborative training, the complementarity between modalities is strengthened, and conflicts are dynamically corrected. For example, when visual information fails due to facial occlusion, the model can rely on audio features to infer emotions, and vice versa. This stage forces the model to treat each modal information equally through balanced multimodal exposure to avoid single dominance; at the end of training (the last 30% stage), the proportion of multimodal data increases to 90%, simulating the complex input distribution of real scenes (such as robots receiving voice, image and environmental sensor signals at the same time), and introducing a dynamic balance optimization strategy to adjust the loss weight according to the real-time prediction confidence and suppress the interference of high-uncertainty modalities. This progressive design has been proven to be able to gradually reduce the modal conflict rate from 18% in the early stage to 6% in the final stage, while improving training efficiency by 30% (compared to the single-stage strategy), ultimately achieving low-error (7%) and highly robust multimodal collaborative reasoning.
[0141] Reference Figure 3 As shown, Figure 3 A schematic diagram of a serial multimodal emotion recognition system for robots provided by the present invention; specifically comprising:
[0142] The data acquisition module is used to obtain the audio data sequence and image data sequence of the current emotional activity, including:
[0143] A microphone is located on the robot head; it is used to record the audio data of the current emotional activity within a preset period of time and construct an audio data sequence of the current emotional activity; wherein, the audio file in the format of WAV or OGG is obtained through four microphones with a sensitivity of 40dB on the robot head;
[0144] The camera is located on the forehead of the robot; it is used to continuously capture multiple images of current emotional activities and construct an image data sequence of the current emotional activities; wherein the camera captures images with a maximum resolution of 1288*968;
[0145] The perception module is connected to the data acquisition module and the instruction text acquisition module, and includes:
[0146] An audio processing unit, used to input a preset audio instruction text and an audio data sequence into a trained audio multimodal model, and output an audio feedback text of the current emotional activity; wherein the audio feedback text includes the volume dynamic range, fundamental frequency contour information and tone classification result of the audio data sequence;
[0147] The image processing unit is connected to the audio processing unit for communication; it is used to align the timestamps of the audio feedback text and the image data sequence, input the trained image multimodal model, and output the emotion recognition result of the current emotional activity and the image feedback text; wherein the image feedback text includes voice response and action response information;
[0148] The interaction module is connected to the perception module in communication; it is used to convert the voice response information in the image feedback text into an audio stream, and obtain the stiffness data, motion angle data and motion time data of each joint of the robot according to the action response information in the image feedback text;
[0149] The voice playback module is located on the left and right sides of the robot head and is connected to the interactive module for communication; it is used to play the audio stream according to the preset volume; wherein the voice playback module has a built-in speaker, and the maximum output power of the speaker is 2W;
[0150] The control module is located on the top of the robot and is in communication with the interaction module; it is used to control the robot to perform corresponding limb movements according to the stiffness data, motion angle data and motion time data of each joint of the robot;
[0151] Wherein, the robot is a Naoqi robot.
[0152] Obviously, the above embodiments are merely examples for the purpose of clear explanation and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the present invention.
Claims
1. A serial multimodal emotion recognition method for robots, characterized in that: include: Obtain the audio data sequence and image data sequence of the current emotional activity; Input the preset audio instruction text and audio data sequence into the trained audio multimodal model, and output the audio feedback text of the current emotional activity; wherein the audio feedback text includes the volume dynamic range, fundamental frequency contour information and tone classification result of the audio data sequence; After aligning the timestamps of the audio feedback text and the image data sequence, the trained image multimodal model is input to output the emotion recognition results of the current emotional activity and the image feedback text; the image feedback text includes voice response and action response information.
2. A serial multimodal emotion recognition method for robots according to claim 1, characterized in that: Based on each audio data sequence in the open source data set and its corresponding audio emotion label, combined with the audio instruction text, an audio-text training set is constructed; wherein the open source data set contains audio data sequences and their corresponding image data sequences; Using the audio-text training set, the initial audio multimodal model is trained by minimizing the first audio loss function to obtain a pre-trained audio multimodal model; Input the audio-text training set into the pre-trained audio multimodal model, extract the audio data feature vector, and output each first audio feedback text; after aligning the timestamps of each first audio feedback text with each image data sequence, combine the image emotion labels corresponding to each image data sequence to construct the first image-text training set; The first image-text training set is used to train the initial image multimodal model by minimizing the first image loss function, thereby obtaining a pre-trained image multimodal model.
3. A serial multimodal emotion recognition method for robots according to claim 2, characterized in that: Based on the cross-modal bidirectional feedback mechanism, the pre-trained audio multimodal model and image multimodal model are trained, including: Input the first audio feedback text and image data sequence after timestamp alignment into the pre-trained image multimodal model, encode the first audio feedback text and image data at the same timestamp respectively through the CLIP architecture dual-tower encoder, and extract the first audio feedback text feature vector and the image data feature vector; Generate a first modified weight matrix based on the image data feature vector extracted by the pre-trained image multimodal model and the audio data feature vector extracted by the pre-trained audio multimodal model; According to the first corrected weight matrix, the emotion probability distribution output by the classification layer of the pre-trained audio multimodal model, and the emotion probability distribution output by the classification layer of the pre-trained image multimodal model, the audio emotion probability distribution output by the classification layer of the pre-trained audio multimodal model is corrected to obtain a first corrected audio emotion probability distribution; Constructing a first reverse feedback loss function based on the first corrected audio emotion probability distribution and the image emotion probability distribution output by the classification layer of the pre-trained image multimodal model; Based on the first reverse feedback loss function and its corresponding weight and the first audio loss function and its corresponding weight, a second audio loss function is constructed; using the audio-text training set, the pre-trained audio multimodal model is trained by minimizing the second audio loss function to obtain an audio multimodal model that has completed collaborative training; Input the audio-text training set into the audio multimodal model after collaborative training, extract the audio data feature vector, and output each second audio feedback text; after aligning the timestamps of each second audio feedback text with each image data sequence, combine the image emotion labels corresponding to each image data sequence to construct the second image-text training set; Based on the audio emotion probability distribution output by the classification layer of the co-trained audio multimodal model, a first positive feedback loss function is constructed; Based on the first forward feedback loss function and its corresponding weights and the first image loss function and its corresponding weights, a second image loss function is constructed; using the second image-text training set, the pre-trained image multimodal model is trained by minimizing the second image loss function to obtain a collaboratively trained image multimodal model.
4. A serial multimodal emotion recognition method for robots according to claim 3, characterized in that: The expression of the first modified weight matrix is: ; in, represents the first modified weight matrix; Represents the audio data feature vector extracted by the pre-trained audio multimodal model; Represents the feature vector of image data extracted by the pre-trained image multimodal model; Represents the Sigmoid activation function; represents the learnable weight matrix; The expression of the first modified audio emotion probability distribution is: ; in, represents the first corrected audio emotion probability distribution; Represents an element-by-element multiplication operation; Represents the audio emotion probability distribution output by the classification layer of the pre-trained audio multimodal model; Represents the image emotion probability distribution output by the classification layer of the pre-trained image multimodal model.
5. The serial multimodal emotion recognition method for robots according to claim 3, characterized in that: The first reverse feedback loss function The expression is: ; in, represents KL divergence; represents the first corrected audio emotion probability distribution; represents the image emotion probability distribution output by the classification layer of the pre-trained image multimodal model; The first positive feedback loss function The expression is: ; in, represents KL divergence; represents the audio emotion probability distribution output by the classification layer of the co-trained audio multimodal model; Represents the image emotion probability distribution output by the classification layer when training the pre-trained image multimodal model.
6. The serial multimodal emotion recognition method for robots according to claim 3, characterized in that: Also includes: Input the second audio feedback text and image data sequence with aligned timestamps into the co-trained image multimodal model, encode the second audio feedback text and image data at the same timestamp respectively through the CLIP architecture dual-tower encoder, and extract the second audio feedback text feature vector and image data feature vector; Generate a second modified weight matrix based on the image data feature vector extracted by the collaboratively trained image multimodal model and the audio data feature vector extracted by the collaboratively trained audio multimodal model; According to the second corrected weight matrix, the emotion probability distribution output by the classification layer of the collaboratively trained audio multimodal model, and the emotion probability distribution output by the classification layer of the collaboratively trained image multimodal model, the audio emotion probability distribution output by the classification layer of the collaboratively trained audio multimodal model is corrected to obtain a second corrected audio emotion probability distribution; Based on the second corrected audio emotion probability distribution and the image emotion probability distribution output by the classification layer of the image multimodal model after collaborative training, a second reverse feedback loss function is constructed; according to the modal confidence range, the corresponding weight of the second reverse feedback loss function is dynamically adjusted to construct a third audio loss function; Using the audio-text training set, the audio multimodal model after the collaborative training is trained by minimizing the third audio loss function to obtain an audio multimodal model that completes the target training; Input the audio-text training set into the target trained audio multimodal model, and output each third audio feedback text; after aligning the timestamps of each third audio feedback text with each image data sequence, combine the image emotion labels corresponding to each image data sequence to construct a third image-text training set; Based on the audio emotion probability distribution output by the classification layer of the target trained audio multimodal model, a second positive feedback loss function is constructed; according to the modal confidence range, the corresponding weight of the second positive feedback loss function is dynamically adjusted to construct a third image loss function; wherein the modal confidence range is 0.1-0.3; The collaboratively trained image multimodal model is trained by minimizing the third image loss function using the third image-text training set to obtain an image multimodal model that completes the target training.
7. A serial multimodal emotion recognition method for robots according to claim 6, characterized in that: In the pre-training stage, the ratio of the audio instruction text replaced by the empty characters in the audio-text training set and the ratio of the audio feedback text replaced by the empty characters in the first image-text training set are both set to the first data ratio; In the collaborative training stage, the proportion of audio instruction text replaced by empty characters in the audio-text training set and the proportion of audio feedback text replaced by empty characters in the second image-text training set are both set to the second data ratio; in the target training stage, the proportion of audio instruction text replaced by empty characters in the audio-text training set and the proportion of audio feedback text replaced by empty characters in the third image-text training set are both set to the third data ratio.
8. The serial multimodal emotion recognition method for robots according to claim 2, characterized in that: The method of using the audio-text training set to train the initial audio multimodal model by minimizing the first audio loss function to obtain the pre-trained audio multimodal model includes: Based on the deep neural network, a teacher audio multimodal model is constructed; the audio data sequence and the audio instruction text are used as the input of the teacher audio multimodal model, and the audio emotion label corresponding to the audio data sequence is used as the output of the teacher audio multimodal model. The teacher audio multimodal model is trained using the audio-text training set to obtain a trained teacher audio multimodal model; Based on the trained teacher audio multimodal model, a student audio multimodal model is constructed as the initial audio multimodal model; a loss function between the teacher audio multimodal model and the student audio multimodal model is constructed as the first audio loss function; using the audio-text training set, the student audio multimodal model is trained by minimizing the first audio loss function to obtain the trained student multimodal model as the pre-trained audio multimodal model.
9. A serial multimodal emotion recognition system for robots, characterized in that: include: A data acquisition module, used to obtain the audio data sequence and image data sequence of the current emotional activity; The perception module is connected to the data acquisition module and the instruction text acquisition module, and includes: An audio processing unit, used to input a preset audio instruction text and an audio data sequence into a trained audio multimodal model, and output an audio feedback text of the current emotional activity; wherein the audio feedback text includes the volume dynamic range, fundamental frequency contour information and tone classification result of the audio data sequence; The image processing unit is communicatively connected with the audio processing unit; it is used to align the timestamps of the audio feedback text and the image data sequence, input the trained image multimodal model, and output the emotion recognition result of the current emotional activity and the image feedback text; wherein the image feedback text includes voice response and action response information.
10. The serial multimodal emotion recognition system for robots according to claim 9, characterized in that: Also includes: An interaction module, communicating with the perception module; It is used to convert the voice response information in the image feedback text into an audio stream, and obtain the stiffness data, motion angle data and motion time data of each joint of the robot according to the action response information in the image feedback text; The voice playback module is located on the robot head and is connected to the interaction module for communication; it is used to play the audio stream at a preset volume; The control module is located on the top of the robot and is in communication with the interaction module. It is used to control the robot to perform corresponding limb movements according to the stiffness data, motion angle data and motion time data of each joint of the robot.
Citation Information
Patent Citations
Emotion recognition method and system based on voice text cross-modal fusion
CN117765981A
Multi-modal sentiment analysis method based on multi-granularity feature comparison and fusion framework
CN117893948A
Unsupervised multi-modal emotion recognition method based on attention aggregation and cross-modal graph fusion
CN119128616A
Multi-modal classroom emotion recognition method and system based on modal adaptive learning
CN119418725A
Cited By
Multi-modal emotion recognition method and system for service-oriented robot
CN120995416A