Emotion recognition method and device based on multi-modal missing data, equipment and medium

By generating a multimodal emotion data sequence set and performing feature extraction and fusion processing, the problems of low recognition accuracy and wasted computing resources of single-modal data are solved, and more efficient emotion recognition is achieved.

CN120654201BActive Publication Date: 2025-11-11HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511129153.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-11
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

In existing technologies, emotion recognition based on only a single modality of data can easily lead to low recognition accuracy, and data loss can easily occur when processing modal data such as text and audio, resulting in a waste of computing resources.

Method used

By generating a multimodal emotion data sequence set, performing feature extraction processing, generating multimodal emotion fusion feature information, constructing a multimodal temporal relationship matrix and a speaker relationship matrix, and using a conditional diffusion generation network and a feature processing network to generate and reconstruct fusion feature information, the final multimodal emotion recognition result is generated.

Benefits of technology

It improves the accuracy of emotion recognition, reduces the waste of computing resources, and reduces the probability of decreased recognition accuracy and repeated recognition due to missing modal data by fusing data from different modalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654201B_ABST
    Figure CN120654201B_ABST
Patent Text Reader

Abstract

This disclosure presents embodiments of an emotion recognition method, apparatus, device, and medium based on missing multimodal data. One specific implementation of the method includes: generating a multimodal emotion data sequence set; performing the following steps: obtaining text modality feature information, audio modality feature information, and video modality feature information corresponding to the multimodal emotion data; generating multimodal emotion fusion feature information corresponding to the multimodal emotion data; constructing a multimodal temporal relation matrix and a speaker relation matrix; generating multimodal graph feature information; generating reconstructed text convolutional features, reconstructed audio convolutional features, and reconstructed video convolutional features; generating reconstructed fusion feature information; generating multimodal emotion recognition results; and sending the generated multimodal emotion recognition results to a user terminal. This implementation can improve the accuracy of emotion recognition and reduce the waste of computational resources during recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed herein relate to the field of computer technology, and more specifically to methods, apparatus, devices, and media for emotion recognition based on multimodal missing data. Background Technology

[0002] Emotion recognition is a technology that identifies a user's emotional category by processing information such as language, speech, and facial expressions. Currently, the common approach to emotion recognition is to process and recognize single-modal data such as text or audio sent by the user terminal to obtain the corresponding emotion recognition result. Then, the emotion recognition result is sent back to the user terminal.

[0003] However, when using the above methods for emotion recognition, the following technical problems often arise:

[0004] Sentiment recognition based on a single modality of data is prone to low accuracy. Furthermore, processing text, audio, and other modalities can lead to data loss, further reducing accuracy and necessitating repeated recognition of multimodal data from the same user terminal, resulting in wasted computational resources. Summary of the Invention

[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0006] Some embodiments of this disclosure propose emotion recognition methods, apparatuses, devices, and media based on multimodal missing data to address one or more of the technical problems mentioned in the background section above.

[0007] In a first aspect, some embodiments of this disclosure provide an emotion recognition method based on missing multimodal data. The method includes: in response to receiving a multimodal statement sequence set sent by a user terminal, generating a multimodal emotion data sequence set based on the multimodal statement sequence set; for each multimodal emotion data in the multimodal emotion data sequence set, performing the following steps: performing feature extraction processing on the multimodal emotion data to obtain text modality feature information, audio modality feature information, and video modality feature information corresponding to the multimodal emotion data; generating multimodal emotion fusion feature information corresponding to the multimodal emotion data based on the text modality feature information, the audio modality feature information, and the video modality feature information; and constructing a multimodal temporal relationship matrix and speaker identification based on the multimodal emotion data. The system generates a relation matrix; based on the aforementioned multimodal emotion fusion feature information, the aforementioned multimodal temporal relation matrix, and the aforementioned speaker relation matrix, it generates multimodal graph feature information; based on the aforementioned multimodal graph feature information, the aforementioned multimodal emotion fusion feature information, and the conditional diffusion generation network in the pre-trained emotion recognition model, it generates reconstructed text convolutional features, reconstructed audio convolutional features, and reconstructed video convolutional features; based on the aforementioned reconstructed text convolutional features, the aforementioned reconstructed audio convolutional features, the aforementioned reconstructed video convolutional features, the aforementioned multimodal graph feature information, and the aforementioned feature processing network in the emotion recognition model, it generates reconstructed fusion feature information; based on the aforementioned reconstructed fusion feature information and the aforementioned classification network in the emotion recognition model, it generates multimodal emotion recognition results; and it sends the generated multimodal emotion recognition results to the aforementioned user terminal.

[0008] Secondly, some embodiments of this disclosure provide an emotion recognition device based on missing multimodal data. The device includes: a generation unit configured to generate a multimodal emotion data sequence set based on a multimodal statement sequence set received from a user terminal; and an execution unit configured to perform the following steps for each multimodal emotion data in the multimodal emotion data sequence set: performing feature extraction processing on the multimodal emotion data to obtain text modality feature information, audio modality feature information, and video modality feature information corresponding to the multimodal emotion data; generating multimodal emotion fusion feature information corresponding to the multimodal emotion data based on the text modality feature information, the audio modality feature information, and the video modality feature information; and constructing a multimodal temporal relationship matrix based on the multimodal emotion data. The system generates a speaker relationship matrix; based on the aforementioned multimodal emotion fusion feature information, the aforementioned multimodal temporal relationship matrix, and the aforementioned speaker relationship matrix, it generates multimodal graph feature information; based on the aforementioned multimodal graph feature information, the aforementioned multimodal emotion fusion feature information, and the conditional diffusion generation network in the pre-trained emotion recognition model, it generates reconstructed text convolutional features, reconstructed audio convolutional features, and reconstructed video convolutional features; based on the aforementioned reconstructed text convolutional features, the aforementioned reconstructed audio convolutional features, the aforementioned reconstructed video convolutional features, the aforementioned multimodal graph feature information, and the aforementioned feature processing network in the emotion recognition model, it generates reconstructed fusion feature information; based on the aforementioned reconstructed fusion feature information and the aforementioned classification network in the emotion recognition model, it generates multimodal emotion recognition results; and a sending unit is configured to send the generated multimodal emotion recognition results to the aforementioned user terminal.

[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0011] The above embodiments of this disclosure have the following beneficial effects: the emotion recognition method based on multimodal missing data in some embodiments of this disclosure can improve the accuracy of emotion recognition and reduce the waste of computational resources during recognition. Specifically, the reason for the low accuracy of emotion recognition and the waste of computational resources during recognition is that emotion recognition based on only a single modality of data is prone to low accuracy. Moreover, when processing modal data such as text and audio, data loss is likely to occur, which can easily lead to low accuracy when performing emotion recognition based on the processed data. This results in the need to call computational resources to repeatedly recognize the multimodal data sent by the same user terminal, causing a waste of computational resources during recognition. Based on this, the emotion recognition method based on multimodal missing data in some embodiments of this disclosure firstly, in response to receiving a multimodal sentence sequence set sent by the user terminal, generates a multimodal emotion data sequence set based on the multimodal sentence sequence set. Thus, the original data to be processed can be obtained. Secondly, for each multimodal sentiment data point in the aforementioned multimodal sentiment data sequence set, the following steps are performed: First, feature extraction processing is performed on the multimodal sentiment data to obtain text modality feature information, audio modality feature information, and video modality feature information corresponding to the multimodal sentiment data. This yields feature vectors corresponding to different modalities in the multimodal sentiment data. Second, based on the aforementioned text modality feature information, audio modality feature information, and video modality feature information, multimodal sentiment fusion feature information corresponding to the aforementioned multimodal sentiment data is generated. This allows for the fusion of feature vectors corresponding to different modalities in the multimodal sentiment data. Then, based on the aforementioned multimodal sentiment data, a multimodal temporal relation matrix and a speaker relation matrix are constructed. This yields the multimodal temporal relation matrix and the speaker relation matrix. Finally, based on the aforementioned multimodal sentiment fusion feature information, the aforementioned multimodal temporal relation matrix, and the aforementioned speaker relation matrix, multimodal graph feature information is generated. This generates multimodal graph feature information. Next, based on the aforementioned multimodal graph feature information, the aforementioned multimodal emotion fusion feature information, and the conditional diffusion generative network in the pre-trained emotion recognition model, reconstructed text convolutional features, reconstructed audio convolutional features, and reconstructed video convolutional features are generated. Then, based on the aforementioned reconstructed text convolutional features, the aforementioned reconstructed audio convolutional features, the aforementioned reconstructed video convolutional features, the aforementioned multimodal graph feature information, and the aforementioned feature processing network in the emotion recognition model, reconstructed fusion feature information is generated. Then, based on the aforementioned reconstructed fusion feature information and the aforementioned classification network in the emotion recognition model, multimodal emotion recognition results are generated. Finally, the generated multimodal emotion recognition results are sent to the aforementioned user terminal.Therefore, the results of various multimodal emotion recognition can be sent to the user terminal. Because the data from multiple modalities sent by the user terminal can be fused, and then emotion recognition can be performed based on the processed data, the accuracy of emotion recognition can be improved. Furthermore, because the feature vectors corresponding to different modalities can be fused first to obtain multimodal emotion fusion feature information, and then emotion recognition can be performed based on the fused features, even if data from one modality is missing, data from other modalities can still provide a data foundation for subsequent emotion recognition. This reduces the probability of low emotion recognition accuracy due to missing modal data, and also reduces the probability of repeatedly using computing resources to recognize the same data due to low accuracy, thus reducing the waste of computing resources. Attached Figure Description

[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0013] Figure 1 This is a flowchart of some embodiments of the emotion recognition method based on multimodal missing data according to the present disclosure;

[0014] Figure 2 This is a schematic diagram of the structure of some embodiments of the emotion recognition device based on multimodal missing data according to the present disclosure;

[0015] Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0016] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0017] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0018] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0019] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0020] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0021] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0022] Figure 1 A flow 100 of some embodiments of a sentiment recognition method based on multimodal missing data according to the present disclosure is shown. This sentiment recognition method based on multimodal missing data includes the following steps:

[0023] Step 101: In response to receiving a multimodal statement sequence set sent by the user terminal, generate a multimodal sentiment data sequence set based on the multimodal statement sequence set.

[0024] In some embodiments, the execution entity (e.g., a computing device) of the emotion recognition method based on missing multimodal data can generate a multimodal emotion data sequence set based on a multimodal statement sequence set received from a user terminal. The user terminal can be a terminal device corresponding to the user. Each multimodal statement in the multimodal statement sequence set can be dialogue data sent by the user, including text data, audio data, and video data. Each multimodal statement in the multimodal statement sequence set can include text data, audio data, and video data. The text data included in each multimodal statement in the multimodal statement sequence set can be text content representing the dialogue content between speakers. The text data in each multimodal statement in the multimodal statement sequence set can include a statement sequence. Each statement in the statement sequence can be text representing the content spoken by a speaker. Each statement in the statement sequence corresponds to a speaker. The audio data included in each multimodal statement in the multimodal statement sequence set can be audio corresponding to the dialogue data between speakers. The video data included in each multimodal statement in the aforementioned multimodal statement sequence set can be the video corresponding to the dialogue data between the speakers. The executing entity can be a server. The aforementioned multimodal sentiment data sequence set can be the data determined by the aforementioned multimodal statement sequence set. Each multimodal sentiment data in the aforementioned multimodal sentiment data sequence set can include text modal sentiment data, audio modal sentiment data, and video modal sentiment data. The aforementioned text modal sentiment data can be the text data in the multimodal statement. The aforementioned text modal sentiment data can include statement sequences. The statement sequences included in the aforementioned text modal sentiment data can be the statement sequences included in the text data of the aforementioned multimodal statement. The aforementioned audio modal sentiment data can be the audio data in the multimodal statement. The aforementioned video modal sentiment data can be the video data in the multimodal statement.

[0025] In practice, for each multimodal statement in the aforementioned multimodal statement sequence set, the text data included in the multimodal statement can be identified as text modal sentiment data, and the statement sequence within the text data can be identified as the statement sequence within the text modal sentiment data. Then, the audio data included in the multimodal statement can be identified as audio modal sentiment data. Next, the video data included in the multimodal statement can be identified as video modal sentiment data. Finally, the text modal sentiment data, the audio modal sentiment data, and the video modal sentiment data can be combined to form multimodal sentiment data.

[0026] Step 102: For each multimodal sentiment data point in the multimodal sentiment data sequence set, perform the following steps:

[0027] Step 1021: Perform feature extraction processing on the multimodal sentiment data to obtain the text modal feature information, audio modal feature information and video modal feature information of the corresponding multimodal sentiment data.

[0028] In some embodiments, the execution entity may perform feature extraction processing on the multimodal sentiment data to obtain text modality feature information, audio modality feature information, and video modality feature information corresponding to the multimodal sentiment data. The text modality feature information may be a feature vector corresponding to the text data included in the multimodal sentiment data. The audio modality feature information may be a feature vector corresponding to the audio data included in the multimodal sentiment data. The video modality feature information may be a feature vector corresponding to the video data included in the multimodal sentiment data.

[0029] In some optional implementations of certain embodiments, the aforementioned execution entity may perform feature extraction processing on the aforementioned multimodal sentiment data through the following steps to obtain text modal feature information, audio modal feature information, and video modal feature information corresponding to the aforementioned multimodal sentiment data:

[0030] The first step is to perform word segmentation on the text modal sentiment data included in the aforementioned multimodal sentiment data, obtaining a text modal word segmentation sequence. Each text modal word in the above text modal word segmentation sequence can be a word from the aforementioned text modal sentiment data. In practice, the executing entity can use a word segmentation tool to perform word segmentation on the text modal sentiment data included in the aforementioned multimodal sentiment data, obtaining a text modal word segmentation sequence. This word segmentation tool can be any tool capable of segmenting text data. For example, the word segmentation tool could be jieba.

[0031] The second step involves encoding the aforementioned text modality segmentation sequence to obtain a text-encoded data sequence. Each text-encoded data point in this sequence can be the data obtained after encoding the text modality segmentation. In practice, the executing entity can input the text modality segmentation sequence into a pre-trained word vector encoding model to obtain the text-encoded data sequence. This word vector encoding model can be a neural network model that takes the text modality segmentation sequence as input and the text-encoded data sequence as output. For example, the word vector encoding model can be a pre-processed Word2Vec model. This preprocessing can be a process of fine-tuning the Word2Vec model using a set of text modality segmentation sequences and a cross-entropy loss function.

[0032] The third step involves feature extraction processing of the aforementioned text-encoded data sequence to obtain text modality feature information. This text modality feature information can be the feature vector corresponding to the text-encoded data sequence. In practice, the executing entity can input the text-encoded data sequence into a pre-trained text feature extraction model to obtain the text modality feature information. This text feature extraction model can be a neural network model that takes the text-encoded data sequence as input and outputs the text modality feature information. For example, the text feature extraction model can be a pre-trained TextCNN (Text Convolutional Neural Network). This pre-training can be a process of fine-tuning TextCNN using a set of text-encoded data sequences and a cross-entropy loss function.

[0033] The fourth step involves resampling the audio modal emotional data included in the aforementioned multimodal emotional data to obtain resampled audio data. This resampled audio data can be the audio modal emotional data after resampling processing. In practice, the executing entity can use a resampling algorithm to resample the audio modal emotional data, converting the sampling rate corresponding to the audio modal emotional data to a preset sampling rate, thus obtaining the resampled audio modal emotional data as the resampled audio data. The preset sampling rate can be a pre-defined sampling rate. For example, the preset sampling rate could be 16kHz.

[0034] Step 5: Based on the resampled audio data, generate frame-level audio feature information corresponding to each frame of the resampled audio data. Each frame-level audio feature information can be a feature vector corresponding to an audio segment of a preset duration in the resampled audio data. The preset duration can be a pre-defined length, for example, 20ms. The audio segments corresponding to the frame-level audio feature information do not overlap. In practice, the resampled audio data can be input into a frame-level feature extraction model to obtain the frame-level audio feature information. The frame-level feature extraction model can be a neural network model that takes the resampled audio data as input and outputs the corresponding frame-level audio feature information. For example, the frame-level feature extraction model can be a pre-processed Wav2Vec model (Waveform-to-Vector). The pre-processing can be a process of fine-tuning the Wav2Vec model using the resampled audio dataset and the cross-entropy loss function.

[0035] The sixth step involves pooling the aforementioned frame-level audio feature information to obtain audio modal feature information. In practice, the executing entity can use average pooling to perform average pooling on the aforementioned frame-level audio feature information to obtain audio modal feature information.

[0036] Step 7: Perform face alignment processing on the video modal sentiment data included in the aforementioned multimodal sentiment data to obtain aligned video data. The aligned video data can be video data obtained by processing each face region in the aforementioned video modal sentiment data. Each face region can be the image region where the face is located in each video frame included in the aforementioned video modal sentiment data.

[0037] In practice, firstly, the aforementioned video modal sentiment data can be input into a face region data generation model to obtain individual face region data. This face region data generation model can be a neural network model that takes the video modal sentiment data as input and outputs the corresponding face region data. For example, the face region data generation model can be a pre-trained multi-task cascaded convolutional neural network. The pre-training process involves fine-tuning the multi-task cascaded convolutional neural network using an annotated video modal sentiment dataset and a cross-entropy loss function. The annotation process involves labeling each facial keypoint in each video frame of the video modal sentiment data. Each facial keypoint can be a key point on the face. For example, a facial keypoint can be the center of the pupil, the tip of the nose, or the corner of the mouth. Each face region data can be the data corresponding to the face region in the video modal sentiment data. Each face region data can include face bounding box data and individual facial keypoint data. Each face region in the aforementioned face region data corresponds to a video frame in the aforementioned video modality sentiment data. The aforementioned face bounding box data can be an array of bounding boxes representing the face regions. This data may include the x-coordinate of the top-left corner, the y-coordinate of the top-left corner, the x-coordinate of the bottom-right corner, and the y-coordinate of the bottom-right corner. The coordinate values ​​in the aforementioned face bounding box data can be coordinate values ​​in the corresponding image coordinate system. This image coordinate system can be a two-dimensional coordinate system constructed with the first pixel in the top-left corner of the video frame as the origin, the horizontal direction to the right of the video frame as the x-coordinate direction, and the vertical direction downwards of the video frame as the y-coordinate direction. Each face keypoint in the aforementioned face keypoint data can be the two-dimensional coordinates of a keypoint in the face region in the corresponding image coordinate system. Each face keypoint in the aforementioned face keypoint data corresponds to a keypoint category. This category can be a label representing the category of the keypoint. For example, the above key point categories can be the center of the left pupil, the center of the right pupil, the tip of the nose, the left corner of the mouth, or the right corner of the mouth.

[0038] Then, for each face region data in the aforementioned face region data, the facial key point data included in the aforementioned face region data can be combined into a matrix as a facial key point matrix according to a preset combination order. The above combination order can be the order in which the facial key point data in the face region data are combined. For example, the above combination order can be {center of left pupil, center of right pupil, tip of nose, left corner of mouth, right corner of mouth}. As an example, when the facial key point data corresponding to the center of left pupil, center of right pupil, and tip of nose are (1,2), (2,3), and (3,4) respectively, when the combination order is {center of left pupil, center of right pupil, tip of nose}, it can be combined into a three-row, two-column matrix [1,2; 2,3; 3,4] as a facial key point matrix.

[0039] Next, the facial landmark matrix and a preset standard facial landmark matrix can be input into the similarity transformation matrix generation function to obtain the landmark transformation matrix. The landmark transformation matrix can be any matrix capable of transforming the facial landmark matrix into the standard facial landmark matrix. The similarity transformation matrix generation function can be any function capable of generating a similarity transformation matrix between two matrices. For example, the similarity transformation matrix generation function can be the `estimateAffinePartial2D` function in OpenCV. The standard facial landmark matrix can be a facial landmark matrix preset by the technician.

[0040] Then, the video frame corresponding to the aforementioned face region data can be determined as the video frame to be processed. Next, the face bounding box data corresponding to the video frame to be processed can be input into an image cropping function to crop the video frame, obtaining the image region corresponding to the face bounding box data in the video frame as the face image data. The image cropping function can be any function capable of cropping image data. For example, the image cropping function can be the `crop` function from the PIL library.

[0041] Then, the aforementioned face image data and the aforementioned keypoint transformation matrix can be input into the transformation function to obtain the transformed face image data, which serves as the target face image data corresponding to the aforementioned face region data. The aforementioned transformation function can be a function capable of linearly mapping the input image. For example, the aforementioned transformation function could be the cv2.warpAffine() function.

[0042] Finally, using a video generation tool, the obtained target face image data can be combined into aligned video data according to the order of the video frames corresponding to each face region data in the aforementioned video modality sentiment data. The video generation tool can be any tool capable of combining images into a video. For example, ffmpeg could be used.

[0043] Step 8: Perform feature extraction processing on the aligned video data to obtain frame-level video feature information. Each frame-level video feature can be a feature vector corresponding to each video frame in the aligned video data. In practice, for each video frame in the aligned video data, the execution entity can use a feature extraction algorithm to perform feature extraction processing on the video frame, obtaining the corresponding feature vector as frame-level video feature information. The feature extraction algorithm can be any algorithm capable of extracting features from video frames. For example, the feature extraction algorithm can be a scale-invariant feature transform algorithm.

[0044] The ninth step involves pooling the aforementioned frame-level video feature information to obtain video modal feature information. In practice, the executing entity can use average pooling to perform average pooling on the aforementioned frame-level video feature information to obtain video modal feature information.

[0045] Step 1022: Based on text modal feature information, audio modal feature information and video modal feature information, generate multimodal sentiment fusion feature information corresponding to the multimodal sentiment data.

[0046] In some embodiments, the executing entity may generate multimodal emotion fusion feature information corresponding to the multimodal emotion data based on the text modal feature information, the audio modal feature information, and the video modal feature information. The multimodal emotion fusion feature information may be a feature vector obtained by processing the text modal feature information, the audio modal feature information, and the video modal feature information.

[0047] In some optional implementations of certain embodiments, the aforementioned execution entity may generate multimodal sentiment fusion feature information corresponding to the aforementioned multimodal sentiment data by means of the following steps based on the aforementioned text modal feature information, the aforementioned audio modal feature information, and the aforementioned video modal feature information:

[0048] The first step involves randomly masking the aforementioned text modal features, audio modal features, and video modal features to obtain masked text features, masked audio features, and masked video features. Specifically, the masked text features can be the masked text modal features, the masked audio features can be the masked audio modal features, and the masked video features can be the masked video modal features.

[0049] The second step involves performing feature fusion processing on the aforementioned masked text feature information, masked audio feature information, and masked video feature information to obtain multimodal sentiment fusion feature information. This multimodal sentiment fusion feature information can be a feature vector obtained by fusing the aforementioned masked text feature information, masked audio feature information, and masked video feature information. In practice, the aforementioned masked text feature information, masked audio feature information, and masked video feature information can be input into a feature fusion model to obtain multimodal sentiment fusion feature information. This feature fusion model can be a neural network model that takes the masked text feature information, masked audio feature information, and masked video feature information as input and outputs the multimodal sentiment fusion feature information. For example, the feature fusion model can be a pre-trained Bi-GRU (Bidirectional Gated Recurrent Unit) model. Pre-training can be the process of fine-tuning the Bi-GRU model using a training dataset and a cross-entropy loss function. Each training data point in the training dataset can be the data used to train the Bi-GRU model. Each training data in the above training dataset may include masked text feature information, masked audio feature information, and masked video feature information.

[0050] In some optional implementations of certain embodiments, the aforementioned execution entity may perform random masking processing on the aforementioned text modal feature information, the aforementioned audio modal feature information, and the aforementioned video modal feature information through the following steps to obtain masked text feature information, masked audio feature information, and masked video feature information:

[0051] The first step involves generating text masking probability data, audio masking probability data, and video masking probability data based on preset text missing probability data, preset audio missing probability data, preset video missing probability data, and preset missing data. Specifically, the text missing probability data can be the missing data rate corresponding to the text modality sentiment data. The audio missing probability data can be the missing data rate corresponding to the audio modality sentiment data. The video missing probability data can be the missing data rate corresponding to the video modality sentiment data. The missing data rate can be the proportion of missing values ​​in the data. The text masking probability data can be the masking ratio corresponding to random masking of the text modality feature information. The audio masking probability data can be the masking ratio corresponding to random masking of the audio modality feature information. The video masking probability data can be the masking ratio corresponding to random masking of the video modality feature information. The missing data can be a pre-set value. For example, the missing data can be 1. In practice, the executing entity can determine the text masking probability data as the difference between the missing data and the text missing probability data. The difference between the aforementioned missing data and the aforementioned audio missing probability data can be used to determine the audio mask probability data. Similarly, the difference between the aforementioned missing data and the aforementioned video missing probability data can be used to determine the video mask probability data.

[0052] The second step involves masking the text modality feature information based on the aforementioned text mask probability data to obtain masked text feature information. In practice, firstly, the executing entity can generate a vector with the same size and shape as the aforementioned text modality feature information, and element values ​​within a preset range, as a random text vector using a vector generation function. This vector generation function can be any function capable of generating random vectors. For example, it could be the `np.random.rand` function. The preset range can be [0,1). Then, for each element in the random text vector, in response to determining that the element value is less than or equal to the aforementioned text mask probability data, 0 can be set as the element value to update the random text vector. In response to determining that the element value is greater than the aforementioned text mask probability data, 1 can be set as the element value to update the random text vector. The updated random text vector can then be multiplied element-wise with the aforementioned text modality feature information to obtain the masked text feature information.

[0053] The third step involves masking the audio modal feature information based on the aforementioned audio mask probability data to obtain masked audio feature information. In practice, firstly, the executing entity can generate a vector with the same size and shape as the audio modal feature information, and element values ​​within the aforementioned preset range, using the aforementioned vector generation function as the audio random vector. Then, for each element in the audio random vector, in response to determining that the element value is less than or equal to the audio mask probability data, 0 can be set as the element value to update the audio random vector. In response to determining that the element value is greater than the audio mask probability data, 1 can be set as the element value to update the audio random vector. The updated audio random vector can then be multiplied element-wise with the aforementioned audio modal feature information to obtain the masked audio feature information.

[0054] The fourth step involves masking the video modal feature information based on the aforementioned video mask probability data to obtain masked video feature information. In practice, firstly, the executing entity can generate a vector with the same size and shape as the video modal feature information, and element values ​​within the aforementioned preset range, using the aforementioned vector generation function as a video random vector. Then, for each element in the video random vector, in response to determining that the element value is less than or equal to the video mask probability data, 0 can be set as the element value to update the video random vector. In response to determining that the element value is greater than the video mask probability data, 1 can be set as the element value to update the video random vector. The updated video random vector can then be multiplied element-wise with the aforementioned video modal feature information to obtain the masked video feature information.

[0055] Step 1023: Based on multimodal sentiment data, construct a multimodal temporal relation matrix and a speaker relation matrix.

[0056] In some embodiments, the executing entity may construct a multimodal temporal relation matrix and a speaker relation matrix based on the multimodal sentiment data. The multimodal temporal relation matrix may be a matrix representing the order of statements in the sentence sequence of the textual modal sentiment data within the multimodal sentiment data. The speaker relation matrix may be a matrix used to represent whether the two speakers corresponding to every two statements in the sentence sequence included in the multimodal sentiment data are the same.

[0057] In practice, for the sentence sequence of text-based sentiment data in the aforementioned multimodal sentiment data, for each sentence in the sequence, its position within the sequence can be determined by its corresponding index number. For example, when a sentence is the first sentence in the sequence, its index number can be 1. Then, the number of sentences in the sequence can be determined as the target number of sentences. Next, an empty matrix with the same number of rows and columns as the target number of sentences can be created as the temporal relationship matrix to be processed. This matrix creation function can be any function capable of creating matrices. For example, it could be the `np.empty()` function. Then, the elements on the main diagonal of the temporal relationship matrix can be set to 0. Finally, for every two sentences in the sequence, the two corresponding index numbers can be combined into two matrix position data. Each matrix position data can be used to represent a position within the matrix. For example, when the two numbers mentioned above are 2 and 3, the combined matrix position data can be (3, 2) and (2, 3). (3, 2) can represent the position corresponding to the third row and second column in the matrix, and (2, 3) can represent the position corresponding to the second row and third column in the matrix. Next, the difference between the two numbers corresponding to the two statements can be determined as the number difference data. Then, the absolute value of the number difference data can be determined as the number absolute value data. Then, in response to determining that the number absolute value data satisfies a preset number data condition, 1 can be determined as the element corresponding to the two matrix position data. Here, the number data condition can be that the number absolute value data equals 1. Then, in response to determining that the number absolute value data does not satisfy the number data condition, 0 can be determined as the element corresponding to the two matrix position data. This process continues until the time series relation matrix to be processed is completely filled, at which point the filled time series relation matrix to be processed can be determined as a multimodal time series relation matrix.

[0058] Next, an empty matrix with the same number of rows and columns as the target statements can be created using the matrix creation function described above. This matrix serves as the speaker matrix to be processed. Then, each element on the main diagonal of this speaker matrix is ​​set to 1. For every two statements in the statement sequence, firstly, the two corresponding index numbers are combined into two matrix position data. Secondly, the two speakers corresponding to the two statements are identified as two target speakers. Then, in response to determining that the two target speakers meet a preset speaker condition, the two elements in the two matrix position data are set to 1. This speaker condition can be that the two target speakers are the same. Then, in response to determining that the two target speakers do not meet the speaker condition, the two elements in the two matrix position data are set to 0. This process continues until the speaker matrix to be processed is completely filled, at which point the filled speaker matrix is ​​determined as the speaker relationship matrix.

[0059] Step 1024: Generate multimodal graph feature information based on multimodal emotion fusion feature information, multimodal temporal relation matrix and speaker relation matrix.

[0060] In some embodiments, the aforementioned executing entity may generate multimodal graph feature information based on the aforementioned multimodal emotion fusion feature information, the aforementioned multimodal temporal relationship matrix, and the aforementioned speaker relationship matrix.

[0061] In some optional implementations of certain embodiments, the aforementioned executing entity can generate multimodal graph feature information based on the aforementioned multimodal emotion fusion feature information, the aforementioned multimodal temporal relationship matrix, and the aforementioned speaker relationship matrix through the following steps:

[0062] The first step is to generate temporal feature information based on the aforementioned multimodal temporal relationship matrix and the aforementioned multimodal sentiment fusion feature information. Specifically, the aforementioned temporal feature information can be a feature vector that integrates the aforementioned multimodal sentiment fusion feature information and is used to characterize the features of each element in the aforementioned multimodal temporal relationship matrix.

[0063] In practice, the aforementioned multimodal temporal relationship matrix and multimodal sentiment fusion feature information can be input into a pre-trained temporal graph fusion network to obtain temporal feature information. The aforementioned temporal graph fusion network can be a neural network that takes the multimodal temporal relationship matrix and multimodal sentiment fusion feature information as input and outputs the temporal feature information. For example, the aforementioned temporal graph fusion network can be a pre-trained graph convolutional neural network. The aforementioned pre-training can be a process of fine-tuning the graph convolutional neural network using a temporal network training dataset and a cross-entropy loss function. Each temporal network training data in the aforementioned temporal network training dataset can be data used to train the graph convolutional neural network. Each temporal network training data in the aforementioned temporal network training dataset can include the multimodal temporal relationship matrix and multimodal sentiment fusion feature information.

[0064] The second step involves generating speaker feature information based on the aforementioned speaker relationship matrix and the aforementioned multimodal emotion fusion feature information. This speaker feature information can be a feature vector that integrates the aforementioned multimodal emotion fusion feature information and represents the features of each element in the aforementioned speaker relationship matrix. In practice, the executing entity can input the aforementioned speaker relationship matrix and the aforementioned multimodal emotion fusion feature information into a pre-trained speaker graph fusion network to obtain the speaker feature information. This speaker graph fusion network can be a neural network that takes the speaker relationship matrix and the multimodal emotion fusion feature information as input and outputs the speaker feature information. For example, the speaker graph fusion network can be a pre-trained graph convolutional neural network. The pre-training process can be a fine-tuning of the graph convolutional neural network using a speaker training dataset and a cross-entropy loss function. Each speaker training data point in the aforementioned speaker training dataset can be used to train the graph convolutional neural network into a speaker graph fusion network. Each speaker training data point in the aforementioned speaker training dataset can include the speaker relationship matrix and the multimodal emotion fusion feature information.

[0065] The third step involves concatenating the aforementioned temporal and speaker features to obtain multimodal graph feature information. This multimodal graph feature information can be a feature vector obtained by concatenating the temporal and speaker features. In practice, the executing entity can input the temporal and speaker features into a feature concatenation function to obtain the multimodal graph feature information. This feature concatenation function can be any function capable of concatenating features. For example, the feature concatenation function could be the `concat()` function.

[0066] Step 1025: Based on multimodal graph feature information, multimodal emotion fusion feature information, and the conditional diffusion generative network in the pre-trained emotion recognition model, generate reconstructed text convolutional features, reconstructed audio convolutional features, and reconstructed video convolutional features.

[0067] In some embodiments, the aforementioned execution entity may generate reconstructed text convolutional features, reconstructed audio convolutional features, and reconstructed video convolutional features based on the aforementioned multimodal graph feature information, the aforementioned multimodal emotion fusion feature information, and the conditional diffusion generation network in the pre-trained emotion recognition model.

[0068] The aforementioned emotion recognition model can be a neural network model that takes multimodal graph feature information and multimodal emotion fusion feature information as input and emotion recognition results as output. The emotion recognition results can be labels used to characterize the emotion categories corresponding to the multimodal graph feature information and multimodal emotion fusion feature information. For example, the emotion recognition results could be "happy," "nervous," or "sad."

[0069] The aforementioned emotion recognition model may include a conditional diffusion generation network, a feature processing network, and a classification network. The conditional diffusion generation network may include a text modal diffusion network, an audio modal diffusion network, a video modal diffusion network, text convolutional layers, audio convolutional layers, and video convolutional layers. Each of these conditional diffusion generation networks corresponds to a diffusion time step. Each diffusion time step can be considered the time step in which the conditional diffusion generation network processes the data. Each diffusion time step corresponds to text Gaussian vector data, audio Gaussian vector data, video Gaussian vector data, sample diffusion text feature information, sample diffusion audio feature information, and sample diffusion video feature information. The text Gaussian vector data can be the Gaussian noise generated by the text modal diffusion network in the conditional diffusion generation network at the corresponding diffusion time step. The audio Gaussian vector data can be the Gaussian noise generated by the audio modal diffusion network in the conditional diffusion generation network at the corresponding diffusion time step. The video Gaussian vector data can be the Gaussian noise generated by the video modal diffusion network in the conditional diffusion generation network at the corresponding diffusion time step. The aforementioned sample diffusion text feature information can be obtained by adding text Gaussian vector data to the feature vector input to the text modality diffusion network. Similarly, the aforementioned sample diffusion audio feature information can be obtained by adding audio Gaussian vector data to the feature vector input to the audio modality diffusion network. Finally, the aforementioned sample diffusion video feature information can be obtained by adding video Gaussian vector data to the feature vector input to the video modality diffusion network.

[0070] In some optional implementations of certain embodiments, the emotion recognition model described above can be trained by the aforementioned execution subject through the following steps:

[0071] The first step is to obtain a sample set. Each sample in the sample set can be used as sample data for training the emotion recognition model. Each sample in the sample set can include sample multimodal feature information, sample multimodal graph feature information, sample multimodal emotion fusion feature information, and sample emotion recognition results. The sample multimodal feature information can be used for training the emotion recognition model. The sample multimodal graph feature information can be used for training the emotion recognition model. The sample multimodal emotion fusion feature information can be used for training the emotion recognition model. The sample multimodal emotion fusion feature information corresponds to data in three modalities: text, audio, and video. The sample multimodal feature information can include sample text feature information, sample audio feature information, and sample video feature information. The sample text feature information can be the text feature information included in the sample multimodal feature information. The sample audio feature information can be the audio feature information included in the sample multimodal feature information. The sample video feature information can be the video feature information included in the sample multimodal feature information. The above sample emotion recognition results can be the actual emotion recognition results corresponding to the samples.

[0072] The second step, based on the sample set, is to perform the following training steps:

[0073] The first sub-step involves, for each sample in at least one sample set, inputting the sample multimodal sentiment fusion feature information and the sample multimodal graph feature information included in the sample into the text modality diffusion network of the conditional diffusion generation network in the initial neural network, to obtain the sample reconstructed text feature information corresponding to the sample. Here, the sample reconstructed text feature information can be a feature vector of data representing the text modality corresponding to the sample multimodal sentiment fusion feature information, obtained after reconstructing the sample multimodal sentiment fusion feature information. The text modality diffusion network can be a neural network that takes the sample multimodal sentiment fusion feature information and the sample multimodal graph feature information as input and the sample reconstructed text feature information as output. For example, the text modality diffusion network can be a pre-trained ScoreNet model. The pre-training can be a process of fine-tuning the ScoreNet model using a text training dataset and a mean squared error loss function. Each text training data in the text training dataset can be data used to train the ScoreNet model to generate sample reconstructed text feature information. Each text training data in the text training dataset can include sample multimodal sentiment fusion feature information and sample multimodal graph feature information.

[0074] The second sub-step involves, for each of the at least one of the aforementioned samples, inputting the reconstructed text feature information corresponding to that sample into the text convolutional layer of the conditional diffusion generative network in the initial neural network, to obtain the reconstructed text convolutional features corresponding to that sample. These reconstructed text convolutional features can be reconstructed text feature information that has undergone convolution processing. The text convolutional layer can be a convolutional layer that takes the reconstructed text feature information as input and the reconstructed text convolutional features as output.

[0075] The third sub-step involves inputting the sample multimodal emotion fusion feature information and the sample multimodal graph feature information included in the sample into the audio modal diffusion network of the conditional diffusion generation network in the initial neural network for each sample, thereby obtaining the sample reconstruction audio feature information corresponding to the sample.

[0076] The aforementioned reconstructed audio feature information can be a feature vector representing the audio modality corresponding to the aforementioned multimodal emotion fusion feature information, obtained after reconstructing the multimodal emotion fusion feature information. The aforementioned audio modality diffusion network can be a neural network that takes the multimodal emotion fusion feature information and the multimodal graph feature information as input and outputs the reconstructed audio feature information. For example, the aforementioned audio modality diffusion network can be a preprocessed ScoreNet model. The aforementioned preprocessing can be a process of fine-tuning the ScoreNet model using an audio training dataset and a mean squared error loss function. Each audio training data point in the aforementioned audio training dataset can be data used to train the ScoreNet model to generate the reconstructed audio feature information. Each audio training data point in the aforementioned audio training dataset can include the multimodal emotion fusion feature information and the multimodal graph feature information.

[0077] The fourth sub-step involves, for each of the at least one of the aforementioned samples, inputting the reconstructed audio feature information corresponding to that sample into the audio convolutional layer of the conditional diffusion generator network in the initial neural network, to obtain the reconstructed audio convolutional features corresponding to that sample. These reconstructed audio convolutional features can be reconstructed audio feature information that has undergone convolution processing. The audio convolutional layer can be a convolutional layer that takes the reconstructed audio feature information as input and outputs the reconstructed audio convolutional features.

[0078] The fifth sub-step involves inputting the sample multimodal emotion fusion feature information and the sample multimodal graph feature information included in the sample into the video modal diffusion network of the conditional diffusion generation network in the initial neural network for each sample, thereby obtaining the sample reconstruction video feature information corresponding to the sample.

[0079] The aforementioned sample-reconstructed video feature information can be a feature vector obtained by reconstructing the aforementioned sample multimodal emotion fusion feature information, used to characterize the video modality corresponding to the aforementioned sample multimodal emotion fusion feature information. The aforementioned video modality diffusion network can be a neural network model that takes sample multimodal emotion fusion feature information and sample multimodal graph feature information as input and outputs sample-reconstructed video feature information. For example, the aforementioned video modality diffusion network can be a preprocessed ScoreNet model. The aforementioned preprocessing can be a process of fine-tuning the ScoreNet model using a video training dataset and a mean squared error loss function. Each video training data in the aforementioned video training dataset can be data used to train the ScoreNet model to generate sample-reconstructed video feature information. Each video training data in the aforementioned video training dataset can include sample multimodal emotion fusion feature information and sample multimodal graph feature information.

[0080] The sixth sub-step involves, for each of the at least one of the aforementioned samples, inputting the sample-reconstructed video feature information corresponding to that sample into the video convolutional layer of the conditional diffusion generator network in the initial neural network, to obtain the sample-reconstructed video convolutional features corresponding to that sample. These sample-reconstructed video convolutional features can be sample-reconstructed video feature information that has undergone convolution processing. The video convolutional layer can be a convolutional layer that takes the sample-reconstructed video feature information as input and outputs the sample-reconstructed video convolutional features.

[0081] The seventh sub-step involves inputting the sample multimodal graph feature information, the sample reconstruction text convolutional feature, the sample reconstruction audio convolutional feature, and the sample reconstruction video convolutional feature of the sample into the feature processing network of the initial neural network for each of the at least one sample mentioned above, to obtain the sample reconstruction fusion feature information.

[0082] The aforementioned sample reconstruction fusion feature information can be a feature vector obtained by fusing the aforementioned sample multimodal map feature information, the sample reconstruction text convolutional features corresponding to the aforementioned samples, the sample reconstruction audio convolutional features, and the sample reconstruction video convolutional features. The aforementioned feature processing network can be a neural network model that takes the sample multimodal map feature information, sample reconstruction text convolutional features, sample reconstruction audio convolutional features, and sample reconstruction video convolutional features as input, and takes the sample reconstruction fusion feature information as output. For example, the aforementioned feature processing network can be a multilayer perceptron (MLP).

[0083] The eighth sub-step involves inputting the reconstructed and fused feature information of each sample corresponding to at least one of the above samples into the classification network in the initial neural network to obtain each emotion recognition result.

[0084] The aforementioned classification network may include a normalization function and an output function. The normalization function may be a normalized exponential function that takes sample reconstruction fusion feature information as input and outputs the sentiment label probability distribution data corresponding to the sample reconstruction fusion feature information. The sentiment label probability distribution data may be the probability distribution data of the sentiment categories corresponding to the sample reconstruction fusion feature information. The sentiment label probability distribution data may include each sentiment label and each label probability data. Each sentiment label and each label probability data corresponds one-to-one. Each sentiment label may be a label used to characterize the sentiment category corresponding to the sample reconstruction fusion feature information. For example, the sentiment label may be "happy," "sad," or "nervous." Each label probability data may be the probability corresponding to the sentiment label.

[0085] The output function described above can be an argmax function that takes sentiment label probability distribution data as input and outputs the sentiment recognition result. The sentiment recognition result is the sentiment label with the highest corresponding label probability data in the aforementioned sentiment label probability distribution data.

[0086] The ninth sub-step generates classification loss data based on the conditional diffusion generative network, the emotion recognition results corresponding to at least one of the above samples, the reconstructed text convolutional features of each sample, the reconstructed audio convolutional features of each sample, the reconstructed video convolutional features of each sample, the emotion recognition results of each sample included in the at least one sample, the multimodal feature information of each sample, and the multimodal graph feature information of each sample. The classification loss data can be the loss value corresponding to the at least one sample.

[0087] The third step involves determining that the classification loss data meets a preset classification loss condition, and then defining the initial neural network as an emotion recognition model. The classification loss condition can be that the classification loss data is less than a preset loss threshold. This loss threshold can be a pre-set value. The specific setting of this loss threshold is not limited here.

[0088] The fourth step involves adjusting the network parameters of the initial neural network in response to the determination that the classification loss data does not meet the classification loss conditions. A sample set is then created using unused samples, and the adjusted initial neural network is used to perform the training steps again. In practice, the execution entity can use the back propagation algorithm (BP algorithm) and gradient descent methods (such as mini-batch gradient descent) to adjust the network parameters of the initial neural network.

[0089] In the process of adopting technical solutions to address the aforementioned technical problems, the following issues often arise:

[0090] When performing emotion recognition using only single-modal data, the accuracy of emotion recognition is likely to be low. This leads to a higher probability of needing to call computing resources to repeatedly recognize the same data, which can easily result in a waste of computing resources during recognition.

[0091] In response to the aforementioned technical problems, the following solution was adopted:

[0092] In some optional implementations of certain embodiments, the aforementioned execution entity can generate classification loss data based on the following steps: a conditional diffusion generation network, the emotion recognition results corresponding to the at least one sample, reconstructed text convolutional features for each sample, reconstructed audio convolutional features for each sample, reconstructed video convolutional features for each sample, the emotion recognition results of each sample included in the at least one sample, multimodal feature information of each sample, and multimodal graph feature information of each sample.

[0093] First, for each of the at least one of the above samples, perform the following steps:

[0094] The first sub-step is to determine the emotion recognition result corresponding to the above sample as the target emotion recognition result.

[0095] The second sub-step involves determining the sentiment recognition results of the samples included in the above samples as the sentiment recognition results of the target samples.

[0096] The third sub-step involves generating recognition result loss data based on the target sentiment recognition result and the target sample sentiment recognition result. This recognition result loss data can be the cross-entropy loss value between the target sentiment recognition result and the target sample sentiment recognition result. In practice, firstly, the sentiment label probability distribution data corresponding to the target sentiment recognition result can be determined as the target probability distribution data. Then, the various label probability data included in the target probability distribution data can be combined into a one-dimensional vector as the label vector. Next, the target sample sentiment recognition result can be one-hot encoded using an encoding function to obtain the one-hot vector corresponding to the target sample sentiment recognition result as the sample label vector. The encoding function can be any function capable of one-hot encoding data. For example, the encoding function could be the `get_dummies()` function in pandas. Finally, the cross-entropy loss value between the label vector and the sample label vector can be determined as the recognition result loss data.

[0097] The fourth sub-step involves determining the sample multimodal feature information included in the aforementioned samples as the target multimodal feature information. This target multimodal feature information may include target text feature information, target audio feature information, and target video feature information. Specifically, the target text feature information can be the sample text feature information included in the aforementioned sample multimodal feature information. The target audio feature information can be the sample audio feature information included in the aforementioned sample multimodal feature information. The target video feature information can be the sample video feature information included in the aforementioned sample multimodal feature information.

[0098] In practice, firstly, the sample audio features included in the aforementioned multimodal feature information can be identified as the target audio feature information. Secondly, the sample audio features included in the aforementioned multimodal feature information can be identified as the target audio feature information. Then, the sample video features included in the aforementioned multimodal feature information can be identified as the target video feature information. Finally, the aforementioned target audio feature information, the aforementioned target video feature information, and the aforementioned target video feature information can be combined to form the target multimodal feature information.

[0099] The fifth sub-step is to determine the difference between the convolutional features of the reconstructed text corresponding to the above samples and the target text feature information in the above target multimodal feature information as the text difference feature information.

[0100] The sixth sub-step is to determine the modulus of the above-mentioned text difference feature information as text feature modulus data.

[0101] The seventh sub-step is to determine the square of the above text feature modulus data as the text feature modulus square data.

[0102] The eighth sub-step is to determine the difference between the reconstructed audio convolutional features of the samples corresponding to the above samples and the target audio feature information in the above target multimodal feature information as the audio difference feature information.

[0103] The ninth sub-step involves determining the modulus of the aforementioned audio difference feature information as audio feature modulus data.

[0104] The tenth sub-step is to determine the square of the above audio feature modulus data as the audio feature modulus square data.

[0105] The eleventh sub-step involves determining the difference between the sample-reconstructed video convolutional features corresponding to the above samples and the target video feature information in the above target multimodal feature information as the video difference feature information.

[0106] The twelfth sub-step involves determining the modulus of the aforementioned video difference feature information as video feature modulus data.

[0107] The thirteenth sub-step is to determine the square of the above video feature modulus data as the video feature modulus square data.

[0108] The fourteenth sub-step involves determining the average value of the above-mentioned text feature modulus square data, the above-mentioned audio feature modulus square data, and the above-mentioned video feature modulus square data as the reconstruction loss data.

[0109] The fifteenth sub-step generates score loss data based on the conditional diffusion generation network and the multimodal graph feature information of the samples included in the above samples. The score loss data can be the loss value corresponding to each data point generated by the conditional diffusion generation network at the above sampling time step.

[0110] The sixteenth sub-step involves generating sample loss data based on the aforementioned recognition result loss data, reconstruction loss data, and score loss data. This sample loss data can be obtained by weighting the recognition result loss data, reconstruction loss data, and score loss data. In practice, firstly, the product of the reconstruction loss data and a first preset weight data can be determined as the weighted reconstruction loss data. The first preset weight data can be the weight data corresponding to the reconstruction loss data. For example, the first preset weight data can be 0.5. Secondly, the product of the score loss data and a second preset weight data can be determined as the weighted score loss data. The second preset weight data can be the weight data corresponding to the score loss data. For example, the second preset weight data can be 0.6. Finally, the sum of the recognition result loss data, the weighted reconstruction loss data, and the weighted score loss data can be determined as the sample loss data.

[0111] The second step is to generate classification loss data based on the generated loss data for each sample. This classification loss data can be the average of the loss data for each sample. In practice, the executing entity can determine the average of the loss data for each sample as the classification loss data.

[0112] The above technical solution and its related content, combined with steps 101 to 103, serve as an inventive point of this disclosure, solving the problem of "waste of computing resources." Factors leading to wasted computing resources often include: when performing emotion recognition using only single-modal data, the accuracy of emotion recognition is easily low, leading to a higher probability of needing to repeatedly recognize the same data using computing resources, easily resulting in wasted computing resources during recognition. Solving these factors can reduce wasted computing resources. To achieve this effect, this disclosure performs the following steps for each of the at least one sample mentioned above: First, the emotion recognition result corresponding to the sample is determined as the target emotion recognition result. Second, the sample emotion recognition results included in the sample are determined as the target sample emotion recognition result. Then, based on the target emotion recognition result and the target sample emotion recognition result, recognition result loss data is generated. Thus, the loss value between the target emotion recognition result and the target sample emotion recognition result can be obtained. Then, the multimodal feature information included in the sample is determined as the target multimodal feature information, wherein the target multimodal feature information includes target text feature information, target audio feature information, and target video feature information. Next, the difference between the reconstructed text convolutional features corresponding to the above samples and the target text features in the above target multimodal feature information is determined as the text difference feature information. Thus, the difference between the reconstructed text convolutional features and the target text features can be obtained. Then, the modulus of the text difference feature information is determined as the text feature modulus data. Thus, the modulus data of the text difference feature information can be obtained. Then, the square of the text feature modulus data is determined as the text feature modulus squared data. Thus, the text feature modulus squared data can be obtained. Next, the difference between the reconstructed audio convolutional features corresponding to the above samples and the target audio features in the above target multimodal feature information is determined as the audio difference feature information. Thus, the difference between the reconstructed audio convolutional features and the target audio features can be obtained. Then, the modulus of the audio difference feature information is determined as the audio feature modulus data. Thus, the audio feature modulus data can be obtained. Then, the square of the audio feature modulus data is determined as the audio feature modulus squared data. Thus, the audio feature modulus squared data can be obtained. Next, the difference between the convolutional features of the reconstructed video corresponding to the above samples and the target video feature information in the above target multimodal feature information is determined as the video difference feature information. Thus, the difference between the convolutional features of the reconstructed video and the target video feature information can be obtained. Then, the modulus of the above video difference feature information is determined as the video feature modulus data. Thus, the modulus data of the video difference feature information can be obtained. Then, the square of the above video feature modulus data is determined as the video feature modulus squared data. Thus, the video feature modulus squared data can be obtained.Next, the average value of the squared modulus data of the text features, the squared modulus data of the audio features, and the squared modulus data of the video features is determined as the reconstruction loss data. Thus, the loss data after feature reconstruction of the three modalities of text, audio, and video data can be obtained. Then, based on the conditional diffusion generation network and the multimodal graph feature information of the samples included in the above samples, score loss data is generated. Thus, the loss value of each data generated by the conditional diffusion generation network can be obtained. Then, based on the recognition result loss data, the reconstruction loss data, and the score loss data, sample loss data is generated. Thus, the loss value of the above samples can be obtained by combining the recognition result loss data, the reconstruction loss data, and the score loss data. Finally, based on the generated sample loss data, classification loss data is generated. Thus, the average loss value of each sample in at least one of the above samples can be obtained. Because the network parameters of the initial neural network can be iteratively adjusted by combining the loss values ​​of the recognition result loss data, the reconstruction loss data, and the score loss data, the recognition accuracy of the trained emotion recognition model when performing emotion recognition on data can be improved. This reduces the probability of multiple recognitions of the same data due to low recognition accuracy, thereby reducing the waste of computational resources.

[0113] In the process of adopting technical solutions to address the aforementioned technical problems, the following issues often arise:

[0114] Emotion recognition based solely on processing unimodal data can easily lead to low accuracy and user satisfaction. It also increases the likelihood of needing to repeatedly recognize the same unimodal data, resulting in wasted computing resources.

[0115] In response to the aforementioned technical problems, the following solution was adopted:

[0116] In some optional implementations of certain embodiments, the aforementioned execution entity can generate score loss data based on the conditional diffusion generation network and the multimodal graph feature information of the samples included in the aforementioned samples through the following steps:

[0117] The first step is to determine the diffusion time step that meets the preset sampling conditions among the above diffusion time steps as the sampling time step. The sampling conditions can be any one of the above diffusion time steps.

[0118] The second step involves generating Gaussian vector data of the text to be processed based on the aforementioned multimodal graph feature information, the aforementioned sampling time step, and the sample diffusion text feature information corresponding to the aforementioned sampling time step. The aforementioned Gaussian vector data of the text to be processed can be Gaussian noise obtained through prediction from the aforementioned sample diffusion text feature information.

[0119] In practice, the aforementioned execution entity can input the sample multimodal graph feature information, the sampling time step, and the sample diffusion text feature information corresponding to the sampling time step into a pre-trained text reward model to obtain Gaussian vector data of the text to be processed. The aforementioned text reward model can be a neural network model that takes the sample multimodal graph feature information, the sampling time step, and the sample diffusion text feature information corresponding to the sampling time step as input and the Gaussian vector data of the text to be processed as output. For example, the aforementioned text reward model can be a pre-processed ScoreNet model. The aforementioned preprocessing can be a process of fine-tuning the ScoreNet model using a text dataset and a cross-entropy loss function. Each text data in the aforementioned text dataset can be data used to train the ScoreNet model to generate the Gaussian vector data of the text to be processed. Each text data in the aforementioned text dataset can include sample multimodal graph feature information, the sampling time step, and the sample diffusion text feature information corresponding to the sampling time step.

[0120] The third step is to determine the weighted text Gaussian vector data by multiplying the Gaussian vector data of the text to be processed by the preset Gaussian weight data. The Gaussian weight data can be a pre-set weight value. For example, the Gaussian weight data can be 0.7.

[0121] The fourth step is to determine the sum of the weighted text Gaussian vector data and the text Gaussian vector data corresponding to the above sampling time step as the summed text Gaussian vector data.

[0122] The fifth step is to determine the modulus of the summed text Gaussian vector data as the text Gaussian modulus data.

[0123] Step 6: Determine the square of the above text Gaussian modulus data as the text Gaussian square data.

[0124] Step 7: Based on the aforementioned sample multimodal graph feature information, the aforementioned sampling time step, and the sample diffusion audio feature information corresponding to the aforementioned sampling time step, generate the audio Gaussian vector data to be processed. The aforementioned audio Gaussian vector data to be processed can be Gaussian noise obtained through prediction, corresponding to the aforementioned sample diffusion audio feature information.

[0125] In practice, the aforementioned multimodal graph feature information, sampling time step, and corresponding sample diffusion audio feature information can be input into a pre-trained audio reward model to obtain the audio Gaussian vector data to be processed. The audio reward model can be a neural network model that takes the multimodal graph feature information, sampling time step, and corresponding sample diffusion audio feature information as input and outputs the audio Gaussian vector data to be processed. For example, the audio reward model can be a pre-processed ScoreNet model. The preprocessing can be a process of fine-tuning the ScoreNet model using the audio dataset and the cross-entropy loss function. Each audio data point in the aforementioned audio dataset can be used to train the ScoreNet model to generate the audio Gaussian vector data to be processed. Each audio data point in the aforementioned audio dataset can include multimodal graph feature information, sampling time step, and corresponding sample diffusion audio feature information.

[0126] Step 8: The product of the audio Gaussian vector data to be processed and the Gaussian weight data is determined as the weighted audio Gaussian vector data.

[0127] Step 9: The sum of the weighted audio Gaussian vector data and the audio Gaussian vector data corresponding to the sampling time step is determined as the summed audio Gaussian vector data.

[0128] Step 10: Determine the modulus of the above summed audio Gaussian vector data as the audio Gaussian modulus data.

[0129] Step 11: Determine the square of the above audio Gaussian modulus data as the audio Gaussian square data.

[0130] Step 12: Based on the aforementioned sample multimodal graph feature information, the aforementioned sampling time step, and the sample diffusion video feature information corresponding to the aforementioned sampling time step, generate Gaussian vector data of the video to be processed. The aforementioned Gaussian vector data of the video to be processed can be Gaussian noise obtained through prediction based on the aforementioned sample diffusion video feature information.

[0131] In practice, the aforementioned multimodal graph feature information, sampling time step, and corresponding sample diffusion video feature information can be input into a pre-trained video reward model to obtain Gaussian vector data of the video to be processed. The video reward model can be a neural network model that takes the multimodal graph feature information, sampling time step, and corresponding sample diffusion video feature information as input and outputs the Gaussian vector data of the video to be processed. For example, the video reward model can be a pre-processed ScoreNet model. The pre-processing can be a process of fine-tuning the ScoreNet model using the video dataset and the cross-entropy loss function. Each video data in the aforementioned video dataset can be used to train the ScoreNet model to generate the Gaussian vector data of the video to be processed. Each video data in the aforementioned video dataset can include multimodal graph feature information, sampling time step, and corresponding sample diffusion video feature information.

[0132] Step 13: The product of the above-mentioned Gaussian vector data of the video to be processed and the above-mentioned Gaussian weight data is determined as the weighted video Gaussian vector data.

[0133] Step fourteen: The sum of the above weighted video Gaussian vector data and the video Gaussian vector data corresponding to the above sampling time step is determined as the summed video Gaussian vector data.

[0134] Step 15: Determine the modulus of the above summed video Gaussian vector data as the video Gaussian modulus data.

[0135] Step sixteen: Determine the square of the above video Gaussian modulus data as the video Gaussian square data.

[0136] Step 17: The average of the above text Gaussian squared data, the above audio Gaussian squared data, and the above video Gaussian squared data is determined as the score loss data.

[0137] The above technical solution and its related content, combined with steps 101 to 103, serve as an inventive point of this disclosure, solving the problem of "waste of computing resources." Factors leading to wasted computing resources often include: processing emotion recognition solely through single-modal data can easily result in low accuracy of the emotion recognition results, low user satisfaction with the results, and a high probability of needing to repeatedly recognize the same single-modal data, thus wasting computing resources. Solving these factors can reduce wasted computing resources. To achieve this, this disclosure first determines the diffusion time step that meets the preset sampling conditions among the above diffusion time steps as the sampling time step. This allows the determination of the time step to be processed. Secondly, based on the sample multimodal graph feature information, the above sampling time step, and the sample diffusion text feature information corresponding to the above sampling time step, Gaussian vector data of the text to be processed is generated. This allows the prediction of the Gaussian noise corresponding to the text modality diffusion network in the conditional diffusion generation network at the above sampling time step. Then, the product of the above-mentioned Gaussian vector data of the text to be processed and the preset Gaussian weight data is determined as the weighted text Gaussian vector data. Therefore, weighted text Gaussian vector data can be obtained. Then, the sum of the weighted text Gaussian vector data and the text Gaussian vector data corresponding to the sampling time step is determined as the summed text Gaussian vector data. Next, the modulus of the summed text Gaussian vector data is determined as the text Gaussian modulus data. Then, the square of the text Gaussian modulus data is determined as the text Gaussian square data. Then, based on the sample multimodal graph feature information, the sampling time step, and the sample diffusion audio feature information corresponding to the sampling time step, audio Gaussian vector data to be processed is generated. Therefore, the Gaussian noise corresponding to the audio modality diffusion network in the conditional diffusion generation network at the sampling time step can be obtained. Then, the product of the audio Gaussian vector data to be processed and the Gaussian weight data is determined as the weighted audio Gaussian vector data. Therefore, weighted audio Gaussian vector data can be obtained. Then, the sum of the weighted audio Gaussian vector data and the audio Gaussian vector data corresponding to the above sampling time step is determined as the summed audio Gaussian vector data. Thus, the summed audio Gaussian vector data can be obtained. Then, the magnitude of the above summed audio Gaussian vector data is determined as the audio Gaussian magnitude data. Then, the square of the above audio Gaussian magnitude data is determined as the audio Gaussian square data. Next, based on the above sample multimodal graph feature information, the above sampling time step, and the sample diffusion video feature information corresponding to the above sampling time step, the Gaussian vector data of the video to be processed is generated. Thus, the Gaussian noise corresponding to the video modal diffusion network in the conditional diffusion generation network at the above sampling time step can be obtained.Then, the product of the aforementioned Gaussian vector data of the video to be processed and the aforementioned Gaussian weight data is determined as the weighted video Gaussian vector data. Thus, the weighted video Gaussian vector data can be obtained. Next, the sum of the aforementioned weighted video Gaussian vector data and the video Gaussian vector data corresponding to the aforementioned sampling time step is determined as the summed video Gaussian vector data. Thus, the summed video Gaussian vector data can be obtained. Then, the modulus of the aforementioned summed video Gaussian vector data is determined as the video Gaussian modulus data. Then, the square of the aforementioned video Gaussian modulus data is determined as the video Gaussian squared data. Finally, the average of the aforementioned text Gaussian squared data, the aforementioned audio Gaussian squared data, and the aforementioned video Gaussian squared data is determined as the score loss data. Thus, by averaging the aforementioned text Gaussian squared data, the aforementioned audio Gaussian squared data, and the aforementioned video Gaussian squared data, the average loss value of the aforementioned text Gaussian squared data, the aforementioned audio Gaussian squared data, and the aforementioned video Gaussian squared data can be obtained. Furthermore, by processing the feature vectors corresponding to the three modalities generated by the text modal diffusion network, audio modal diffusion network, and video modal diffusion network in the conditional diffusion generator network at any diffusion time step, the Gaussian noise corresponding to the above diffusion time step can be inferred. Then, by comparing the inferred Gaussian noise with the Gaussian noise actually generated by the conditional diffusion generator network at the above diffusion time step, the loss values ​​corresponding to the feature vectors of the three modalities can be obtained. Then, the initial neural network can be iteratively optimized based on the loss values ​​corresponding to the feature vectors of the three modalities. Therefore, the accuracy of the trained emotion recognition model in performing emotion recognition on data can be improved, the probability of repeated recognition of the same data due to low recognition accuracy can be reduced, and thus the waste of computing resources can be reduced.

[0138] Step 1026: Based on the reconstructed text convolutional features, reconstructed audio convolutional features, reconstructed video convolutional features, multimodal graph feature information, and the feature processing network in the emotion recognition model, generate reconstructed fusion feature information.

[0139] In some embodiments, the execution entity can generate reconstructed fusion feature information based on the reconstructed text convolutional features, the reconstructed audio convolutional features, the reconstructed video convolutional features, the multimodal graph feature information, and the feature processing network in the emotion recognition model. The reconstructed fusion feature information can be a feature vector obtained by fusing the reconstructed text convolutional features, the reconstructed audio convolutional features, the reconstructed video convolutional features, and the multimodal graph feature information. In practice, the reconstructed text convolutional features, the reconstructed audio convolutional features, the reconstructed video convolutional features, and the multimodal graph feature information can be input into the feature processing network in the emotion recognition model to obtain the reconstructed fusion feature information.

[0140] Step 1027: Based on the reconstruction of the classification network in the fusion feature information and the emotion recognition model, generate multimodal emotion recognition results.

[0141] In some embodiments, the executing entity can generate multimodal emotion recognition results based on the reconstructed fusion feature information and the classification network in the emotion recognition model. The multimodal emotion recognition results can be labels representing the emotion categories corresponding to the multimodal emotion data. In practice, the reconstructed fusion feature information can be input into the classification network in the emotion recognition model to obtain the emotion recognition results. Then, the emotion recognition results can be identified as multimodal emotion recognition results.

[0142] Step 103: Send the generated multimodal emotion recognition results to the user terminal.

[0143] In some embodiments, the aforementioned executing entity may send the generated multimodal emotion recognition results to the aforementioned user terminal.

[0144] The above embodiments of this disclosure have the following beneficial effects: the emotion recognition method based on multimodal missing data in some embodiments of this disclosure can improve the accuracy of emotion recognition and reduce the waste of computational resources during recognition. Specifically, the reason for the low accuracy of emotion recognition and the waste of computational resources during recognition is that emotion recognition based on only a single modality of data is prone to low accuracy. Moreover, when processing modal data such as text and audio, data loss is likely to occur, which can easily lead to low accuracy when performing emotion recognition based on the processed data. This results in the need to call computational resources to repeatedly recognize the multimodal data sent by the same user terminal, causing a waste of computational resources during recognition. Based on this, the emotion recognition method based on multimodal missing data in some embodiments of this disclosure firstly, in response to receiving a multimodal sentence sequence set sent by the user terminal, generates a multimodal emotion data sequence set based on the multimodal sentence sequence set. Thus, the original data to be processed can be obtained. Secondly, for each multimodal sentiment data point in the aforementioned multimodal sentiment data sequence set, the following steps are performed: First, feature extraction processing is performed on the multimodal sentiment data to obtain text modality feature information, audio modality feature information, and video modality feature information corresponding to the multimodal sentiment data. This yields feature vectors corresponding to different modalities in the multimodal sentiment data. Second, based on the aforementioned text modality feature information, audio modality feature information, and video modality feature information, multimodal sentiment fusion feature information corresponding to the aforementioned multimodal sentiment data is generated. This allows for the fusion of feature vectors corresponding to different modalities in the multimodal sentiment data. Then, based on the aforementioned multimodal sentiment data, a multimodal temporal relation matrix and a speaker relation matrix are constructed. This yields the multimodal temporal relation matrix and the speaker relation matrix. Finally, based on the aforementioned multimodal sentiment fusion feature information, the aforementioned multimodal temporal relation matrix, and the aforementioned speaker relation matrix, multimodal graph feature information is generated. This generates multimodal graph feature information. Next, based on the aforementioned multimodal graph feature information, the aforementioned multimodal emotion fusion feature information, and the conditional diffusion generative network in the pre-trained emotion recognition model, reconstructed text convolutional features, reconstructed audio convolutional features, and reconstructed video convolutional features are generated. Then, based on the aforementioned reconstructed text convolutional features, the aforementioned reconstructed audio convolutional features, the aforementioned reconstructed video convolutional features, the aforementioned multimodal graph feature information, and the aforementioned feature processing network in the emotion recognition model, reconstructed fusion feature information is generated. Then, based on the aforementioned reconstructed fusion feature information and the aforementioned classification network in the emotion recognition model, multimodal emotion recognition results are generated. Finally, the generated multimodal emotion recognition results are sent to the aforementioned user terminal.Therefore, the results of various multimodal emotion recognition can be sent to the user terminal. Because the data from multiple modalities sent by the user terminal can be fused, and then emotion recognition can be performed based on the processed data, the accuracy of emotion recognition can be improved. Furthermore, because the feature vectors corresponding to different modalities can be fused first to obtain multimodal emotion fusion feature information, and then emotion recognition can be performed based on the fused features, even if data from one modality is missing, data from other modalities can still provide a data foundation for subsequent emotion recognition. This reduces the probability of low emotion recognition accuracy due to missing modal data, and also reduces the probability of repeatedly using computing resources to recognize the same data due to low accuracy, thus reducing the waste of computing resources.

[0145] Further reference Figure 2 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of an emotion recognition device based on multimodal missing data. These device embodiments are similar to... Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.

[0146] like Figure 2As shown, an emotion recognition device 200 based on multimodal missing data in some embodiments includes: a generation unit 201, an execution unit 202, and a sending unit 203. The generation unit 201 is configured to generate a multimodal emotion data sequence set based on a multimodal statement sequence set received from a user terminal. The execution unit 202 is configured to perform the following steps for each multimodal emotion data in the multimodal emotion data sequence set: perform feature extraction processing on the multimodal emotion data to obtain text modality feature information, audio modality feature information, and video modality feature information corresponding to the multimodal emotion data; generate multimodal emotion fusion feature information corresponding to the multimodal emotion data based on the text modality feature information, the audio modality feature information, and the video modality feature information; construct a multimodal temporal relationship matrix and a speaker relationship matrix based on the multimodal emotion data; and perform multimodal emotion fusion feature information based on the multimodal emotion data. Multimodal graph feature information is generated by fusing feature information, the aforementioned multimodal temporal relation matrix, and the aforementioned speaker relation matrix. Based on the aforementioned multimodal graph feature information, the aforementioned multimodal emotion fusion feature information, and the conditional diffusion generation network in the pre-trained emotion recognition model, reconstructed text convolutional features, reconstructed audio convolutional features, and reconstructed video convolutional features are generated. Based on the aforementioned reconstructed text convolutional features, the aforementioned reconstructed audio convolutional features, the aforementioned reconstructed video convolutional features, the aforementioned multimodal graph feature information, and the aforementioned feature processing network in the aforementioned emotion recognition model, reconstructed fusion feature information is generated. Based on the aforementioned reconstructed fusion feature information and the aforementioned classification network in the aforementioned emotion recognition model, a multimodal emotion recognition result is generated. The sending unit 203 is configured to send the generated multimodal emotion recognition results to the aforementioned user terminal.

[0147] It is understandable that the units described in the emotion recognition device 200 based on multimodal missing data are related to the reference. Figure 1 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method are also applicable to the emotion recognition device 200 based on multimodal missing data and the units contained therein, and will not be repeated here.

[0148] The following is for reference. Figure 3 It shows a schematic diagram of the structure of an electronic device (such as a computing device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0149] like Figure 3As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0150] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.

[0151] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.

[0152] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0153] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0154] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: in response to receiving a multimodal statement sequence set sent by a user terminal, generate a multimodal sentiment data sequence set based on the multimodal statement sequence set; for each multimodal sentiment data in the multimodal sentiment data sequence set, perform the following steps: perform feature extraction processing on the multimodal sentiment data to obtain text modality feature information, audio modality feature information, and video modality feature information corresponding to the multimodal sentiment data; generate multimodal sentiment fusion feature information corresponding to the multimodal sentiment data based on the text modality feature information, the audio modality feature information, and the video modality feature information; and construct a multimodal temporal relationship matrix based on the multimodal sentiment data. The system generates a speaker relationship matrix; based on the aforementioned multimodal emotion fusion feature information, the aforementioned multimodal temporal relationship matrix, and the aforementioned speaker relationship matrix, it generates multimodal graph feature information; based on the aforementioned multimodal graph feature information, the aforementioned multimodal emotion fusion feature information, and the conditional diffusion generation network in the pre-trained emotion recognition model, it generates reconstructed text convolutional features, reconstructed audio convolutional features, and reconstructed video convolutional features; based on the aforementioned reconstructed text convolutional features, the aforementioned reconstructed audio convolutional features, the aforementioned reconstructed video convolutional features, the aforementioned multimodal graph feature information, and the aforementioned feature processing network in the aforementioned emotion recognition model, it generates reconstructed fusion feature information; based on the aforementioned reconstructed fusion feature information and the aforementioned classification network in the aforementioned emotion recognition model, it generates multimodal emotion recognition results; and it sends the generated multimodal emotion recognition results to the aforementioned user terminal.

[0155] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0156] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0157] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including a generation unit, an execution unit, and a transmission unit. The names of these units do not necessarily limit the specific unit; for example, a generation unit may also be described as a "unit that generates a set of multimodal sentiment data sequences."

[0158] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0159] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A sentiment recognition method based on multimodal missing data, comprising: In response to receiving a multimodal statement sequence set sent by a user terminal, a multimodal sentiment data sequence set is generated based on the multimodal statement sequence set; For each multimodal sentiment data point in the multimodal sentiment data sequence set, perform the following steps: Feature extraction processing is performed on the multimodal sentiment data to obtain text modal feature information, audio modal feature information, and video modal feature information corresponding to the multimodal sentiment data; Based on the text modal feature information, the audio modal feature information, and the video modal feature information, multimodal emotion fusion feature information corresponding to the multimodal emotion data is generated; Based on the aforementioned multimodal sentiment data, a multimodal temporal relation matrix and a speaker relation matrix are constructed; Based on the multimodal emotion fusion feature information, the multimodal temporal relationship matrix, and the speaker relationship matrix, multimodal graph feature information is generated; Based on the multimodal graph feature information, the multimodal emotion fusion feature information, and the conditional diffusion generation network in the pre-trained emotion recognition model, reconstructed text convolutional features, reconstructed audio convolutional features, and reconstructed video convolutional features are generated. The emotion recognition model is trained through the following steps: Obtain a sample set, wherein each sample in the sample set includes sample multimodal feature information, sample multimodal graph feature information, sample multimodal emotion fusion feature information, and sample emotion recognition result; Based on the sample set, perform the following training steps: For each sample in at least one sample in the sample set, the multimodal emotion fusion feature information and the multimodal graph feature information of the sample are input into the initial neural network to obtain the sample reconstruction text convolutional feature, sample reconstruction audio convolutional feature, sample reconstruction video convolutional feature, sample reconstruction fusion feature information and emotion recognition result corresponding to the sample. The initial neural network includes a conditional diffusion generation network, a feature processing network and a classification network. Based on the conditional diffusion generative network, the emotion recognition results corresponding to the at least one sample, the reconstructed text convolutional features of each sample, the reconstructed audio convolutional features of each sample, the reconstructed video convolutional features of each sample, the emotion recognition results of each sample included in the at least one sample, the multimodal feature information of each sample, and the multimodal graph feature information of each sample, classification loss data is generated. In response to determining that the classification loss data meets the preset classification loss conditions, the initial neural network is determined as an emotion recognition model; Based on the reconstructed text convolutional features, the reconstructed audio convolutional features, the reconstructed video convolutional features, the multimodal graph feature information, and the feature processing network in the emotion recognition model, reconstructed fusion feature information is generated; Based on the reconstructed fusion feature information and the classification network in the emotion recognition model, multimodal emotion recognition results are generated; The generated multimodal emotion recognition results are sent to the user terminal.

2. The method according to claim 1, wherein, The step of generating multimodal emotion fusion feature information corresponding to the multimodal emotion data based on the text modal feature information, the audio modal feature information, and the video modal feature information includes: Random masking is performed on the text modal feature information, the audio modal feature information, and the video modal feature information to obtain masked text feature information, masked audio feature information, and masked video feature information; The masked text feature information, the masked audio feature information, and the masked video feature information are subjected to feature fusion processing to obtain multimodal emotion fusion feature information.

3. The method according to claim 2, wherein, The random masking process performed on the text modal feature information, the audio modal feature information, and the video modal feature information to obtain masked text feature information, masked audio feature information, and masked video feature information includes: Based on preset text missing probability data, preset audio missing probability data, preset video missing probability data, and preset missing data, generate text mask probability data, audio mask probability data, and video mask probability data. Based on the text mask probability data, the text modal feature information is masked to obtain masked text feature information; Based on the audio mask probability data, the audio modal feature information is masked to obtain masked audio feature information; Based on the video mask probability data, the video modal feature information is masked to obtain masked video feature information.

4. The method according to claim 1, wherein, The generation of multimodal graph feature information based on the multimodal emotion fusion feature information, the multimodal temporal relationship matrix, and the speaker relationship matrix includes: Based on the multimodal temporal relationship matrix and the multimodal emotion fusion feature information, temporal feature information is generated; Based on the speaker relationship matrix and the multimodal emotion fusion feature information, speaker feature information is generated; The temporal feature information and the speaker feature information are subjected to feature concatenation processing to obtain multimodal graph feature information.

5. The method according to claim 1, wherein, The multimodal sentiment data includes text-based sentiment data, audio-based sentiment data, and video-based sentiment data; and the feature extraction process performed on the multimodal sentiment data to obtain corresponding text modal feature information, audio modal feature information, and video modal feature information includes: The text modal sentiment data, which includes the multimodal sentiment data, is segmented into words to obtain a text modal word segmentation sequence; The text modality segmentation sequence is encoded to obtain a text encoded data sequence; The text encoded data sequence is subjected to feature extraction processing to obtain text modal feature information; The audio modal emotional data included in the multimodal emotional data is resampled to obtain resampled audio data; Based on the resampled audio data, generate frame-level audio feature information corresponding to each frame of the resampled audio data; The audio feature information at each frame level is pooled to obtain audio modal feature information; Face alignment processing is performed on the video modal emotion data included in the multimodal emotion data to obtain aligned video data; The aligned video data is subjected to feature extraction processing to obtain video feature information at each frame level; The video feature information at each frame level is pooled to obtain video modal feature information.

6. An emotion recognition device based on multimodal missing data, comprising: The generation unit is configured to generate a multimodal sentiment data sequence set based on a multimodal statement sequence set received from a user terminal in response to such a set. The execution unit is configured to perform the following steps for each multimodal sentiment data in the multimodal sentiment data sequence set: perform feature extraction processing on the multimodal sentiment data to obtain text modal feature information, audio modal feature information, and video modal feature information corresponding to the multimodal sentiment data; and generate multimodal sentiment fusion feature information corresponding to the multimodal sentiment data based on the text modal feature information, the audio modal feature information, and the video modal feature information. Based on the aforementioned multimodal sentiment data, a multimodal temporal relation matrix and a speaker relation matrix are constructed; Based on the multimodal emotion fusion feature information, the multimodal temporal relation matrix, and the speaker relation matrix, multimodal graph feature information is generated. Based on the multimodal graph feature information, the multimodal emotion fusion feature information, and the conditional diffusion generation network in the pre-trained emotion recognition model, reconstructed text convolutional features, reconstructed audio convolutional features, and reconstructed video convolutional features are generated. The emotion recognition model is trained through the following steps: obtaining a sample set, wherein each sample in the sample set includes sample multimodal feature information, sample multimodal graph feature information, sample multimodal emotion fusion feature information, and sample emotion recognition result; based on the sample set, performing the following training steps: for each sample in at least one sample in the sample set, inputting the sample multimodal emotion fusion feature information and the sample multimodal graph feature information included in the sample into an initial neural network to obtain the sample reconstructed text convolutional features, sample reconstructed audio convolutional features, and sample reconstructed video convolutional features corresponding to the sample. The process involves reconstructing and fusing feature information and emotion recognition results from samples. The initial neural network includes a conditional diffusion generation network, a feature processing network, and a classification network. Based on the conditional diffusion generation network, each emotion recognition result corresponding to the at least one sample, reconstructed text convolutional features of each sample, reconstructed audio convolutional features of each sample, reconstructed video convolutional features of each sample, emotion recognition results of each sample included in the at least one sample, multimodal feature information of each sample, and multimodal graph feature information of each sample, classification loss data is generated. In response to determining that the classification loss data meets preset classification loss conditions, the initial neural network is identified as an emotion recognition model. Based on the reconstructed text convolutional features, the reconstructed audio convolutional features, the reconstructed video convolutional features, the multimodal graph feature information, and the feature processing network in the emotion recognition model, reconstructed and fused feature information is generated. Based on the reconstructed and fused feature information and the classification network in the emotion recognition model, multimodal emotion recognition results are generated. The sending unit is configured to send the generated multimodal emotion recognition results to the user terminal.

7. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 5.

8. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 5.