Emotion recognition method and device based on multi-modal missing data, equipment and medium

The emotion recognition method with multimodal missing data solves the problems of low recognition accuracy of single modal data and waste of computing resources. Through multimodal feature extraction and fusion processing, the accuracy of emotion recognition is improved and resource waste is reduced.

CN120654201AActive Publication Date: 2025-09-16HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS +1

Patent Information

Application Number
CN202511129153.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-09-16
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

In the existing technology, emotion recognition using only single modal data can easily lead to low recognition accuracy, and data loss is prone to occur when processing modal data such as text and audio, resulting in waste of computing resources and repeated recognition.

Method used

The emotion recognition method with multimodal missing data is adopted. By extracting features from the multimodal sentence sequence set, a multimodal emotion data sequence is generated, and a multimodal temporal relationship matrix and a speaker relationship matrix are constructed. The conditional diffusion generation network and the feature processing network are used to generate and reconstruct the fused feature information, and finally the multimodal emotion recognition result is generated.

Benefits of technology

The accuracy of emotion recognition is improved, the waste of computing resources is reduced, and by fusing data from different modalities, the decline in recognition accuracy and repeated calls to computing resources caused by missing modal data are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654201A_ABST
    Figure CN120654201A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an emotion recognition method and device based on multi-modal missing data, equipment and a medium. A specific embodiment of the method comprises the following steps: generating a multi-modal emotion data sequence set; executing the following steps: obtaining text modal feature information, audio modal feature information and video modal feature information corresponding to the multi-modal emotion data; multi-modal emotion fusion feature information corresponding to the multi-modal emotion data is generated; constructing a multi-modal time sequence relation matrix and a speaker relation matrix; multi-modal diagram feature information is generated; generating reconstructed text convolution features, reconstructed audio convolution features and reconstructed video convolution features; generating reconstruction fusion feature information; generating a multi-modal emotion recognition result; and sending each generated multi-modal emotion recognition result to the user terminal. According to the embodiment, the accuracy during emotion recognition can be improved, and waste of computing resources during recognition can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of computer technology, and in particular to a method, apparatus, device, and medium for emotion recognition based on multimodal missing data. Background Art

[0002] Emotion recognition is a technology that identifies a user's emotions by processing information such as their speech, voice, and facial expressions. Currently, emotion recognition typically involves processing and recognizing single-modal data, such as text or audio, sent by a user terminal to generate an emotion recognition result. This result is then sent to the user terminal.

[0003] However, when using the above method for emotion recognition, the following technical problems often occur: Emotion recognition based solely on a single modality can easily lead to low recognition accuracy. Furthermore, when processing modal data such as text and audio, data loss is prone to occur, which in turn leads to low accuracy when performing emotion recognition based on the processed data. This requires computing resources to repeatedly recognize multimodal data sent from the same user terminal, resulting in wasted computing resources during recognition. Summary of the Invention

[0004] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0005] Some embodiments of the present disclosure propose emotion recognition methods, devices, equipment and media based on multimodal missing data to solve one or more of the technical problems mentioned in the above background technology section.

[0006] In the first aspect, some embodiments of the present disclosure provide an emotion recognition method based on multimodal missing data, the method comprising: in response to receiving a multimodal sentence sequence set sent by a user terminal, generating a multimodal emotion data sequence set based on the multimodal sentence sequence set; for each multimodal emotion data in the multimodal emotion data sequence set, performing the following steps: performing feature extraction processing on the multimodal emotion data to obtain text modal feature information, audio modal feature information and video modal feature information corresponding to the multimodal emotion data; based on the text modal feature information, the audio modal feature information and the video modal feature information, generating multimodal emotion fusion feature information corresponding to the multimodal emotion data; constructing a multimodal temporal relationship matrix and a speaker matrix based on the multimodal emotion data. relationship matrix; based on the above-mentioned multimodal emotion fusion feature information, the above-mentioned multimodal temporal relationship matrix and the above-mentioned speaker relationship matrix, generate multimodal graph feature information; based on the above-mentioned multimodal graph feature information, the above-mentioned multimodal emotion fusion feature information and the conditional diffusion generation network in the pre-trained emotion recognition model, generate reconstructed text convolution features, reconstructed audio convolution features and reconstructed video convolution features; based on the above-mentioned reconstructed text convolution features, the above-mentioned reconstructed audio convolution features, the above-mentioned reconstructed video convolution features, the above-mentioned multimodal graph feature information and the feature processing network in the above-mentioned emotion recognition model, generate reconstructed fusion feature information; based on the above-mentioned reconstructed fusion feature information and the classification network in the above-mentioned emotion recognition model, generate multimodal emotion recognition results; and send the generated multimodal emotion recognition results to the above-mentioned user terminal.

[0007] In a second aspect, some embodiments of the present disclosure provide an emotion recognition device based on multimodal missing data, the device comprising: a generation unit, configured to generate a multimodal emotion data sequence set based on the multimodal sentence sequence set in response to receiving a multimodal sentence sequence set sent by a user terminal; an execution unit, configured to perform the following steps for each multimodal emotion data in the multimodal emotion data sequence set: performing feature extraction processing on the multimodal emotion data to obtain text modal feature information, audio modal feature information and video modal feature information corresponding to the multimodal emotion data; generating multimodal emotion fusion feature information corresponding to the multimodal emotion data based on the text modal feature information, the audio modal feature information and the video modal feature information; constructing a multimodal temporal relationship matrix based on the multimodal emotion data. and speaker relationship matrix; based on the above-mentioned multimodal emotion fusion feature information, the above-mentioned multimodal temporal relationship matrix and the above-mentioned speaker relationship matrix, generate multimodal graph feature information; based on the above-mentioned multimodal graph feature information, the above-mentioned multimodal emotion fusion feature information and the conditional diffusion generation network in the pre-trained emotion recognition model, generate reconstructed text convolution features, reconstructed audio convolution features and reconstructed video convolution features; based on the above-mentioned reconstructed text convolution features, the above-mentioned reconstructed audio convolution features, the above-mentioned reconstructed video convolution features, the above-mentioned multimodal graph feature information and the feature processing network in the above-mentioned emotion recognition model, generate reconstructed fusion feature information; based on the above-mentioned reconstructed fusion feature information and the classification network in the above-mentioned emotion recognition model, generate multimodal emotion recognition results; the sending unit is configured to send the generated various multimodal emotion recognition results to the above-mentioned user terminal.

[0008] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0009] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation of the first aspect is implemented.

[0010] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: through the emotion recognition method based on multimodal missing data of some embodiments of the present disclosure, the accuracy of emotion recognition can be improved, and the waste of computing resources during recognition can be reduced. Specifically, the reason for the low accuracy of emotion recognition and the waste of computing resources during recognition is that only using a single modal data for emotion recognition can easily lead to low recognition accuracy. And when processing modal data such as text and audio, data missing and the like are likely to occur, which in turn can easily lead to low recognition accuracy when performing emotion recognition based on the processed data, resulting in the need to call computing resources for multiple repeated recognition of multimodal data sent by the same user terminal, resulting in waste of computing resources during recognition. Based on this, the emotion recognition method based on multimodal missing data of some embodiments of the present disclosure, first, in response to receiving a multimodal sentence sequence set sent by a user terminal, generates a multimodal emotion data sequence set based on the above-mentioned multimodal sentence sequence set. In this way, the original data that needs to be processed can be obtained. Secondly, for each multimodal emotion data in the above-mentioned multimodal emotion data sequence set, the following steps are performed: First, feature extraction processing is performed on the above-mentioned multimodal emotion data to obtain text modal feature information, audio modal feature information and video modal feature information corresponding to the above-mentioned multimodal emotion data. In this way, feature vectors corresponding to different modalities in the multimodal emotion data can be obtained. Secondly, based on the above-mentioned text modal feature information, the above-mentioned audio modal feature information and the above-mentioned video modal feature information, multimodal emotion fusion feature information corresponding to the above-mentioned multimodal emotion data is generated. In this way, feature vectors corresponding to different modalities in the multimodal emotion data can be fused. Then, based on the above-mentioned multimodal emotion data, a multimodal temporal relationship matrix and a speaker relationship matrix are constructed. In this way, a multimodal temporal relationship matrix and a speaker relationship matrix can be obtained. Then, based on the above-mentioned multimodal emotion fusion feature information, the above-mentioned multimodal temporal relationship matrix and the above-mentioned speaker relationship matrix, multimodal graph feature information is generated. In this way, multimodal graph feature information can be generated. Then, based on the above-mentioned multimodal graph feature information, the above-mentioned multimodal emotion fusion feature information and the conditional diffusion generation network in the pre-trained emotion recognition model, reconstructed text convolution features, reconstructed audio convolution features and reconstructed video convolution features are generated. In this way, reconstructed text convolution features, reconstructed audio convolution features and reconstructed video convolution features can be generated. Then, based on the above-mentioned reconstructed text convolution features, the above-mentioned reconstructed audio convolution features, the above-mentioned reconstructed video convolution features, the above-mentioned multimodal graph feature information and the feature processing network in the above-mentioned emotion recognition model, reconstructed fusion feature information is generated. In this way, reconstructed fusion feature information can be generated. Then, based on the above-mentioned reconstructed fusion feature information and the classification network in the above-mentioned emotion recognition model, a multimodal emotion recognition result is generated. In this way, a multimodal emotion recognition result can be generated. Finally, the generated multimodal emotion recognition results are sent to the above-mentioned user terminal.Thus, each multimodal emotion recognition result can be sent to the user terminal. Also, because the data of multiple modalities sent by the user terminal can be fused and processed, and then emotion recognition can be performed based on the processed data, the recognition accuracy during emotion recognition can be improved. Also, because the feature vectors corresponding to different modal data can be fused first to obtain multimodal emotion fusion feature information, and then emotion recognition can be performed based on the fused features, when the data of one modality is missing, the data of each modality other than the above modality can still provide a data basis for subsequent emotion recognition, thereby reducing the probability of low recognition accuracy of emotion recognition due to missing modal data, reducing the probability of calling computing resources to repeatedly recognize the same data due to low recognition accuracy, and reducing the waste of computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0012] Figure 1 is a flowchart of some embodiments of the emotion recognition method based on multimodal missing data according to the present disclosure; Figure 2 is a schematic structural diagram of some embodiments of an emotion recognition device based on multimodal missing data according to the present disclosure; Figure 3 It is a structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0013] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0014] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.

[0015] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0016] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0017] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0018] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0019] Figure 1 The process 100 of some embodiments of the emotion recognition method based on multimodal missing data according to the present disclosure is shown. The emotion recognition method based on multimodal missing data includes the following steps: Step 101: In response to receiving a multimodal sentence sequence set sent by a user terminal, a multimodal emotion data sequence set is generated based on the multimodal sentence sequence set.

[0020] In some embodiments, an execution entity (e.g., a computing device) of a method for emotion recognition based on multimodal missing data may generate a multimodal emotion data sequence set based on a multimodal sentence sequence set in response to receiving a multimodal sentence sequence set sent by a user terminal. The user terminal may be a terminal device corresponding to a user. Each multimodal sentence in the multimodal sentence sequence set may be conversation data sent by the user, including text data, audio data, and video data. Each multimodal sentence in the multimodal sentence sequence set may include text data, audio data, and video data. The text data included in each multimodal sentence in the multimodal sentence sequence set may be text content representing the content of the conversation between each speaker. The text data in each multimodal sentence in the multimodal sentence sequence set may include a sentence sequence. Each sentence in the sentence sequence may be text representing the content of a speaker's speech. Each sentence in the sentence sequence corresponds to a speaker. The audio data included in each multimodal sentence in the multimodal sentence sequence set may be audio corresponding to the conversation data between each speaker. The video data included in each multimodal sentence in the above-mentioned multimodal sentence sequence set may be the video corresponding to the conversation data between the speakers. The above-mentioned execution subject may be a server. The above-mentioned multimodal emotion data sequence set may be the data determined by the above-mentioned multimodal sentence sequence set. Each multimodal emotion data in the above-mentioned multimodal emotion data sequence set may include text modal emotion data, audio modal emotion data and video modal emotion data. The above-mentioned text modal emotion data may be the text data in the multimodal sentence. The above-mentioned text modal emotion data may include a sentence sequence. The sentence sequence included in the above-mentioned text modal emotion data may be the sentence sequence included in the text data in the above-mentioned multimodal sentence. The above-mentioned audio modal emotion data may be the audio data in the multimodal sentence. The above-mentioned video modal emotion data may be the video data in the multimodal sentence.

[0021] In practice, for each multimodal sentence in the above-mentioned multimodal sentence sequence set, the text data included in the above-mentioned multimodal sentence can be determined as text modal emotion data, and the sentence sequence in the text data included in the above-mentioned multimodal sentence can be determined as a sentence sequence in text modal emotion data. Then, the audio data included in the above-mentioned multimodal sentence can be determined as audio modal emotion data. Then, the video data included in the above-mentioned multimodal sentence can be determined as video modal emotion data. Finally, the above-mentioned text modal emotion data, the above-mentioned audio modal emotion data and the above-mentioned video modal emotion data can be combined into multimodal emotion data.

[0022] Step 102: for each multimodal emotion data in the multimodal emotion data sequence set, perform the following steps: Step 1021 , performing feature extraction processing on the multimodal emotion data to obtain text modal feature information, audio modal feature information, and video modal feature information corresponding to the multimodal emotion data.

[0023] In some embodiments, the execution entity may perform feature extraction processing on the multimodal emotion data to obtain text modal feature information, audio modal feature information, and video modal feature information corresponding to the multimodal emotion data. The text modal feature information may be a feature vector corresponding to the text data included in the multimodal emotion data. The audio modal feature information may be a feature vector corresponding to the audio data included in the multimodal emotion data. The video modal feature information may be a feature vector corresponding to the video data included in the multimodal emotion data.

[0024] In some optional implementations of some embodiments, the execution entity may perform feature extraction processing on the multimodal emotion data through the following steps to obtain text modal feature information, audio modal feature information, and video modal feature information corresponding to the multimodal emotion data: The first step is to perform word segmentation processing on the text modal emotional data included in the above multimodal emotional data to obtain a text modal word segmentation sequence. Wherein, each text modal word segmentation in the above text modal word segmentation sequence can be a word in the above text modal emotional data. In practice, the above execution subject can perform word segmentation processing on the text modal emotional data included in the above multimodal emotional data through a word segmentation tool to obtain a text modal word segmentation sequence. Wherein, the above word segmentation tool can be a tool that can segment text data. For example, the above word segmentation tool can be jieba.

[0025] The second step is to encode the above-mentioned text modal word segmentation sequence to obtain a text encoding data sequence. Among them, each text encoding data in the above-mentioned text encoding data sequence can be the data obtained after encoding the text modal word segmentation. In practice, the above-mentioned execution entity can input the above-mentioned text modal word segmentation sequence into a pre-trained word vector encoding model to obtain a text encoding data sequence. Among them, the above-mentioned word vector encoding model can be a neural network model that takes the text modal word segmentation sequence as input and the text encoding data sequence as output. For example, the above-mentioned word vector encoding model can be a preprocessed Word2Vec model (Word to Vector). The above-mentioned preprocessing can be a process of fine-tuning the Word2Vec model using a text modal word segmentation sequence set and a cross-entropy loss function.

[0026] The third step is to perform feature extraction processing on the above-mentioned text encoding data sequence to obtain text modal feature information. The above-mentioned text modal feature information can be the feature vector corresponding to the above-mentioned text encoding data sequence. In practice, the above-mentioned execution entity can input the above-mentioned text encoding data sequence into a pre-trained text feature extraction model to obtain text modal feature information. The above-mentioned text feature extraction model can be a neural network model that takes the text encoding data sequence as input and takes the text modal feature information as output. For example, the above-mentioned text feature extraction model can be a pre-trained TextCNN (Text Convolutional Neural Network, convolutional neural network for text). The above-mentioned pre-training can be a process of fine-tuning TextCNN using a set of text encoding data sequences and a cross-entropy loss function.

[0027] The fourth step is to resample the audio modal emotion data included in the multimodal emotion data to obtain resampled audio data. The resampled audio data may be the audio modal emotion data after resampling. In practice, the execution subject may resample the audio modal emotion data through a resampling algorithm to convert the sampling rate corresponding to the audio modal emotion data to a preset sampling rate, and obtain the resampled audio modal emotion data as the resampled audio data. The preset sampling rate may be a pre-set sampling rate. For example, the preset sampling rate may be 16kHz.

[0028] In the fifth step, based on the resampled audio data, frame-level audio feature information corresponding to the resampled audio data is generated. Each frame-level audio feature information in the frame-level audio feature information may be a feature vector corresponding to an audio segment of a preset duration in the resampled audio data. The preset duration may be a pre-set duration. For example, the preset duration may be 20ms. The audio segments corresponding to the frame-level audio feature information do not overlap. In practice, the resampled audio data may be input into a frame-level feature extraction model to obtain the frame-level audio feature information. The frame-level feature extraction model may be a neural network model that takes the resampled audio data as input and outputs the frame-level audio feature information corresponding to the resampled audio data. For example, the frame-level feature extraction model may be a preprocessed Wav2Vec model (Waveform-to-Vector). The preprocessing may be a process of fine-tuning the Wav2Vec model using a resampled audio dataset and a cross-entropy loss function.

[0029] In step 6, the above-mentioned frame-level audio feature information is pooled to obtain audio modal feature information. In practice, the above-mentioned execution subject can perform average pooling processing on the above-mentioned frame-level audio feature information by an average pooling method to obtain audio modal feature information.

[0030] In the seventh step, face alignment processing is performed on the video modality emotion data included in the multimodal emotion data to obtain aligned video data. The aligned video data may be video data obtained by processing each face region in the video modality emotion data. Each of the face regions may be an image region where a face is located in each video frame included in the video modality emotion data.

[0031] In practice, first, the above-mentioned video modality emotion data can be input into the face region data generation model to obtain each face region data. The above-mentioned face region data generation model can be a neural network model that takes the video modality emotion data as input and takes each face region data corresponding to the video modality emotion data as output. For example, the above-mentioned face region data generation model can be a pre-trained multi-task cascade convolutional neural network. The above-mentioned pre-training can be a process of fine-tuning the multi-task cascade convolutional neural network using the labeled video modality emotion data set and the cross-entropy loss function. The above-mentioned labeling can be a process of labeling each face key point of the face region in each video frame of the video modality emotion data. Each face key point in the above-mentioned face key points can be a key point in the face. For example, the above-mentioned face key point can be the center of the pupil, the tip of the nose or the corner of the mouth. Each face region data in the above-mentioned face region data can be data corresponding to the face region in the video modality emotion data. Each face region data in the above-mentioned face region data can include face bounding box data and each face key point data. Each face region data in the above-mentioned face region data corresponds to a video frame in the above-mentioned video modality emotion data. The above-mentioned face bounding box data can be an array of bounding boxes corresponding to the face region. The above-mentioned face bounding box data can include the horizontal coordinate of the upper left corner of the bounding box, the vertical coordinate of the upper left corner of the bounding box, the horizontal coordinate of the lower right corner of the bounding box, and the vertical coordinate of the lower right corner of the bounding box. The coordinate values ​​in the above-mentioned face bounding box data can be coordinate values ​​in the image coordinate system corresponding to the face bounding box data. The above-mentioned image coordinate system can be a two-dimensional coordinate system constructed with the first pixel point in the upper left corner of the video frame of the video modality emotion data as the coordinate origin, the horizontal right direction of the video frame as the horizontal coordinate direction, and the vertical downward direction of the video frame as the vertical coordinate direction. Each face key point data in the above-mentioned face key point data can be the two-dimensional coordinate of the key point in the face region in the corresponding image coordinate system. Each face key point data in the above-mentioned face key point data corresponds to a key point category. The above-mentioned key point category can be a label for characterizing the category of the key point. For example, the key point category may be the center of the left pupil, the center of the right pupil, the tip of the nose, the left corner of the mouth, or the right corner of the mouth.

[0032] Then, for each face region data in the above-mentioned face region data, the various face key point data included in the above-mentioned face region data can be combined into a matrix as a face key point matrix according to a preset combination order. The above-mentioned combination order can be the order in which the various face key point data in the face region data are combined. For example, the above-mentioned combination order can be {left eye pupil center, right eye pupil center, nose tip, left mouth corner, right mouth corner}. As an example, when the face key point data corresponding to the left eye pupil center, the right eye pupil center and the nose tip are (1,2), (2,3), and (3,4) respectively, and the combination order is {left eye pupil center, right eye pupil center, nose tip}, they can be combined into a three-row and two-column matrix [1,2; 2,3; 3,4] as the face key point matrix.

[0033] Then, the facial key point matrix and the preset standard facial key point matrix can be input into a similarity transformation matrix generation function to obtain a key point transformation matrix. The key point transformation matrix can be a matrix capable of similarity transforming the facial key point matrix into the standard facial key point matrix. The similarity transformation matrix generation function can be a function capable of generating a similarity transformation matrix between two matrices. For example, the similarity transformation matrix generation function can be the estimateAffinePartial2D function in OpenCV. The standard facial key point matrix can be a facial key point matrix preset by a technician.

[0034] The video frame corresponding to the facial region data can then be determined as the video frame to be processed. The facial bounding box data corresponding to the video frame to be processed can then be input into an image cropping function to crop the video frame to be processed, obtaining the image region corresponding to the facial bounding box data in the video frame to be processed as the facial image data. The image cropping function can be any function capable of cropping image data. For example, the image cropping function can be the crop function of the PIL library.

[0035] Then, the facial image data and the key point transformation matrix can be input into a conversion function to obtain the converted facial image data as the target facial image data corresponding to the facial region data. The conversion function can be a function that can perform a linear mapping on the input image. For example, the conversion function can be the cv2.warpAffine() function.

[0036] Finally, a video generation tool can be used to combine the obtained target facial image data into video data as aligned video data, according to the order of the video frames corresponding to the facial region data in the video modality emotion data. The video generation tool can be a tool capable of combining images into a video. For example, the video generation tool can be ffmpeg.

[0037] In the eighth step, feature extraction processing is performed on the aligned video data to obtain frame-level video feature information. Each frame-level video feature information in the aligned video data may be a feature vector corresponding to each video frame in the aligned video data. In practice, for each video frame in the aligned video data, the execution entity may perform feature extraction processing on the video frame using a feature extraction algorithm to obtain a feature vector corresponding to the video frame as frame-level video feature information. The feature extraction algorithm may be an algorithm capable of extracting features from video frames. For example, the feature extraction algorithm may be a scale-invariant feature transformation algorithm.

[0038] In the ninth step, the above-mentioned frame-level video feature information is pooled to obtain the video modality feature information. In practice, the above-mentioned execution subject can perform average pooling processing on the above-mentioned frame-level video feature information by an average pooling method to obtain the video modality feature information.

[0039] Step 1022: Generate multimodal emotion fusion feature information corresponding to the multimodal emotion data based on the text modal feature information, the audio modal feature information, and the video modal feature information.

[0040] In some embodiments, the execution entity may generate multimodal emotion fusion feature information corresponding to the multimodal emotion data based on the text modal feature information, the audio modal feature information, and the video modal feature information. The multimodal emotion fusion feature information may be a feature vector obtained by processing the text modal feature information, the audio modal feature information, and the video modal feature information.

[0041] In some optional implementations of some embodiments, the execution entity may generate multimodal emotion fusion feature information corresponding to the multimodal emotion data based on the text modal feature information, the audio modal feature information, and the video modal feature information through the following steps: In the first step, random masking is performed on the text modal feature information, the audio modal feature information, and the video modal feature information to obtain masked text feature information, masked audio feature information, and masked video feature information. The masked text feature information may be masked text modal feature information. The masked audio feature information may be masked audio modal feature information. The masked video feature information may be masked video modal feature information.

[0042] In the second step, feature fusion processing is performed on the masked text feature information, the masked audio feature information, and the masked video feature information to obtain multimodal emotion fusion feature information. The multimodal emotion fusion feature information may be a feature vector obtained by fusing the masked text feature information, the masked audio feature information, and the masked video feature information. In practice, the masked text feature information, the masked audio feature information, and the masked video feature information may be input into a feature fusion model to obtain multimodal emotion fusion feature information. The feature fusion model may be a neural network model that takes the masked text feature information, the masked audio feature information, and the masked video feature information as input and outputs the multimodal emotion fusion feature information. For example, the feature fusion model may be a pre-trained Bidirectional Gated Recurrent Unit (BGRU) model. The pre-training may be a process of fine-tuning the Bi-GRU model using a training dataset and a cross-entropy loss function. Each training data point in the training dataset may be data used to train the Bi-GRU model. Each training data in the above training data set may include masked text feature information, masked audio feature information, and masked video feature information.

[0043] In some optional implementations of some embodiments, the execution entity may perform random masking on the text modality feature information, the audio modality feature information, and the video modality feature information through the following steps to obtain masked text feature information, masked audio feature information, and masked video feature information: The first step is to generate text mask probability data, audio mask probability data, and video mask probability data based on preset text missing probability data, preset audio missing probability data, preset video missing probability data, and preset missing data. The text missing probability data may be the data missing rate corresponding to the text modal emotion data. The audio missing probability data may be the data missing rate corresponding to the audio modal emotion data. The video missing probability data may be the data missing rate corresponding to the video modal emotion data. The data missing rate may be the ratio of missing values ​​in the data to the total data. The text mask probability data may be the mask ratio corresponding to random masking of the text modal feature information. The audio mask probability data may be the mask ratio corresponding to random masking of the audio modal feature information. The video mask probability data may be the mask ratio corresponding to random masking of the video modal feature information. The missing data may be a preset value. For example, the missing data may be 1. In practice, the execution entity may determine the difference between the missing data and the text missing probability data as the text mask probability data. The difference between the missing data and the audio missing probability data may be determined as audio mask probability data. The difference between the missing data and the video missing probability data may be determined as video mask probability data.

[0044] The second step is to mask the text modal feature information based on the text mask probability data to obtain masked text feature information. In practice, first, the execution subject can generate a vector with the same size as the text modal feature information and an element value between the preset interval as a text random vector through a vector generation function. The vector generation function can be a function that can generate a random vector. For example, the vector generation function can be the np.random.rand function. The preset interval can be [0,1). Then, for each element included in the text random vector, in response to determining that the element value corresponding to the element is less than or equal to the text mask probability data, 0 can be determined as the element value corresponding to the element to update the text random vector. In response to determining that the element value corresponding to the element is greater than the text mask probability data, 1 can be determined as the element value corresponding to the element to update the text random vector. The updated text random vector can be element-by-element multiplied by the text modal feature information to obtain masked text feature information.

[0045] The third step is to mask the audio modal feature information based on the audio mask probability data to obtain masked audio feature information. In practice, first, the execution subject can generate a vector with the same size as the audio modal feature information and element values ​​between the preset interval as an audio random vector through the vector generation function. Then, for each element of the various elements included in the audio random vector, in response to determining that the element value corresponding to the element is less than or equal to the audio mask probability data, 0 can be determined as the element value corresponding to the element to update the audio random vector. In response to determining that the element value corresponding to the element is greater than the audio mask probability data, 1 can be determined as the element value corresponding to the element to update the audio random vector. The updated audio random vector can be multiplied element by element with the audio modal feature information to obtain masked audio feature information.

[0046] The fourth step is to perform mask processing on the video modal feature information based on the above-mentioned video mask probability data to obtain masked video feature information. In practice, first, the above-mentioned execution subject can generate a vector with the same size as the above-mentioned video modal feature information and an element value between the above-mentioned preset interval as a video random vector through the above-mentioned vector generation function. Then, for each element of the various elements included in the above-mentioned video random vector, in response to determining that the element value corresponding to the above-mentioned element is less than or equal to the above-mentioned video mask probability data, 0 can be determined as the element value corresponding to the above-mentioned element to update the video random vector. In response to determining that the element value corresponding to the above-mentioned element is greater than the above-mentioned video mask probability data, 1 can be determined as the element value corresponding to the above-mentioned element to update the video random vector. The updated video random vector can be multiplied element by element with the above-mentioned video modal feature information to obtain masked video feature information.

[0047] Step 1023: construct a multimodal temporal relationship matrix and a speaker relationship matrix based on the multimodal emotion data.

[0048] In some embodiments, the execution entity may construct a multimodal temporal relationship matrix and a speaker relationship matrix based on the multimodal emotion data. The multimodal temporal relationship matrix may be a matrix representing the order of each sentence in the sentence sequence of the text modal emotion data in the multimodal emotion data. The speaker relationship matrix may be a matrix representing whether the two speakers corresponding to each two sentences in the sentence sequence included in the multimodal emotion data are the same.

[0049] In practice, for the sentence sequence of text modal sentiment data in the above-mentioned multimodal sentiment data, for each sentence in the above-mentioned sentence sequence, the position of the above-mentioned sentence in the sentence sequence can be determined as the number corresponding to the above-mentioned sentence. For example, when a sentence is the first sentence in the above-mentioned sentence sequence, the number corresponding to the above-mentioned sentence can be 1. Then, the number of each sentence in the above-mentioned sentence sequence can be determined as the number of target sentences. Then, an empty matrix with the same number of rows as the number of the above-mentioned target sentences and the same number of columns as the above-mentioned target sentences can be created through a matrix creation function as a time series relationship matrix to be processed. Among them, the above-mentioned matrix creation function can be a function that can create a matrix. For example, the above-mentioned matrix creation function can be the np.empty() function. Then, each element on the main diagonal of the above-mentioned time series relationship matrix to be processed can be determined as 0. Then, for every two sentences in the above-mentioned sentence sequence, first, the two numbered numbers corresponding to the above-mentioned two sentences can be combined into two matrix position data. Among them, each matrix position data in the above-mentioned two matrix position data can be data used to characterize the position in the matrix. For example, when the two numbered numbers are 2 and 3 respectively, the two matrix position data obtained by the combination can be (3, 2) and (2, 3). (3, 2) can represent the position corresponding to the third row and second column in the matrix, and (2, 3) can represent the position corresponding to the second row and third column in the matrix. Secondly, the difference between the two numbered numbers corresponding to the two statements can be determined as the numbered difference data. Then, the absolute value of the numbered difference data can be determined as the numbered absolute value data. Then, in response to determining that the numbered absolute value data meets the preset numbered data condition, 1 can be determined as the element corresponding to the two matrix position data. The numbered data condition can be that the numbered absolute value data is equal to 1. Then, in response to determining that the numbered absolute value data does not meet the numbered data condition, 0 can be determined as the element corresponding to the two matrix position data. This process can be deduced in this way until the time series relationship matrix to be processed is completely filled, and the filled time series relationship matrix to be processed can be determined as a multimodal time series relationship matrix.

[0050] Then, an empty matrix with the same number of rows and columns as the target sentences can be created using the matrix creation function as the speaker matrix to be processed. Each element on the main diagonal of the speaker matrix to be processed can be set to 1. For each two sentences in the sentence sequence, first, the two numbers corresponding to the two sentences can be combined into two matrix position data. Secondly, the two speakers corresponding to the two sentences can be determined as two target speakers. Then, in response to determining that the two target speakers meet a preset speaker condition, the values ​​of the two elements in the two matrix position data can be set to 1. The speaker condition can be that the two target speakers are the same. Then, in response to determining that the two target speakers do not meet the speaker condition, the values ​​of the two elements in the two matrix position data can be set to 0. This process continues in this manner until the speaker matrix to be processed is completely filled, and the filled speaker matrix to be processed can be determined as the speaker relationship matrix.

[0051] Step 1024: Generate multimodal graph feature information based on the multimodal emotion fusion feature information, the multimodal temporal relationship matrix, and the speaker relationship matrix.

[0052] In some embodiments, the execution entity may generate multimodal graph feature information based on the multimodal emotion fusion feature information, the multimodal temporal relationship matrix, and the speaker relationship matrix.

[0053] In some optional implementations of some embodiments, the execution entity may generate multimodal graph feature information based on the multimodal emotion fusion feature information, the multimodal temporal relationship matrix, and the speaker relationship matrix through the following steps: The first step is to generate temporal feature information based on the multimodal temporal relationship matrix and the multimodal emotion fusion feature information. The temporal feature information may be a feature vector that is integrated with the multimodal emotion fusion feature information and is used to characterize the characteristics of each element in the multimodal temporal relationship matrix.

[0054] In practice, the multimodal temporal relationship matrix and the multimodal emotion fusion feature information can be input into a pre-trained temporal graph fusion network to obtain temporal feature information. The temporal graph fusion network can be a neural network that takes the multimodal temporal relationship matrix and the multimodal emotion fusion feature information as input and the temporal feature information as output. For example, the temporal graph fusion network can be a pre-trained graph convolutional neural network. The pre-training can be a process of fine-tuning the graph convolutional neural network using a temporal network training data set and a cross-entropy loss function. Each temporal network training data in the temporal network training data set can be data for training a graph convolutional neural network. Each temporal network training data in the temporal network training data set can include a multimodal temporal relationship matrix and multimodal emotion fusion feature information.

[0055] The second step is to generate speaker feature information based on the speaker relationship matrix and the multimodal emotion fusion feature information. The speaker feature information may be a feature vector that fuses the multimodal emotion fusion feature information and is used to characterize the characteristics of each element in the speaker relationship matrix. In practice, the execution entity may input the speaker relationship matrix and the multimodal emotion fusion feature information into a pre-trained speaker graph fusion network to obtain the speaker feature information. The speaker graph fusion network may be a neural network that takes the speaker relationship matrix and the multimodal emotion fusion feature information as input and outputs the speaker feature information. For example, the speaker graph fusion network may be a pre-trained graph convolutional neural network. The pre-training may be a process of fine-tuning the graph convolutional neural network using a speaker training dataset and a cross-entropy loss function. Each speaker training data in the speaker training dataset may be data used to train the graph convolutional neural network into a speaker graph fusion network. Each speaker training data in the speaker training dataset may include a speaker relationship matrix and multimodal emotion fusion feature information.

[0056] The third step is to perform feature concatenation processing on the above-mentioned time series feature information and the above-mentioned speaker feature information to obtain multimodal graph feature information. The above-mentioned multimodal graph feature information can be a feature vector obtained by concatenating the above-mentioned time series feature information and the above-mentioned speaker feature information. In practice, the above-mentioned execution entity can input the above-mentioned time series feature information and the above-mentioned speaker feature information into a feature concatenation function to obtain multimodal graph feature information. The above-mentioned feature concatenation function can be a function that can concatenate features. For example, the above-mentioned feature concatenation function can be a concat() function.

[0057] Step 1025: Generate reconstructed text convolution features, reconstructed audio convolution features, and reconstructed video convolution features based on the multimodal graph feature information, the multimodal emotion fusion feature information, and the conditional diffusion generation network in the pre-trained emotion recognition model.

[0058] In some embodiments, the above-mentioned execution entity can generate reconstructed text convolution features, reconstructed audio convolution features and reconstructed video convolution features based on the above-mentioned multimodal graph feature information, the above-mentioned multimodal emotion fusion feature information and the conditional diffusion generation network in the pre-trained emotion recognition model.

[0059] The emotion recognition model may be a neural network model that takes multimodal graph feature information and multimodal emotion fusion feature information as input and outputs emotion recognition results. The emotion recognition results may be labels representing the emotion categories corresponding to the multimodal graph feature information and multimodal emotion fusion feature information. For example, the emotion recognition results may be "happy," "nervous," or "sad."

[0060] The emotion recognition model may include a conditional diffusion generative network, a feature processing network, and a classification network. The conditional diffusion generative network may include a text modality diffusion network, an audio modality diffusion network, a video modality diffusion network, a text convolution layer, an audio convolution layer, and a video convolution layer. The conditional diffusion generative network corresponds to various diffusion time steps. Each of these diffusion time steps may correspond to a time step when the conditional diffusion generative network processes data. Each of these diffusion time steps may correspond to text Gaussian vector data, audio Gaussian vector data, video Gaussian vector data, sample diffusion text feature information, sample diffusion audio feature information, and sample diffusion video feature information. The text Gaussian vector data may be the Gaussian noise generated by the text modality diffusion network in the conditional diffusion generative network at the corresponding diffusion time step. The audio Gaussian vector data may be the Gaussian noise generated by the audio modality diffusion network in the conditional diffusion generative network at the corresponding diffusion time step. The video Gaussian vector data may be the Gaussian noise generated by the video modality diffusion network in the conditional diffusion generative network at the corresponding diffusion time step. The sample diffusion text feature information may be a feature vector obtained by adding text Gaussian vector data to a feature vector input into a text modal diffusion network. The sample diffusion audio feature information may be a feature vector obtained by adding audio Gaussian vector data to a feature vector input into an audio modal diffusion network. The sample diffusion video feature information may be a feature vector obtained by adding video Gaussian vector data to a feature vector input into a video modal diffusion network.

[0061] In some optional implementations of some embodiments, the emotion recognition model may be trained by the execution subject through the following steps: The first step is to obtain a sample set. Each sample in the sample set may be sample data used to train an emotion recognition model. Each sample in the sample set may include sample multimodal feature information, sample multimodal graph feature information, sample multimodal emotion fusion feature information, and a sample emotion recognition result. The sample multimodal feature information may be multimodal feature information used to train the emotion recognition model. The sample multimodal graph feature information may be multimodal graph feature information used to train the emotion recognition model. The sample multimodal emotion fusion feature information may be multimodal emotion fusion feature information used to train the emotion recognition model. The sample multimodal emotion fusion feature information may correspond to data in three modalities: text, audio, and video. The sample multimodal feature information may include sample text feature information, sample audio feature information, and sample video feature information. The sample text feature information may be text feature information included in the sample multimodal feature information. The sample audio feature information may be audio feature information included in the sample multimodal feature information. The sample video feature information may be video feature information included in the sample multimodal feature information. The above sample emotion recognition results may be real emotion recognition results corresponding to the samples.

[0062] In the second step, based on the sample set, the following training steps are performed: In the first sub-step, for each of at least one sample in the sample set, the sample multimodal emotion fusion feature information and the sample multimodal graph feature information included in the sample are input into the text modality diffusion network of the conditional diffusion generative network in the initial neural network to obtain sample reconstructed text feature information corresponding to the sample. The sample reconstructed text feature information may be a feature vector obtained by reconstructing the sample multimodal emotion fusion feature information and used to characterize the data of the text modality corresponding to the sample multimodal emotion fusion feature information. The text modality diffusion network may be a neural network that takes the sample multimodal emotion fusion feature information and the sample multimodal graph feature information as input and outputs the sample reconstructed text feature information. For example, the text modality diffusion network may be a pre-trained ScoreNet model. The pre-training may be a process of fine-tuning the ScoreNet model using a text training dataset and a mean squared error loss function. Each text training data in the text training dataset may be data used to train the ScoreNet model to generate sample reconstructed text feature information. Each text training data in the text training dataset may include sample multimodal emotion fusion feature information and sample multimodal graph feature information.

[0063] In a second sub-step, for each of the at least one sample, the sample-reconstructed text feature information corresponding to the sample is input into a text convolution layer of a conditional diffusion generative network in the initial neural network to obtain a sample-reconstructed text convolution feature corresponding to the sample. The sample-reconstructed text convolution feature may be sample-reconstructed text feature information that has undergone convolution processing. The text convolution layer may be a convolution layer that takes the sample-reconstructed text feature information as input and outputs the sample-reconstructed text convolution feature.

[0064] In the third sub-step, for each of the at least one sample mentioned above, the sample multimodal emotion fusion feature information and the sample multimodal graph feature information included in the sample are input into the audio modal diffusion network of the conditional diffusion generation network in the initial neural network to obtain the sample reconstructed audio feature information corresponding to the sample.

[0065] Among them, the above-mentioned sample reconstructed audio feature information can be a feature vector of data obtained after reconstructing the above-mentioned sample multimodal emotion fusion feature information, which is used to characterize the audio modality corresponding to the above-mentioned sample multimodal emotion fusion feature information. The above-mentioned audio modal diffusion network can be a neural network that takes sample multimodal emotion fusion feature information and sample multimodal graph feature information as input and takes sample reconstructed audio feature information as output. For example, the above-mentioned audio modal diffusion network can be a preprocessed ScoreNet model. The above-mentioned preprocessing can be a process of fine-tuning the ScoreNet model using an audio training data set and a mean square error loss function. Each audio training data in the above-mentioned audio training data set can be data used to train the ScoreNet model to generate sample reconstructed audio feature information. Each audio training data in the above-mentioned audio training data set can include sample multimodal emotion fusion feature information and sample multimodal graph feature information.

[0066] In a fourth sub-step, for each of the at least one sample, the sample-reconstructed audio feature information corresponding to the sample is input into the audio convolution layer of the conditional diffusion generative network in the initial neural network to obtain a sample-reconstructed audio convolution feature corresponding to the sample. The sample-reconstructed audio convolution feature may be sample-reconstructed audio feature information that has undergone convolution processing. The audio convolution layer may be a convolution layer that takes the sample-reconstructed audio feature information as input and outputs the sample-reconstructed audio convolution feature.

[0067] The fifth sub-step is to input the sample multimodal emotion fusion feature information and the sample multimodal graph feature information included in the above sample into the video modal diffusion network of the conditional diffusion generation network in the initial neural network for each sample of the above at least one sample, and obtain the sample reconstructed video feature information corresponding to the above sample.

[0068] Among them, the above-mentioned sample reconstructed video feature information can be a feature vector of data obtained after reconstructing the above-mentioned sample multimodal emotion fusion feature information, which is used to characterize the video modality corresponding to the above-mentioned sample multimodal emotion fusion feature information. The above-mentioned video modality diffusion network can be a neural network model that takes sample multimodal emotion fusion feature information and sample multimodal graph feature information as input and takes sample reconstructed video feature information as output. For example, the above-mentioned video modality diffusion network can be a preprocessed ScoreNet model. The above-mentioned preprocessing can be a process of fine-tuning the ScoreNet model using a video training data set and a mean square error loss function. Each video training data in the above-mentioned video training data set can be data used to train the ScoreNet model to generate sample reconstructed video feature information. Each video training data in the above-mentioned video training data set can include sample multimodal emotion fusion feature information and sample multimodal graph feature information.

[0069] In a sixth sub-step, for each of the at least one sample, the sample-reconstructed video feature information corresponding to the sample is input into the video convolution layer of the conditional diffusion generative network in the initial neural network to obtain a sample-reconstructed video convolution feature corresponding to the sample. The sample-reconstructed video convolution feature may be sample-reconstructed video feature information that has undergone convolution processing. The video convolution layer may be a convolution layer that takes the sample-reconstructed video feature information as input and outputs the sample-reconstructed video convolution feature.

[0070] The seventh sub-step is to input the sample multimodal graph feature information included in the above-mentioned at least one sample, the sample reconstructed text convolution feature, the sample reconstructed audio convolution feature and the sample reconstructed video convolution feature corresponding to the above-mentioned sample into the feature processing network of the initial neural network to obtain the sample reconstruction fusion feature information.

[0071] The sample reconstruction fusion feature information may be a feature vector obtained by fusing the sample multimodal graph feature information, the sample reconstructed text convolutional features corresponding to the sample, the sample reconstructed audio convolutional features, and the sample reconstructed video convolutional features. The feature processing network may be a neural network model that takes the sample multimodal graph feature information, the sample reconstructed text convolutional features, the sample reconstructed audio convolutional features, and the sample reconstructed video convolutional features as input and outputs the sample reconstruction fusion feature information. For example, the feature processing network may be a multilayer perceptron (MLP).

[0072] In the eighth sub-step, the reconstructed fusion feature information of each sample corresponding to the at least one sample is input into the classification network in the initial neural network to obtain various emotion recognition results.

[0073] In which, the above-mentioned classification network may include a normalization function and an output function. The above-mentioned normalization function may be a normalized exponential function that takes the sample reconstruction fusion feature information as input and takes the emotion label probability distribution data corresponding to the sample reconstruction fusion feature information as output. The above-mentioned emotion label probability distribution data may be the probability distribution data of the emotion category corresponding to the sample reconstruction fusion feature information. The above-mentioned emotion label probability distribution data may include each emotion label and each label probability data. The above-mentioned each emotion label and the above-mentioned each label probability data correspond one to one. Each of the above-mentioned emotion labels may be a label for characterizing the emotion category corresponding to the above-mentioned sample reconstruction fusion feature information. For example, the above-mentioned emotion label may be "happy", "sad", or "nervous". Each of the above-mentioned label probability data may be the probability corresponding to the emotion label.

[0074] The output function may be an argmax function that takes the emotion label probability distribution data as input and outputs the emotion recognition result. The emotion recognition result is the emotion label with the largest corresponding label probability data in the emotion label probability distribution data.

[0075] A ninth sub-step is generating classification loss data based on the conditional diffusion generative network, each emotion recognition result corresponding to the at least one sample, reconstructed text convolution features of each sample, reconstructed audio convolution features of each sample, reconstructed video convolution features of each sample, each emotion recognition result included in the at least one sample, multimodal feature information of each sample, and multimodal graph feature information of each sample. The classification loss data may be a loss value corresponding to the at least one sample.

[0076] In the third step, in response to determining that the classification loss data satisfies a preset classification loss condition, the initial neural network is determined to be a sentiment recognition model. The classification loss condition may be that the classification loss data is less than a preset loss threshold. The loss threshold may be a pre-set value. The specific setting of the loss threshold is not limited herein.

[0077] In step 4, in response to determining that the classification loss data does not meet the classification loss condition, the network parameters of the initial neural network are adjusted, and a sample set is formed using unused samples, and the training step is performed again using the adjusted initial neural network. In practice, the execution entity may use a back propagation algorithm (BP algorithm) and a gradient descent method (e.g., a mini-batch gradient descent algorithm) to adjust the network parameters of the initial neural network.

[0078] In the process of adopting technical solutions to solve the above technical problems, the following problems often arise: When emotion recognition is performed only through single-modal data, it is easy to lead to low accuracy of emotion recognition, which in turn leads to a high probability of needing to call computing resources for repeated recognition of the same data, which easily causes a waste of computing resources during recognition.

[0079] Faced with the above technical problems, we decided to adopt the following solutions: In some optional implementations of some embodiments, the execution entity may generate classification loss data based on the conditional diffusion generation network, the emotion recognition results corresponding to the at least one sample, the reconstructed text convolution features of each sample, the reconstructed audio convolution features of each sample, the reconstructed video convolution features of each sample, the emotion recognition results of each sample included in the at least one sample, the multimodal feature information of each sample, and the multimodal graph feature information of each sample through the following steps: In the first step, for each of the at least one sample, perform the following steps: In the first sub-step, the emotion recognition result corresponding to the above sample is determined as the target emotion recognition result.

[0080] The second sub-step is to determine the sample emotion recognition result included in the above sample as the target sample emotion recognition result.

[0081] The third sub-step is to generate recognition result loss data based on the above-mentioned target emotion recognition result and the above-mentioned target sample emotion recognition result. The above-mentioned recognition result loss data can be the cross-entropy loss value between the above-mentioned target emotion recognition result and the above-mentioned target sample emotion recognition result. In practice, first, the emotion label probability distribution data corresponding to the above-mentioned target emotion recognition result can be determined as the target probability distribution data. Then, the individual label probability data included in the above-mentioned target probability distribution data can be combined into a one-dimensional vector as a label vector. Then, the above-mentioned target sample emotion recognition result can be subjected to one-hot encoding processing by an encoding function to obtain a one-hot vector corresponding to the above-mentioned target sample emotion recognition result as a sample label vector. The above-mentioned encoding function can be a function that can perform one-hot encoding on data. For example, the above-mentioned encoding function can be the get_dummies() function in pandas. Finally, the cross-entropy loss value between the above-mentioned label vector and the above-mentioned sample label vector can be determined as the recognition result loss data.

[0082] In a fourth sub-step, the sample multimodal feature information included in the sample is determined as target multimodal feature information. The target multimodal feature information may include target text feature information, target audio feature information, and target video feature information. The target text feature information may be the sample text feature information included in the sample multimodal feature information. The target audio feature information may be the sample audio feature information included in the sample multimodal feature information. The target video feature information may be the sample video feature information included in the sample multimodal feature information.

[0083] In practice, first, the sample audio feature information included in the sample multimodal feature information can be determined as the target audio feature information. Second, the sample audio feature information included in the sample multimodal feature information can be determined as the target audio feature information. Then, the sample video feature information included in the sample multimodal feature information can be determined as the target video feature information. Finally, the target audio feature information, the target audio feature information, and the target video feature information can be combined to form the target multimodal feature information.

[0084] In a fifth sub-step, the difference between the sample reconstructed text convolution feature corresponding to the sample and the target text feature information in the target multimodal feature information is determined as text difference feature information.

[0085] The sixth sub-step is to determine the modulus of the above text difference feature information as text feature modulus data.

[0086] The seventh sub-step is to determine the square of the text feature modulus data as the text feature modulus square data.

[0087] In an eighth sub-step, the difference between the sample-reconstructed audio convolution feature corresponding to the sample and the target audio feature information in the target multimodal feature information is determined as audio difference feature information.

[0088] In a ninth sub-step, the modulus of the audio difference feature information is determined as audio feature modulus data.

[0089] In a tenth sub-step, the square of the audio feature modulus data is determined as the audio feature modulus square data.

[0090] In the eleventh sub-step, the difference between the sample reconstructed video convolution feature corresponding to the above sample and the target video feature information in the above target multimodal feature information is determined as video difference feature information.

[0091] In the twelfth sub-step, the modulus of the video difference feature information is determined as video feature modulus data.

[0092] In the thirteenth sub-step, the square of the video feature modulus data is determined as the video feature modulus square data.

[0093] In the fourteenth sub-step, an average value of the text feature modulus square data, the audio feature modulus square data, and the video feature modulus square data is determined as reconstruction loss data.

[0094] In a fifteenth sub-step, based on the conditional diffusion generative network and the multimodal graph feature information of the sample, score loss data is generated. The score loss data may be the loss value corresponding to each data generated by the conditional diffusion generative network at the sampling time step.

[0095] The sixteenth sub-step is to generate sample loss data based on the above-mentioned recognition result loss data, the above-mentioned reconstruction loss data and the above-mentioned score loss data. The above-mentioned sample loss data can be the data obtained by weighted processing of the above-mentioned recognition result loss data, the above-mentioned reconstruction loss data and the above-mentioned score loss data. In practice, first, the product of the above-mentioned reconstruction loss data and the first preset weight data can be determined as weighted reconstruction loss data. The above-mentioned first preset weight data can be the weight data corresponding to the above-mentioned reconstruction loss data. For example, the above-mentioned first preset weight data can be 0.5. Secondly, the product of the above-mentioned score loss data and the second preset weight data can be determined as weighted score loss data. The above-mentioned second preset weight data can be the weight data corresponding to the above-mentioned score loss data. For example, the above-mentioned second preset weight data can be 0.6. Then, the sum of the above-mentioned recognition result loss data, the above-mentioned weighted reconstruction loss data and the above-mentioned weighted score loss data can be determined as sample loss data.

[0096] The second step is to generate classification loss data based on the generated sample loss data. The classification loss data may be the average value of the sample loss data. In practice, the execution entity may determine the average value of the sample loss data as the classification loss data.

[0097] The above technical solution and its related contents, combined with steps 101 to 103, serve as an inventive feature of an embodiment of the present disclosure, solving the problem of "wasted computing resources." Factors that often lead to wasted computing resources are as follows: When emotion recognition is performed only on single-modal data, the accuracy of emotion recognition is often low, which in turn leads to a high probability of requiring computing resources to repeatedly recognize the same data, which easily wastes computing resources during recognition. If these factors are resolved, computing resource waste can be reduced. To achieve this effect, the present disclosure performs the following steps for each of the at least one sample: First, the emotion recognition result corresponding to the sample is determined as the target emotion recognition result. Second, the sample emotion recognition result included in the sample is determined as the target sample emotion recognition result. Then, recognition result loss data is generated based on the target emotion recognition result and the target sample emotion recognition result. This allows the loss value between the target emotion recognition result and the target sample emotion recognition result to be obtained. Then, the sample multimodal feature information included in the sample is determined as the target multimodal feature information, where the target multimodal feature information includes target text feature information, target audio feature information, and target video feature information. Then, the difference between the sample-reconstructed text convolution feature corresponding to the above sample and the target text feature information in the above target multimodal feature information is determined as text difference feature information. Thus, the difference between the sample-reconstructed text convolution feature and the above target text feature information can be obtained. Then, the modulus of the above text difference feature information is determined as text feature modulus data. Thus, the modulus data of the above text difference feature information can be obtained. Then, the square of the above text feature modulus data is determined as text feature modulus squared data. Thus, the text feature modulus squared data can be obtained. Then, the difference between the sample-reconstructed audio convolution feature corresponding to the above sample and the target audio feature information in the above target multimodal feature information is determined as audio difference feature information. Thus, the difference between the sample-reconstructed audio convolution feature and the above target audio feature information can be obtained. Then, the modulus of the above audio difference feature information is determined as audio feature modulus data. Thus, the audio feature modulus data can be obtained. Then, the square of the above audio feature modulus data is determined as audio feature modulus squared data. Thus, the audio feature modulus squared data can be obtained. Then, the difference between the sample-reconstructed video convolution feature corresponding to the sample and the target video feature information in the target multimodal feature information is determined as video difference feature information. Thus, the difference between the sample-reconstructed video convolution feature and the target video feature information is obtained. Then, the modulus of the video difference feature information is determined as video feature modulus data. Thus, the modulus data of the video difference feature information is obtained. Then, the square of the video feature modulus data is determined as video feature modulus squared data. Thus, the video feature modulus squared data is obtained.Then, the average of the squared modulus data of the text features, the squared modulus data of the audio features, and the squared modulus data of the video features is determined as reconstruction loss data. This yields feature reconstruction loss data for the three modalities of text, audio, and video. Then, based on the conditional diffusion generative network and the multimodal graph feature information included in the sample, score loss data is generated. This yields loss values ​​for each data item generated by the conditional diffusion generative network. Then, based on the recognition loss data, the reconstruction loss data, and the score loss data, sample loss data is generated. This yields a loss value for the sample item by combining the recognition loss data, the reconstruction loss data, and the score loss data. Finally, based on the generated sample loss data, classification loss data is generated. This yields an average loss value for each sample in the at least one sample item. Because the network parameters of the initial neural network can be iteratively adjusted by combining the loss values ​​of the recognition loss data, the reconstruction loss data, and the score loss data, the recognition accuracy of the trained emotion recognition model can be improved, thereby reducing the probability of multiple recognitions of the same data item due to low recognition accuracy, and thus reducing the waste of computing resources.

[0098] In the process of adopting technical solutions to solve the above technical problems, the following problems often arise: Performing emotion recognition only by processing unimodal data can easily lead to low accuracy of the emotion recognition results, low user satisfaction with the emotion recognition results, and a high probability of needing to call computing resources for repeated recognition of the same unimodal data, resulting in a waste of computing resources.

[0099] Faced with the above technical problems, we decided to adopt the following solutions: In some optional implementations of some embodiments, the execution entity may generate score loss data based on the conditional diffusion generation network and the sample multimodal graph feature information included in the sample through the following steps: In the first step, a diffusion time step that satisfies a preset sampling condition among the above diffusion time steps is determined as a sampling time step, wherein the above sampling condition may be that the diffusion time step is any one of the above diffusion time steps.

[0100] The second step is to generate Gaussian vector data of the text to be processed based on the sample multimodal graph feature information, the sampling time step, and the sample diffuse text feature information corresponding to the sampling time step. The Gaussian vector data of the text to be processed may be Gaussian noise corresponding to the sample diffuse text feature information obtained through prediction.

[0101] In practice, the execution entity may input the sample multimodal graph feature information, the sampling time step, and the sample diffusion text feature information corresponding to the sampling time step into a pre-trained text reward model to obtain the Gaussian vector data of the text to be processed. The text reward model may be a neural network model that takes the sample multimodal graph feature information, the sampling time step, and the sample diffusion text feature information corresponding to the sampling time step as input, and takes the Gaussian vector data of the text to be processed as output. For example, the text reward model may be a preprocessed ScoreNet model. The preprocessing may be a process of fine-tuning the ScoreNet model using a text dataset and a cross-entropy loss function. Each text data in the text dataset may be data used to train the ScoreNet model to generate the Gaussian vector data of the text to be processed. Each text data in the text dataset may include sample multimodal graph feature information, a sampling time step, and the sample diffusion text feature information corresponding to the sampling time step.

[0102] The third step is to determine the product of the above-mentioned Gaussian vector data of the text to be processed and the preset Gaussian weight data as the weighted text Gaussian vector data. Wherein, the above-mentioned Gaussian weight data can be a preset weight value. For example, the above-mentioned Gaussian weight data can be 0.7.

[0103] In the fourth step, the sum of the weighted text Gaussian vector data and the text Gaussian vector data corresponding to the sampling time step is determined as the summed text Gaussian vector data.

[0104] In the fifth step, the modulus of the above-mentioned summed text Gaussian vector data is determined as text Gaussian modulus data.

[0105] In the sixth step, the square of the above text Gaussian modulus data is determined as the text Gaussian square data.

[0106] In step 7, based on the sample multimodal graph feature information, the sampling time step, and the sample diffuse audio feature information corresponding to the sampling time step, the audio Gaussian vector data to be processed is generated. The audio Gaussian vector data to be processed may be Gaussian noise corresponding to the sample diffuse audio feature information obtained through prediction.

[0107] In practice, the above-mentioned sample multimodal graph feature information, the above-mentioned sampling time step and the sample diffusion audio feature information corresponding to the above-mentioned sampling time step can be input into a pre-trained audio reward model to obtain the audio Gaussian vector data to be processed. The above-mentioned audio reward model can be a neural network model that takes the sample multimodal graph feature information, the sampling time step and the sample diffusion audio feature information corresponding to the sampling time step as input, and takes the audio Gaussian vector data to be processed as output. For example, the above-mentioned audio reward model can be a preprocessed ScoreNet model. The above-mentioned preprocessing can be a process of fine-tuning the ScoreNet model using an audio dataset and a cross-entropy loss function. Each audio data in the above-mentioned audio dataset can be data used to train the ScoreNet model to generate the audio Gaussian vector data to be processed. Each audio data in the above-mentioned audio dataset can include sample multimodal graph feature information, a sampling time step and the sample diffusion audio feature information corresponding to the sampling time step.

[0108] In the eighth step, the product of the to-be-processed audio Gaussian vector data and the Gaussian weight data is determined as weighted audio Gaussian vector data.

[0109] In the ninth step, the sum of the weighted audio Gaussian vector data and the audio Gaussian vector data corresponding to the sampling time step is determined as the summed audio Gaussian vector data.

[0110] In the tenth step, the modulus of the summed audio Gaussian vector data is determined as the audio Gaussian modulus data.

[0111] In the eleventh step, the square of the audio Gaussian modulus data is determined as the audio Gaussian square data.

[0112] In step 12, based on the sample multimodal image feature information, the sampling time step, and the sample diffusion video feature information corresponding to the sampling time step, Gaussian vector data of the video to be processed is generated. The Gaussian vector data of the video to be processed may be Gaussian noise corresponding to the sample diffusion video feature information obtained through prediction.

[0113] In practice, the above-mentioned sample multimodal graph feature information, the above-mentioned sampling time step, and the sample diffusion video feature information corresponding to the above-mentioned sampling time step can be input into a pre-trained video reward model to obtain the Gaussian vector data of the video to be processed. The above-mentioned video reward model can be a neural network model that takes the sample multimodal graph feature information, the sampling time step, and the sample diffusion video feature information corresponding to the sampling time step as input, and takes the Gaussian vector data of the video to be processed as output. For example, the above-mentioned video reward model can be a preprocessed ScoreNet model. The above-mentioned preprocessing can be a process of fine-tuning the ScoreNet model using a video dataset and a cross-entropy loss function. Each video data in the above-mentioned video dataset can be data used to train the ScoreNet model to generate the Gaussian vector data of the video to be processed. Each video data in the above-mentioned video dataset can include sample multimodal graph feature information, a sampling time step, and the sample diffusion video feature information corresponding to the sampling time step.

[0114] In the thirteenth step, the product of the above-mentioned Gaussian vector data of the video to be processed and the above-mentioned Gaussian weight data is determined as weighted video Gaussian vector data.

[0115] In the fourteenth step, the sum of the weighted video Gaussian vector data and the video Gaussian vector data corresponding to the sampling time step is determined as the summed video Gaussian vector data.

[0116] In the fifteenth step, the modulus of the summed video Gaussian vector data is determined as the video Gaussian modulus data.

[0117] In step 16, the square of the video Gaussian modulus data is determined as the video Gaussian square data.

[0118] In step 17, the average value of the text Gaussian square data, the audio Gaussian square data, and the video Gaussian square data is determined as the score loss data.

[0119] The above technical solution and its related contents, combined with steps 101 to 103, serve as an inventive feature of an embodiment of the present disclosure, solving the problem of "wasted computing resources." Factors that often lead to wasted computing resources include: performing emotion recognition solely on unimodal data can easily lead to low accuracy in the resulting emotion recognition results, low user satisfaction with the results, and a high probability of requiring computing resources to repeatedly recognize the same unimodal data, resulting in wasted computing resources. Addressing these factors can reduce computing resource waste. To achieve this, the present disclosure first identifies the diffusion time steps that meet preset sampling conditions among the aforementioned diffusion time steps as sampling time steps. This allows the determination of the time steps to be processed. Secondly, based on the aforementioned sample multimodal graph feature information, the aforementioned sampling time steps, and the sample diffusion text feature information corresponding to the aforementioned sampling time steps, Gaussian vector data for the text to be processed is generated. This allows the prediction of the Gaussian noise corresponding to the text modality diffusion network in the conditional diffusion generation network at the aforementioned sampling time step. Then, the product of the aforementioned Gaussian vector data for the text to be processed and preset Gaussian weight data is determined as weighted text Gaussian vector data. Thus, weighted text Gaussian vector data can be obtained. Then, the sum of the above-mentioned weighted text Gaussian vector data and the text Gaussian vector data corresponding to the above-mentioned sampling time step is determined as the summed text Gaussian vector data. Thus, summed text Gaussian vector data can be obtained. Then, the modulus of the above-mentioned summed text Gaussian vector data is determined as text Gaussian modulus data. Thus, text Gaussian modulus data can be obtained. Then, the square of the above-mentioned text Gaussian modulus data is determined as text Gaussian square data. Thus, text Gaussian square data can be obtained. Then, based on the above-mentioned sample multimodal graph feature information, the above-mentioned sampling time step and the sample diffusion audio feature information corresponding to the above-mentioned sampling time step, the audio Gaussian vector data to be processed is generated. Thus, at the above-mentioned sampling time step, the Gaussian noise corresponding to the audio modal diffusion network in the conditional diffusion generation network can be obtained. Then, the product of the above-mentioned audio Gaussian vector data to be processed and the above-mentioned Gaussian weight data is determined as weighted audio Gaussian vector data. Thus, weighted audio Gaussian vector data can be obtained. Then, the sum of the weighted audio Gaussian vector data and the audio Gaussian vector data corresponding to the sampling time step is determined as the summed audio Gaussian vector data. Thus, the summed audio Gaussian vector data can be obtained. Then, the modulus of the summed audio Gaussian vector data is determined as the audio Gaussian modulus data. Then, the square of the audio Gaussian modulus data is determined as the audio Gaussian square data. Then, based on the sample multimodal graph feature information, the sampling time step, and the sample diffusion video feature information corresponding to the sampling time step, the video Gaussian vector data to be processed is generated. Thus, the Gaussian noise corresponding to the video modal diffusion network in the conditional diffusion generation network at the sampling time step can be obtained.Then, the product of the above-mentioned video Gaussian vector data to be processed and the above-mentioned Gaussian weight data is determined as weighted video Gaussian vector data. Thus, weighted video Gaussian vector data can be obtained. Then, the sum of the above-mentioned weighted video Gaussian vector data and the video Gaussian vector data corresponding to the above-mentioned sampling time step is determined as summed video Gaussian vector data. Thus, summed video Gaussian vector data can be obtained. Then, the modulus of the above-mentioned summed video Gaussian vector data is determined as video Gaussian modulus data. Then, the square of the above-mentioned video Gaussian modulus data is determined as video Gaussian square data. Finally, the average value of the above-mentioned text Gaussian square data, the above-mentioned audio Gaussian square data and the above-mentioned video Gaussian square data is determined as score loss data. Thus, the average loss value of the above-mentioned text Gaussian square data, the above-mentioned audio Gaussian square data and the above-mentioned video Gaussian square data can be obtained by averaging the above-mentioned text Gaussian square data, the above-mentioned audio Gaussian square data and the above-mentioned video Gaussian square data. This is also because the characteristic vectors corresponding to the three modalities generated by the text modal diffusion network, audio modal diffusion network and video modal diffusion network in the conditional diffusion generation network at any diffusion time step can be processed, and then the Gaussian noises corresponding to the above diffusion time step can be inferred. Then, by comparing the inferred Gaussian noises with the Gaussian noises actually generated by the conditional diffusion generation network at the above diffusion time step, the loss values ​​corresponding to the characteristic vectors of the three modalities can be obtained. Then, the initial neural network is iteratively optimized according to the loss values ​​corresponding to the characteristic vectors of the three modalities. Therefore, the accuracy of the emotion recognition model after training for emotion recognition of data can be improved, the probability of repeated recognition of the same data due to low recognition accuracy can be reduced, and the waste of computing resources can be reduced.

[0120] Step 1026: Generate reconstructed fusion feature information based on the reconstructed text convolution features, the reconstructed audio convolution features, the reconstructed video convolution features, the multimodal graph feature information, and the feature processing network in the emotion recognition model.

[0121] In some embodiments, the execution entity may generate reconstructed fusion feature information based on the reconstructed text convolution features, the reconstructed audio convolution features, the reconstructed video convolution features, the multimodal graph feature information, and the feature processing network in the emotion recognition model. The reconstructed fusion feature information may be a feature vector obtained by fusing the reconstructed text convolution features, the reconstructed audio convolution features, the reconstructed video convolution features, and the multimodal graph feature information. In practice, the reconstructed text convolution features, the reconstructed audio convolution features, the reconstructed video convolution features, and the multimodal graph feature information may be input into the feature processing network in the emotion recognition model to obtain the reconstructed fusion feature information.

[0122] Step 1027: Generate a multimodal emotion recognition result based on the reconstructed fusion feature information and the classification network in the emotion recognition model.

[0123] In some embodiments, the execution entity may generate a multimodal emotion recognition result based on the reconstructed fusion feature information and the classification network in the emotion recognition model. The multimodal emotion recognition result may be a label representing the emotion category corresponding to the multimodal emotion data. In practice, the reconstructed fusion feature information may be input into the classification network in the emotion recognition model to obtain an emotion recognition result. The emotion recognition result may then be determined as a multimodal emotion recognition result.

[0124] Step 103: Send the generated multimodal emotion recognition results to the user terminal.

[0125] In some embodiments, the execution entity may send the generated multimodal emotion recognition results to the user terminal.

[0126] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: through the emotion recognition method based on multimodal missing data of some embodiments of the present disclosure, the accuracy of emotion recognition can be improved, and the waste of computing resources during recognition can be reduced. Specifically, the reason for the low accuracy of emotion recognition and the waste of computing resources during recognition is that only using a single modal data for emotion recognition can easily lead to low recognition accuracy. And when processing modal data such as text and audio, data missing and the like are likely to occur, which in turn can easily lead to low recognition accuracy when performing emotion recognition based on the processed data, resulting in the need to call computing resources for multiple repeated recognition of multimodal data sent by the same user terminal, resulting in waste of computing resources during recognition. Based on this, the emotion recognition method based on multimodal missing data of some embodiments of the present disclosure, first, in response to receiving a multimodal sentence sequence set sent by a user terminal, generates a multimodal emotion data sequence set based on the above-mentioned multimodal sentence sequence set. In this way, the original data that needs to be processed can be obtained. Secondly, for each multimodal emotion data in the above-mentioned multimodal emotion data sequence set, the following steps are performed: First, feature extraction processing is performed on the above-mentioned multimodal emotion data to obtain text modal feature information, audio modal feature information and video modal feature information corresponding to the above-mentioned multimodal emotion data. In this way, feature vectors corresponding to different modalities in the multimodal emotion data can be obtained. Secondly, based on the above-mentioned text modal feature information, the above-mentioned audio modal feature information and the above-mentioned video modal feature information, multimodal emotion fusion feature information corresponding to the above-mentioned multimodal emotion data is generated. In this way, feature vectors corresponding to different modalities in the multimodal emotion data can be fused. Then, based on the above-mentioned multimodal emotion data, a multimodal temporal relationship matrix and a speaker relationship matrix are constructed. In this way, a multimodal temporal relationship matrix and a speaker relationship matrix can be obtained. Then, based on the above-mentioned multimodal emotion fusion feature information, the above-mentioned multimodal temporal relationship matrix and the above-mentioned speaker relationship matrix, multimodal graph feature information is generated. In this way, multimodal graph feature information can be generated. Then, based on the above-mentioned multimodal graph feature information, the above-mentioned multimodal emotion fusion feature information and the conditional diffusion generation network in the pre-trained emotion recognition model, reconstructed text convolution features, reconstructed audio convolution features and reconstructed video convolution features are generated. In this way, reconstructed text convolution features, reconstructed audio convolution features and reconstructed video convolution features can be generated. Then, based on the above-mentioned reconstructed text convolution features, the above-mentioned reconstructed audio convolution features, the above-mentioned reconstructed video convolution features, the above-mentioned multimodal graph feature information and the feature processing network in the above-mentioned emotion recognition model, reconstructed fusion feature information is generated. In this way, reconstructed fusion feature information can be generated. Then, based on the above-mentioned reconstructed fusion feature information and the classification network in the above-mentioned emotion recognition model, a multimodal emotion recognition result is generated. In this way, a multimodal emotion recognition result can be generated. Finally, the generated multimodal emotion recognition results are sent to the above-mentioned user terminal.Thus, each multimodal emotion recognition result can be sent to the user terminal. Also, because the data of multiple modalities sent by the user terminal can be fused and processed, and then emotion recognition can be performed based on the processed data, the recognition accuracy during emotion recognition can be improved. Also, because the feature vectors corresponding to different modal data can be fused first to obtain multimodal emotion fusion feature information, and then emotion recognition can be performed based on the fused features, when the data of one modality is missing, the data of each modality other than the above modality can still provide a data basis for subsequent emotion recognition, thereby reducing the probability of low recognition accuracy of emotion recognition due to missing modal data, reducing the probability of calling computing resources to repeatedly recognize the same data due to low recognition accuracy, and reducing the waste of computing resources.

[0127] Further references Figure 2 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of an emotion recognition device based on multimodal missing data. These device embodiments are similar to Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.

[0128] like Figure 2As shown, some embodiments of the emotion recognition device 200 based on multimodal missing data include: a generating unit 201, an executing unit 202 and a sending unit 203. The generating unit 201 is configured to generate a multimodal emotion data sequence set based on the multimodal sentence sequence set in response to receiving a multimodal sentence sequence set sent by a user terminal; the executing unit 202 is configured to perform the following steps for each multimodal emotion data in the multimodal emotion data sequence set: perform feature extraction processing on the multimodal emotion data to obtain text modal feature information, audio modal feature information and video modal feature information corresponding to the multimodal emotion data; based on the text modal feature information, the audio modal feature information and the video modal feature information, generate multimodal emotion fusion feature information corresponding to the multimodal emotion data; based on the multimodal emotion data, construct a multimodal temporal relationship matrix and a speaker relationship matrix; based on the multimodal emotion The feature information, the multimodal temporal relationship matrix and the speaker relationship matrix are integrated to generate multimodal graph feature information; based on the multimodal graph feature information, the multimodal emotion fusion feature information and the conditional diffusion generation network in the pre-trained emotion recognition model, reconstructed text convolution features, reconstructed audio convolution features and reconstructed video convolution features are generated; based on the reconstructed text convolution features, the reconstructed audio convolution features, the reconstructed video convolution features, the multimodal graph feature information and the feature processing network in the emotion recognition model, reconstructed fusion feature information is generated; based on the reconstructed fusion feature information and the classification network in the emotion recognition model, a multimodal emotion recognition result is generated; the sending unit 203 is configured to send the generated multimodal emotion recognition results to the user terminal.

[0129] It can be understood that the units recorded in the emotion recognition device 200 based on multimodal missing data are similar to the reference Figure 1 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the emotion recognition device 200 based on multimodal missing data and the units contained therein, and will not be repeated here.

[0130] Reference below Figure 3 , which shows a structural diagram of an electronic device (such as a computing device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0131] like Figure 3As shown, electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 302 or programs loaded from a storage device 308 into a random access memory (RAM) 303. RAM 303 also stores various programs and data required for the operation of electronic device 300. Processing device 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to bus 304.

[0132] Typically, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Figure 3 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead. Figure 3 Each block shown in the figure may represent one device, or may represent multiple devices as needed.

[0133] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from a network via the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the method of some embodiments of the present disclosure are performed.

[0134] It should be noted that the computer-readable medium described in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable storage media may include, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In some embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. Furthermore, in some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. This propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wire, optical cable, RF (radio frequency), or any suitable combination thereof.

[0135] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0136] The above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or it may exist independently without being assembled into the electronic device. The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: generates a multimodal emotion data sequence set based on the above-mentioned multimodal sentence sequence set in response to receiving a multimodal sentence sequence set sent by the user terminal; for each multimodal emotion data in the above-mentioned multimodal emotion data sequence set, performs the following steps: performs feature extraction processing on the above-mentioned multimodal emotion data to obtain text modal feature information, audio modal feature information and video modal feature information corresponding to the above-mentioned multimodal emotion data; generates multimodal emotion fusion feature information corresponding to the above-mentioned multimodal emotion data based on the above-mentioned text modal feature information, the above-mentioned audio modal feature information and the above-mentioned video modal feature information; constructs a multimodal temporal relationship matrix based on the above-mentioned multimodal emotion data. and speaker relationship matrix; based on the above-mentioned multimodal emotion fusion feature information, the above-mentioned multimodal temporal relationship matrix and the above-mentioned speaker relationship matrix, generate multimodal graph feature information; based on the above-mentioned multimodal graph feature information, the above-mentioned multimodal emotion fusion feature information and the conditional diffusion generation network in the pre-trained emotion recognition model, generate reconstructed text convolution features, reconstructed audio convolution features and reconstructed video convolution features; based on the above-mentioned reconstructed text convolution features, the above-mentioned reconstructed audio convolution features, the above-mentioned reconstructed video convolution features, the above-mentioned multimodal graph feature information and the feature processing network in the above-mentioned emotion recognition model, generate reconstructed fusion feature information; based on the above-mentioned reconstructed fusion feature information and the classification network in the above-mentioned emotion recognition model, generate multimodal emotion recognition results; and send the generated multimodal emotion recognition results to the above-mentioned user terminal.

[0137] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0138] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0139] The units described in some embodiments of the present disclosure may be implemented in software or in hardware. The described units may also be provided in a processor. For example, they may be described as: a processor comprising a generation unit, an execution unit, and a sending unit. The names of these units do not, in some cases, constitute a limitation on the units themselves. For example, the generation unit may also be described as a “unit for generating a set of multimodal emotion data sequences.”

[0140] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0141] The above descriptions are merely some preferred embodiments of the present disclosure and illustrate the underlying technical principles. Those skilled in the art should understand that the scope of the invention encompassed by the embodiments of the present disclosure is not limited to technical solutions formed by specific combinations of the aforementioned technical features. It also encompasses other technical solutions formed by any combination of the aforementioned technical features or their equivalents, without departing from the aforementioned inventive concept. For example, a technical solution formed by replacing the aforementioned features with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A method for emotion recognition based on multimodal missing data, comprising: In response to receiving a multimodal sentence sequence set sent by a user terminal, generating a multimodal emotion data sequence set based on the multimodal sentence sequence set; For each multimodal emotion data in the multimodal emotion data sequence set, the following steps are performed: Performing feature extraction processing on the multimodal emotion data to obtain text modal feature information, audio modal feature information, and video modal feature information corresponding to the multimodal emotion data; Generating multimodal emotion fusion feature information corresponding to the multimodal emotion data based on the text modality feature information, the audio modality feature information, and the video modality feature information; Based on the multimodal emotion data, constructing a multimodal temporal relationship matrix and a speaker relationship matrix; generating multimodal graph feature information based on the multimodal emotion fusion feature information, the multimodal temporal relationship matrix, and the speaker relationship matrix; Generate reconstructed text convolution features, reconstructed audio convolution features, and reconstructed video convolution features based on the multimodal graph feature information, the multimodal emotion fusion feature information, and a conditional diffusion generation network in a pre-trained emotion recognition model; Generate reconstructed fusion feature information based on the reconstructed text convolution feature, the reconstructed audio convolution feature, the reconstructed video convolution feature, the multimodal graph feature information and the feature processing network in the emotion recognition model; Generating a multimodal emotion recognition result based on the reconstructed fusion feature information and the classification network in the emotion recognition model; The generated multimodal emotion recognition results are sent to the user terminal.

2. The method according to claim 1, wherein The generating of multimodal emotion fusion feature information corresponding to the multimodal emotion data based on the text modal feature information, the audio modal feature information, and the video modal feature information includes: Performing random masking processing on the text modality feature information, the audio modality feature information, and the video modality feature information to obtain masked text feature information, masked audio feature information, and masked video feature information; Feature fusion processing is performed on the masked text feature information, the masked audio feature information, and the masked video feature information to obtain multimodal emotion fusion feature information.

3. The method according to claim 2, wherein: The randomly masking the text modality feature information, the audio modality feature information, and the video modality feature information to obtain masked text feature information, masked audio feature information, and masked video feature information includes: Generate text mask probability data, audio mask probability data, and video mask probability data based on preset text missing probability data, preset audio missing probability data, preset video missing probability data, and preset missing data; Based on the text mask probability data, masking the text modal feature information to obtain masked text feature information; Based on the audio mask probability data, masking the audio modal feature information to obtain masked audio feature information; Based on the video mask probability data, mask processing is performed on the video modality feature information to obtain masked video feature information.

4. The method according to claim 1, wherein The generating of multimodal graph feature information based on the multimodal emotion fusion feature information, the multimodal temporal relationship matrix, and the speaker relationship matrix includes: Generate time series feature information based on the multimodal time series relationship matrix and the multimodal emotion fusion feature information; generating speaker feature information based on the speaker relationship matrix and the multimodal emotion fusion feature information; Feature splicing processing is performed on the time series feature information and the speaker feature information to obtain multimodal graph feature information.

5. The method according to claim 1, wherein The multimodal emotion data includes text modal emotion data, audio modal emotion data, and video modal emotion data; and the feature extraction processing of the multimodal emotion data to obtain text modal feature information, audio modal feature information, and video modal feature information corresponding to the multimodal emotion data includes: Performing word segmentation processing on the text modality emotion data included in the multimodal emotion data to obtain a text modality word segmentation sequence; Encoding the text modal word segmentation sequence to obtain a text encoding data sequence; Performing feature extraction processing on the text encoding data sequence to obtain text modal feature information; Resampling the audio modality emotion data included in the multimodal emotion data to obtain resampled audio data; Based on the resampled audio data, generating frame-level audio feature information corresponding to the resampled audio data; Performing pooling processing on the respective frame-level audio feature information to obtain audio modal feature information; Performing face alignment processing on the video modality emotion data included in the multimodal emotion data to obtain aligned video data; Performing feature extraction processing on the aligned video data to obtain video feature information at each frame level; Pooling is performed on the frame-level video feature information to obtain video modality feature information.

6. The method according to claim 1, wherein The conditional diffusion generation network includes a text modality diffusion network, an audio modality diffusion network, a video modality diffusion network, a text convolution layer, an audio convolution layer, and a video convolution layer; and the emotion recognition model is trained by the following steps: Acquire a sample set, wherein each sample in the sample set includes sample multimodal feature information, sample multimodal graph feature information, sample multimodal emotion fusion feature information, and a sample emotion recognition result, and the sample multimodal feature information includes sample text feature information, sample audio feature information, and sample video feature information; Based on the sample set, the following training steps are performed: For each sample in at least one sample in the sample set, inputting the sample multimodal emotion fusion feature information and the sample multimodal graph feature information included in the sample into the text modality diffusion network of the conditional diffusion generation network in the initial neural network to obtain sample reconstructed text feature information corresponding to the sample; For each sample of the at least one sample, inputting the sample reconstructed text feature information corresponding to the sample into the text convolution layer of the conditional diffusion generation network in the initial neural network to obtain the sample reconstructed text convolution feature corresponding to the sample; For each of the at least one sample, inputting the sample multimodal emotion fusion feature information and the sample multimodal graph feature information included in the sample into the audio modal diffusion network of the conditional diffusion generation network in the initial neural network to obtain sample reconstructed audio feature information corresponding to the sample; For each sample of the at least one sample, inputting the sample-reconstructed audio feature information corresponding to the sample into the audio convolution layer of the conditional diffusion generation network in the initial neural network to obtain the sample-reconstructed audio convolution feature corresponding to the sample; For each of the at least one sample, inputting the sample multimodal emotion fusion feature information and the sample multimodal graph feature information included in the sample into the video modality diffusion network of the conditional diffusion generation network in the initial neural network to obtain sample reconstructed video feature information corresponding to the sample; For each sample of the at least one sample, inputting sample reconstructed video feature information corresponding to the sample into a video convolution layer of a conditional diffusion generation network in the initial neural network to obtain a sample reconstructed video convolution feature corresponding to the sample; For each of the at least one sample, inputting the sample multimodal graph feature information included in the sample, the sample reconstructed text convolution feature, the sample reconstructed audio convolution feature, and the sample reconstructed video convolution feature corresponding to the sample into a feature processing network of the initial neural network to obtain sample reconstruction fusion feature information; Inputting the reconstructed fusion feature information of each sample corresponding to the at least one sample into the classification network in the initial neural network to obtain each emotion recognition result; Generate classification loss data based on the conditional diffusion generative network, each emotion recognition result corresponding to the at least one sample, each sample reconstructed text convolution feature, each sample reconstructed audio convolution feature, each sample reconstructed video convolution feature, each sample emotion recognition result included in the at least one sample, each sample multimodal feature information, and each sample multimodal graph feature information; In response to determining that the classification loss data satisfies a preset classification loss condition, determining the initial neural network as a sentiment recognition model; In response to determining that the classification loss data does not meet the classification loss condition, adjusting the network parameters of the initial neural network, and using unused samples to form a sample set, and performing the training step again using the adjusted initial neural network.

7. An emotion recognition device based on multimodal missing data, comprising: a generating unit configured to generate a multimodal emotion data sequence set based on the multimodal sentence sequence set in response to receiving the multimodal sentence sequence set sent by the user terminal; The execution unit is configured to perform the following steps for each multimodal emotion data in the multimodal emotion data sequence set: performing feature extraction processing on the multimodal emotion data to obtain text modal feature information, audio modal feature information, and video modal feature information corresponding to the multimodal emotion data; and generating multimodal emotion fusion feature information corresponding to the multimodal emotion data based on the text modal feature information, the audio modal feature information, and the video modal feature information; Based on the multimodal emotion data, constructing a multimodal temporal relationship matrix and a speaker relationship matrix; Based on the multimodal emotion fusion feature information, the multimodal temporal relationship matrix and the speaker relationship matrix, multimodal graph feature information is generated; based on the multimodal graph feature information, the multimodal emotion fusion feature information and the conditional diffusion generation network in the pre-trained emotion recognition model, reconstructed text convolution features, reconstructed audio convolution features and reconstructed video convolution features are generated; based on the reconstructed text convolution features, the reconstructed audio convolution features, the reconstructed video convolution features, the multimodal graph feature information and the feature processing network in the emotion recognition model, reconstructed fusion feature information is generated; Generating a multimodal emotion recognition result based on the reconstructed fusion feature information and the classification network in the emotion recognition model; The sending unit is configured to send the generated multimodal emotion recognition results to the user terminal.

8. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A computer-readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Multi-modal emotion recognition method based on graph convolutional network

    CN116229225A

  • Multi-modal sentiment analysis method based on diffusion attention fusion and Fourier reconstruction

    CN117953513A

  • Video behavior recognition method, device and equipment based on multi-mode large model fine tuning

    CN119495127A

  • Multi-dimensional character relationship discovery method based on multi-modal information fusion

    CN119760641A

  • Text user sentiment analysis method and system based on multiple modes and AI

    CN120216700A

Cited By

  • Multi-modal emotion recognition model training method and device, equipment and medium

    CN122045965A