An emotion recognition method, device, system, and storage medium

By annotating and fusing video data, and combining visual and textual contextual semantic information, the problem of insufficient information utilization in cross-modal fusion dialogue emotion recognition is solved, achieving more accurate emotion recognition and generalization capabilities.

CN115690875BActive Publication Date: 2026-05-19GUILIN UNIV OF ELECTRONIC TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUILIN UNIV OF ELECTRONIC TECH
Filing Date
2022-10-19
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies fail to fully utilize multimodal contextual semantic information in cross-modal fusion dialogue sentiment recognition, resulting in an inability to effectively identify dialogue sentiment in complex contexts.

Method used

By labeling video data and dividing it into training and testing sets, a fusion analysis is performed to build an emotion recognition model. The model is trained using role vectors and fusion features, the loss function is calculated and the model parameters are updated, and finally, the model is tested to recognize emotions.

Benefits of technology

By effectively combining visual and textual contextual semantic information, the accuracy of emotion recognition is improved, and it also has a certain degree of generalization ability and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115690875B_ABST
    Figure CN115690875B_ABST
Patent Text Reader

Abstract

The application provides an emotion recognition method, device, system and storage medium, and belongs to the field of video recognition, and the method comprises the following steps: labeling video data to obtain labeled video data; dividing the labeled video data and a role vector into a video training set and a video test set according to a preset proportion; performing fusion analysis on the labeled video data to obtain fusion features; and training an emotion recognition model by using the role vector and the fusion features to obtain trained features. The application effectively combines the context semantic information of vision and text, can connect the context and auxiliary questions to learn speaker-specific features, improves the accuracy of emotion recognition, and has certain generalization ability, and has good reliability and robustness in other emotion recognition tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates primarily to the field of video recognition technology, specifically to an emotion recognition method, device, system, and storage medium. Background Technology

[0002] Currently, cross-modal fusion dialogue sentiment recognition has become a hot topic. However, mainstream methods do not fully utilize the contextual semantic information of multimodal interactions and fail to pay attention to the transmission of emotions among speakers during the dialogue, resulting in an inability to effectively recognize dialogue sentiment in complex contexts. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide an emotion recognition method, device, system and storage medium to address the shortcomings of the prior art.

[0004] The technical solution of the present invention to solve the above-mentioned technical problems is as follows: An emotion recognition method, comprising the following steps:

[0005] Import multiple video data sets and annotate each video data set to obtain multiple annotated video data sets for each video data set; each annotated video data set corresponds to a character vector.

[0006] All the labeled video data are divided into a video training set and a video test set according to a preset ratio;

[0007] The fusion analysis is performed on each of the labeled video data in the video training set to obtain the fusion features of each of the labeled video data in the video training set.

[0008] An emotion recognition model is constructed by training the emotion recognition model using the role vectors and fusion features corresponding to each of the annotated video data in the video training set, thereby obtaining the trained features of each of the annotated video data in the video training set.

[0009] Import the real labels corresponding to the training features of each labeled video data in the video training set, and calculate the loss function based on the training features and real labels of all labeled video data in the video training set to obtain the loss function.

[0010] The parameters of the emotion recognition model are updated according to the loss function to obtain the updated emotion recognition model.

[0011] The updated emotion recognition model is tested using the video test set to obtain emotion recognition results.

[0012] Another technical solution of the present invention to solve the above-mentioned technical problems is as follows: An emotion recognition device, comprising:

[0013] The annotation module is used to import multiple video data and annotate each of the video data to obtain multiple annotated video data for each video data; wherein each of the annotated video data corresponds to a character vector.

[0014] The partitioning module is used to divide all the labeled video data into a video training set and a video test set according to a preset ratio.

[0015] The fusion analysis module is used to perform fusion analysis on each of the labeled video data in the video training set to obtain the fusion features of each of the labeled video data in the video training set.

[0016] The model training module is used to construct an emotion recognition model. It trains the emotion recognition model using the role vectors and fusion features corresponding to each labeled video data in the video training set, and obtains the trained features of each labeled video data in the video training set.

[0017] The loss function calculation module is used to import the real labels corresponding to the training features of each labeled video data in the video training set, and calculate the loss function based on the training features and real labels of all labeled video data in the video training set to obtain the loss function.

[0018] The parameter update module is used to update the parameters of the emotion recognition model according to the loss function to obtain the updated emotion recognition model.

[0019] The recognition result acquisition module is used to test the updated emotion recognition model using the video test set to obtain the emotion recognition result.

[0020] Based on the above-mentioned emotion recognition method, the present invention also provides an emotion recognition system.

[0021] Another technical solution of the present invention to solve the above-mentioned technical problems is as follows: an emotion recognition system, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the emotion recognition method described above is implemented.

[0022] Based on the above-mentioned emotion recognition method, the present invention also provides a computer-readable storage medium.

[0023] Another technical solution of the present invention to solve the above-mentioned technical problems is as follows: a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the emotion recognition method as described above.

[0024] The beneficial effects of this invention are as follows: Annotated video data is obtained by labeling video data; the labeled video data is divided into a video training set and a video test set according to a preset ratio; fusion analysis of the labeled video data yields fusion features; training of the emotion recognition model using role vectors and fusion features yields trained features; a loss function is calculated based on the trained features and the loss function of the real labels; the parameters of the emotion recognition model are updated according to the loss function to obtain an updated emotion recognition model; and the updated emotion recognition model is tested using the video test set to obtain the emotion recognition result. This effectively combines visual and textual contextual semantic information, connecting context with auxiliary questions to learn speaker-specific features, improving the accuracy of emotion recognition, and exhibiting a certain generalization ability. It also demonstrates good reliability and robustness in other emotion recognition tasks. Attached Figure Description

[0025] Figure 1 A flowchart illustrating an emotion recognition method provided in an embodiment of the present invention;

[0026] Figure 2 This is a block diagram of an emotion recognition device provided in an embodiment of the present invention. Detailed Implementation

[0027] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0028] Figure 1 This is a flowchart illustrating an emotion recognition method provided in an embodiment of the present invention.

[0029] like Figure 1 As shown, an emotion recognition method includes the following steps:

[0030] Import multiple video data sets and annotate each video data set to obtain multiple annotated video data sets for each video data set; each annotated video data set corresponds to a character vector.

[0031] All the labeled video data are divided into a video training set and a video test set according to a preset ratio;

[0032] The fusion analysis is performed on each of the labeled video data in the video training set to obtain the fusion features of each of the labeled video data in the video training set.

[0033] An emotion recognition model is constructed by training the emotion recognition model using the role vectors and fusion features corresponding to each of the annotated video data in the video training set, thereby obtaining the trained features of each of the annotated video data in the video training set.

[0034] Import the real labels corresponding to the training features of each labeled video data in the video training set, and calculate the loss function based on the training features and real labels of all labeled video data in the video training set to obtain the loss function.

[0035] The parameters of the emotion recognition model are updated according to the loss function to obtain the updated emotion recognition model.

[0036] The updated emotion recognition model is tested using the video test set to obtain emotion recognition results.

[0037] Preferably, the preset ratio can be 3:7.

[0038] It should be understood that video data is collected and preprocessed to obtain video dialogue segments (i.e., the annotated video data).

[0039] It should be understood that the video segments (i.e., the annotated video data) are annotated.

[0040] Specifically, the video dialogue segments (i.e., the annotated video data) are labeled with roles and corresponding emotion tags to obtain a set (U,E) = {(U1,E1),(U2,E2),...,(U... N E N )},(U i E i )={(u (i,1) ,e (i,1) ),(u (i,2) ,e (i,2) ),…,(u (i,N) ,e (i,N) )}

[0041] e (i,j) ∈{joy, sadness, fear, anger, surprise, disgust}

[0042] (U i E i ) represents the set of characters and corresponding emotions in the i-th group of dialogue video segments (i.e., the labeled video data), u(i,j) ,e (i,j) This represents the character and corresponding emotion speaking in the j-th round of the i-th group of dialogue video clips (i.e., the annotated video data).

[0043] It should be understood that the trained model (i.e., the updated emotion recognition model) is tested with test set samples (i.e., the video test set) to achieve dialogue emotion recognition results.

[0044] In the above embodiments, labeled video data is obtained by annotating the video data. The labeled video data is then divided into a video training set and a video test set according to a preset ratio. Fusion features are obtained by fusion analysis of the labeled video data. The emotion recognition model is trained using role vectors and fusion features to obtain trained features. A loss function is calculated based on the loss function of the trained features and the real labels. The parameters of the emotion recognition model are updated according to the loss function to obtain an updated emotion recognition model. The updated emotion recognition model is tested using the video test set to obtain the emotion recognition result. This effectively combines the contextual semantic information of vision and text, and can connect the context with auxiliary questions to learn speaker-specific features, thereby improving the accuracy of emotion recognition. It also has a certain generalization ability and good reliability and robustness in other emotion recognition tasks.

[0045] Optionally, as an embodiment of the present invention, the process of performing fusion analysis on each of the labeled video data in the video training set to obtain the fusion features of each of the labeled video data in the video training set includes:

[0046] Face detection processing is performed on each of the annotated video data in the video training set to obtain a face image sequence for each of the annotated video data in the video training set.

[0047] Text extraction processing is performed on each of the annotated video data in the video training set to obtain the text sequence of each of the annotated video data in the video training set;

[0048] The face image sequence and text sequence of each labeled video data in the video training set are encoded and analyzed using an encoder to obtain the image features and text features of each labeled video data in the video training set.

[0049] The image features and text features of each labeled video data in the video training set are weighted and concatenated to obtain the fusion features of each labeled video data in the video training set.

[0050] It should be understood that video dialogue segments (i.e., each of the labeled video data in the video training set) are imported, and the corresponding face image sequences and text sequences are extracted.

[0051] It should be understood that the image sequence and text sequence are encoded and fused to obtain fused features.

[0052] Specifically, by performing face detection on the videos (i.e., each of the annotated video data in the video training set), a sequence of face images for each dialogue segment is extracted.

[0053] Let i represent the set of face image sequences of the i-th group of dialogue video segments.

[0054] Let represent the sequence of facial images of the speaker in the j-th round of the i-th dialogue video segment.

[0055] Specifically, by extracting text from the video subtitles (annotated segments), the text sequences corresponding to the speaker's turn in the video are extracted.

[0056] This represents the set of subtitle text for the i-th group of dialogue video clips.

[0057] This represents the subtitle text sequence of the speaker in the j-th round of the i-th dialogue video segment.

[0058] It should be understood that the features (i.e., the image features) and The text features are weighted and concatenated to obtain the fused feature f. (i,j) .

[0059] In the above embodiments, the fusion analysis of the labeled video data yields fusion features, which effectively combine the contextual semantic information of vision and text. This allows the context to be connected with auxiliary questions to learn speaker-specific features, thereby improving the accuracy of emotion recognition.

[0060] Optionally, as an embodiment of the present invention, the process of using an encoder to encode and analyze the face image sequences and text sequences of each labeled video data in the video training set, and correspondingly obtaining the image features and text features of each labeled video data in the video training set, includes:

[0061] One-dimensional convolution processing is performed on the face image sequence and text sequence of each labeled video data in the video training set to obtain the initial image features and initial text features of each labeled video data in the video training set.

[0062] The initial image features and initial text features of each labeled video data in the video training set are respectively processed by the trigonometric function encoding algorithm to obtain the image position information and text position information of each labeled video data in the video training set.

[0063] The image location information and text location information of each labeled video data in the video training set are encoded using an encoder to obtain the image features and text features of each labeled video data in the video training set.

[0064] It should be understood that the encoding module (i.e. the encoder) consists of two structurally independent transformer encoders, which are used to encode the image sequence (i.e. the face image sequence) and the text sequence, respectively.

[0065] It should be understood that a transformer encoder can consist of N identical sub-decoders cascaded together, where each sub-decoder consists of a head attention layer and a feedforward layer.

[0066] Specifically, the input sequence (i.e., the facial image sequence) and (That is, the text sequence) is fed into the input encoding layer and subjected to a one-dimensional convolution operation to obtain... Where k∈{v,t}.

[0067] Specifically, the aforementioned (i.e., the initial image features or the initial text features) are positionally encoded to obtain the position information of each frame image and each word in the sequence. The encoding method uses trigonometric function encoding, and the specific formula is as follows:

[0068]

[0069]

[0070] It should be understood that location information (i.e., the image location information or the text location information) is input into the transformer encoder for encoding, and then processed by Q... k K k and V k The image features are obtained from vectors. and text (i.e., the text features).

[0071] In the above embodiments, the encoder is used to analyze the encoding of the face image sequence and the text sequence to obtain the corresponding image features and text features. This allows us to obtain the position information of each frame image and each word in the sequence, effectively combining the contextual semantic information of vision and text, and improving the accuracy of emotion recognition.

[0072] Optionally, as an embodiment of the present invention, the emotion recognition model includes a BERT encoder and a fully connected layer;

[0073] The process of constructing an emotion recognition model, which involves training the emotion recognition model using the character vectors and fusion features corresponding to each of the annotated video data in the video training set, to obtain the trained features of each of the annotated video data in the video training set, includes:

[0074] The weighted concatenation of the character vectors and fusion features corresponding to each labeled video data in the video training set is performed to obtain the concatenated features of each labeled video data in the video training set.

[0075] Import the problem task vector, and then perform a weighted concatenation of the problem task vector with the concatenated features of each of the annotated video data in the video training set to obtain the features to be encoded for each of the annotated video data in the video training set.

[0076] The BERT encoder is used to encode the features to be encoded for each of the labeled video data in the video training set, thereby obtaining the classification features for each of the labeled video data in the video training set.

[0077] The fully connected layer is used to map the classification features of each labeled video data in the video training set to obtain the post-training features of each labeled video data in the video training set.

[0078] It should be understood that the fused features are input into the emotion recognition model to obtain the output result (i.e., the trained features).

[0079] Specifically, firstly, regarding f (i,j) (i.e., the fusion feature) is used for role embedding u (i,j) (i.e., the character vector), to obtain (f) (i,j) ,u (i,j) (i.e., the cascaded features); secondly, construct a question-answering task T. (i,j) ={u (i,j)How are you feeling? |i∈N,j∈n i} (i.e., the problem task vector); then, (f (i,j) ,u (i,j) ,T (i,j) (i.e., the feature to be encoded) is input into the BERT encoder to obtain the classification feature h. (i,j) Finally, h (i,j) (i.e., the classification features) are input into the fully connected layer for mapping, resulting in the output r. (i,j) (i.e., the post-training features).

[0080] In the above embodiments, the training of the emotion recognition model is carried out using role vectors and fused features to obtain trained features, which improves the accuracy of emotion recognition and has a certain generalization ability. It also has good reliability and robustness in other emotion recognition tasks.

[0081] Optionally, as an embodiment of the present invention, the process of calculating the loss function based on the post-training features of all the labeled video data in the video training set and the real labels to obtain the loss function includes:

[0082] Based on the first equation, the loss function is calculated using the post-training features and real labels of all the labeled video data in the video training set, resulting in the loss function. The first equation is:

[0083]

[0084] Specifically, e′ (i,j) =argmax(P (i,j) ),

[0085] Specifically,

[0086] Specifically,

[0087] Where Loss is the loss function, e′ (i,j) Let e ​​be the sentiment prediction label for the j-th labeled video data of the i-th video data. (i,j) Let n be the true label of the j-th labeled video data of the i-th video data, and N be the total number of video data. i P represents the total number of labeled video data. (i,j) Z represents the sentiment probability of the j-th labeled video data for the i-th video data. (i,j) For the j-th labeled video data of the i-th video data, and All are learnable parameters, r (i,j)The features of the j-th labeled video data are the post-training features of the i-th video data.

[0088] It should be understood that the entire model is trained on all dialogue samples in the training set (i.e., the video training set) by calculating the loss function.

[0089] Specifically, the loss function adopted is the classification cross-entropy loss function L, where,

[0090] Z (i,j) It is a normalization term, which can be calculated using a forward and backward algorithm. The specific calculation method is as follows:

[0091]

[0092] P (i,j) The sentiment probability is calculated using the softmax layer.

[0093]

[0094] e′ (i,j) It is based on P (i,j) The sentiment prediction label is obtained from the sentiment probability with the highest probability.

[0095] e′ (i,j) =argmax(P (i,j) ),

[0096]

[0097] W a and b a It is a learnable parameter, e (i,j) These are the corresponding real tags.

[0098] In the above embodiments, the loss function is calculated based on the loss function of the trained features and the real labels according to the first formula. It effectively combines the contextual semantic information of vision and text, and can connect the context with auxiliary questions to learn speaker-specific features, thereby improving the accuracy of emotion recognition. It also has a certain generalization ability and good reliability and robustness in other emotion recognition tasks.

[0099] Optionally, as another embodiment of the present invention, the present invention uses the transformer framework combined with a gated full fusion module to solve the problem of selecting highly useful features for learning, so that the entire framework can generate event descriptions with a holistic understanding of the video, and obtain more accurate description results.

[0100] Optionally, as another embodiment of the present invention, the present invention collects video data and preprocesses the video data to obtain video dialogue segments; annotates the video segments; imports the video dialogue segments and extracts corresponding face image sequences and text sequences; inputs the image sequences and text sequences into an encoding module for encoding to obtain corresponding video features and text features; fuses the video features and text features to obtain corresponding fused features; inputs the fused features into an emotion recognition model to obtain output results; trains the entire model on the training set data through loss function calculation and backpropagation; and tests the trained model with test set data to achieve dialogue emotion recognition. The present invention effectively combines visual and textual contextual semantic information, connecting context with auxiliary questions to learn speaker-specific features, improving the accuracy of emotion recognition, and possessing a certain generalization ability, exhibiting good reliability and robustness in other emotion recognition tasks.

[0101] Figure 2 This is a block diagram of an emotion recognition device provided in an embodiment of the present invention.

[0102] Alternatively, as another embodiment of the present invention, such as Figure 2 As shown, an emotion recognition device includes:

[0103] The annotation module is used to import multiple video data and annotate each of the video data to obtain multiple annotated video data for each video data; wherein each of the annotated video data corresponds to a character vector.

[0104] The partitioning module is used to divide all the labeled video data into a video training set and a video test set according to a preset ratio.

[0105] The fusion analysis module is used to perform fusion analysis on each of the labeled video data in the video training set to obtain the fusion features of each of the labeled video data in the video training set.

[0106] The model training module is used to construct an emotion recognition model. It trains the emotion recognition model using the role vectors and fusion features corresponding to each labeled video data in the video training set, and obtains the trained features of each labeled video data in the video training set.

[0107] The loss function calculation module is used to import the real labels corresponding to the training features of each labeled video data in the video training set, and calculate the loss function based on the training features and real labels of all labeled video data in the video training set to obtain the loss function.

[0108] The parameter update module is used to update the parameters of the emotion recognition model according to the loss function to obtain the updated emotion recognition model.

[0109] The recognition result acquisition module is used to test the updated emotion recognition model using the video test set to obtain the emotion recognition result.

[0110] Optionally, as an embodiment of the present invention, the fusion analysis module is specifically used for:

[0111] Face detection processing is performed on each of the annotated video data in the video training set to obtain a face image sequence for each of the annotated video data in the video training set.

[0112] Text extraction processing is performed on each of the annotated video data in the video training set to obtain the text sequence of each of the annotated video data in the video training set;

[0113] The face image sequence and text sequence of each labeled video data in the video training set are encoded and analyzed using an encoder to obtain the image features and text features of each labeled video data in the video training set.

[0114] The image features and text features of each labeled video data in the video training set are weighted and concatenated to obtain the fusion features of each labeled video data in the video training set.

[0115] Optionally, as an embodiment of the present invention, the fusion analysis module is specifically used for:

[0116] One-dimensional convolution processing is performed on the face image sequence and text sequence of each labeled video data in the video training set to obtain the initial image features and initial text features of each labeled video data in the video training set.

[0117] The initial image features and initial text features of each labeled video data in the video training set are respectively processed by the trigonometric function encoding algorithm to obtain the image position information and text position information of each labeled video data in the video training set.

[0118] The image location information and text location information of each labeled video data in the video training set are encoded using an encoder to obtain the image features and text features of each labeled video data in the video training set.

[0119] Optionally, another embodiment of the present invention provides an emotion recognition system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the emotion recognition method as described above. This system can be a computer or similar system.

[0120] Optionally, another embodiment of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the emotion recognition method as described above.

[0121] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0122] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described apparatus and unit can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0123] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0124] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of the present invention, depending on actual needs.

[0125] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0126] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0127] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An emotion recognition method, characterized in that, Includes the following steps: Import multiple video data sets and annotate each video data set to obtain multiple annotated video data sets for each video data set; each annotated video data set corresponds to a character vector. All the labeled video data are divided into a video training set and a video test set according to a preset ratio; The fusion analysis is performed on each of the labeled video data in the video training set to obtain the fusion features of each of the labeled video data in the video training set. An emotion recognition model is constructed by training the emotion recognition model using the role vectors and fusion features corresponding to each of the annotated video data in the video training set, thereby obtaining the trained features of each of the annotated video data in the video training set. Import the real labels corresponding to the training features of each labeled video data in the video training set, and calculate the loss function based on the training features and real labels of all labeled video data in the video training set to obtain the loss function. The parameters of the emotion recognition model are updated according to the loss function to obtain the updated emotion recognition model. The updated emotion recognition model is tested using the video test set to obtain emotion recognition results; The process of performing fusion analysis on each of the labeled video data in the video training set to obtain the fusion features of each of the labeled video data in the video training set includes: Face detection processing is performed on each of the annotated video data in the video training set to obtain a face image sequence for each of the annotated video data in the video training set. Text extraction processing is performed on each of the annotated video data in the video training set to obtain the text sequence of each of the annotated video data in the video training set; The face image sequence and text sequence of each labeled video data in the video training set are encoded and analyzed using an encoder to obtain the image features and text features of each labeled video data in the video training set. The image features and text features of each labeled video data in the video training set are weighted and concatenated to obtain the fusion features of each labeled video data in the video training set.

2. The emotion recognition method according to claim 1, characterized in that, The process of using an encoder to encode and analyze the face image sequences and text sequences of each labeled video data in the video training set, and obtaining the corresponding image features and text features of each labeled video data in the video training set, includes: One-dimensional convolution processing is performed on the face image sequence and text sequence of each labeled video data in the video training set to obtain the initial image features and initial text features of each labeled video data in the video training set. The initial image features and initial text features of each labeled video data in the video training set are respectively processed by the trigonometric function encoding algorithm to obtain the image position information and text position information of each labeled video data in the video training set. The image location information and text location information of each labeled video data in the video training set are encoded using an encoder to obtain the image features and text features of each labeled video data in the video training set.

3. The emotion recognition method according to claim 1, characterized in that, The emotion recognition model includes a BERT encoder and a fully connected layer; The process of constructing an emotion recognition model, which involves training the emotion recognition model using the character vectors and fusion features corresponding to each of the annotated video data in the video training set, to obtain the trained features of each of the annotated video data in the video training set, includes: The weighted concatenation of the character vectors and fusion features corresponding to each labeled video data in the video training set is performed to obtain the concatenated features of each labeled video data in the video training set. Import the problem task vector, and then perform a weighted concatenation of the problem task vector with the concatenated features of each of the annotated video data in the video training set to obtain the features to be encoded for each of the annotated video data in the video training set. The BERT encoder is used to encode the features to be encoded for each of the labeled video data in the video training set, thereby obtaining the classification features for each of the labeled video data in the video training set. The fully connected layer is used to map the classification features of each labeled video data in the video training set to obtain the post-training features of each labeled video data in the video training set.

4. The emotion recognition method according to claim 1, characterized in that, The process of calculating the loss function based on the post-training features and ground truth labels of all labeled video data in the video training set includes: Based on the first equation, the loss function is calculated using the post-training features and real labels of all the labeled video data in the video training set, resulting in the loss function. The first equation is: , In particular, , In particular, , In particular, , in, For loss function, For the first The first video data Sentiment prediction labels for labeled video data For the first The first video data The true labels of the annotated video data The total number of video data. This represents the total number of labeled video data. For the first The first video data The sentiment probability of labeled video data For the first The first video data The normalization term for the labeled video data. , , and All are learnable parameters. For the first The first video data Post-training features of labeled video data.

5. An emotion recognition device, characterized in that, include: The annotation module is used to import multiple video data and annotate each of the video data to obtain multiple annotated video data for each of the video data. Each of the annotated video data points corresponds to a character vector; The partitioning module is used to divide all the labeled video data into a video training set and a video test set according to a preset ratio. The fusion analysis module is used to perform fusion analysis on each of the labeled video data in the video training set to obtain the fusion features of each of the labeled video data in the video training set. The model training module is used to construct an emotion recognition model. It trains the emotion recognition model using the role vectors and fusion features corresponding to each labeled video data in the video training set, and obtains the trained features of each labeled video data in the video training set. The loss function calculation module is used to import the real labels corresponding to the training features of each labeled video data in the video training set, and calculate the loss function based on the training features and real labels of all labeled video data in the video training set to obtain the loss function. The parameter update module is used to update the parameters of the emotion recognition model according to the loss function to obtain the updated emotion recognition model. The recognition result acquisition module is used to test the updated emotion recognition model using the video test set to obtain the emotion recognition result; The fusion analysis module is specifically used for: Face detection processing is performed on each of the annotated video data in the video training set to obtain a face image sequence for each of the annotated video data in the video training set. Text extraction processing is performed on each of the annotated video data in the video training set to obtain the text sequence of each of the annotated video data in the video training set; The face image sequence and text sequence of each labeled video data in the video training set are encoded and analyzed using an encoder to obtain the image features and text features of each labeled video data in the video training set. The image features and text features of each labeled video data in the video training set are weighted and concatenated to obtain the fusion features of each labeled video data in the video training set.

6. The emotion recognition device according to claim 5, characterized in that, The fusion analysis module is specifically used for: One-dimensional convolution processing is performed on the face image sequence and text sequence of each labeled video data in the video training set to obtain the initial image features and initial text features of each labeled video data in the video training set. The initial image features and initial text features of each labeled video data in the video training set are respectively processed by the trigonometric function encoding algorithm to obtain the image position information and text position information of each labeled video data in the video training set. The image location information and text location information of each labeled video data in the video training set are encoded using an encoder to obtain the image features and text features of each labeled video data in the video training set.

7. An emotion recognition system, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the emotion recognition method as described in any one of claims 1 to 4.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the emotion recognition method as described in any one of claims 1 to 4.