Multi-modal emotion recognition method based on image and character early fusion

By adopting early fusion methods of images and text in multimodal emotion recognition, combining pre-trained models and multi-sample loss training, the shortcomings of traditional methods in image-text paired feature analysis are solved, and efficient and accurate emotion recognition effect is achieved.

CN120123985APending Publication Date: 2025-06-10XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510283481.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Traditional multimodal emotion recognition methods ignore the use of image-text paired features, resulting in insufficient accuracy and efficiency when analyzing emotional information in image-text paired features.

Method used

A multimodal emotion recognition method based on early fusion of images and text is adopted, and the pre-trained visual and language transformer models ViLT and VAuLT are integrated into early, and combined with multi-sample loss training, the recognition ability of the model is improved.

Benefits of technology

It realizes the rapid and accurate identification of emotional information in image-text paired features, improves the performance and generalization capabilities of the model, and is suitable for the analysis of large-scale multimodal emotion data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123985A_ABST
    Figure CN120123985A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal emotion recognition method based on image and character early fusion. The method comprises the following steps: preprocessing a multi-modal emotion data set to obtain a data sequence of the multi-modal emotion data set; establishing an interactive fusion model based on image and character multi-modal emotion data by using a data sequence of the multi-modal emotion data set; performing multi-sample loss training on an interactive fusion model based on image and character multi-modal emotion data by using a data sequence of the multi-modal emotion data set; and inputting test data by using the interactive fusion model based on the image and character multi-modal emotion data to obtain a multi-modal emotion recognition result, and analyzing model performance. According to the method, the ability of visual capture and context information understanding is improved, the model is allowed to adapt to knowledge in a specific field, generalization is improved through early fusion, and the method is superior to an existing method in the aspects of sentiment analysis and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence, computer vision, and affective computing, and particularly relates to a multimodal emotion recognition method based on early fusion of images and text. Background Art

[0002] Emotions and feelings are crucial for understanding and predicting human behavior, affecting all aspects of decision-making, communication, and social interaction. In the context of multimodal data (such as images and text) from social media, the analysis of emotions and feelings provides profound insights into aspects such as cultural diversity, mental health, and consumer behavior. The work tasks of multimodal data analysis have put forward increasingly high requirements, and traditional multimodal emotion recognition ignores the use of image-text paired features. There is a greater need for a fast and reliable method to analyze the emotion information in image-text paired features. Summary of the Invention

[0003] To solve the deficiencies of existing image-text multimodal emotion recognition methods, the purpose of the present invention is to provide a multimodal emotion recognition method based on early fusion of images and text. Through the early fusion of two image and text models and training with multi-sampled dropout (MSD), the purpose of accurately recognizing the emotional information of image-text pairs is achieved.

[0004] To achieve the above purpose, the present invention adopts the following technical solutions:

[0005] A multimodal emotion recognition method based on early fusion of images and text, comprising the following steps:

[0006] S1. Preprocess the multimodal emotion dataset to obtain the data sequence of the multimodal emotion dataset;

[0007] S2. Use the data sequence of the multimodal emotion dataset to establish an interactive fusion model based on image and text multimodal emotion data;

[0008] S3: Use the data sequence of the multimodal emotion dataset to perform multi-sampled dropout training on the interactive fusion model based on image and text multimodal emotion data;

[0009] S4: Use the interactive fusion model based on image and text multimodal emotion data, input test data to obtain the multimodal emotion recognition result and analyze the model performance.

[0010] In S1, preprocessing the multimodal emotion dataset to obtain the data sequence of the multimodal emotion dataset includes the following steps:

[0011] S11: The publicly available multimodal benchmark dataset MSED is selected as the source of the multimodal emotion dataset. The dataset contains 9,190 image, title, and description combinations. The dataset is divided into 6127 training instances, 1021 validation instances, and 2042 test instances.

[0012] S12: performing face detection, face normalization and data enhancement preprocessing operations on the images in the multimodal benchmark dataset MSED in sequence to obtain an image data sequence;

[0013] S13: Perform noise removal and text standardization preprocessing on the titles and descriptions in the multimodal benchmark dataset MSED to obtain title and description data sequences.

[0014] In S2, an interactive fusion model based on image and text multimodal emotion data is established using the data sequence of the multimodal emotion data set, as follows:

[0015] S21: using the pre-trained vision and language transformer model ViLT, the pre-trained vision and language model ViLT uses the bert-base-uncased word segmenter to process the title and description data sequence to obtain a text input vector, using the ViT-B / 32 transformer model to process the input image to obtain an image input vector, using the image input vector and the text input vector to splice into an early fusion vector in the order of picture and text, and processing the early fusion vector with a transformer encoder and a pooling layer;

[0016] S22: The pre-trained vision and language transformer model VAuLT is used. The vision and language transformer model VAuLT uses a modified version of the vision and language transformer model ViLT. Compared with the vision and language transformer model ViLT, the vision and language transformer model VAuLT includes an additional text encoder. The bertweet-base word segmenter is used to process the input text data sequence to obtain the text input vector. The remaining parameters are the same as the pre-trained vision and language model ViLT described in S21. The input vector is processed by the transformer encoder and the pooling layer.

[0017] S23: Connect the pooling layer outputs of the pre-trained vision and augmented language transformer model VAuLT and the pre-trained vision and language transformer model ViLT to obtain the feature vector of the early fusion vector, and connect the fully connected layers to build an interactive fusion model based on image and text multimodal emotion data.

[0018] S21, comprising the following steps:

[0019] S211: The bert-base-uncased tokenizer selects the title and description data sequences obtained by noise removal and text standardization preprocessing, regards the text in the title and description as a word sequence, encodes the word sequence to obtain a title vector and a description vector composed of multiple word vectors, splices the sequence head CLS as a word vector at the beginning of the title vector as a new title vector, splices the sequence segmentation SEP as a word vector at the beginning and end of the description vector as a new title vector, and inserts the position code into the first position of each word vector code according to the order of the words in the original text, and finally obtains the text input vector in the order of sequence head, title vector, sequence segmentation, description vector, and sequence segmentation;

[0020] S212: The ViT-B / 32 transformer model selects an image data sequence obtained by performing preprocessing operations of face detection, face normalization and data enhancement, in which the image size is 612*408*3. The image is divided into blocks of a fixed size, each block is flattened and linearly mapped to obtain a block vector, the sequence header CLS is spliced ​​in front of the block vector, and the position code is spliced ​​to the first bit of each block vector code to obtain the image input sequence.

[0021] The bertweet-base tokenizer uses a tokenizer optimized for Twitter data. The bertweet-base tokenizer is more adaptable to noisy words and abbreviations on social media.

[0022] In S3, the interactive fusion model based on image and text multimodal emotion data is trained with multi-sample loss by using the data sequence of the multimodal emotion data set, as follows:

[0023] S31: Select a training instance of the multimodal benchmark dataset MSED and input it into the interactive fusion model based on image and text multimodal emotion data to obtain a feature vector of an early fusion vector, and use three random loss layers with different masks to perform random loss processing on the feature vector of the early fusion vector to obtain a loss feature vector;

[0024] S33: inputting the loss feature vectors into the fully connected layer respectively, and sharing the weight parameters of the fully connected layer to obtain the fully connected layer output;

[0025] S34: Apply the rectified linear unit ReLU activation function to calculate the activation value of the fully connected layer output, and use the binary cross entropy loss with logits to calculate the loss value of each fully connected layer output;

[0026] S35: Summarize all loss values ​​of the loss feature vector and calculate their average, update the weights according to the average loss value and repeat the above training.

[0027] In S4, the interactive fusion model based on image and text multimodal emotion data is used to input test data to obtain multimodal emotion recognition results and analyze model performance, as follows:

[0028] S41: Select a test instance of the multimodal benchmark dataset MSED and input it into the interactive fusion model based on image and text multimodal emotion data described in S2 to obtain an output result;

[0029] S42: Compare the output results with the standard results of the test set to calculate standard evaluation indicators of the interactive fusion model based on image and text multimodal emotion data, including precision P, recall R and F1 score to evaluate model performance.

[0030] Compared with the existing invention, the present invention has the following advantages:

[0031] 1. The pre-trained vision and language transformer model ViLT and the pre-trained vision and enhanced language transformer model VAuLT are used to build an interactive fusion model based on image and text multimodal emotion data. Convolution calculation is not used. When the data set is large, the response speed is faster and the accuracy is higher;

[0032] 2. An interactive fusion model based on multimodal sentiment data of images and texts, using a unified architecture of BERT-based text encoder and ViT image encoder in its backbone. It helps the model recognize long-term dependencies of context and establish internal relationships between images and context;

[0033] 3. A multi-sample loss training strategy is adopted to make the training speed of the base model faster and improve the generalization, thereby improving the overall performance of the model. Binary cross entropy loss with logits is used to calculate the loss of each repeated feature vector. Training is accelerated by evaluating multiple models simultaneously, thereby promoting faster convergence. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 Flow chart of the present invention DETAILED DESCRIPTION

[0035] The following describes the embodiments of the present application, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0036] likeFigure 1 As shown, the present invention provides a multimodal emotion recognition method based on early fusion of images and texts, comprising the following steps:

[0037] S1. Preprocessing the multimodal emotion dataset to obtain a data sequence of the multimodal emotion dataset;

[0038] S2, using the data sequence of the multimodal emotion data set to establish an interactive fusion model based on image and text multimodal emotion data;

[0039] S3: using the data sequence of the multimodal emotion data set to perform multi-sample loss training on the interactive fusion model based on image and text multimodal emotion data;

[0040] S4: Using the interactive fusion model based on image and text multimodal emotion data, input test data to obtain multimodal emotion recognition results and analyze model performance.

[0041] In S1, preprocessing the multimodal emotion dataset to obtain a data sequence of the multimodal emotion dataset includes the following steps:

[0042] S11: The publicly available multimodal benchmark dataset MSED is selected as the source of the multimodal emotion dataset. The dataset contains 9,190 image, title, and description combinations. The dataset is divided into 6127 training instances, 1021 validation instances, and 2042 test instances.

[0043] S12: performing face detection, face normalization and data enhancement preprocessing operations on the images in the multimodal benchmark dataset MSED in sequence to obtain an image data sequence;

[0044] S13: Perform noise removal and text standardization preprocessing on the titles and descriptions in the multimodal benchmark dataset MSED to obtain title and description data sequences.

[0045] In the embodiment, S1 is as follows:

[0046] The publicly available multimodal benchmark dataset MSED is selected to train and test the model of the present invention. MSED contains 9,190 text-image combinations collected from various social media (including Twitter, Getty Image and Flickr). The dataset contains 6127 training instances, 1021 verification instances and 2042 test instances. The images are mainly subjected to geometric transformations, including flipping, rotation, cropping, deformation, scaling and other operations to achieve image enhancement. The Viola-Jones object detector proposed by Viola and Jones in 2001 is used to extract faces from complex images. When preprocessing the description and title data, the description and title are first cleaned to remove extra spaces, line breaks and irrelevant content such as comments. Subsequently, a natural language processing system is used to detect and correct grammatical errors to ensure smooth sentences. Then, irrelevant words are screened and eliminated through a vocabulary or word embedding technology. Further, an N-gram model is used to identify and delete repeated words. At the same time, ambiguous words are eliminated with the help of word meaning analysis. Finally, the text of the description and title is formatted to strictly comply with specific format specifications.

[0047] In S2, an interactive fusion model based on image and text multimodal emotion data is established using the data sequence of the multimodal emotion data set, as follows:

[0048] S21: using the pre-trained vision and language transformer model ViLT, the pre-trained vision and language model ViLT uses the bert-base-uncased word segmenter to process the title and description data sequence to obtain a text input vector, using the ViT-B / 32 transformer model to process the input image to obtain an image input vector, using the image input vector and the text input vector to splice into an early fusion vector in the order of picture and text, and processing the early fusion vector with a transformer encoder and a pooling layer;

[0049] S22: The pre-trained vision and language transformer model VAuLT is used. The vision and language transformer model VAuLT uses a modified version of the vision and language transformer model ViLT. Compared with the vision and language transformer model ViLT, the vision and language transformer model VAuLT includes an additional text encoder. The bertweet-base word segmenter is used to process the input text data sequence to obtain the text input vector. The remaining parameters are the same as the pre-trained vision and language model ViLT described in S21. The input vector is processed by the transformer encoder and the pooling layer.

[0050] S23: Connect the pooling layer outputs of the pre-trained vision and augmented language transformer model VAuLT and the pre-trained vision and language transformer model ViLT to obtain the feature vector of the early fusion vector, and connect the fully connected layers to build an interactive fusion model based on image and text multimodal emotion data.

[0051] S21, comprising the following steps:

[0052] S211: The bert-base-uncased tokenizer selects the title and description data sequences obtained by removing noise and preprocessing text standardization, regards the text in the title and description as a word sequence, encodes the word sequence to obtain a title vector and a description vector composed of multiple word vectors, splices the sequence head CLS as a word vector at the beginning of the title vector as a new title vector, splices the sequence segmentation SEP as a word vector at the beginning and end of the description vector as a new title vector, and inserts the position code into the first position of each word vector code according to the order of the words in the original text, and finally obtains the text input vector in the order of sequence head, title vector, sequence segmentation, description vector, and sequence segmentation;

[0053] S212: The ViT-B / 32 transformer model selects an image data sequence obtained by performing preprocessing operations of face detection, face normalization and data enhancement, in which the image size is 612*408*3. The image is divided into blocks of a fixed size, each block is flattened and linearly mapped to obtain a block vector, the sequence header CLS is spliced ​​in front of the block vector, and the position code is spliced ​​to the first bit of each block vector code to obtain the image input sequence.

[0054] S22, comprising the following steps:

[0055] S221: The bertweet-base tokenizer uses a tokenizer optimized for Twitter data. The bertweet-base tokenizer is more adaptable to noisy words and abbreviations on social media.

[0056] In the embodiment, S2 is as follows:

[0057] ViLT uses bert-base-uncased tokenizer to process input text and ViT-B / 32 transformer model to process input images. It uses weights pre-trained on ImageNet, including 12 transformer layers, with a multi-layer perceptron (MLP) size of 3072, a hidden layer size of 768, a patch size of 32, 12 attention heads, and the vilt-b32-mlm checkpoint of the ViLT model implemented by HuggingFace.

[0058] The VAuLT model uses a modified version of the ViLT architecture that contains an additional text encoder module. VAuLT solves this problem by replacing the language embedding input of ViLT with language features from LLMs with more language data varieties, which can also be selected to better meet the needs of downstream applications. To tokenize the input text, the bertweet-base tokenizer is used, using the vilt-b32-mlm checkpoint and vinai / bertweet-base checkpoint of the ViLT model using HuggingFace.

[0059] In S3, the interactive fusion model based on image and text multimodal emotion data is trained with multi-sample loss by using the data sequence of the multimodal emotion data set, as follows:

[0060] S31: Select a training instance of the multimodal benchmark dataset MSED and input it into the interactive fusion model based on image and text multimodal emotion data to obtain a feature vector of an early fusion vector, and use three random loss layers with different masks to perform random loss processing on the feature vector of the early fusion vector to obtain a loss feature vector;

[0061] S32: inputting the loss feature vectors into the fully connected layer respectively, and sharing the weight parameters of the fully connected layer to obtain the fully connected layer output;

[0062] S33: Apply the rectified linear unit ReLU activation function to calculate the activation value of the fully connected layer output, and use the binary cross entropy loss with logits to calculate the loss value of each fully connected layer output;

[0063] S34: Summarize all loss values ​​of the loss feature vector and calculate their average, update the weights according to the average loss value and repeat the above training.

[0064] In the embodiment, S3 is as follows:

[0065] During training, three random loss layers with different sample sizes are used, each with a different mask. The feature vectors of the multimodal early fusion representation are replicated after applying the random loss layer, and the weights are shared between these replicated fully connected layers. In addition, the rectified linear unit (ReLU) activation function is applied to calculate the activation value, and then the binary cross entropy loss with logits (BCEWithLogitsLoss) is used to calculate the loss of each repeated feature vector.

[0066]

[0067] Where: LOSS is the loss value, y is the real data, the value is 0 or 1, is the predicted value of the model, ranging from (0,1) (the output value of ReLU), and log is the natural logarithm;

[0068] To calculate the final loss, the losses of all samples are aggregated and their average is calculated. The weight parameters are updated according to the loss values ​​of the samples, and the above steps are repeated multiple times.

[0069] In addition, in order to enhance the learning and adaptation ability of the interactive fusion model based on image and text multimodal emotion data to specific knowledge, several hyperparameters (including training batch size, test batch size, learning rate, random loss and number of training rounds) were fine-tuned, and the grid search method based on the validation dataset was used to select the best hyperparameters. The hyperparameter search space is shown in the table:

[0070]

[0071]

[0072] By fine-tuning the interactive fusion model based on image and text multimodal emotion data, the model can learn and adapt to knowledge in specific fields, improving the performance of model recognition;

[0073] In S4, the interactive fusion model based on image and text multimodal emotion data is used to input test data to obtain multimodal emotion recognition results and analyze model performance, as follows:

[0074] S41: Select a test instance of the multimodal benchmark dataset MSED and input it into the interactive fusion model based on image and text multimodal emotion data described in S2 to obtain an output result;

[0075] S42: Compare the output results with the standard results of the test set to calculate standard evaluation indicators of the interactive fusion model based on image and text multimodal emotion data, including precision P, recall rate R and F1 score to evaluate model performance.

[0076] In the embodiment, S4 is as follows:

[0077] In order to verify the effectiveness of the model, the test examples of the multimodal benchmark dataset MSED were selected as input to the model. The output results were compared with the standard results to calculate the precision P, recall R and F1 score of the output results. The precision refers to the proportion of samples predicted by the model to be positive that are actually positive. The formula is:

[0078]

[0079] Recall rate refers to the proportion of samples that are actually positive that are correctly predicted as positive by the model. The formula is:

[0080]

[0081] The F1 score is the harmonic mean of precision and recall, and the formula is:

[0082]

[0083] In addition, the calculated standard evaluation indicators were compared with the current advanced benchmark methods (Multimodal Transformers, M3GAT, BERT+ResNet). The results showed that the interactive fusion model based on image and text multimodal emotion data proposed in the present invention is significantly better than the current benchmark methods.

[0084] Finally, the whole process is expressed as follows:

[0085] T=[t CLS ;t 1 ;…;t M ;t SEP ;c 1 ;…c M ;t SEP ]

[0086] I=[p CLS ;p 1 ;…;p M ]

[0087] [T; I]

[0088] V pool =(V m (V pr [T; I]))

[0089] V Apool =(VA m (VA pr [T; t]))

[0090]

[0091] MSD(D 0 (IFV)

[0092] O=Avg.(FC(MSD(D 0 (IFV)

[0093] Where T and I represent the text input vector and image input vector, T contains the title vector and the description vector, t CLS Indicates the sequence header of the title, t i Represents each word in the title, t SEP is the sequence separation between title and description, c i Represents each word in the description, p in I CLS Indicates the sequence header of the picture, p i represents each block of the input image, [T; I} represents the early fusion vector;

[0094] V pr represents the tramsformer encoder processing of ViLT, V pool Represents the pooling layer output of ViLT, VA pr represents the tramsformer encoder processing of VAuLT, VA Pool represents the pooling layer output of VAuLT, IFV represents the feature vector of the early fusion vector that combines the pooling layer outputs, Represents the concatenation of two vectors;

[0095] MSD(D 0 The feature vector (IFV) representing the early fusion vector is input into MSD for multi-sample loss training. The output of MSD is connected to the fully connected layer FC, and the obtained outputs are averaged to obtain the final prediction O of this method;

[0096] The implementation of a multimodal emotion recognition method based on early fusion of images and texts in the present invention is conducive to improving the robustness and generalization ability of recognition, can adapt to knowledge in specific fields, reduce training time loss, improve operation accuracy, and save time costs, thereby reducing training time costs, and has important application value for quickly and accurately identifying emotions.

[0097] The above description of the present invention is provided to enable any person of ordinary skill in the art to implement or use the present invention. Various modifications to the present invention are obvious to those of ordinary skill in the art, and the general principles defined herein may be applied to other variations without departing from the scope of protection of the present invention. Therefore, the present invention is not limited to the examples and designs described herein, but is consistent with the widest range of principles and novel features disclosed herein.

Claims

1. A multimodal emotion recognition method based on early fusion of images and texts, characterized in that: The following steps are involved: S1. Preprocessing the multimodal emotion dataset to obtain a data sequence of the multimodal emotion dataset; S2, using the data sequence of the multimodal emotion data set to establish an interactive fusion model based on image and text multimodal emotion data; S3: using the data sequence of the multimodal emotion data set to perform multi-sample loss training on the interactive fusion model based on image and text multimodal emotion data; S4: Using the interactive fusion model based on image and text multimodal emotion data, input test data to obtain multimodal emotion recognition results and analyze model performance.

2. The multimodal emotion recognition method based on early fusion of image and text according to claim 1, characterized in that: In S1, preprocessing the multimodal emotion dataset to obtain a data sequence of the multimodal emotion dataset includes the following steps: S11: The publicly available multimodal benchmark dataset MSED is selected as the source of the multimodal emotion dataset. The dataset contains 9,190 image, title, and description combinations. The dataset is divided into 6127 training instances, 1021 validation instances, and 2042 test instances. S12: performing face detection, face normalization and data enhancement preprocessing operations on the images in the multimodal benchmark dataset MSED in sequence to obtain an image data sequence; S13: Perform noise removal and text standardization preprocessing on the titles and descriptions in the multimodal benchmark dataset MSED to obtain title and description data sequences.

3. The multimodal emotion recognition method based on early fusion of images and text according to claim 1, characterized in that: In S2, an interactive fusion model based on image and text multimodal emotion data is established using the data sequence of the multimodal emotion data set, as follows: S21: using the pre-trained vision and language transformer model ViLT, the pre-trained vision and language model ViLT uses the bert-base-uncased word segmenter to process the title and description data sequence to obtain a text input vector, using the ViT-B / 32 transformer model to process the input image to obtain an image input vector, using the image input vector and the text input vector to splice into an early fusion vector in the order of picture and text, and processing the early fusion vector with a transformer encoder and a pooling layer; S22: The pre-trained vision and language transformer model VAuLT is used. The vision and language transformer model VAuLT uses a modified version of the vision and language transformer model ViLT. Compared with the vision and language transformer model ViLT, the vision and language transformer model VAuLT includes an additional text encoder. The bertweet-base word segmenter is used to process the input text data sequence to obtain the text input vector. The remaining parameters are the same as the pre-trained vision and language model ViLT described in S21. The input vector is processed by the transformer encoder and the pooling layer. S23: Connect the pooling layer outputs of the pre-trained vision and augmented language transformer model VAuLT and the pre-trained vision and language transformer model ViLT to obtain the feature vector of the early fusion vector, and connect the fully connected layers to build an interactive fusion model based on image and text multimodal emotion data.

4. The multimodal emotion recognition method based on early fusion of images and text according to claim 3 is characterized in that: The specific steps of S21 are as follows: S211: The bert-base-uncased tokenizer selects the title and description data sequences obtained by removing noise and preprocessing text standardization, regards the text in the title and description as a word sequence, encodes the word sequence to obtain a title vector and a description vector composed of multiple word vectors, splices the sequence head CLS as a word vector at the beginning of the title vector as a new title vector, splices the sequence segmentation SEP as a word vector at the beginning and end of the description vector as a new title vector, and inserts the position code into the first position of each word vector code according to the order of the words in the original text, and finally obtains the text input vector in the order of sequence head, title vector, sequence segmentation, description vector, and sequence segmentation; S212: The ViT-B / 32 transformer model selects an image data sequence obtained by performing preprocessing operations of face detection, face normalization and data enhancement, in which the image size is 612*408*3. The image is divided into blocks of a fixed size, each block is flattened and linearly mapped to obtain a block vector, the sequence header CLS is spliced ​​in front of the block vector, and the position code is spliced ​​to the first bit of each block vector code to obtain the image input sequence.

5. The multimodal emotion recognition method based on early fusion of images and text according to claim 3 is characterized in that: The bertweet-base tokenizer uses a tokenizer optimized for Twitter data. The bertweet-base tokenizer is more adaptable to noisy words and abbreviations on social media.

6. The multimodal emotion recognition method based on early fusion of images and text according to claim 1, characterized in that: In S3, the interactive fusion model based on image and text multimodal emotion data is trained with multi-sample loss by using the data sequence of the multimodal emotion data set, as follows: S31: Select a training instance of the multimodal benchmark dataset MSED and input it into the interactive fusion model based on image and text multimodal emotion data to obtain a feature vector of an early fusion vector, and use three random loss layers with different masks to perform random loss processing on the feature vector of the early fusion vector to obtain a loss feature vector; S32: inputting the loss feature vectors into the fully connected layer respectively, and sharing the weight parameters of the fully connected layer to obtain the fully connected layer output; S33: Apply the rectified linear unit ReLU activation function to calculate the activation value of the fully connected layer output, and use the binary cross entropy loss with logits to calculate the loss value of each fully connected layer output; S34: Summarize all loss values ​​of the loss feature vector and calculate their average, update the weights according to the average loss value and repeat the above training.

7. The multimodal emotion recognition method based on early fusion of images and text according to claim 1, characterized in that: In S4, the interactive fusion model based on image and text multimodal emotion data is used to input test data to obtain multimodal emotion recognition results and analyze model performance, as follows: S41: Select a test instance of the multimodal benchmark dataset MSED and input it into the interactive fusion model based on image and text multimodal emotion data described in S2 to obtain an output result; S42: Compare the output results with the standard results of the test set to calculate standard evaluation indicators of the interactive fusion model based on image and text multimodal emotion data, including precision P, recall R and F1 score to evaluate model performance.