A multi-modal sentiment analysis system and method based on circuit generation

Through a multimodal sentiment analysis system based on loop generation, feature extraction and mutual attention alignment processing are used to solve the problem of information missing in multimodal sentiment analysis, achieve efficient sentiment analysis in the absence of information, and improve the robustness and accuracy of the model.

CN116861181BActive Publication Date: 2025-10-10BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310706053.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-14
Publication Date
2025-10-10
Estimated Expiration
2043-06-14

AI Technical Summary

Technical Problem

Existing multimodal sentiment analysis methods cannot effectively express multimodal information at the feature level fusion, and the decision-level fusion cost is high and contains a lot of redundant information, resulting in a decrease in model effectiveness when multimodal information is missing.

Method used

A multimodal sentiment analysis system based on loop generation is adopted. Through feature extraction, loop generation feature fusion and mutual attention alignment processing, combined with multi-task sentiment analysis, accurate analysis of multimodal features is achieved.

Benefits of technology

It reduces the decline in model effectiveness when multimodal information is missing, improves the robustness and accuracy of sentiment analysis, makes full use of unimodal feature information, and improves the overall performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116861181B_ABST
    Figure CN116861181B_ABST
Patent Text Reader

Abstract

The application provides a multi-modal sentiment analysis system based on loop generation, comprising a feature extraction subsystem, which is used for storing multi-modal original data, performing feature extraction on the multi-modal original data, and obtaining multi-modal features; a loop generation feature fusion subsystem, which is used for fusing the multi-modal features to obtain multi-modal fusion features, and generating modes to obtain multi-modal generated features; fusing the multi-modal features, the multi-modal generated features and the multi-modal fusion features to obtain mixed features; performing mutual attention alignment processing on the multi-modal generated features and the multi-modal features to obtain reinforced single-modal features; and a sentiment analysis subsystem, which is used for performing sentiment analysis on the mixed features and the reinforced single-modal features. Through the method provided by the application, accurate analysis of multi-modal sentiment is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision. Background Art

[0002] Sentiment analysis, also known as opinion mining and subjective analysis, is aimed at obtaining corresponding sentiment analysis results for a text that expresses human emotions through computer algorithm analysis and mining. The modality of information refers to the source or form of information, such as text, video, voice, etc. Multimodal data usually refers to data that includes multiple modalities, in which there are often various information associations between different modalities. In the real world, information generally exists in a multimodal manner, and a single modal information often cannot fully describe things. With the rapid development of social networks, people can express their emotions and opinions through pictures, text and videos on the Internet. Multiple modalities can complement each other. Analyzing emotions in multimodal data is an important issue in the field of sentiment analysis. The fusion of multimodal information features is the key to the multimodal sentiment analysis model. Currently, there are two main forms of multimodal sentiment analysis methods, feature-level fusion and decision-level fusion:

[0003] Feature-level fusion is also called early fusion. This solution integrates features immediately after feature extraction, usually by simply connecting their representations. Mainly, the feature vectors of each modality, such as text feature vectors, image feature vectors, etc., are fused into a multimodal feature vector through a feature fusion unit, and then a decision analysis is performed on the combined features to output the sentiment analysis results. Feature-level fusion fuses multimodal information by connecting multimodal features, which can achieve the effect of multimodal fusion with a smaller amount of computation. However, the features extracted from multimodal information come from different feature spaces and have large differences in time and semantic dimensions. Simply connecting features cannot well express multimodal information, and when there are omissions or noise in the multimodal information, the model effect will drop sharply.

[0004] Decision-level fusion, also known as late-stage fusion, performs independent analysis on each modality after feature extraction. A separate analysis model is established for each modality, and the analysis results are fused into a decision vector to obtain the final decision. Decision-level fusion uses a method that performs a separate analysis on each modality before fusion. This method can map multimodal features into the same space, but requires separate training for each classifier, resulting in high model training costs. It also amplifies redundant information between modalities, covering the truly effective information. Summary of the Invention

[0005] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.

[0006] To this end, the purpose of the present invention is to propose a multimodal sentiment analysis system based on loop generation for accurate analysis of multimodal sentiment.

[0007] To achieve the above objectives, the first embodiment of the present invention proposes a multimodal sentiment analysis system based on loop generation, comprising:

[0008] A feature extraction subsystem is used to store multimodal raw data and perform feature extraction on the multimodal raw data to obtain multimodal features;

[0009] A loop generation feature fusion subsystem is used to fuse the multimodal features to obtain a multimodal fusion feature, and to generate a modality to obtain a multimodal generation feature; to fuse the multimodal feature with the multimodal generation feature and the multimodal fusion feature to obtain a hybrid feature; and to perform mutual attention alignment processing on the multimodal generation feature and the multimodal feature to obtain an enhanced single-modal feature;

[0010] The sentiment analysis subsystem is used to perform sentiment analysis on the mixed features and the enhanced unimodal features.

[0011] In addition, the multimodal sentiment analysis system based on loop generation according to the above embodiment of the present invention may also have the following additional technical features:

[0012] Furthermore, in one embodiment of the present invention, the multimodal original data includes images, audio, and text, and the multimodal features include image features, audio features, and text features.

[0013] Furthermore, in one embodiment of the present invention, the feature extraction subsystem includes:

[0014] A multimodal feature database module, used for storing the multimodal original data and the multimodal features;

[0015] The feature extraction module is used to extract features from the multimodal raw data to obtain multimodal features.

[0016] Furthermore, in one embodiment of the present invention, the loop generation feature fusion subsystem includes:

[0017] The loop generation module is used to fuse the features of different modalities to obtain multimodal fusion features, and to generate modalities to obtain multimodal generation features;

[0018] A feature fusion module, configured to fuse the multimodal features, the multimodal generated features, and the multimodal fusion features;

[0019] A single-modal reinforcement module is configured to construct a single-modal reinforced Transformer by using a mutual attention mechanism, perform mutual attention alignment processing on the multi-modal generated feature and the multi-modal feature, and obtain a reinforced single-modal feature.

[0020] Further, in an embodiment of the present application, the loop generation module is further configured to:

[0021] input the text modal feature X1 and the image feature X2 into the loop generation module;

[0022] generate a vector X'1 in a video modal feature space from the text modal feature X1 by using an encoder, and generate a vector X'2 in a text modal feature space from the image modal feature X2 by using the encoder;

[0023] perform error constraint on the X'1 and the X1 so that the X'1 and the X1 have minimum error, and perform error constraint on the X'2 and the X2 so that the X'2 and the X2 have minimum error;

[0024] select a feature vector of a middle layer of the encoder as a mixed vector X3, X4 and output the mixed vector X3, X4 to the feature fusion module.

[0025] Further, in an embodiment of the present application, the sentiment analysis subsystem further comprises:

[0026] a multi-task module configured to perform sentiment analysis on the mixed feature and the reinforced single-modal feature, and obtain a sentiment analysis result;

[0027] a user interaction module configured to provide an uploading and accessing information interface to a user, and display the sentiment analysis result to the user.

[0028] To achieve the above object, a second aspect embodiment of the present application provides a multi-modal sentiment analysis method based on loop generation, comprising:

[0029] obtain multi-modal original data, perform feature extraction on the multi-modal original data, and obtain multi-modal features;

[0030] fuse the multi-modal features to obtain multi-modal fusion features, perform modal generation to obtain multi-modal generated features, fuse the multi-modal features, the multi-modal generated features and the multi-modal fusion features to obtain mixed features, and perform mutual attention alignment processing on the multi-modal generated features and the multi-modal features to obtain reinforced single-modal features;

[0031] perform sentiment analysis on the mixed features and the reinforced single-modal features.

[0032] To achieve the above-mentioned purpose, the third aspect of the present invention proposes a computer device, characterized in that it includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements a multimodal sentiment analysis system based on loop generation as described above.

[0033] To achieve the above-mentioned purpose, the fourth aspect of the present invention proposes a computer-readable storage medium on which a computer program is stored, characterized in that when the computer program is executed by a processor, a multimodal sentiment analysis system based on loop generation as described above is implemented.

[0034] The multimodal sentiment analysis system based on loop generation proposed in an embodiment of the present invention utilizes a loop generation method to establish a multimodal feature fusion module, combines multimodal information, and obtains a feature fusion module that can reduce the decline in effect when multimodal information is missing, thereby realizing accurate analysis of multimodal sentiment. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0036] Figure 1 A schematic diagram of the process of a multimodal sentiment analysis system based on loop generation provided by an embodiment of the present invention.

[0037] Figure 2 Schematic diagram of a multimodal feature database of a multimodal sentiment analysis system based on loop generation provided by an embodiment of the present invention.

[0038] Figure 3 A schematic flow chart of a loop generation feature fusion subsystem of a multimodal sentiment analysis system based on loop generation provided in an embodiment of the present invention.

[0039] Figure 4 A schematic diagram of the flow of a loop generation module of a multimodal sentiment analysis system based on loop generation provided in an embodiment of the present invention.

[0040] Figure 5 A flow chart of a feature fusion mechanism provided by an embodiment of the present invention.

[0041] Figure 6 A schematic diagram of the multi-task module flow of a multimodal sentiment analysis system based on loop generation provided by an embodiment of the present invention.

[0042] Figure 7 A flowchart of a multimodal sentiment analysis method based on loop generation provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0043] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.

[0044] The following describes a multimodal sentiment analysis system based on loop generation according to an embodiment of the present invention with reference to the accompanying drawings.

[0045] Figure 1 A schematic diagram of the process of a multimodal sentiment analysis system based on loop generation provided by an embodiment of the present invention.

[0046] like Figure 1 As shown, the multimodal sentiment analysis system based on loop generation includes:

[0047] The feature extraction subsystem is used to store multimodal raw data and perform feature extraction on the multimodal raw data to obtain multimodal features;

[0048] Furthermore, in one embodiment of the present invention, the multimodal original data includes images, audio, and text, and the multimodal features include image features, audio features, and text features.

[0049] Furthermore, in one embodiment of the present invention, the feature extraction subsystem includes:

[0050] Multimodal feature database module, used to store multimodal raw data and multimodal features;

[0051] The feature extraction module is used to extract features from multimodal raw data to obtain multimodal features.

[0052] Specifically, the feature extraction subsystem provides data support for the loop generation feature fusion subsystem and the sentiment analysis subsystem. On the one hand, it provides complete data query, and on the other hand, it provides features extracted from information of different modalities for feature fusion and sentiment analysis.

[0053] The feature extraction subsystem includes a multimodal feature database to store multimodal raw data such as images, audio, and text, as well as the extracted multimodal feature data. The feature extraction subsystem also includes a feature extraction model to extract features of different modalities. Figure 2 shown.

[0054] The multimodal feature database mainly stores multimodal original data such as images, audio, text, and the extracted multimodal feature data.

[0055] The multi-modal raw data refers to image, audio and text data obtained after preliminary processing of the video. The image data is obtained by extracting specific frames from the video, the audio data is extracted from the video, and the text data is obtained by processing the audio through a text transcription model. The user can provide video data by himself, the system can automatically process data, and the user can also change the processing result by himself.

[0056] The multi-modal feature data refers to feature data obtained after feature extraction of the raw data, which is generally represented as a multi-dimensional data vector.

[0057] The feature extraction model can extract features from the original multi-modal data, providing a data basis for the downstream feature fusion system and the sentiment analysis system.

[0058] Among them, the image features can be extracted through a pre-trained model such as ResNet or ViT. According to the different languages used, the text data can be extracted through a pre-trained model such as BERT, BERT-CHINESE or ERNIE.

[0059] The loop generation feature fusion subsystem is used to fuse the multi-modal features to obtain multi-modal fusion features, and to generate the modal to obtain multi-modal generation features; the multi-modal features, the multi-modal generation features and the multi-modal fusion features are fused to obtain mixed features; the multi-modal generation features and the multi-modal features are processed through mutual attention alignment to obtain reinforced single-modal features;

[0060] Further, in an embodiment of the present application, the loop generation feature fusion subsystem comprises:

[0061] The loop generation module is used to fuse the features of different modalities to obtain multi-modal fusion features, and to generate the modal to obtain multi-modal generation features;

[0062] The feature fusion module is used to fuse the multi-modal features, the multi-modal generation features and the multi-modal fusion features;

[0063] The single-modal reinforcement module is used to construct a single-modal reinforcement Transformer using a mutual attention mechanism, to process the multi-modal generation features and the multi-modal features through mutual attention alignment, and to obtain reinforced single-modal features.

[0064] Further, in an embodiment of the present application, the loop generation module is further used to:

[0065] The text modal feature X1 and the image feature X2 are input into the loop generation module;

[0066] The encoder converts the text modality feature X1 into a vector X′1 in the video modality feature space, and converts the image modality feature X2 into a vector X′2 in the text modality feature space;

[0067] Perform error constraints on X′1 and X1 so that the error between X′1 and X1 is minimized; perform error constraints on X′2 and X2 so that the error between X′2 and X2 is minimized;

[0068] The feature vectors of the middle layer of the encoder are selected as mixed vectors X3 and X4 and output to the feature fusion module.

[0069] Specifically, the loop generation feature fusion subsystem is responsible for interacting with the feature extraction subsystem. It achieves multimodal feature fusion by fusing feature vectors from different feature spaces and different modalities and providing the fusion results to the sentiment analysis subsystem. In addition, the loop generation method strengthens the model's robustness to modality loss. The flow chart of the loop generation feature fusion subsystem is shown below. Figure 3 shown.

[0070] The present invention designs a feature fusion subsystem based on loop generation, which includes a loop generation module, a single-modality enhancement module, and a feature fusion module.

[0071] The loop generation module is responsible for the initial fusion of features of different modes and the generation of modes. After obtaining features from the feature extraction module,

[0072] In order to solve the problem of sensitivity to modality loss in current multimodal sentiment analysis models, this solution uses a loop generation module to process multimodality. The flow chart of the loop generation module is as follows: Figure 4 shown.

[0073] To map different modalities to different feature spaces, the loop generation module uses only the encoder structure in the Transformer to encode features, removing the decoder structure. An error constraint module is used to control the error between the generated features and the original features.

[0074] This module first generates the missing modal information, and then treats the generated modal as a certain original feature for processing, thereby reconstructing the missing modal information to a certain extent.

[0075] Figure 4 The middle process is the loop generation module of the text (T) modality and the image (V) modality. The process is as follows:

[0076] 1) Input text modality features X1 and image features X2 to the loop generation module;

[0077] 2) The encoder converts the text modality feature X1 into a vector X′1 in the video modality feature space, and converts the image modality feature X2 into a vector X′2 in the text modality feature space;

[0078] 3) Error constraints are applied to the features of X′1 and X1, and the error between the generated features and the original features is minimized by training the model; the same operation is performed on X2 and X′2;

[0079] 4) The multi-layer encoder can be regarded as a fusion process between modalities. The feature vectors of the middle layer are selected as the mixed vectors X3 and X4 and output to the feature fusion module.

[0080] 5) The features X′1 and X′2 generated after training can be considered to contain some information from other modalities, and also have a certain amount of information redundancy. In order to make the information from different modalities produce positive effects, the original features and generated features need to be input into the single-modal enhancement module.

[0081] The feature fusion module is responsible for further fusing the multimodal original features obtained in the feature loop generation module with the multimodal generated features and the multimodal fusion features. To achieve this goal, this module establishes a co-attention module and a BiLSTM module.

[0082] The mutual attention module adopts the mutual attention mechanism to align and process the feature information of different modalities, thereby achieving complementarity between information from different modalities and making full use of multimodal feature information.

[0083] like Figure 5 The following is a flowchart of the feature fusion mechanism. To enhance the model's ability to memorize context, the model integrates multiple BiLSTM modules to achieve feature information fusion.

[0084] The generated features obtained by the loop generation module can be considered to have obtained information from different modes. However, there will also be some redundant information between different modes. The information fusion between multiple modes may cause the superposition of redundant information, which in turn covers the valid information in the original mode.

[0085] In order to improve the effective information of unimodal features, this scheme sets up a unimodal enhancement module to remove redundant information in the features.

[0086] The mutual attention mechanism is adopted to construct a unimodal enhanced Transformer, and the generated features are aligned with the original features through mutual attention to obtain enhanced unimodal features, which are provided to the sentiment analysis subsystem for processing.

[0087] The sentiment analysis subsystem is used to perform sentiment analysis on mixed features and enhanced unimodal features.

[0088] Further, in an embodiment of the present application, the sentiment analysis subsystem further comprises:

[0089] a multi-task module for performing sentiment analysis on the mixed features and reinforced single-modal features to obtain sentiment analysis results;

[0090] a user interaction module for providing an upload and access information interface to the user and displaying the sentiment analysis results to the user.

[0091] Specifically, the processing flow of the user interaction module includes the multi-task module and the user interaction module.

[0092] The present application adopts a multi-task mechanism, and through the multi-task module, not only multi-modal complete information is subjected to sentiment analysis, but also single-modal information is subjected to separate sentiment analysis, so as to realize the universality of the sentiment analysis task, and the multi-task mechanism is also helpful to the improvement of the overall effect of the whole model, and the robustness of the model is improved as a whole.

[0093] As shown in Figure 6 Human emotions include various classification methods, and common methods include classification into two categories of positive and negative, or classification into five or seven different emotions according to different emotions. In order to classify multi-modal mixed information, a multi-classification module is arranged in this module to classify the information into a certain emotion.

[0094] The user interaction module is responsible for interacting with the user, providing an upload and access information interface to the user, and displaying the sentiment analysis results to the user.

[0095] The user uploads multi-modal information requiring sentiment analysis through the interactive interface, and after uploading, the module stores the multi-modal information data into the database, and after the sentiment analysis is completed, the generated analysis results are updated into the database.

[0096] The user can also query the sentiment analysis results through the interactive interface, and the interactive interface displays the results of sentiment analysis of different modalities and the results of sentiment analysis of all multi-modal information fusion to the user.

[0097] The multi-modal sentiment analysis system based on loop generation provided in the embodiment of the present application utilizes the loop generation method to establish a multi-modal feature fusion module, combines multi-modal information, obtains a feature fusion module capable of reducing the effect of decline when multi-modal information is missing, and realizes accurate analysis of multi-modal sentiment.

[0098] Compared with the existing technology, the advantages of the present invention are as follows: 1) The existing solution uses a feature-level fusion approach to complete multimodal mixing, which is less effective when multimodal information is missing. It requires users to provide all modal information, which increases the difficulty of use and the requirements for information quality. This solution utilizes a feature fusion model based on loop generation, which effectively supplements user information from an objective perspective and significantly improves the robustness and analytical capabilities of sentiment analysis. In the test of the model, in the case of modal missing, the performance of the model proposed in this proposal is less degraded than other models. 2) The existing solution only uses feature information after multimodal fusion for sentiment analysis, and does not use single-modal feature information for multi-task sentiment analysis. This proposal uses a multi-task mechanism and a single-modal enhancement mechanism to fully consider the effective information and redundant information in a single modality, improve the effect of single-modal information, and improve the overall effect of the model and the overall robustness of the model. In the test of the model, the model proposed in this proposal has a certain performance improvement compared with other models.

[0099] Figure 7 A schematic diagram of a multimodal sentiment analysis method based on loop generation provided in an embodiment of the present invention.

[0100] like Figure 7 As shown, the multimodal sentiment analysis method based on loop generation includes:

[0101] S101: Acquire multimodal raw data, and perform feature extraction on the multimodal raw data to obtain multimodal features;

[0102] S102: fusing the multimodal features to obtain a multimodal fusion feature, and generating a modality to obtain a multimodal generation feature; fusing the multimodal feature with the multimodal generation feature and the multimodal fusion feature to obtain a hybrid feature; performing mutual attention alignment processing on the multimodal generation feature and the multimodal feature to obtain an enhanced unimodal feature;

[0103] S103: Perform sentiment analysis on the mixed features and the enhanced unimodal features.

[0104] To achieve the above-mentioned purpose, the third aspect of the present invention proposes a computer device, which is characterized in that it includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the multimodal sentiment analysis system based on loop generation as described above.

[0105] To achieve the above-mentioned purpose, the fourth aspect of the present invention proposes a computer-readable storage medium on which a computer program is stored, characterized in that when the computer program is executed by a processor, the multimodal sentiment analysis system based on loop generation as described above is implemented.

[0106] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0107] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0108] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limiting the present invention. A person skilled in the art may change, modify, replace, and modify the above embodiments within the scope of the present invention.

Claims

1. A multimodal sentiment analysis system based on loop generation, characterized in that: include: A feature extraction subsystem, configured to store multimodal raw data and perform feature extraction on the multimodal raw data to obtain multimodal features; wherein the multimodal raw data includes images, audio, and text, and the multimodal features include image features, audio features, and text features; The loop generation feature fusion subsystem includes: a loop generation module for fusing features of different modalities to obtain multimodal fusion features, and performing modal generation to obtain multimodal generation features; a feature fusion module for fusing the multimodal features with the multimodal generation features and the multimodal fusion features; a single-modal enhancement module for constructing a single-modal enhancement Transformer using a mutual attention mechanism, performing mutual attention alignment processing on the multimodal generation features and the multimodal features to obtain enhanced single-modal features; the loop generation module is further used to: input text modal features X1 and image features X2 into the loop generation module; generate a vector in the video modal feature space from the text modal feature X1 through an encoder , the image feature X2 generates a vector in the text modality feature space ; The error constraint is performed with X1 so that The error with X1 is the smallest; The error constraint is performed with the X2 so that the The error with X2 is the smallest; the feature vector of the middle layer of the encoder is selected as the mixed vector X3 and X4 and output to the feature fusion module; The sentiment analysis subsystem is used to perform sentiment analysis on the fusion features and the enhanced unimodal features.

2. The system according to claim 1, wherein: The feature extraction subsystem includes: A multimodal feature database module, used for storing the multimodal original data and the multimodal features; The feature extraction module is used to extract features from the multimodal raw data to obtain multimodal features.

3. The system according to claim 1, wherein: The sentiment analysis subsystem further includes: a multi-task module, configured to perform sentiment analysis on the mixed features and the enhanced unimodal features to obtain sentiment analysis results; The user interaction module is used to provide the user with an interface for uploading and accessing information, and to display the sentiment analysis results to the user.

4. A multimodal sentiment analysis method based on loop generation applied to the system of claim 1, characterized in that: include: Acquiring multimodal raw data, and performing feature extraction on the multimodal raw data to obtain multimodal features; fusing the multimodal features to obtain multimodal fusion features, and generating modalities to obtain multimodal generation features; fusing the multimodal feature with the multimodal generated feature and the multimodal fusion feature to obtain a hybrid feature; Performing mutual attention alignment processing on the multimodal generated features and the multimodal features to obtain enhanced unimodal features; Sentiment analysis is performed on the mixed features and the enhanced unimodal features.

5. A computer device, characterized in that: The system comprises a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the system realizes the multimodal sentiment analysis system based on loop generation as described in any one of claims 1 to 3.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multimodal sentiment analysis system based on loop generation according to any one of claims 1 to 3 is implemented.