Public opinion sentiment analysis method based on multiple modes
Through the feature-level and decision-level fusion method combined with iterative attention mechanism, sentiment analysis of multimodal data is solved, and the problem of multimodal data integration and interaction relationship neglect in the existing technology is achieved, achieving higher accuracy and robustness of sentiment analysis.
Patent Information
- Application Number
- CN202411969603.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-27
AI Technical Summary
The existing multimodal emotion analysis methods are difficult to effectively integrate multimodal data such as text, images and videos, resulting in limited accuracy of emotion recognition and neglecting the complex interaction between different modes.
The emotional characteristics of text, image and audio and video data are fused by the feature-level fusion method and the decision-level fusion method, and the fusion results are optimized through the iterative attention mechanism to analyze the emotional consistency between the modals to adjust the feature weight.
It improves the accuracy and robustness of sentiment analysis, can capture the potential correlation of multimodal data more comprehensively, and enhances the comprehensive discrimination ability of complex emotions.
Smart Images

Figure CN120045991A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data processing, and particularly relates to a multi-modal public opinion sentiment analysis method. Background Art
[0002] With the wide popularity of social media and online platforms, people often use multi-modal forms such as text, images, and videos to express their emotions when expressing emotions. This diverse expression makes the emotional information richer and more complex, beyond the scope of traditional single-modal sentiment analysis. For example, forms of expression such as irony, exaggeration, and metaphor often rely on the mutual cooperation between multiple modalities, and it is difficult to accurately identify the emotional information contained therein by analyzing text or images alone. Therefore, how to effectively integrate multi-modal data has become an important challenge in current sentiment analysis research.
[0003] Existing multi-modal sentiment analysis methods usually rely on simple feature fusion techniques, such as directly splicing features of different modalities. However, this method is prone to introducing information redundancy and noise, resulting in limited accuracy of emotion recognition. At the same time, it ignores the complex interaction relationships between different modalities, such as the context dependence and emotional consistency between text and images. This limitation has prompted researchers to explore deeper modal interaction modeling methods to more comprehensively capture the potential correlations of multi-modal data, thereby improving the accuracy and robustness of sentiment analysis. Summary of the Invention
[0004] The present application provides a multi-modal public opinion sentiment analysis method, which can more comprehensively capture the potential correlations of multi-modal data, thereby improving the accuracy and robustness of sentiment analysis.
[0005] In a first aspect of the present application, a multi-modal public opinion sentiment analysis method is provided. The method includes:
[0006] Preprocess the multi-modal data collected by the public opinion information collection platform to obtain various processed data, and the various processed data include text data, image data, and audio-visual data;
[0007] For different types of the processed data, adopt different feature extraction methods to obtain different modal sentiment features corresponding to the processed data, where the text data corresponds to the first sentiment feature, the image data corresponds to the second sentiment feature, and the audio-visual data corresponds to the third sentiment feature;
[0008] Adopt a feature-level fusion method and a decision-level fusion method to fuse the first sentiment feature, the second sentiment feature, and the third sentiment feature, and further optimize the fusion result through an iterative attention mechanism to obtain a fusion feature;
[0009] Analyze the emotional consistency among the first emotional feature, the second emotional feature, and the third emotional feature, and analyze the impact of the first emotional feature, the second emotional feature, and the third emotional feature on the overall emotion. Adjust the weights of the fusion features, and output the public opinion emotion of the multimodal data according to the adjustment results.
[0010] Based on the above technical solutions, preferably, the feature-level fusion method and the decision-level fusion method are used to fuse the first emotional feature, the second emotional feature, and the third emotional feature, and the fusion result is further optimized through an iterative attention mechanism to obtain a fusion feature, which specifically includes:
[0011] Project the first emotional feature, the second emotional feature, and the third emotional feature into a unified feature space through a mapping layer to obtain the first feature corresponding to the first emotional feature, the second feature corresponding to the second emotional feature, and the third feature corresponding to the third emotional feature;
[0012] Generate attention weights corresponding to the first feature, the second feature, and the third feature according to the spatial dimensions of the first feature, the second feature, and the third feature;
[0013] Use a multi-scale channel attention module to combine the first feature, the second feature, and the third feature according to the attention weights corresponding to the first feature, the second feature, and the third feature to obtain a preliminary fusion feature;
[0014] Feed the preliminary fusion feature back into the first emotional feature, the second emotional feature, and the third emotional feature, and perform fusion again through the multi-scale channel attention module. Adjust the attention weights during the multi-layer iteration process, gradually optimize the fusion result, and output the fusion feature when the iteration process reaches a preset convergence condition.
[0015] Based on the above technical solutions, preferably, the feature-level fusion method and the decision-level fusion method are used to fuse the first emotional feature, the second emotional feature, and the third emotional feature, and the fusion result is further optimized through an iterative attention mechanism to obtain a fusion feature, which specifically further includes:
[0016] Input the first emotional feature into a deep learning classification network and output a text emotion classification result;
[0017] Input the second emotional feature into a convolutional neural network or a visual feature extraction model and output an image emotion classification result;
[0018] Input the third emotional feature into a multimodal audio-visual network and output an audio-visual emotion classification result;
[0019] According to the priori knowledge design rules, determine the probability distributions of the text sentiment classification result, the image sentiment classification result, and the audio-visual sentiment classification result;
[0020] According to the probability distributions of the text sentiment classification result, the image sentiment classification result, and the audio-visual sentiment classification result, perform weighted summation on the sentiment classification result, the image sentiment classification result, and the audio-visual sentiment classification result to obtain a fused probability distribution;
[0021] Use a fully connected layer to perform a non-linear transformation on the fused probability distribution to generate a preliminary fused feature of a higher dimension;
[0022] Dynamically adjust the weights corresponding to the probability distribution through spatial attention and channel attention, and further optimize the preliminary fused feature to obtain an optimized fused feature;
[0023] Feed the optimized fused feature back to the preliminary fused feature, dynamically adjust the weights corresponding to the probability distribution by spatial attention and channel attention again, and perform multiple iterations. When the iteration process reaches a preset convergence condition, output the fused feature.
[0024] Based on the above technical solutions, preferably, for different types of the processed data, different feature extraction methods are adopted to obtain different modal sentiment features corresponding to the processed data, specifically including:
[0025] Vectorize the text data to generate sentence vectors;
[0026] Repeat the sentence vectors through a preset number of convolutional layers and pooling layers in sequence, and then connect a fully connected layer to obtain an intermediate output;
[0027] Use the intermediate output as the input of the Softmax layer, output the sentiment feature of the text data, and obtain the first sentiment feature.
[0028] Based on the above technical solutions, preferably, for different types of the processed data, different feature extraction methods are adopted to obtain different modal sentiment features corresponding to the processed data, specifically further including:
[0029] Encode the image data into a convolutional feature map through a convolutional neural network;
[0030] Adopt a multi-instance learning method to extract the global image feature of the convolutional feature map;
[0031] Adopt an object detection method to extract the rectangular regions where each independent object in the image data is located to obtain the local image feature of the image data;
[0032] Input the global image feature and the local image feature into the attention model, and dynamically adjust the model's attention degree to different image regions by calculating the weighted average value, so as to obtain the processed global feature corresponding to the global image feature and the processed local feature corresponding to the local image feature;
[0033] Input the processed global feature and the processed local feature into the bidirectional LSTM for sequence modeling. Among them, the first LSTM layer of the bidirectional LSTM processes the temporal dependence of the processed global feature and the processed local feature, and generates an intermediate vector representing the dynamic emotional feature of the image;
[0034] The second LSTM layer of the bidirectional LSTM further refines the generation of the image description according to the intermediate vector, and takes the output of the second LSTM layer as the visual feature of the image to obtain the second emotional feature.
[0035] On the basis of the above technical solutions, preferably, for different types of the processed data, different feature extraction methods are adopted to obtain different modal emotional features corresponding to the processed data. Specifically, it further includes:
[0036] Take the audio-visual FBank feature in the audio-visual data as the input of the convolutional neural network. The convolutional neural network mines the correlation between dimensions and extracts the audio-visual local feature of the audio-visual data;
[0037] Input the audio-visual local feature into the time-delay neural network expansion module to further model the time dynamic information of the audio-visual local feature and obtain the dynamic audio-visual feature;
[0038] Input the dynamic audio-visual feature into the attention mechanism. The attention mechanism automatically assigns attention weights and focuses on the key features related to emotions in the dynamic audio-visual feature;
[0039] Input the key feature into the time-delay neural network expansion module again. The time-delay neural network expansion module extracts the emotional feature in the audio-visual data through the key feature to obtain the third emotional feature.
[0040] On the basis of the above technical solutions, preferably, analyze the emotional consistency among the first emotional feature, the second emotional feature, and the third emotional feature, and analyze the influence of the first emotional feature, the second emotional feature, and the third emotional feature on the overall emotion, adjust the weight of the fusion feature, and output the public opinion emotion of the multi-modal data according to the adjustment result. Specifically, it includes:
[0041] Construct positive and negative sample pairs using contrastive learning, and calculate the within-modal consistency scores within the first sentiment feature, the second sentiment feature, and the third sentiment feature respectively;
[0042] Construct cross-modal positive and negative sample pairs, and calculate the between-modal consistency scores between the first sentiment feature, the second sentiment feature, and the third sentiment feature respectively;
[0043] According to multiple within-modal consistency scores and multiple between-modal consistency scores, calculate the feature-level fusion weights corresponding to the first sentiment feature, the second sentiment feature, and the third sentiment feature respectively;
[0044] Adjust the fused feature obtained by the feature-level fusion method through the feature-level fusion weights to obtain the first fused feature;
[0045] Adjust the fused feature obtained by the decision-level fusion method through the feature-level fusion weights to obtain the second fused feature;
[0046] Generate the final fused feature through the first fused feature and the second fused feature;
[0047] Use a deep learning model to classify the final fused feature, and output the public opinion sentiment category and its confidence level.
[0048] In the second aspect of the present application, a multi-modal public opinion sentiment analysis device is provided. The device includes a processing module, an extraction module, a fusion module, and an output module, where:
[0049] The processing module is used to preprocess the multi-modal data collected by the public opinion information collection platform to obtain various processed data. The various processed data include text data, image data, and audio-visual data;
[0050] The extraction module is used to adopt different feature extraction methods for different types of the processed data to obtain the sentiment features of different modalities corresponding to the processed data. Among them, the text data corresponds to the first sentiment feature, the image data corresponds to the second sentiment feature, and the audio-visual data corresponds to the third sentiment feature;
[0051] The fusion module is used to fuse the first sentiment feature, the second sentiment feature, and the third sentiment feature by using the feature-level fusion method and the decision-level fusion method, and further optimize the fusion result through the iterative attention mechanism to obtain the fused feature;
[0052] The output module is used to analyze the emotional consistency among the first emotional feature, the second emotional feature, and the third emotional feature, analyze the influence of the first emotional feature, the second emotional feature, and the third emotional feature on the overall emotion, adjust the weights of the fusion features, and output the public opinion emotion of the multimodal data according to the adjustment result.
[0053] Based on the above technical solutions, preferably, the processing module is used to project the first emotional feature, the second emotional feature, and the third emotional feature into a unified feature space through a mapping layer to obtain the first feature corresponding to the first emotional feature, the second feature corresponding to the second emotional feature, and the third feature corresponding to the third emotional feature.
[0054] The extraction module is used to generate attention weights corresponding to the first feature, the second feature, and the third feature according to the spatial dimensions of the first feature, the second feature, and the third feature.
[0055] The fusion module is used to use a multi-scale channel attention module to combine the first feature, the second feature, and the third feature according to the attention weights corresponding to the first feature, the second feature, and the third feature to obtain a preliminary fusion feature.
[0056] The fusion module is used to feedback the preliminary fusion feature into the first emotional feature, the second emotional feature, and the third emotional feature, and perform fusion again through the multi-scale channel attention module. Adjust the attention weights during the multi-layer iteration process, gradually optimize the fusion result, and output the fusion feature when the iteration process reaches a preset convergence condition.
[0057] Based on the above technical solutions, preferably, the output module is used to input the first emotional feature into a deep learning classification network and output a text emotion classification result.
[0058] The extraction module is used to input the second emotional feature into a convolutional neural network or a visual feature extraction model and output an image emotion classification result.
[0059] The output module is used to input the third emotional feature into a multimodal audio-visual network and output an audio-visual emotion classification result.
[0060] The processing module is used to design rules according to prior knowledge to determine the probability distributions of the text emotion classification result, the image emotion classification result, and the audio-visual emotion classification result.
[0061] The processing module is configured to perform a weighted sum of the text sentiment classification result, the image sentiment classification result, and the audio-visual sentiment classification result according to the probability distributions thereof, so as to obtain a fused probability distribution;
[0062] The processing module is configured to perform a non-linear transformation on the fused probability distribution using a fully connected layer to generate a preliminary fused feature of a higher dimension;
[0063] The processing module is configured to further optimize the preliminary fused feature by dynamically adjusting the weights corresponding to the probability distribution through spatial attention and channel attention, so as to obtain an optimized fused feature;
[0064] The fusion module is configured to feedback the optimized fused feature to the preliminary fused feature, dynamically adjust the weights corresponding to the probability distribution again through spatial attention and channel attention, and perform multiple iterations. When a preset convergence condition is reached during the iteration process, the fused feature is output.
[0065] Based on the above technical solutions, preferably, the processing module is configured to vectorize the text data to generate sentence vectors;
[0066] The processing module is configured to sequentially repeat the sentence vectors through a preset number of convolutional layers and pooling layers, and then connect a fully connected layer to obtain an intermediate output;
[0067] The output module is configured to use the intermediate output as the input of a Softmax layer, output the sentiment feature of the text data, and obtain the first sentiment feature.
[0068] Based on the above technical solutions, preferably, the processing module is configured to encode the image data into a convolutional feature map through a convolutional neural network;
[0069] The processing module is configured to extract the global image feature of the convolutional feature map by using a multi-instance learning method;
[0070] The processing module is configured to use an object detection method to extract the rectangular regions where each independent object in the image data is located, and obtain the local image feature of the image data;
[0071] The processing module is configured to input the global image feature and the local image feature into an attention model, and dynamically adjust the attention degree of the model to different image regions by calculating a weighted average value, so as to obtain a processed global feature corresponding to the global image feature and a processed local feature corresponding to the local image feature;
[0072] The processing module is configured to input the processed global feature and the processed local feature into a bidirectional LSTM for sequence modeling. Among them, the first LSTM layer of the bidirectional LSTM processes the temporal dependence of the processed global feature and the processed local feature, and generates an intermediate vector representing the dynamic emotional feature of the image;
[0073] The processing module is configured to further refine the generation of the image description according to the intermediate vector by the second LSTM layer of the bidirectional LSTM, and use the output of the second LSTM layer as the visual feature of the image to obtain the second emotional feature.
[0074] Based on the above technical solutions, preferably, the extraction module is configured to use the audio-visual FBank features in the audio-visual data as the input of a convolutional neural network. The convolutional neural network mines the correlation between dimensions and extracts the audio-visual local features of the audio-visual data;
[0075] The processing module is configured to input the audio-visual local features into a time-delay neural network expansion module to further model the time dynamics information of the audio-visual local features and obtain dynamic audio-visual features;
[0076] The processing module is configured to input the dynamic audio-visual features into an attention mechanism. The attention mechanism automatically assigns attention weights and focuses on the key features related to emotions in the dynamic audio-visual features;
[0077] The processing module is configured to input the key features into the time-delay neural network expansion module again. The time-delay neural network expansion module extracts the emotional features in the audio-visual data through the key features to obtain the third emotional feature.
[0078] Based on the above technical solutions, preferably, the processing module is configured to construct positive and negative sample pairs using contrastive learning, and calculate the intra-modal consistency scores within the first emotional feature, the second emotional feature, and the third emotional feature respectively;
[0079] The processing module is configured to construct cross-modal positive and negative sample pairs and calculate the inter-modal consistency scores between the first emotional feature, the second emotional feature, and the third emotional feature respectively;
[0080] The processing module is configured to calculate the feature-level fusion weights corresponding to the first emotional feature, the second emotional feature, and the third emotional feature respectively according to the multiple intra-modal consistency scores and the multiple inter-modal consistency scores;
[0081] The processing module is configured to adjust the fused feature obtained by the feature-level fusion method through the feature-level fusion weight to obtain a first fused feature;
[0082] The fusion module is configured to adjust the fused feature obtained by the decision-level fusion method through the feature-level fusion weight to obtain a second fused feature;
[0083] The fusion module is configured to generate a final fused feature based on the first fused feature and the second fused feature;
[0084] The output module is configured to classify the final fused feature using a deep learning model and output the public opinion sentiment category and its confidence level.
[0085] In a third aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. Both the user interface and the network interface are used to communicate with other devices. The processor is configured to execute the instructions stored in the memory to enable the electronic device to execute the method described in any one of the above.
[0086] In a fourth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores instructions that, when executed, perform the method described in any one of the above.
[0087] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0088] 1. In the present application, different feature extraction methods are adopted for data of different modalities such as text, images, and audio / video, which can effectively capture the unique emotional information of each data type. Secondly, the combination of feature-level fusion and decision-level fusion optimizes the fusion result through an iterative attention mechanism to ensure that the emotional features of each modality play their due roles in the final output. More importantly, emotional consistency analysis is introduced. By analyzing and adjusting the emotional consistency within and between modalities, the weights of each modality can be dynamically adjusted, enabling the fused feature to more accurately reflect the contribution of different modalities to the overall emotion, thereby achieving a more accurate grasp of the emotion of multi-modal data. This comprehensive processing method improves the robustness of the model when facing complex and diverse public opinion data, can effectively handle the differences and noises between different modalities, and enhances the reliability and stability of the overall emotion analysis.
[0089] 2. By combining the feature-level fusion and decision-level fusion methods and continuously optimizing the fusion results with the iterative attention mechanism, the accuracy and robustness of multi-modal sentiment analysis are significantly improved. First, feature-level fusion ensures the effective integration of sentiment features from different modalities, enhancing the complementarity of information across modalities. Second, decision-level fusion sums up the sentiment classification results of each modality with weights and rationally allocates the contributions of each modality by combining prior knowledge, thereby enhancing the model's comprehensive discrimination ability for complex sentiments. The iterative attention mechanism optimizes the fused features by dynamically focusing on the most important information regions through multiple adjustments of feature weights, thus improving the accuracy, stability of the overall sentiment analysis and its adaptability to different sentiment data.
[0090] 3. By analyzing the sentiment consistency among different modalities and dynamically adjusting the fusion weights of the features of each modality in combination with the intra-modal and inter-modal consistency scores, the precise fusion of multi-modal data is achieved. This application can effectively capture the consistency and differences in the sentiment information of each modality, avoid the over-influence or neglect of a certain modality in sentiment analysis, and improve the accuracy and robustness of multi-modal sentiment analysis. At the same time, by finely adjusting the fusion weights at the feature level and decision level, the performance of the fused features is optimized, making the final sentiment classification more accurate, with higher confidence, and better reflecting the comprehensive sentiment state of multi-modal data. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] Figure 1 is a schematic flowchart of a multi-modal public opinion sentiment analysis method disclosed in an embodiment of the present application;
[0092] Figure 2 is a schematic block diagram of a multi-modal public opinion sentiment analysis device disclosed in an embodiment of the present application;
[0093] Figure 3 is a schematic structural diagram of an electronic device disclosed in an embodiment of the present application.
[0094] Description of the reference numerals: 201, processing module; 202, extraction module; 203, fusion module; 204, output module; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0095] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.
[0096] In the description of the embodiments of the present application, words such as "for example" or "for instance" are used to give examples, illustrations, or explanations. Any embodiment or design solution described as "for example" or "for instance" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "for example" or "for instance" is intended to present relevant concepts in a specific manner.
[0097] In the description of the embodiments of the present application, the term "a plurality of" means two or more. For example, a plurality of systems means two or more systems, and a plurality of screen terminals means two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0098] Multimodal sentiment analysis faces the challenge of integrating multimodal data such as text, images, and videos. Traditional feature concatenation methods are prone to introducing redundancy and noise, and neglect the complex interaction relationships between modalities, making it difficult to accurately identify emotional expressions such as sarcasm and metaphor. To improve the accuracy and robustness of sentiment analysis, researchers are working on exploring deeper modal interaction modeling methods.
[0099] This embodiment discloses a multimodal-based public opinion sentiment analysis method. Refer to Figure 1 , and it includes the following steps S110 - S140:
[0100] S110, preprocess the multimodal data collected by the public opinion information collection platform to obtain various processed data.
[0101] A multimodal-based public opinion sentiment analysis method disclosed in the embodiments of the present application is applied to a server. The server includes, but is not limited to, electronic devices such as mobile phones, tablets, wearable devices, and PCs (Personal Computers), and can also be a background server running a multimodal-based public opinion sentiment analysis method. The server can be implemented by an independent server or a server cluster composed of multiple servers.
[0102] In the public opinion information collection platform, the specific implementation plan for preprocessing multi-modal data includes: First, perform word segmentation, stop word removal, syntax parsing, and sentiment word extraction on the collected text data to generate structured text features; Second, perform denoising and normalization on the image data, and extract key features, such as object detection, scene classification, or sentiment labels; For audio-visual data, first perform audio-visual separation and frame extraction, and then extract features from the audio-visual data, such as intonation, speech rate, and sentiment waveform analysis, perform key frame analysis and multi-object tracking on the video data, and extract emotional features such as expressions and actions. Finally, uniformly format all processed data to lay the foundation for subsequent multi-modal sentiment analysis.
[0103] S120, for different types of processed data, adopt different feature extraction methods to obtain different modal sentiment features corresponding to the processed data.
[0104] In a possible implementation manner, for different types of processed data, adopt different feature extraction methods to obtain different modal sentiment features corresponding to the processed data, specifically including: vectorize the text data to generate sentence vectors; sequentially pass the sentence vectors through a preset number of convolutional layers and pooling layers, and after completion, connect a fully connected layer to obtain an intermediate output; use the intermediate output as the input of the Softmax layer, output the sentiment features of the text data, and obtain the first sentiment feature.
[0105] Specifically, first, preprocess the collected text data, including word segmentation, stop word removal, and normalization operations. Then, use a pre-trained word embedding model, such as Word2Vec, GloVe, or BERT, to convert the text into a word vector representation. Set a fixed dimension according to the sentence length, and maintain consistency by padding or truncating. Then, perform weighted average, concatenation, or context encoding on the word vectors in the sentence, such as RNN, LSTM, or Transformer, to generate a fixed-dimension sentence vector representation.
[0106] Input the generated sentence vectors into a convolutional neural network (CNN). The sentence vectors sequentially pass through multiple preset convolutional layers to extract local context features. Each convolutional layer uses multiple groups of convolutional kernels to capture semantic features of different granularities. After the convolutional operation, the ReLU activation function is used to increase the non-linear characteristics, and then a pooling layer, such as max pooling or average pooling, is added to reduce the dimension, retain important features, and reduce redundancy. The stacking of multiple convolutional and pooling layers helps to capture deep text semantic information.
[0107] The pooled feature map is flattened into a one-dimensional vector and input into a fully connected layer for further processing. The fully connected layer maps high-dimensional features to a lower-dimensional space and enhances the expressive power of features through activation functions (such as ReLU or Tanh). The intermediate output not only retains the key features extracted by the convolutional layer but also further strengthens the aggregation of emotional features.
[0108] The intermediate output is passed to the Softmax layer to calculate the probability distribution of different emotional categories (such as positive, negative, neutral, etc.). The Softmax function maps the input to the interval [0, 1] and ensures that the sum of the probabilities of all categories is 1, thus achieving emotional classification. The output emotional feature is the first emotional feature of the text data and is used for subsequent analysis or multimodal emotional fusion.
[0109] In a possible implementation, different feature extraction methods are adopted for different types of processed data to obtain emotional features of different modalities corresponding to the processed data. Specifically, it further includes: encoding image data into a convolutional feature map through a convolutional neural network; extracting the global image features of the convolutional feature map using the multi-instance learning method; using the object detection method to extract the rectangular regions where each independent object is located in the image data to obtain the local image features of the image data; inputting the global image local features and the local image features into the attention model, and dynamically adjusting the model's attention degree to different image regions by calculating the weighted average value, so as to obtain the processed global features corresponding to the global image features and the processed local features corresponding to the local image features; inputting the processed global features and the processed local features into a bidirectional LSTM for sequence modeling, where the first LSTM layer of the bidirectional LSTM processes the temporal dependencies of the processed global features and the processed local features to generate an intermediate vector representing the dynamic emotional features of the image; the second LSTM layer of the bidirectional LSTM further refines the generation of the image description according to the intermediate vector, and takes the output of the second LSTM layer as the visual feature of the image to obtain the second emotional feature.
[0110] Specifically, first, a pre-trained convolutional neural network such as ResNet or VGG is used to process the image data. The input image is passed through multiple convolutional and pooling operations to extract its deep visual features and encode them into a convolutional feature map. These feature maps retain important information of the image, such as texture, edge, and shape features, providing a basis for subsequent processing.
[0111] The global features in the convolutional feature map are extracted using the multi-instance learning method. The multi-instance learning method models the overall feature distribution of the image to capture the representative global emotional information in the image. These global features reflect the overall emotional features and semantic background of the entire image.
[0112] Use object detection algorithms such as SSD or YOLO to perform object detection on image data, and extract the rectangular regions (ROIs) where independent objects are located in the image. Each rectangular region represents a local area of the image. By further extracting the features of these regions, the local emotional features of the image are obtained. These local features can reflect important information in specific regions, such as facial expressions or specific emotional markers.
[0113] Input the global features and local features into the attention model. The attention model calculates through weighted average to dynamically adjust the model's attention degree to different image regions. The model will automatically assign weights to focus on the regions with more significant emotional information, thereby generating optimized processed global features and processed local features.
[0114] Input the optimized processed global features and processed local features into a bidirectional LSTM network. The first layer of LSTM mainly models the temporal dependencies of these features to generate an intermediate vector, which synthesizes the dynamic emotional information of global and local features. The bidirectional LSTM can capture the forward and backward dependencies of features, thus better expressing the dynamic emotional features of the image.
[0115] The second layer of LSTM further processes the intermediate vector to refine the generation of the image description and improve the expression of emotional features by combining context information. Finally, the output of the second layer of LSTM is used as the visual emotional feature of the image, that is, the second emotional feature of the image, for subsequent multi-modal fusion and emotion classification.
[0116] In a possible implementation manner, for different types of processed data, different feature extraction methods are adopted to obtain different modal emotional features corresponding to the processed data. Specifically, it further includes: using the audio-visual FBank features in the audio-visual data as the input of the convolutional neural network, and the convolutional neural network mines the correlations between dimensions to extract the audio-visual local features of the audio-visual data; inputting the audio-visual local features into the time-delay neural network expansion module to further model the time dynamic information of the audio-visual local features to obtain dynamic audio-visual features; inputting the dynamic audio-visual features into the attention mechanism, and the attention mechanism automatically assigns attention weights to focus on the key features related to emotions in the dynamic audio-visual features; inputting the key features into the time-delay neural network expansion module again, and the time-delay neural network expansion module extracts the emotional features in the audio-visual data through the key features to obtain the third emotional feature.
[0117] Specifically, the audio-visual data is converted into FBank features, i.e., filter bank features, and used as the input of the convolutional neural network in the form of a two-dimensional tensor. The FBank features contain the spectral information of the audio-visual signal and have a two-dimensional correlation in time and frequency. The convolutional layer of the CNN operates on the FBank features to extract their local features, and at the same time, the activation function is combined to mine the correlation between dimensions. After convolution, a pooling layer such as max pooling or average pooling is used to reduce the dimension of the features and reduce redundant data, thereby generating local audio-visual features that reflect the emotional information in local time segments.
[0118] The extracted local audio-visual features are input into the time-delay neural network expansion module. This module can capture the temporal dynamic changes of the features and model the dependencies of the audio-visual signal in the time series. The time-delay neural network expansion module consists of multiple fully connected layers, delay layers, and activation layers. By extracting features with different time spans in the time series layer by layer, it captures the temporal dynamic patterns of the audio-visual data. The final output is the dynamic audio-visual features, which reflect the temporal characteristics related to emotions in the audio-visual signal.
[0119] Then, the dynamic audio-visual features are input into the attention mechanism module, and by automatically assigning attention weights, the key features related to emotions are located. The multi-head attention mechanism or self-attention mechanism is used to weight the input dynamic features to generate a weight matrix. The model highlights the features highly relevant to the emotional information according to the weight matrix, while suppressing noise and irrelevant features, thereby improving the emotional representation ability of the features and extracting the key features.
[0120] The key features extracted by the attention mechanism are input into the time-delay neural network expansion module again for deep modeling. This process further optimizes the key features and enhances the expression ability of the emotional information. At this stage, the time-delay neural network expansion module performs multi-layer non-linear mapping on the input features and combines the context information to mine the potential emotional patterns of the key features. The output result is the emotional features of the audio-visual data, obtaining the third emotional feature, which provides the input for multi-modal emotion analysis.
[0121] S130, the first emotional feature, the second emotional feature, and the third emotional feature are fused using the feature-level fusion method and the decision-level fusion method, and the fusion result is further optimized through the iterative attention mechanism to obtain the fusion features.
[0122] In a possible implementation, a feature-level fusion method and a decision-level fusion method are used to fuse the first emotional feature, the second emotional feature, and the third emotional feature, and the fusion result is further optimized through an iterative attention mechanism to obtain a fusion feature. Specifically, it includes: projecting the first emotional feature, the second emotional feature, and the third emotional feature into a unified feature space through a mapping layer to obtain the first feature corresponding to the first emotional feature, the second feature corresponding to the second emotional feature, and the third feature corresponding to the third emotional feature; generating attention weights corresponding to the first feature, the second feature, and the third feature according to the spatial dimensions of the first feature, the second feature, and the third feature; using a multi-scale channel attention module to combine the first feature, the second feature, and the third feature according to the attention weights corresponding to the first feature, the second feature, and the third feature to obtain a preliminary fusion feature; feeding the preliminary fusion feature back into the first emotional feature, the second emotional feature, and the third emotional feature, and performing fusion again through the multi-scale channel attention module, adjusting the attention weights during the multi-layer iterative process, gradually optimizing the fusion result, and outputting the fusion feature when the iterative process reaches a preset convergence condition.
[0123] Specifically, a set of mapping layers such as a fully connected layer or a convolutional layer are used to project the first emotional feature, the second emotional feature, and the third emotional feature of different modalities from their respective original spaces into the same unified feature space. During the projection process, the feature data is standardized and normalized to ensure that different modality features have similar scale and distribution characteristics. The result of the projection is to obtain the first feature, the second feature, and the third feature respectively, and these features have the same dimension, which is convenient for subsequent operations.
[0124] By calculating the spatial dimensions of the features such as width, height, depth, and channel dimension, the corresponding attention weights are generated. For the generation of spatial dimension attention, global average pooling or max pooling operations can be specifically used to extract spatial features to obtain the importance distribution of each region. Or through Softmax or normalization operations, these distributions are mapped into spatial attention weights. Then, the channel correlation is extracted, and the importance of each channel is generated through a self-attention mechanism or weight learning. The attention weights are used to control the information contribution of different modality features in a specific region or channel.
[0125] Using a multi-scale channel attention module (MS-CAM), combining the first feature, the second feature, and the third feature and their corresponding attention weights, global and local features are extracted through multi-scale convolution operations, and the three modality features are weighted and combined. This step maps the features of each modality to the fusion feature space, generates a preliminary fusion feature containing multi-modal information, and at the same time retains the key semantic information of each modality.
[0126] Feed the preliminary fusion features back into the original features, namely the first emotional feature, the second emotional feature, and the third emotional feature, as the input of the multi-scale channel attention module for the next round of fusion. In this process, by dynamically adjusting the attention weights, the model gradually optimizes the contribution ratio of each modality feature, reduces the interference of invalid information, and strengthens the collaborative relationship between modalities, thereby improving the quality of the fusion features.
[0127] During the multi-layer iteration process, by monitoring the changes in the fusion results, set preset convergence conditions, such as the loss function no longer decreasing significantly or the feature distribution tending to be stable. When the convergence conditions are met, output the finally optimized fusion features, which integrate the important emotional information of all modalities and can be used for subsequent emotional classification or analysis tasks.
[0128] In a possible implementation manner, a feature-level fusion method and a decision-level fusion method are used to fuse the first emotional feature, the second emotional feature, and the third emotional feature, and the fusion result is further optimized through an iterative attention mechanism to obtain fusion features. Specifically, it further includes: input the first emotional feature into a deep learning classification network to output a text emotional classification result; input the second emotional feature into a convolutional neural network or a visual feature extraction model to output an image emotional classification result; input the third emotional feature into a multi-modal audio-visual network to output an audio-visual emotional classification result; design rules according to prior knowledge to determine the probability distributions of the text emotional classification result, the image emotional classification result, and the audio-visual emotional classification result; perform weighted summation on the text emotional classification result, the image emotional classification result, and the audio-visual emotional classification result according to their probability distributions to obtain a fused probability distribution; use a fully connected layer to perform a non-linear transformation on the fused probability distribution to generate a preliminary fusion feature with a higher dimension; further optimize the preliminary fusion feature by dynamically adjusting the weights corresponding to the probability distribution through spatial attention and channel attention to obtain an optimized fusion feature; feed the optimized fusion feature back to the preliminary fusion feature, dynamically adjust the weights corresponding to the probability distribution by spatial attention and channel attention again, and perform multiple iterations. When the preset convergence conditions are reached during the iteration process, output the fusion features.
[0129] Specifically, input the first emotional feature into a deep learning classification network. Usually, a network architecture based on models such as Transformer or LSTM is used for the emotional classification task. After the network is trained, it will output the emotional classification result of the text, which represents the emotional category corresponding to the text content and its corresponding probability distribution.
[0130] Input the second emotional feature into a convolutional neural network or a dedicated visual feature extraction model, such as ResNet, VGG, etc. Through the extraction of image features and the learning of convolutional layers, the model outputs the emotional classification result of the image. This result represents the emotional category of the image and its probability distribution, reflecting the emotional information of the image content.
[0131] Input the third emotional feature into a multi-modal audio-visual network, which is usually composed of models such as convolutional neural networks and long short-term memory networks for processing sequential audio-visual data. The network outputs the audio-visual emotional classification result, representing the emotional category of the audio-visual content and its probability distribution.
[0132] Design a set of rules based on prior knowledge to determine the probability distributions of the text emotional classification result, the image emotional classification result, and the audio-visual emotional classification result. These rules can be based on the weights, credibility, or other external conditions of each modality. For example, according to the accuracy of emotional classification, the importance of the modality, or the context information, determine the weighted ratio of different modality classification results.
[0133] Perform a weighted sum of the probability distributions of the text, image, and audio-visual emotional classification results obtained according to the rules. Through weighting, the model can integrate the emotional information of each modality and adjust its influence according to the weights of different modalities to obtain a fused probability distribution, representing the final emotional prediction after integrating the modality information.
[0134] Use a fully connected layer to perform a non-linear transformation on the fused probability distribution, usually through activation functions such as ReLU, Sigmoid, etc. for feature dimension elevation and non-linear expression. At this time, the fused features are mapped from a low-dimensional space to a higher-dimensional space to generate preliminary fused features, facilitating subsequent optimization and adjustment.
[0135] Use spatial attention mechanisms and channel attention mechanisms to dynamically adjust the importance of each part in the fused features according to the attention weights. Spatial attention helps to focus on important regions, and channel attention helps to optimize the influence of different feature channels. By weighted adjustment of feature contributions, the preliminary fused features are further optimized.
[0136] Feed the optimized fused features back into the preliminary fused features and perform multiple iterative optimizations through the dynamic adjustment of spatial attention and channel attention. In each round of the iterative process, judge whether the preset convergence condition is reached according to the changes in features and the convergence progress of the model, such as the changes in the loss function and the feature stability. When the iteration reaches the preset convergence condition, output the final optimized fused features, which integrate the emotional information of all modalities and provide a basis for the final emotional recognition result.
[0137] S140, analyze the emotional consistency among the first emotional feature, the second emotional feature, and the third emotional feature, analyze the impact of the first emotional feature, the second emotional feature, and the third emotional feature on the overall emotion, adjust the weights of the fusion features, and output the public opinion emotion of the multimodal data according to the adjustment result.
[0138] In a possible implementation manner, analyze the emotional consistency among the first emotional feature, the second emotional feature, and the third emotional feature, analyze the impact of the first emotional feature, the second emotional feature, and the third emotional feature on the overall emotion, adjust the weights of the fusion features, and output the public opinion emotion of the multimodal data according to the adjustment result. Specifically, use contrastive learning to construct positive and negative sample pairs, and calculate the intra-modal consistency scores within the first emotional feature, the second emotional feature, and the third emotional feature respectively; construct cross-modal positive and negative sample pairs, and calculate the inter-modal consistency scores among the first emotional feature, the second emotional feature, and the third emotional feature respectively; according to multiple intra-modal consistency scores and multiple inter-modal consistency scores, calculate the feature-level fusion weights corresponding to the first emotional feature, the second emotional feature, and the third emotional feature respectively; through the feature-level fusion weights, adjust the fusion features obtained by the feature-level fusion method to obtain the first fusion feature; through the feature-level fusion weights, adjust the fusion features obtained by the decision-level fusion method to obtain the second fusion feature; generate the final fusion feature through the first fusion feature and the second fusion feature; use a deep learning model to classify the final fusion feature and output the public opinion emotion category and its confidence.
[0139] Specifically, using contrastive learning, first construct positive and negative sample pairs, which are emotional features from the same modality. For example, for the text emotional feature of the first emotional feature, select samples with similar emotions as positive samples and samples with different emotions as negative samples. By calculating the similarity of each pair of samples within the same modality, the intra-modal consistency score is obtained. These scores reflect the consistency of emotions within this modality, that is, the intensity of emotional consistency within the modality features.
[0140] By constructing cross-modal positive and negative sample pairs, select emotional features from different modalities (such as text, image, and audio-video), and compare their emotional consistency. For example, the emotional feature of the text can be matched with the emotional feature of the image or audio-video, and the consistency score of the cross-modal emotion of the positive and negative sample pairs is calculated. This step helps to measure the emotional consistency between different modalities and ensures that the emotional information between multimodal data is coordinated.
[0141] Based on the calculated multiple intra-modal consistency scores and inter-modal consistency scores, fusion weights are assigned to each emotional feature of text, image, and audio-video respectively. Modal features with high consistency are given higher weights, while those with low consistency are given lower weights. Through these consistency scores, the contribution ratio of each modal feature in the final emotional fusion can be adjusted, thereby optimizing the feature fusion process.
[0142] For the first emotional feature, the calculation process of its corresponding feature-level fusion weight is as follows:
[0143]
[0144] Where, W T is the feature-level fusion weight corresponding to the first emotional feature, C T is the intra-modal consistency score corresponding to the first emotional feature, C T-inter is the inter-modal consistency score corresponding to the first emotional feature, W I is the feature-level fusion weight corresponding to the second emotional feature, C I is the intra-modal consistency score corresponding to the second emotional feature, C I-inter is the inter-modal consistency score corresponding to the second emotional feature, W V is the feature-level fusion weight corresponding to the third emotional feature, C V is the intra-modal consistency score corresponding to the third emotional feature, C V-inter is the inter-modal consistency score corresponding to the third emotional feature, and λ is a tuning parameter used to balance the influence of intra-modal consistency and inter-modal consistency on the weight.
[0145] Furthermore, using the obtained feature-level fusion weights, the fusion features obtained by the feature-level fusion method are adjusted. Specifically, the weights of each modal feature are adjusted according to the importance of each modal feature, so that the fused features can better reflect the intra-modal and inter-modal consistency, thereby optimizing the fusion result. These adjustments can make the final fused features more representative and distinguishable.
[0146] Similar to the feature-level fusion step, using the fusion weights obtained in the above step, the fusion features obtained by the decision-level fusion method are adjusted. By adjusting the weighted ratio of the modal classification results, it is ensured that the influence of each modal classification result can match its emotional consistency, thereby optimizing the fusion effect at the decision-making level.
[0147] Generate the final fused feature by combining the first fused feature and the second fused feature. This fused feature is an optimized multi-modal fused feature that takes into account the intra-modal and inter-modal consistency scores as well as the adjustment of feature-level and decision-level fusion weights, and has higher sentiment recognition accuracy. The final fused feature is usually a high-dimensional vector that contains sentiment information from multiple modalities such as text, images, and audio-video. These fused features have undergone modal consistency analysis and feature fusion optimization, and can effectively represent the sentiment states of each modality. Input the final fused feature into a deep learning model such as a fully connected neural network, CNN, or LSTM for sentiment classification. After being trained, the model will output the category of public opinion sentiment (such as positive, negative, neutral) and the confidence of this classification result as the final sentiment analysis result.
[0148] Selecting an appropriate deep learning architecture usually depends on the type of input features and the specific requirements of the task. If the fused features mainly include text and audio-video sequence information, LSTM or Transformer can be selected. If image features dominate the fused features, CNN can be considered.
[0149] Normalize or standardize the fused features so that the neural network can learn more efficiently. For sentiment classification problems, common loss functions include Cross-Entropy Loss. This loss function is applicable to multi-class sentiment classification problems, and the model optimizes the parameters by minimizing the loss function. Use the training dataset (including labels and fused features) to train the model, usually using the Backpropagation algorithm and optimization algorithms (such as Adam, SGD, etc.) for parameter updates.
[0150] After training is completed, use the model to infer (predict) new fused features. The model will output the sentiment classification result and its corresponding confidence. The sentiment classification result is usually one of multiple categories, such as "positive", "negative", and "neutral". The confidence refers to the confidence of the model in each category, usually calculated through the softmax layer. It represents the probability of each sentiment category.
[0151] Suppose we have a sentiment analysis task with the goal of judging public opinion sentiment based on text, image, and audio-video data, and the sentiment categories include "positive", "negative", and "neutral". The text sentiment feature is obtained by performing sentiment analysis on a piece of text (such as "This movie is really great!"), resulting in the text sentiment feature F text . Image sentiment feature: Perform sentiment analysis on the movie cover image to obtain the image sentiment feature F image . Audio-video sentiment feature: Obtain the audio-video sentiment feature F by analyzing the audio-video (such as the tone and emotion of the speech) in the movie traileraudio By using the feature-level fusion method, the emotional features of these modalities are fused to obtain the final fused feature F final = Ftext , F image , F audio . F final is input into a deep learning model. The model generates a prediction of an emotional category based on the training data, such as "positive". Suppose the output result of the model is that the confidence of the "positive" category is 0.85, the confidence of the "negative" category is 0.10, and the confidence of the "neutral" category is 0.05. Therefore, the final classification result is the "positive" category with a confidence of 85%.
[0152] This embodiment also discloses a multi-modal public opinion sentiment analysis device. Referring to Figure 2 , the device includes a processing module 201, an extraction module 202, a fusion module 203, and an output module 204, where:
[0153] The processing module 201 is configured to preprocess the multi-modal data collected by the public opinion information collection platform to obtain various processed data, and the various processed data include text data, image data, and audio-video data.
[0154] The extraction module 202 is configured to adopt different feature extraction methods for different types of processed data to obtain different modal emotional features corresponding to the processed data, where the text data corresponds to the first emotional feature, the image data corresponds to the second emotional feature, and the audio-video data corresponds to the third emotional feature.
[0155] The fusion module 203 is configured to fuse the first emotional feature, the second emotional feature, and the third emotional feature by using the feature-level fusion method and the decision-level fusion method, and further optimize the fusion result through an iterative attention mechanism to obtain a fused feature.
[0156] The output module 204 is configured to analyze the emotional consistency among the first emotional feature, the second emotional feature, and the third emotional feature, and analyze the influence of the first emotional feature, the second emotional feature, and the third emotional feature on the overall emotion, adjust the weight of the fused feature, and output the public opinion sentiment of the multi-modal data according to the adjustment result.
[0157] In a possible implementation manner, the processing module 201 is configured to project the first emotional feature, the second emotional feature, and the third emotional feature into a unified feature space through a mapping layer to obtain the first feature corresponding to the first emotional feature, the second feature corresponding to the second emotional feature, and the third feature corresponding to the third emotional feature.
[0158] An extraction module 202, configured to generate attention weights corresponding to the first feature, the second feature, and the third feature according to the spatial dimensions of the first feature, the second feature, and the third feature.
[0159] A fusion module 203, configured to use a multi-scale channel attention module to combine the first feature, the second feature, and the third feature according to the attention weights corresponding to the first feature, the second feature, and the third feature, so as to obtain a preliminary fusion feature.
[0160] The fusion module 203 is configured to feedback the preliminary fusion feature into the first sentiment feature, the second sentiment feature, and the third sentiment feature, and perform fusion again through the multi-scale channel attention module, adjust the attention weights during the multi-layer iteration process, gradually optimize the fusion result, and output the fusion feature when the iteration process reaches a preset convergence condition.
[0161] In a possible implementation manner, an output module 204 is configured to input the first sentiment feature into a deep learning classification network and output a text sentiment classification result.
[0162] The extraction module 202 is configured to input the second sentiment feature into a convolutional neural network or a visual feature extraction model and output an image sentiment classification result.
[0163] The output module 204 is configured to input the third sentiment feature into a multi-modal audio-visual network and output an audio-visual sentiment classification result.
[0164] A processing module 201 is configured to design rules according to prior knowledge to determine the probability distributions of the text sentiment classification result, the image sentiment classification result, and the audio-visual sentiment classification result.
[0165] The processing module 201 is configured to perform weighted summation on the sentiment classification result, the image sentiment classification result, and the audio-visual sentiment classification result according to the probability distributions of the text sentiment classification result, the image sentiment classification result, and the audio-visual sentiment classification result to obtain a fusion probability distribution.
[0166] The processing module 201 is configured to perform a non-linear transformation on the fusion probability distribution using a fully connected layer to generate a preliminary fusion feature of a higher dimension.
[0167] The processing module 201 is configured to further optimize the preliminary fusion feature by dynamically adjusting the weights corresponding to the probability distribution through spatial attention and channel attention to obtain an optimized fusion feature.
[0168] The fusion module 203 is configured to feedback the optimized fusion feature to the preliminary fusion feature, dynamically adjust the weights corresponding to the probability distribution by spatial attention and channel attention again, and perform multiple iterations, and output the fusion feature when the iteration process reaches a preset convergence condition.
[0169] In a possible implementation, the processing module 201 is configured to vectorize text data to generate sentence vectors.
[0170] The processing module 201 is configured to sequentially pass the sentence vectors through a preset number of convolutional layers and pooling layers, and then connect a fully connected layer to obtain an intermediate output.
[0171] The output module 204 is configured to use the intermediate output as the input to the Softmax layer, output the sentiment features of the text data, and obtain the first sentiment feature.
[0172] In a possible implementation, the processing module 201 is configured to encode image data into a convolutional feature map through a convolutional neural network.
[0173] The processing module 201 is configured to extract the global image features of the convolutional feature map using the multi-instance learning method.
[0174] The processing module 201 is configured to use the object detection method to extract the rectangular regions where each independent object is located in the image data, and obtain the local image features of the image data.
[0175] The processing module 201 is configured to input the global image features and the local image features into the attention model, and dynamically adjust the model's attention to different image regions by calculating the weighted average, so as to obtain the processed global features corresponding to the global image features and the processed local features corresponding to the local image features.
[0176] The processing module 201 is configured to input the processed global features and the processed local features into a bidirectional LSTM for sequence modeling. Among them, the first LSTM layer of the bidirectional LSTM processes the temporal dependencies of the processed global features and the processed local features, and generates an intermediate vector representing the dynamic sentiment features of the image.
[0177] The processing module 201 is configured to further refine the generation of the image description according to the intermediate vector by the second LSTM layer of the bidirectional LSTM, and use the output of the second LSTM layer as the visual features of the image to obtain the second sentiment feature.
[0178] In a possible implementation, the extraction module 202 is configured to use the audio-visual FBank features in the audio-visual data as the input to the convolutional neural network, and the convolutional neural network mines the correlations between dimensions to extract the local audio-visual features of the audio-visual data.
[0179] The processing module 201 is configured to input the local audio-visual features into the time-delay neural network expansion module to further model the time dynamic information of the local audio-visual features, and obtain the dynamic audio-visual features.
[0180] The processing module 201 is configured to input dynamic audio-visual features into the attention mechanism, and the attention mechanism automatically assigns attention weights to focus on the key features related to emotions in the dynamic audio-visual features.
[0181] The processing module 201 is configured to input the key features into the time-delay neural network extension module again, and the time-delay neural network extension module extracts the emotional features from the audio-visual data through the key features to obtain the third emotional feature.
[0182] In a possible implementation manner, the processing module 201 is configured to construct positive and negative sample pairs by using contrastive learning, and calculate the intra-modal consistency scores within the first emotional feature, the second emotional feature, and the third emotional feature respectively.
[0183] The processing module 201 is configured to construct cross-modal positive and negative sample pairs, and calculate the inter-modal consistency scores between the first emotional feature, the second emotional feature, and the third emotional feature respectively.
[0184] The processing module 201 is configured to calculate the feature-level fusion weights corresponding to the first emotional feature, the second emotional feature, and the third emotional feature respectively according to the multiple intra-modal consistency scores and the multiple inter-modal consistency scores.
[0185] The processing module 201 is configured to adjust the fusion feature obtained by the feature-level fusion method through the feature-level fusion weights to obtain the first fusion feature.
[0186] The fusion module 203 is configured to adjust the fusion feature obtained by the decision-level fusion method through the feature-level fusion weights to obtain the second fusion feature.
[0187] The fusion module 203 is configured to generate the final fusion feature through the first fusion feature and the second fusion feature.
[0188] The output module 204 is configured to classify the final fusion feature by using a deep learning model and output the public opinion emotion category and its confidence level.
[0189] It should be noted that when the device provided in the above embodiment realizes its functions, only the above-mentioned division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be elaborated here.
[0190] This embodiment also discloses an electronic device, referring to Figure 3, the electronic device may include: at least one processor 301, at least one communication bus 302, a user interface 303, a network interface 304, and at least one memory 305.
[0191] Among them, the communication bus 302 is used to realize the connection and communication between these components.
[0192] Among them, the user interface 303 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 303 may further include a standard wired interface and a wireless interface.
[0193] Among them, the network interface 304 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0194] Among them, the processor 301 may include one or more processing cores. The processor 301 connects various parts within the entire server through various interfaces and lines, and executes various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 305, and by calling the data stored in the memory 305. Optionally, the processor 301 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 301 may integrate one or several combinations of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor 301 and may be implemented separately through a single chip.
[0195] Among them, the memory 305 may include a Random Access Memory (RAM), or may also include a Read-Only Memory. Optionally, the memory includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 305 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area can store the data involved in the above-mentioned method embodiments. Optionally, the memory 305 may also be at least one storage device located far from the aforementioned processor 301. The memory 305, as a computer storage medium, may include an operating system, a network communication module, a user interface 303 module, and an application program for a multi-modal public opinion sentiment analysis method.
[0196] In Figure 3 In the electronic device shown, the user interface 303 is mainly used to provide an input interface for the user to obtain the data input by the user; and the processor 301 can be used to call the application program for a multi-modal public opinion sentiment analysis method stored in the memory 305. When executed by one or more processors 301, the electronic device executes the method as described in one or more of the above embodiments.
[0197] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0198] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0199] In several embodiments provided in this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some service interfaces. The indirect couplings or communication connections of devices or units can be in electrical or other forms.
[0200] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0201] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0202] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory 305 and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of this application. And the aforementioned memory 305 includes: various media such as USB flash drives, mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0203] This application also discloses a computer-readable storage medium, and the computer-readable storage medium stores instructions. When executed by one or more processors 301, it causes the electronic device to execute the method as described in one or more of the above embodiments.
[0204] The above are only exemplary embodiments of the present disclosure, and the scope of the present disclosure cannot be limited thereby. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. Other embodiments of the present disclosure will be readily envisioned by those skilled in the art after considering the specification and the practice of the disclosed truth. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.
Claims
1. A multimodal public opinion sentiment analysis method, characterized in that: The method comprises: Preprocessing the multimodal data collected by the public opinion information collection platform to obtain a variety of processed data, wherein the various processed data include text data, image data, and audio and video data; For different types of processed data, different feature extraction methods are adopted to obtain emotional features of different modalities corresponding to the processed data, wherein the text data corresponds to a first emotional feature, the image data corresponds to a second emotional feature, and the audio and video data corresponds to a third emotional feature; The first emotion feature, the second emotion feature and the third emotion feature are fused by using a feature-level fusion method and a decision-level fusion method, and the fusion result is further optimized by an iterative attention mechanism to obtain a fusion feature; Analyze the emotional consistency among the first emotional feature, the second emotional feature, and the third emotional feature, and analyze the impact of the first emotional feature, the second emotional feature, and the third emotional feature on the overall emotion, adjust the weight of the fusion feature, and output the public opinion emotion of the multimodal data according to the adjustment result.
2. According to the multimodal public opinion sentiment analysis method of claim 1, it is characterized in that: The method of using a feature-level fusion method and a decision-level fusion method to fuse the first emotion feature, the second emotion feature, and the third emotion feature, and further optimizing the fusion result through an iterative attention mechanism to obtain a fusion feature, specifically includes: Projecting the first emotion feature, the second emotion feature, and the third emotion feature into a unified feature space through a mapping layer to obtain a first feature corresponding to the first emotion feature, a second feature corresponding to the second emotion feature, and a third feature corresponding to the third emotion feature; Generate attention weights corresponding to the first feature, the second feature, and the third feature according to the spatial dimensions of the first feature, the second feature, and the third feature; Using a multi-scale channel attention module, combining the first feature, the second feature, and the third feature according to the attention weights corresponding to the first feature, the second feature, and the third feature to obtain a preliminary fusion feature; The preliminary fusion feature is fed back to the first emotion feature, the second emotion feature and the third emotion feature, and fused again through the multi-scale channel attention module. The attention weight is adjusted in the multi-layer iteration process, the fusion result is gradually optimized, and the fusion feature is output when the iterative process reaches a preset convergence condition.
3. According to the multimodal public opinion sentiment analysis method of claim 1, it is characterized in that: The method of using a feature-level fusion method and a decision-level fusion method to fuse the first emotion feature, the second emotion feature, and the third emotion feature, and further optimizing the fusion result through an iterative attention mechanism to obtain a fusion feature, specifically includes: Inputting the first sentiment feature into a deep learning classification network, and outputting a text sentiment classification result; Inputting the second emotion feature into a convolutional neural network or a visual feature extraction model, and outputting an image emotion classification result; Inputting the third emotion feature into a multimodal audio and video network, and outputting an audio and video emotion classification result; Design rules based on prior knowledge to determine the probability distribution of the text sentiment classification result, the image sentiment classification result, and the audio and video sentiment classification result; According to the probability distribution of the text emotion classification result, the image emotion classification result and the audio and video emotion classification result, weighted summing the emotion classification result, the image emotion classification result and the audio and video emotion classification result is performed to obtain a fusion probability distribution; Using a fully connected layer to perform nonlinear transformation on the fused probability distribution to generate a higher-dimensional preliminary fused feature; Dynamically adjusting the weight corresponding to the probability distribution through spatial attention and channel attention, further optimizing the preliminary fusion feature, and obtaining an optimized fusion feature; The optimized fusion feature is fed back to the preliminary fusion feature, and the spatial attention and channel attention are used to dynamically adjust the weight corresponding to the probability distribution, and multiple iterations are performed. When the iterative process reaches a preset convergence condition, the fusion feature is output.
4. The method for analyzing public opinion sentiment based on multimodality according to claim 1, characterized in that: The method of using different feature extraction methods for different types of processed data to obtain emotional features of different modes corresponding to the processed data specifically includes: Vectorizing the text data to generate sentence vectors; The sentence vector is repeatedly passed through a preset number of convolutional layers and pooling layers, and then connected to a fully connected layer to obtain an intermediate output; The intermediate output is used as the input of the Softmax layer, and the sentiment feature of the text data is output to obtain the first sentiment feature.
5. The method for analyzing public opinion sentiment based on multimodality according to claim 1, characterized in that: The method of using different feature extraction methods for different types of processed data to obtain emotional features of different modes corresponding to the processed data specifically includes: Encoding the image data into a convolutional feature map through a convolutional neural network; A multi-instance learning method is used to extract the global image features of the convolutional feature map; Using a target detection method to extract the rectangular area where each independent object in the image data is located, to obtain the local image features of the image data; Input the image global features and the image local features into the attention model, and dynamically adjust the model's attention to different image regions by calculating the weighted average, so as to obtain the processed global features corresponding to the image global features and the processed local features corresponding to the image local features; Inputting the processed global features and the processed local features into a bidirectional LSTM for sequence modeling, wherein a first LSTM layer of the bidirectional LSTM processes the temporal dependency of the processed global features and the processed local features to generate an intermediate vector representing the dynamic emotional features of the image; The second LSTM layer of the bidirectional LSTM further refines the generation of the image description according to the intermediate vector, and uses the output of the second LSTM layer as the visual feature of the image to obtain the second emotional feature.
6. The method for analyzing public opinion based on multimodality according to claim 1, characterized in that: The method of using different feature extraction methods for different types of processed data to obtain emotional features of different modes corresponding to the processed data specifically includes: The audio and video FBank features in the audio and video data are used as the convolutional neural network input, and the convolutional neural network mines the correlation between dimensions and extracts the audio and video local features of the audio and video data; Inputting the local audio and video features into a time delay neural network expansion module, further modeling the temporal dynamic information of the local audio and video features, and obtaining dynamic audio and video features; Inputting the dynamic audio and video features into an attention mechanism, the attention mechanism automatically assigns attention weights and focuses on key features of the dynamic audio and video features related to emotions; The key feature is input again into the time delay neural network extension module, and the time delay neural network extension module extracts the emotional features in the audio and video data through the key feature to obtain the third emotional feature.
7. The method for analyzing public opinion based on multimodality according to claim 1, characterized in that: The analyzing the emotional consistency among the first emotional feature, the second emotional feature, and the third emotional feature, and analyzing the influence of the first emotional feature, the second emotional feature, and the third emotional feature on the overall emotion, adjusting the weight of the fusion feature, and outputting the public opinion emotion of the multimodal data according to the adjustment result, specifically includes: Constructing positive and negative sample pairs by contrastive learning, and calculating intra-modal consistency scores within the first emotional feature, the second emotional feature, and the third emotional feature respectively; Constructing cross-modal positive and negative sample pairs, and respectively calculating inter-modal consistency scores between the first emotion feature, the second emotion feature, and the third emotion feature; According to the plurality of intra-modality consistency scores and the plurality of inter-modality consistency scores, respectively calculating feature-level fusion weights corresponding to the first emotion feature, the second emotion feature, and the third emotion feature; By using the feature-level fusion weight, adjusting the fusion feature obtained by the feature-level fusion method to obtain a first fusion feature; By using the feature-level fusion weight, adjusting the fusion feature obtained by the decision-level fusion method to obtain a second fusion feature; Generate a final fusion feature through the first fusion feature and the second fusion feature; The final fusion features are classified using a deep learning model to output the public opinion sentiment category and its confidence.
8. A multimodal public opinion sentiment analysis device, characterized in that: The device comprises a processing module (201), an extraction module (202), a fusion module (203) and an output module (204), wherein: The processing module (201) is used to pre-process the multimodal data collected by the public opinion information collection platform to obtain a variety of processed data, wherein the various processed data include text data, image data, and audio and video data; The extraction module (202) is used to adopt different feature extraction methods for different types of processed data to obtain emotional features of different modes corresponding to the processed data, wherein the text data corresponds to a first emotional feature, the image data corresponds to a second emotional feature, and the audio and video data corresponds to a third emotional feature; The fusion module (203) is used to fuse the first emotion feature, the second emotion feature and the third emotion feature by using a feature-level fusion method and a decision-level fusion method, and further optimize the fusion result by an iterative attention mechanism to obtain a fusion feature; The output module (204) is used to analyze the emotional consistency between the first emotional feature, the second emotional feature and the third emotional feature, and analyze the impact of the first emotional feature, the second emotional feature and the third emotional feature on the overall emotion, adjust the weight of the fusion feature, and output the public opinion emotion of the multimodal data according to the adjustment result.
9. An electronic device, characterized in that: The electronic device comprises a processor (301), a communication bus (302), a user interface (303), a network interface (304) and a memory (305), wherein the memory (305) is used to store instructions, the user interface (303) and the network interface (304) are both used to communicate with other devices, the communication bus (302) is used to realize connection and communication between components in the electronic device, and the processor (301) is used to execute the instructions stored in the memory (305) so that the electronic device executes the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 7 is performed.
Citation Information
Cited By
Multi-modal convergence media generation system and method based on dynamic emotion map
CN121030019A