An Emergency Sentiment Analysis Method Integrating Theme and Multimodality
Through the integration of themes and multimodal emergencies sentiment analysis methods, neural theme modeling and BiLSTM, textCNN and other technologies, the emotional characteristics of text and pictures in Weibo comments were extracted, and the difficulties in emotional analysis of multimodal data in the existing technology were solved, achieving more accurate sentiment analysis effects.
Patent Information
- Application Number
- CN202211003522.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-19
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-08-19
AI Technical Summary
It is difficult for the prior art to accurately analyze the emotions of comment content in platforms such as Weibo, especially in multimodal data such as text, pictures and emojis. The close correlation between themes and emotions complicates emotional analysis.
The emergencies sentiment analysis method that integrates the theme and multimodal is adopted to integrate external knowledge through neural theme modeling, pre-train and fine-tune neural theme models, and calculate the theme emotion distribution in combination with the emotion dictionary. Then, the emotional characteristics of the text and pictures are extracted by the fusion method of BiLSTM, textCNN and attention mechanism, and weighted averages according to the graphic correlation coefficients, and finally the emotional values are fusion at the model result level.
It has achieved more accurate and effective sentiment analysis of comment content on platforms such as Weibo, and can better deal with sentiment analysis problems in multimodal data.
Smart Images

Figure CN115392232B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, relates to sentiment analysis in network events, and particularly relates to a method for sentiment analysis of emergencies integrating topics and multi-modalities. Background Art
[0002] Sentiment analysis is a task in the field of natural language processing, also known as sentiment tendency analysis, opinion extraction, opinion mining, sentiment mining, subjective analysis, etc. It is a process of analyzing, processing, summarizing, and reasoning on subjective texts with emotional colors.
[0003] With the rapid development of the network, social networks have become the main platforms for the spread of network public opinion, and similar websites or platforms such as Weibo, as important media for the spread of network public opinion, enable users to express their views anytime and anywhere. Different from traditional text data, these comment data are redundant, containing text, pictures, videos, and a large amount of unique information such as website or platform-specific emoticons. At the same time, the text sentiment is also closely related to its discussion topic, which all lead to great difficulties in sentiment analysis of comment content in websites or platforms such as Weibo. Summary of the Invention
[0004] In order to overcome the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a method for sentiment analysis of emergencies integrating topics and multi-modalities, in order to more accurately perform sentiment analysis on comment content in websites or platforms such as Weibo.
[0005] In order to achieve the above purpose, the technical solution adopted by the present invention is:
[0006] A method for sentiment analysis of emergencies integrating topics and multi-modalities, comprising the following steps:
[0007] Step 1, incorporate external knowledge into neural topic modeling, pre-train the modeled neural topic model on a large corpus, then fine-tune it on the target dataset, and then calculate the topic sentiment distribution using a sentiment dictionary, and further obtain the sentiment tendency of each comment in the dataset; the data in the large corpus and the target dataset include text and pictures; based on this neural topic model, use the comment to be analyzed as input to obtain its sentiment value M 1 ;
[0008] Step 2, perform text and picture correlation analysis on the comment to be analyzed to obtain the text-picture correlation coefficient μ;
[0009] Step 3, adopt a method integrating BiLSTM, textCNN, and attention mechanism to extract the sentiment features of text and pictures to obtain the text sentiment value and the picture sentiment value According to the text-picture correlation coefficient μ for and perform a weighted average operation to obtain the graphic and text sentiment value
[0010] Step 4. Combine the sentiment value M 1 and the sentiment value M 2 at the model result level to obtain the final sentiment value M of the comment.
[0011] In one embodiment, in Step 1, the external knowledge is knowledge related to the theme learned during pre-training of the neural topic model and can be reused during fine-tuning on the target dataset; it is incorporated by the neural topic model through pre-training.
[0012] In one embodiment, the neural topic model adopts an encoder-decoder architecture. The BoW model is used to process the text in the dataset to obtain x. The encoder takes x ∈ R v as input, and the topic distribution of the text in the dataset is t ∈ R k , where v is the vocabulary size and k is the topic number of the topic distribution t; the decoder reconstructs the original document; the encoder is a stack composed of N + 1 MLP layers. From bottom to top, the first N layers have the same structure, and each layer has four sub-layers: Dropout, Linear, BatchNorm, and LeakyReLU. The last layer is a Dropout sub-layer and a linear transformation, followed by a Softmax layer. The decoder has the same architecture as the encoder.
[0013] In one embodiment, the encoder receives x ∈ R v as input and infers its topic distribution t ∈ R k , and then, the decoder reconstructs the original document from t. During the process, the dropout probability of each layer of the encoder and the negative slope of the LeakyReLU sub-layer are set to obtain the reconstruction loss: l rec (x, t) = -E(x log t);
[0014] where t has the same size m as x, and the topic distribution obtained by the neural topic model is adjusted by minimizing the maximum mean discrepancy with the Dirichlet distribution P. The formula is as follows:
[0015]
[0016] The total training objective is: l = l rec (x, t) + r · λ · l MMD (t, t′)
[0017] t′ is a topic distribution randomly sampled from P, k() is the information diffusion function, and i and j take values from 1 to m;
[0018] r is a hyperparameter used to balance l rec and l MMD , Normalization is performed using the second norm, and b(N + 1) is the bias term before the Softmax sublayer of the encoder. is the derivative operator.
[0019] In one embodiment, the large corpus is the DBPedia dataset, and the target dataset is the dataset of the competition for identifying the sentiment of Internet users during the epidemic in CCIR2020; the neural topic model is trained once on the DBPedia dataset to complete pre-training; then, fine-tuning is completed on the dataset of the competition for identifying the sentiment of Internet users during the epidemic in CCIR2020.
[0020] In one embodiment, the fine-tuning starts from the pre-trained model, randomly re-initializes the parameters at the last layer of the encoder and the first layer of the decoder, obtains the sentiment values of each topic according to the sentiment dictionary, and further obtains the sentiment value M of the entire comment based on the topic. 1 .
[0021] In one embodiment, for step 2, a method that combines BiLSTM and the attention mechanism is used for text and image correlation analysis. The method is as follows:
[0022] First, the text and the image are processed. The Glove method is used to convert the text into a text matrix, and the tools provided by the Vision platform in Google Cloud Platform are used to extract the image tags and represent them as the same word matrix as the text.
[0023] Then, two independent BiLSTMs are used to receive the image tags and the text tags respectively, and the BiLSTM represents the image and the text as vectors of the same dimension.
[0024] Finally, the image vector and the text vector are feature concatenated as the input of the fully connected layer, and finally the text-image correlation coefficient μ is output through the softmax layer.
[0025] In one embodiment, in step 3, for the text classification method based on the BiLSTM-Attention-textCNN hybrid neural network, the text in the comment to be analyzed is mapped into a vector through the word embedding layer, and the BiLSTM network is used to learn the representation of the previous context and the next context of the words in the comment to be analyzed, so as to obtain a deeper semantic vector of the current word; an attention model is established to calculate the probability weights of each word vector, so that the words with larger weights receive more attention, and these words that receive more attention are often crucial for the classification task; the vector output after the attention mechanism is connected to the pooling layer, and k-max pooling is performed to retain the first k words with larger weights; connect the textCNN network to extract features and output the text sentiment value
[0026] For the image classification method based on the BiLSTM-Attention-CNN hybrid neural network, the labels of the images in the comment to be analyzed are extracted and represented as the same matrix as the text. The BiLSTM is used to extract the image features, and an attention mechanism is established to select the features. Then, the vector output after the attention mechanism is connected to the pooling layer, and k-max pooling is performed to retain the first k words with larger weights; connect the CNN network to extract features and output the image sentiment value
[0027] In one embodiment, in step 4, the fusion calculation formula for the final sentiment value M is as follows:
[0028]
[0029] The proportion δ is determined by adjusting the parameters during the model training process.
[0030] Since network comments often contain specific information such as text and pictures, and at the same time, the comment sentiment is closely related to the discussion topic. Therefore, compared with the existing sentiment analysis methods, the present invention performs sentiment analysis through the theme and multi-modal fusion analysis method, and is more likely to handle the sentiment analysis of network comment data. Description of the Drawings
[0031] Figure 1 It is a schematic diagram of the architecture of the neural theme model.
[0032] Figure 2 It is a structure diagram of the text-image correlation.
[0033] Figure 3 It is a schematic diagram of text sentiment extraction.
[0034] Figure 4 It is a schematic diagram of sentiment feature extraction.
[0035] Figure 5 It is a schematic diagram of the structure for obtaining the comment sentiment polarity.
[0036] Figure 6 is the matching of text comments in the embodiments of the present invention Figure 1 。
[0037] Figure 7 is the matching of text comments in the embodiments of the present invention Figure 2 。 Detailed implementation manners
[0038] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0039] The present invention is an emergency sentiment analysis method that integrates themes and multi-modalities, which combines themes and multi-modalities for emergency sentiment analysis.
[0040] In the present invention, multi-modal refers to text and pictures. Since it revolves around emergencies, there will obviously be themes. The present invention obtains themes from the comments in the dataset, and then combines multi-modalities to obtain their emotions.
[0041] The implementation of the present invention includes the following steps:
[0042] Step 1, incorporate external knowledge into neural topic modeling, pre-train the modeled neural topic model on a large corpus, then fine-tune it on the target dataset, and then calculate the topic emotion distribution using an emotion dictionary, so as to obtain the emotion tendency of each comment in the dataset. Based on this neural topic model, with the comment to be analyzed as the input, obtain its emotion value M 1 ;
[0043] In the present invention, external knowledge is knowledge related to the theme learned during the pre-training of the neural topic model and can be reused during the fine-tuning on the target dataset, which is incorporated by the neural topic model through pre-training.
[0044] The large corpus used in the present invention is the DBPedia dataset, and the target dataset is the dataset of the competition for identifying the emotions of Internet users during the epidemic in CCIR2020. The target dataset contains Weibo comment data related to the epidemic. In this embodiment, the sentiment analysis is to analyze the sentiment polarity of each comment, that is, negative or positive. It should be noted that the data in the large corpus and the target dataset both include text, pictures, expressions, etc. The present invention only uses text and pictures, and only text is used in step 1. The neural topic model of the present invention is trained once on the DBPedia dataset to complete pre-training; then it is fine-tuned on the target dataset.
[0045] Reference Figure 1, the neural topic model established in the present invention adopts an encoder-decoder architecture. First, the BoW model is used to process the text in the datasets (DBPedia dataset and target dataset) to obtain x, and the topic distribution of the text in the dataset is t ∈ R k , where v is the vocabulary size, k is the topic number of the topic distribution t, the encoder is a stack composed of N + 1 MLP layers. From bottom to top, the first N layers have the same structure, and each layer has four sub-layers: Dropout, Linear, BatchNorm, and LeakyReLU. The last layer is a Dropout sub-layer and a linear transformation, followed by a Softmax layer. The decoder reconstructs the original document, and it has the same architecture as the encoder.
[0046] Specifically, the encoder receives x ∈ R v as input and infers its topic distribution t ∈ R k , then, the decoder reconstructs the original document from t. During the process, the dropout probability of each layer of the encoder and the negative slope of the LeakyReLU sub-layer are set to obtain the reconstruction loss: l rec (x, t) = -E(x log t);
[0047] where t has the same size m as x, and the topic distribution obtained by the neural topic model adjusts the topic by minimizing the maximum mean discrepancy between the Dirichlet distribution P. The formula is as follows:
[0048]
[0049] The total training objective is: l = l rec (x, t) + r · λ · l MMD (t, t′)
[0050] t′ is the topic distribution randomly sampled from P, k() is the information diffusion function, i and j take values from 1 to m; r is a hyperparameter used to balance l rec and l MMD , The second norm is used for normalization, b(N + 1) is the bias term before the Softmax sub-layer of the encoder, is the derivative operator.
[0051] The fine-tuning in this step starts from the pre-trained neural topic model, but the parameters are randomly re-initialized at the last layer of the encoder and the first layer of the decoder. According to the sentiment dictionary, the sentiment value of each topic is obtained, and then the sentiment value M of the entire comment based on the topic is obtained 1 .
[0052] Step 2, for a certain comment to be analyzed, perform text and image correlation analysis to obtain the text-image correlation coefficient μ.
[0053] Structural reference for text and image relevance Figure 2 As shown, it can be expressed as relevant or irrelevant. If it is irrelevant, only sentiment analysis of the text is performed; if it is relevant, comprehensive analysis is performed.
[0054] In this step, a method that combines BiLSTM and attention mechanism is used for text and image relevance analysis. The method is as follows:
[0055] First, the text and image are processed. The Glove method is used to convert the text into a text matrix, and the tools provided by the Vision platform in Google Cloud Platform are used to extract image tags and represent them as the same word matrix as the text.
[0056] Then, two independent BiLSTMs are used to receive the image tags and text tags respectively, and the BiLSTM represents the image and text as vectors of the same dimension.
[0057] Finally, the image vector and text vector are feature concatenated and used as the input of the fully connected layer (that is, the fully connected layer formed by concatenating the text and image feature vectors). Finally, the text-image correlation coefficient μ is output through the softmax layer.
[0058] Step 3: Use a method that combines BiLSTM, textCNN, and attention mechanism to extract the sentiment features of the text and image, and obtain the text sentiment value and the image sentiment value According to the text-image correlation μ and perform weighted average operation to obtain the text-image sentiment value M 2 .
[0059] Relevant:
[0060] Irrelevant:
[0061] In this step, the function of the attention mechanism to automatically learn and calculate the contribution of the input data to the output data is used to make the extraction of sentiment features more reasonable.
[0062] Specifically, refer to Figure 3, a text classification method based on a BiLSTM-Attention-textCNN hybrid neural network, maps the text in the comment to be analyzed into a vector through a word embedding layer, uses the BiLSTM network to learn the context representation and the next context representation of the words in the comment to be analyzed, and obtains a deeper semantic vector of the current word; establishes an attention model to calculate the probability weights of each word vector, so that the words with larger weights receive more attention, and these words that receive more attention are often the key words for the classification task; the vector output after the attention mechanism is connected to the pooling layer, performs k-max pooling, and retains the first k words with larger weights; connects the textCNN network to extract features and output the text sentiment value
[0063] Reference Figure 4 , an image classification method based on a BiLSTM-Attention-CNN hybrid neural network, extracts the labels of the images in the comment to be analyzed and represents them as the same matrix as the text, uses BiLSTM to extract image features, establishes an attention mechanism to select the features, and then connects the vector output after the attention mechanism to the pooling layer, performs k-max pooling, and retains the first k words with larger weights; connects the CNN network to extract features and output the image sentiment value
[0064] Step 4, reference Figure 5 , the sentiment value M 1 and the sentiment value M 2 are fused at the model result level to obtain the final sentiment value M of the comment, and its fusion calculation formula is as follows:
[0065]
[0066] The proportion δ is determined by parameter tuning during the model training process.
[0067] In a specific embodiment of the present invention, the text comment data is intercepted as follows:
[0068] 4456427143652010,01 / 02 23:21, mimiko sweetheart, I had a minor fever for the first time on January 2nd. I hope this year will be safe, healthy, and happy. 2 Shijiazhuang · Hebei University of Science and Technology?
[0069] Figure 6 And Figure 7 is the picture accompanying this text comment.
[0070] The text of the target data set is passed through a neural topic model and a sentiment dictionary to obtain the sentiment value of the topic.
[0071] This comment accumulates and then averages the topic sentiment values according to the topic. The sentiment value of the topic included in this comment is m 1,m 2 ,…,m n , then the sentiment value M derived from the theme 1 =(m 1 +m 2 +…+m n ) / n.
[0072] Input this data into Figure 2 's correlation model to obtain the correlation coefficient μ between the text and the picture in this comment. Input it into Figure 3 Figure 4 's text-picture sentiment value extraction model to obtain the sentiment values of the text and the picture and Then the sentiment value obtained from the text and picture is:
[0073] Take the sentiment value M 1 and the sentiment value M 2 Calculate the final sentiment value M of the comment according to the weight δ. Then the sentiment value of this comment is:
[0074] M>0 Positive; M = 0 Neutral; M<0 Negative.
Claims
1. A method for analyzing the sentiment of emergencies by integrating themes and multi-modality. It is characterized in that The steps include: Step 1, incorporate external knowledge into neural topic modeling, pre-train the obtained neural topic model on a large corpus, then fine-tune it on the target dataset, and then calculate the topic sentiment distribution using a sentiment dictionary to obtain the sentiment tendency of each comment in the dataset; the data in the large corpus and the target dataset includes text and pictures; based on this neural topic model, take the comment to be analyzed as input to obtain its sentiment value M 1 ; where the neural topic model adopts an encoder-decoder architecture, uses the BoW model to process the text in the dataset to obtain x, and the encoder takes x ∈ R v as input, and the topic distribution of the text in the dataset is t ∈ R k , where ν is the vocabulary size and k is the topic number of the topic distribution t; the decoder reconstructs the original document; the encoder is a stack composed of N + 1 MLP layers. From bottom to top, the first N layers have the same structure, and each layer has four sub-layers: Dropout, Linear, BatchNorm, and LeakyReLU. The last layer is a Dropout sub-layer and a linear transformation, followed by a Softmax layer. The decoder has the same architecture as the encoder; The encoder receives \(x\in\mathbb{R}\) v as input and infers its topic distribution \(t\in\mathbb{R}\) k , and then, the decoder reconstructs the original document from \(t\). In the process, by setting the exit probability of each layer of the encoder and the negative slope of the LeakyReLU sublayer, the reconstruction loss is obtained: \(l\) rec (x, t)=-E(x log t); Among them, t and x have the same size m, and the topic distribution obtained by the neural topic model is adjusted by minimizing the maximum mean difference between it and the Dirichlet distribution P, as follows: The overall training objective is: l = l rec (x, t) + r·λ·l MMD (t, t′) t′ is the topic distribution randomly drawn from P, k() is the information diffusion function, and i and j take values from 1 to m; r is a hyperparameter used to balance l rec and l MMD , Normalization is performed using the second norm, and b(N + 1) is the bias term before the Softmax sublayer of the encoder, is the derivative operator; Step 2: For the comment to be analyzed, perform text and image correlation analysis to obtain the image-text correlation coefficient μ; Step 3: Use the method of fusing BiLSTM, textCNN and attention mechanism to extract the emotional features of text and pictures, and obtain the text emotional value and the picture emotional value According to the text-picture correlation coefficient μ and perform a weighted average operation to obtain the text-picture emotional value Step 4, combine sentiment value M 1 and sentiment value M 2 at the model result level to obtain the final sentiment value M of the comment.
2. According to the method for analyzing the emotion of emergencies integrating themes and multi-modalities as described in claim 1, It is characterized in that In step 1, the external knowledge is the knowledge related to the topic learned during pre-training of the neural topic model, which can be reused when fine-tuning on the target dataset; it is incorporated into the neural topic model through pre-training.
3. According to the method for analyzing the emotion of emergencies integrating themes and multi-modalities as claimed in claim 1, It is characterized in that The large corpus is the DBPedia dataset, and the target dataset is the dataset of the Internet users' emotion recognition competition during the CCIR2020 epidemic. The neural topic model is trained once on the DBPedia dataset to complete pre-training. Then, fine-tuning is completed on the dataset of the Internet users' emotion recognition competition during the CCIR2020 epidemic.
4. The method for analyzing the emotion of emergencies integrating themes and multi-modality according to claim 1 or 3, It is characterized in that The fine-tuning starts from a pre-trained model, randomly re-initializes the parameters at the last layer of the encoder and the first layer of the decoder, obtains the sentiment values of each topic according to the sentiment dictionary, and then obtains the sentiment value M of the entire comment based on the topic. 1 .
5. According to the method for analyzing the emotion of emergencies integrating themes and multi-modalities as claimed in claim 1, It is characterized in that In step 2, a method combining BiLSTM and attention mechanism is used to perform text and image correlation analysis, as follows: First, the text and images are processed. The Glove method is used to convert the text into a text matrix. The image labels are extracted using the tools provided by the Vision platform in Google Cloud Platform and expressed as a word matrix that is the same as the text. Then, two independent BiLSTMs are used to accept image labels and text labels respectively, and the images and texts are represented as vectors of the same dimension through BiLSTM. Finally, the image vector and text vector are concatenated as the input of the fully connected layer, and the image-text correlation coefficient μ is finally output through the softmax layer.
6. According to the method for analyzing the emotion of emergencies integrating themes and multi-modalities as claimed in claim 1, It is characterized in that In step 3, for the text classification method based on the BiLSTM-Attention-textCNN hybrid neural network, the text in the comment to be analyzed is mapped into a vector through the word embedding layer. The BiLSTM network is used to learn the above-mentioned representation and the below-mentioned representation of the words in the comment to be analyzed, so as to obtain a deeper semantic vector of the current word. An attention model is established to calculate the probability weights of each word vector, so that the words with larger weights receive more attention. These words that receive more attention are often crucial for the classification task. The vector output after the attention mechanism is connected to the pooling layer, and k-max pooling is performed to retain the first k words with larger weights. Connect the textCNN network to extract features and output the text sentiment value A method for image classification based on a hybrid neural network of BiLSTM-Attention-CNN extracts tags from the images in the comments to be analyzed and represents them as matrices identical to the text, uses BiLSTM to extract image features, establishes an attention mechanism to select features, then connects the vector output by the attention mechanism to a pooling layer, performs k-max pooling to retain the top k words with larger weights; connects to a CNN network to extract features and outputs the image sentiment value 7. According to the method for analyzing the emotion of emergencies integrating themes and multi-modalities as claimed in claim 1, It is characterized in that In step 4, the fusion calculation formula of the final sentiment value M is as follows: The weight δ is determined by adjusting parameters during model training.
Citation Information
Patent Citations
Telecom fraud detection system and method based on user privacy protection
CN106851633A
Cross-modal image semantic extraction method, system and device and medium
CN111144410A