Emotion Mining Method and Device under Public Events
By analyzing public events-related release information and comment information, determining content categories and analyzing emotional intensity, the problems of weak emotional interpretation and low accuracy in the prior art are solved, and higher accuracy and explanatory emotions are achieved.
Patent Information
- Application Number
- CN202111580392.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-22
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-12-22
AI Technical Summary
The prior art is weak in the excavation of public event sentiment and has low accuracy.
By obtaining public events-related public information and comment information, determining content categories, and quantitatively analyzing emotions intensity based on comment information, enhancing the interpretability of emotions.
It improves the accuracy and interpretability of emotional mining and can more accurately reflect public emotional changes.
Smart Images

Figure CN114490925B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method and device for emotion mining under public events. Background Art
[0002] With the rise of online social networks, the public more often chooses to obtain and share information related to public events on social media. Based on social media, it is possible to understand the public's emotions regarding public events. In related technologies, there are some studies applying natural language processing techniques to mine the public's emotions regarding public events. However, most of them are based on manual tags of words and focus on identifying the positive and negative polarities of emotions, or use deep learning models to attempt to solve the binary classification problem of emotions, with relatively weak interpretability of emotions and inaccurate emotion mining. Summary of the Invention
[0003] The present invention provides a method and device for emotion mining under public events, to solve the defect in the prior art of relatively weak interpretability of emotions and inaccurate emotion mining, and to achieve enhanced interpretability of emotions, thereby improving the accuracy of emotion mining.
[0004] The present invention provides a method for emotion mining under public events, including:
[0005] Obtaining each piece of release information corresponding to a target public event and each piece of comment information of the release information;
[0006] For each piece of the release information, determining the content category to which the release information belongs;
[0007] For each content category, based on each piece of comment information of each piece of the release information under the content category, determining the emotion intensity information of the content category in each emotion category.
[0008] The present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method for emotion mining under public events as described in any one of the above are implemented.
[0009] The present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for emotion mining under public events as described in any one of the above are implemented.
[0010] The present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the method for emotion mining under public events as described in any one of the above are implemented.
[0011] The emotion mining method under public events provided by the present invention determines the characteristics of the content category to which each piece of release information belongs qualitatively by obtaining each piece of release information corresponding to the target public event and each piece of comment information of the release information. For each content category, based on each piece of comment information of each piece of release information under the content category, the emotion intensity information of the content category in each emotion category is determined quantitatively, enhancing the interpretability of emotions, thereby improving the accuracy of emotion mining. Description of the Drawings
[0012] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0013] Figure 1 It is one of the schematic flowcharts of the emotion mining method under public events provided by the present invention;
[0014] Figure 2 It is another schematic flowchart of the emotion mining method under public events provided by the present invention;
[0015] Figure 3 It is a schematic diagram of the first box plot provided by the present invention;
[0016] Figure 4 It is a schematic diagram of the second box plot provided by the present invention;
[0017] Figure 5a It is one of the schematic diagrams of the line chart provided by the present invention;
[0018] Figure 5b It is one of the schematic diagrams of the line chart provided by the present invention;
[0019] Figure 6a It is one of the schematic diagrams of the undirected graph provided by the present invention;
[0020] Figure 6b It is one of the schematic diagrams of the undirected graph provided by the present invention;
[0021] Figure 6c It is one of the schematic diagrams of the undirected graph provided by the present invention;
[0022] Figure 7 It is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Embodiments
[0023] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts belong to the scope of protection of the present invention.
[0024] With the rise of online social networks, the public more often chooses to obtain and share information related to public events on social media. Based on social media, the public's emotions regarding public events can be understood. In related technologies, there are some studies applying natural language processing technologies to mine the public's emotions regarding public events. However, most of them are based on manual labels of words and focus on identifying the positive and negative polarities of emotions, or use deep learning models to attempt to solve the binary classification problem of emotions, with relatively weak interpretability of emotions and inaccurate emotion mining.
[0025] Therefore, the embodiments of the present invention provide an emotion mining method under public events, which can determine the content category of the published information, the emotion category and emotion intensity information of the comment information induced by the published information under this content category based on the published information and the public's comment information on social media under public events, thereby improving the interpretability of emotions and making emotion mining more accurate. The following will introduce the solution provided by the embodiments of the present invention in detail through examples.
[0026] Figure 1 is one of the schematic flowcharts of the emotion mining method under public events provided by the present invention. As Figure 1 shown, this embodiment provides an emotion mining method under public events. This method can be executed by a terminal or a combination of software and / or hardware therein. The terminal can be, but is not limited to, a PC, a tablet computer, etc., or can also be executed by a server or a combination of software and / or hardware therein. This method at least includes the following steps:
[0027] Step 110, obtain each piece of published information corresponding to the target public event and each piece of comment information on the published information.
[0028] The target public event is the public event for which emotion mining needs to be performed currently. Exemplarily, it can be a sudden public event, such as a sudden public health event, a sudden public safety event, a sudden natural disaster event, etc. Of course, it can also be a non-sudden public event, that is, a frequently occurring public event.
[0029] In practical applications, after a target public event occurs, various pieces of information related to the target public event can be released through social media. Based on this, the released information can be obtained. The public can then comment on the released information to express their thoughts. Based on this, the comment information can be obtained. The released information and comment information here can be text information. In implementation, the released information and the comment information on the released information within a period of time after the occurrence of the target public event can be collected in advance to perform sentiment mining on the target public event.
[0030] Step 120: For each piece of the released information, determine the content category to which the released information belongs.
[0031] In practical applications, for the released information of a target public event, multiple content categories can be preset. Taking a sudden health event as an example, exemplarily, eight content categories can be set: release, traffic, domestic, documentary, policy, international, emotion, and diagnosis and treatment. Among them, the content category of release can include the information released by the news, such as the information of a press conference. The content category of traffic can include the traffic announcements corresponding to the target public event. The content category of domestic can include the domestic progress of the target public event. The content category of documentary can include the social documentary corresponding to the target public event, such as various social work arrangements and other information. The content category of policy can include the handling policies corresponding to the target public event. The content category of international can include the international progress of the target public event. The content category of emotion can include the information of emotional support corresponding to the target public event. The content category of diagnosis and treatment can include the patient diagnosis and treatment information corresponding to the target public event. For each piece of released information, the content category to which the released information belongs can be determined, thereby completing the classification of the released information.
[0032] Step 130: For each content category, based on the comment information of each piece of the released information under the content category, determine the emotion intensity information of the content category in each emotion category.
[0033] In practical applications, multiple emotion categories can be preset. Exemplarily, eight emotion categories can be set: expectation, happiness, trust, fear, surprise, sadness, disgust, and anger. The comment information of each piece of the released information under a content category can reflect the emotions induced by the public in the content category. Further, it can reflect the emotions in each emotion category. Based on this, in this step, based on the comment information of each piece of the released information under the content category, the emotion intensity information of the content category in each emotion category can be determined. The emotion intensity information of this emotion category is the quantitative information of the emotion intensity of the content category in this emotion category.
[0034] In this embodiment, by obtaining each piece of release information corresponding to the target public event and each piece of comment information of the release information, for each piece of the release information, the characteristics of the content category to which the release information belongs are qualitatively determined, and for each content category, based on each piece of the comment information of each piece of the release information under the content category, the emotional intensity information of the content category in each emotional category is quantitatively determined, enhancing the interpretability of emotions, thereby improving the accuracy of emotion mining.
[0035] Based on the above embodiments, for each piece of the release information, determining the content category to which the release information belongs, as Figure 2 shown, its specific implementation manner may include:
[0036] Step 210, input all the release information into a word segmentation tool to obtain a first word segmentation result output by the word segmentation tool.
[0037] Among them, the word segmentation dictionary of the word segmentation tool is supplemented with the proprietary new words corresponding to the target public event and / or the colloquial feature words in the social media, so that the word segmentation tool can output the proprietary new words and / or the colloquial feature words. Here, the proprietary new words can be the proprietary words corresponding to the target public event (i.e., a specific corpus background), that is, the terms used in the target public event, such as medical terms. Among them, the colloquial feature words can be the common colloquial expressions in the social media for the target public event, for example, "rest in peace". By adding proprietary new words and colloquial feature words to the word segmentation dictionary of the word segmentation tool, the dictionary is enriched, so that the words related to the target public event can be output more accurately, realizing the preprocessing of the release information.
[0038] In practical applications, if the target public event is for emotion mining for the first time, it is necessary to supplement the above proprietary new words and colloquial feature words to the word segmentation dictionary of the word segmentation tool, and then input all the release information into the word segmentation tool. If the target public event is not for emotion mining for the first time, since the above proprietary new words and colloquial feature words have been supplemented during the first emotion mining, all the release information can be directly input into the word segmentation tool.
[0039] Step 220, obtain the word vector space obtained by converting the first word segmentation result.
[0040] Step 230, obtain the probability that each word vector in the dimension-reduced word vector space belongs to each content category after soft clustering of the word vectors.
[0041] Soft clustering means that data is assigned to various categories with a certain probability. The algorithms for soft clustering include Gaussian Mixed Model (GMM), such as Fuzzy c-Means, and so on.
[0042] Taking eight content categories of release, traffic, domestic, documentary, policy, international, emotion, and diagnosis and treatment as examples, each word vector has a probability belonging to each content category.
[0043] Step 240: For each of the said release information, obtain the corresponding word vectors of the release information from the dimension-reduced word vector space, and determine the content category to which the release information belongs based on the probabilities that the corresponding word vectors of the release information belong to each of the content categories.
[0044] In this embodiment, by segmenting all release information and converting it into a word vector space, after soft clustering of each word vector in the obtained dimension-reduced word vector space, the probability that each word vector belongs to each content category can be obtained. Based on the probabilities that the corresponding word vectors of a release information belong to each content category, the content category to which a release information belongs can be accurately determined. In addition, since the data volume after segmenting all release information is large and the dimension of the converted word vector space is also high, the processing efficiency is low. However, after the word vector space is dimension-reduced, the processing efficiency can be greatly improved.
[0045] Based on any of the above embodiments, there are various specific implementation manners for determining the content category to which the release information belongs based on the probabilities that the corresponding word vectors of the release information belong to each of the content categories. The following lists two of the ways.
[0046] Way 1: For each of the content categories, calculate the average value of the probabilities that the corresponding word vectors of the release information belong to the content category; determine the content category with the highest average value as the content category to which the release information belongs. Taking eight content categories of release, traffic, domestic, documentary, policy, international, emotion, and diagnosis and treatment as examples, the corresponding word vectors of release information A include word vector a, word vector b, and word vector c. For the content category of release, the probabilities that word vector a belongs to release, word vector b belongs to release, and word vector c belongs to release can be averaged, and then by analogy, for each of the other content categories, the probabilities that the corresponding word vectors of the release information belong to the content category can be averaged. Finally, determine the content category with the highest average value as the content category to which the release information belongs.
[0047] The method of this embodiment can quickly determine the content category to which the release information belongs.
[0048] Method 2: Cluster the word vectors corresponding to the release information. Based on the clustering category with the largest number of included word vectors, for each of the content categories, calculate the average probability that the word vectors in the clustering category belong to the content category. Determine the content category to which the release information belongs as the content category with the highest average value.
[0049] Still taking the eight content categories of release, transportation, domestic, documentary, policy, international, emotion, and diagnosis and treatment as an example, the word vectors corresponding to the release information A include word vector a, word vector b, and word vector c. Cluster word vector a, word vector b, and word vector c. If word vector a and word vector b are classified into one category and word vector c is in a separate category, based on the clustering category where word vector a and word vector b are located, using the result of soft clustering as prior information, for the content category of transportation, calculate the average probability that word vector a belongs to transportation and the probability that word vector b belongs to transportation. Then, by analogy, for each of the other content categories, calculate the average probability that word vector a and word vector b belong to the content category. Finally, determine the content category with the highest average value as the content category to which the release information belongs.
[0050] The clustering here can be hard clustering. Hard clustering divides the data precisely into a certain category. For example, it can be the K-Means algorithm.
[0051] In the method of this embodiment, after soft clustering, using the result of soft clustering as prior information, perform clustering again. Word vectors belonging to the same content category are more likely to be classified into one category, and the clustering category with the largest number of included word vectors can better reflect the main content of the release information. Based on this, determining the category to which the release information belongs is more accurate.
[0052] Based on any of the above embodiments, before obtaining the word vector space obtained by converting the first word segmentation result, the above method may further include: obtaining different values of the hyperparameters used by the representation learning model; for each value of the hyperparameters, based on the representation learning model corresponding to the value, respectively convert the first word segmentation result to obtain a word vector space; based on the user's first input, select the word vector space obtained by the representation learning model corresponding to the value that meets the requirements.
[0053] Representation learning is used to convert the original data into a form that can be effectively exploited by machine learning. In this embodiment, the first word segmentation result is converted into a word vector space and saved through the representation learning model for subsequent processing.
[0054] Exemplarily, representation learning techniques for representing a learning model may be distributed word representation methods such as Brown Cluster, latent semantic analysis, Word2vec, or context-based word representation methods, etc. Among them, Word2vec is short for word to vector, which is a group of related models for generating word vectors.
[0055] Here, the first input of the user may be a selection operation for selecting the value that meets the requirements.
[0056] In this embodiment, by providing the user with the results of multiple hyperparameters for the user to select appropriate hyperparameters, the sensitivity analysis of the hyperparameters is thus realized.
[0057] In practical applications, if the target public event is for emotion mining for the first time, a word vector space obtained by a representation learning model corresponding to the value of the hyperparameter that meets the requirements can be selected based on the method provided in this embodiment. Further, the value of the selected hyperparameter that meets the requirements is saved. If the target public event is not for emotion mining for the first time, the first word segmentation result can be directly converted into a word vector space based on the representation learning model corresponding to the value of the selected hyperparameter that has been saved.
[0058] It should be noted that words with a word frequency lower than a certain threshold in the first word segmentation result can be deleted because when determining the hyperparameter, the influence of rare low-frequency words needs to be reduced and semantic connections need to be enhanced.
[0059] Based on any of the above embodiments, before obtaining the probability that each word vector in the reduced-dimensional word vector space belongs to each content category after soft clustering, it further includes: obtaining a preset number of clustering numbers; for each of the clustering numbers, performing soft clustering on each word vector in the reduced-dimensional word vector space according to the clustering number to obtain the probability that each word vector belongs to each content category and the result evaluation index of the soft clustering; based on the second input of the user, selecting the probability that each word vector belongs to each content category corresponding to the clustering number that meets the requirements.
[0060] Among them, the result evaluation index of soft clustering is used to evaluate the quality of the clustering result. For example, it may include indicators such as the Akaike information criterion (AIC), the Bayesian Information Criterion (BIC), and the silhouette coefficient.
[0061] Among them, the second input of the user may be a selection operation for selecting the clustering number that meets the requirements.
[0062] In this embodiment, the soft clustering results with multiple numbers of clusters and the result evaluation indexes are provided for the user. The user can name the content categories for each number of clusters based on each category in the soft clustering results of that number of clusters, and select the number of clusters with better naming explanations for the content categories. In this way, a more interpretable content classification result is obtained by combining the result evaluation indexes of clustering and manual experience verification.
[0063] In practical applications, if the target public event is for emotion mining for the first time, the probability that each word vector corresponding to the number of clusters meeting the requirements belongs to each content category can be selected based on the method provided in this embodiment. Further, the selected number of clusters meeting the requirements is saved. If the target public event is not for emotion mining for the first time, the word vectors in the dimension-reduced word vector space can be directly soft-clustered according to the saved selected number of clusters.
[0064] Based on any of the above embodiments, for each content category, the words included in the published information in this content category and their corresponding word frequencies are determined, and a word cloud is generated based on the determined words and their corresponding word frequencies. Exemplarily, a word cloud can be generated based on a set number of words with the highest word frequencies. The word cloud can also be visually displayed. The word cloud can visually highlight the keywords with higher frequencies in the text. In this way, it is convenient for the user to directly understand the main information in a content category.
[0065] Based on any of the above embodiments, within each preset unit time when the target public event evolves over time, for each content category, the number of published information in this content category and the percentage of the number of published information in this content category to the total number of published information in all content categories are determined. The preset unit time can be a day, a week, etc. A visual chart can also be generated and displayed based on the number of published information in this content category and the percentage of the number of published information in this content category to the total number of published information in all content categories. In this embodiment, the visualization of the published information in the content category is convenient for the user to understand the information publishing situation of each content category.
[0066] Based on any of the above embodiments, before soft-clustering the word vectors in the dimension-reduced word vector space according to the number of clusters, the above method may further include: performing principal component dimensionality reduction on the word vectors in the word vector space based on principal component analysis, and the dimension after dimensionality reduction is determined based on the third input of the user. The principal component analysis (PCA) method is a commonly used dimensionality reduction method, and the dimension after dimensionality reduction needs to be determined in advance, that is, the dimensionality reduction dimension.
[0067] The determination of the dimension reduction dimension needs to enable the variance value of each principal component retained after dimension reduction to account for at least a preset ratio of the total variance value of each component before dimension reduction. For example, the preset ratio is 70%.
[0068] In practical applications, a default set dimension reduction dimension can be provided, or an input control for the dimension reduction dimension can be provided for the user to input the dimension reduction dimension. First, the principal component dimension reduction can be performed on each word vector in the word vector space according to the default set dimension reduction dimension. The user can also adjust and update the dimension reduction dimension based on the dimension reduction result, and then perform the principal component dimension reduction on each word vector in the word vector space according to the adjusted dimension reduction dimension. A suitable dimension reduction dimension can be selected from the default set dimension reduction dimension and the adjusted dimension reduction dimension based on the third input of the user. The third input of the user is a selection operation of the dimension reduction dimension.
[0069] In this embodiment, a suitable dimension reduction dimension can be selected according to actual needs to meet the actual needs of the user.
[0070] If the target public event is for emotion mining for the first time, a suitable dimension reduction dimension can be selected from the default set dimension reduction dimension and the adjusted dimension reduction dimension. Further, the selected dimension reduction dimension is saved. If the target public event is not for emotion mining for the first time, the principal component dimension reduction can be directly performed on each word vector in the word vector space based on the saved selected dimension reduction dimension.
[0071] Based on any of the above embodiments, for each of the content categories, based on each of the comment information of the release information under the content category, to determine the emotion intensity information of the content category in each emotion category, the specific implementation manner may include:
[0072] First, all the comment information is input into a word segmentation tool to obtain a second word segmentation result output by the word segmentation tool.
[0073] Then, for each of the content categories, under the content category, for each of the comment information, the words included in the comment information are obtained from the second word segmentation result, the words included in the comment information are compared with the words in a pre-constructed emotion dictionary, and based on the comparison result, the words included in the comment information that exist in the emotion dictionary are determined. For each emotion discriminant model corresponding to each emotion category, the words included in the comment information that exist in the emotion dictionary are input into the emotion discriminant model corresponding to the emotion category to obtain the emotion intensity value of the comment information in the emotion category output by the emotion discriminant model corresponding to the emotion category. Based on the emotion intensity values of all the comment information under the content category in each emotion category, the emotion intensity information of the content category in each emotion category is determined;
[0074] Among them, the emotion discrimination model is trained based on the emotion dictionary and the emotion intensity value labels of each word in the emotion dictionary for each emotion category.
[0075] The emotion intensity value of an emotion category here is the quantization value of the emotion intensity of the emotion category.
[0076] In practical applications, an emotion discrimination model can be pre-trained for each emotion category. Taking the eight emotion categories of anticipation, happiness, trust, fear, surprise, sadness, disgust, and anger as an example, corresponding emotion discrimination models need to be pre-trained for each of the eight emotion categories to obtain eight emotion discrimination models.
[0077] In this embodiment, through the pre-trained emotion discrimination model, the emotion intensity value of each comment information in the emotion category can be obtained quickly and accurately. Furthermore, based on the emotion intensity values of all the comment information in each emotion category under the content category, the emotion intensity information of the content category in each emotion category can be determined quickly.
[0078] Based on any of the above embodiments, for each piece of published information, under each emotion category, the emotion intensity values of all the comment information of the published information in the emotion category are aggregated to obtain an aggregated value. When aggregating, the emotion intensity values of each comment information in the emotion category can be summed, or weighted sum can be performed according to the popularity of the comment information. Among them, the popularity of the comment information can be the total number of likes and comments corresponding to the comment information. The popularity of each comment information is normalized, that is, processed into a value between 0 and 1, and the normalized popularity of each comment information is used as the weight of each comment information, so as to perform weighted sum.
[0079] Furthermore, the above method may further include: within each preset unit time, under each emotion category, the emotion intensity values of all the comment information of the target public event in the emotion category are aggregated to obtain an aggregated value, and the percentage of the aggregated value in the total of the aggregated values under each emotion category is determined. Based on the aggregated values and percentages under each emotion category within each preset unit time, a visual chart representing the change of the aggregated values and the percentage fluctuation of each emotion category in the time series is generated and displayed.
[0080] Among them, the emotion dictionary is constructed in the following way:
[0081] Step 1: Obtain each comment information sample. Each comment information sample is marked with the emotional intensity value of the comment information sample in each emotional category. Use the emotional intensity value of the comment information sample in each emotional category as the emotional intensity value of the word segmentation in the comment information sample in each emotional category. Under each emotional category, based on the emotional intensity values of the word segmentation in each comment information sample in the emotional category, determine the sum of the emotional intensity values of the same word segmentation in the emotional category. Based on the word segmentation with the largest preset percentage of the sum of the emotional intensity values of the emotional category, obtain the emotional feature word set of the emotional category.
[0082] Each of the comment information samples can be a part of the comment information selected from all the release information corresponding to the target public event, or can be obtained through other means. Each comment information sample is marked with the emotional intensity value on each emotional category. For example, a comment information sample is "I believe in national peace and people's security". The emotional intensity value in the emotional category of trust is 3, and the emotional intensity values in other emotional categories such as fear are 0. Correspondingly, the word segmentation "believe" in "I believe in national peace and people's security" has an emotional intensity value of 3 in the emotional category of trust and an emotional intensity value of 0 in other emotional categories such as fear. Then, under the emotional category of trust, determine the sum of the emotional intensity values of "believe" in all comment information samples. Sort the emotional intensity values of the word segmentation in each emotional category from high to low, and select the word segmentation with the largest preset percentage of the sum to obtain the emotional feature word set of the emotional category. The preset percentage can be set according to actual needs. Exemplarily, it is 40%-60%, for example, it can be 50%.
[0083] The word segmentation with a higher ranking of the emotional intensity value of the emotional category is more representative. Therefore, these word segmentations can be used as the emotional feature words of the emotional category.
[0084] Step 2: Based on the word segmentation in each comment information sample, obtain the high-frequency word set based on the preset number of word segmentations with the highest word frequencies.
[0085] Specifically, based on the term frequency–inverse document frequency (TF-IDF), obtain the preset number of word segmentations with the highest word frequencies. The preset number can be set according to actual needs. Exemplarily, it is 250-450, for example, it can be 300.
[0086] Step 3: Based on the word segmentation in each comment information sample, determine at least one combined word. Based on the at least one combined word, obtain the combined word set. The combined word includes at least two word segmentations with an associated relationship obtained based on the association analysis algorithm.
[0087] Specifically, at least one combined word can be determined based on association analysis algorithms, such as the Apriori algorithm and the Frequent Pattern growth (FP-growth) algorithm. The Apriori algorithm is a frequent item set algorithm for mining association rules. Among them, the hyperparameters adopted by the association analysis algorithm can be determined with reference to related technologies and will not be elaborated here.
[0088] Among them, at least two word segments with an association relationship can refer to at least two word segments that usually appear simultaneously. For example, the two word segments "pay tribute to" and "angels in white" generally appear simultaneously to form "pay tribute to angels in white", and their correlation is relatively strong. Based on the association analysis algorithm, it can be analyzed that they are two word segments with an association relationship and form a combined word.
[0089] Step four, obtain the emotion dictionary based on the emotion feature word sets of each emotion category, the high-frequency word set, and the combined word set.
[0090] Among them, the emotion feature word sets of each emotion category and the high-frequency word set can form a single-word dictionary, and each word in the single-word dictionary includes one word segment. The combined word set can form a multi-word dictionary, and each word in the multi-word dictionary includes at least two word segments.
[0091] If the target public event is for emotion mining for the first time, the above emotion dictionary needs to be constructed. If the target public event is not for the first time in emotion mining, the already constructed emotion dictionary can be directly used.
[0092] The emotion discrimination model of the above emotion category is trained in the following way:
[0093] Randomly select a part from each comment information sample as the training set, and the other part as the test set; train the initial emotion discrimination model of the emotion category based on the training set, where the initial emotion discrimination model of the emotion category takes the elements (i.e., words) in the emotion dictionary as independent variables and the emotion intensity value as the dependent variable; test the emotion discrimination model of the emotion category based on the test set.
[0094] In implementation, the training and testing can be carried out based on the method of k-fold cross-validation. Specifically, each comment information sample can be divided into k parts, where k - 1 parts are used as the training set and 1 part is used as the test set. The value of k is a positive integer greater than 1 and can be set according to actual needs. Exemplarily, it is 5 - 20, for example, it can be 10.
[0095] The emotion discrimination model can be a linear regression model, a Least Absolute Shrinkage and Selection Operator (LASSO) regression model, an elastic regression model, etc. One or more regression evaluation metrics can also be combined, such as Root Mean Squared Error (RMSE), Mean Absolute Error (MAE), R-Squared, etc., to select the parameters of the model and verify the effectiveness of the model.
[0096] Similarly, if the target public event is for emotion mining for the first time, the above-mentioned emotion discrimination model needs to be constructed. If the target public event is not for emotion mining for the first time, the already constructed emotion discrimination model can be directly applied.
[0097] Based on any of the above embodiments, the emotion intensity information of the content category in the emotion category includes at least one of the following:
[0098] The first piece of information includes the emotion intensity value of the content category in the emotion category, and the emotion intensity value of the content category in the emotion category is the sum of the emotion intensity values of all the comment information under the content category in the emotion category.
[0099] The second piece of information includes the percentage of the emotion intensity value of the content category in the emotion category to the sum of the emotion intensity values of the content category in each emotion category.
[0100] The third piece of information includes at least one of the average value, median, upper quartile, lower quartile, maximum value, and minimum value of the emotion intensity values of each of the release information of the content category in the emotion category; wherein, the emotion intensity value of the release information in the emotion category is the sum of the emotion intensity values of all the comment information of the release information.
[0101] Among them, the average value of the emotion intensity values of each of the release information of the content category in the emotion category is also the emotion intensity value induced by all the comment information of a single release information of the content category.
[0102] The fourth piece of information includes, at each stage of the evolution of the target public event over time, the emotion intensity value of the content category in the emotion category during the stage; wherein, the emotion intensity value of the content category in the emotion category during the stage is the sum of the emotion intensity values of each of the release information of the content category in the emotion category during the stage.
[0103] In practical applications, users can autonomously pre-define the specific stages of the evolution of the target public event over time by combining the event cycle theory with event evolution indicators (such as key event nodes, the number of confirmed cases, disaster losses, and investment in disaster relief and reconstruction in a public health emergency). For example, three stages, namely the onset stage, the containment stage, and the recovery stage, can be defined. Each content category has an emotional intensity value for each emotional category within each stage.
[0104] The fifth piece of information includes, for each stage of the evolution of the target public event over time, at least one of the average value, median value, upper quartile value, lower quartile value, maximum value, and minimum value of the emotional intensity values of each piece of the published information in the content category within the stage in the emotional category; wherein, the emotional intensity value of the published information in the emotional category within the stage is the sum of the emotional intensity values of all the comment information of the published information within the stage.
[0105] Among them, the average value of the emotional intensity values of each piece of the published information in the content category within the stage is the emotional intensity value induced by a single piece of published information within the stage under the content category.
[0106] The sixth piece of information includes, for each stage of the evolution of the target public event over time, the emotional intensity value of the content category in the emotional category for each preset unit duration within the stage; the emotional intensity value of the content category in the emotional category for the preset unit duration within the stage is the sum of the emotional intensity values of each piece of the published information in the content category within the preset unit duration within the stage.
[0107] Taking the onset stage as an example, assuming the preset unit duration is one day and the onset stage includes 20 days, there are corresponding emotional intensity values for each emotional category every day.
[0108] The seventh piece of information includes, for each stage of the evolution of the target public event over time, at least one of the average value, median value, upper quartile value, lower quartile value, maximum value, and minimum value of the emotional intensity values of the content category in the emotional category for each preset unit duration within the stage.
[0109] Still taking the onset stage as an example, assuming the preset unit duration is one day and the onset stage includes 20 days, and the emotional category is trust, there are emotional intensity values of trust every day, so there are 20 emotional intensity values of trust. Based on these 20 emotional intensity values of trust, at least one of the average value, median value, upper quartile value, lower quartile value, maximum value, and minimum value is obtained.
[0110] The eighth piece of information includes, for each stage of the evolution of the target public event over time, at least one of the average value, median value, upper quartile value, lower quartile value, maximum value, and minimum value of the emotional intensity values of the single piece of the release information in the content category within each preset unit time period in the stage. The emotional intensity value of the single piece of the release information in the content category within the preset unit time period in the stage is the average value of the emotional intensity values of the release information in the content category within the preset unit time period in the stage.
[0111] Still taking the attack stage as an example, assuming that the preset unit time period is one day, the attack stage includes 20 days, the emotional category is trust, and there is a single piece of release information with an emotional intensity value in trust every day. Then there are 20 single pieces of release information with emotional intensity values in trust. Based on these 20 single pieces of release information with emotional intensity values in trust, at least one of the average value, median value, upper quartile value, lower quartile value, maximum value, and minimum value is obtained.
[0112] In this embodiment, by statistically analyzing the emotional intensity values of the comment information in each emotional category, various pieces of information in the emotional intensity information of the content category in each emotional category are obtained, so as to explain the emotional intensity of the content category in each emotional category from various perspectives, and the interpretability of emotions is stronger.
[0113] Based on any of the above embodiments, the above method may further include: generating a visual chart based on the emotional intensity information of the content category in each emotional category. The generated chart may include a line chart, a bar chart, a box plot, or an undirected graph, etc. In implementation, an appropriate chart can be selected according to the characteristics of the emotional intensity information. Through the visual chart, the emotional intensity information of the content category in each emotional category can be visually displayed, facilitating users to intuitively understand the emotional intensity information of the public.
[0114] Based on any of the above embodiments, for generating a visual chart based on the emotional intensity information of the content category in each emotional category, the specific implementation manner may include:
[0115] Step 1. If the emotional intensity information of the emotional categories includes the median, upper quartile, lower quartile, maximum value, and minimum value in the seventh information, based on the median, upper quartile, lower quartile, maximum value, and minimum value in the seventh information included in the emotional intensity information of each emotional category for the content category, generate a first box plot. The box plot gets its name because of its box-like shape. Among them, the median, upper quartile, lower quartile, maximum value, and minimum value in one stage of the seventh information included in the emotional intensity information of an emotional category correspond to a box in the first box plot. The first box plot includes the boxes for each stage corresponding to each emotional category, and is used to represent the overall emotional distribution per preset unit duration at each stage of the evolution of the target public event over time.
[0116] Exemplarily, each stage of the evolution of the target public event over time includes an outbreak period, a containment period, and a recovery period, the preset unit duration is one day, and the emotional categories include eight emotional categories: anticipation, joy, trust, fear, surprise, sadness, disgust, and anger. As Figure 3 shown in the first box plot, it includes the emotional intensity distributions of 8 groups of outbreak periods, containment periods, and recovery periods. In the illustrated state, from left to right, they correspond one by one to the eight emotional categories of anticipation, joy, trust, fear, surprise, sadness, disgust, and anger, reflecting the overall daily emotional distribution at each stage of the target public event.
[0117] Step 2. If the emotional intensity information of the emotional categories includes the median, upper quartile, lower quartile, maximum value, and minimum value in the eighth information, based on the median, upper quartile, lower quartile, maximum value, and minimum value in the eighth information included in the emotional intensity information of each emotional category for the content category, generate a second box plot. Among them, the median, upper quartile, lower quartile, maximum value, and minimum value in one stage of the eighth information included in the emotional intensity information of an emotional category correspond to a box in the first box plot. The second box plot includes the boxes for each stage corresponding to each emotional category, and is used to represent the emotional distribution of each single piece of published information per preset unit duration at each stage of the evolution of the target public event over time.
[0118] As Figure 4 shown in the second box plot, it includes the emotional intensity distributions of 8 groups of outbreak periods, containment periods, and recovery periods. In the illustrated state, from left to right, they correspond one by one to the eight emotional categories of anticipation, joy, trust, fear, surprise, sadness, disgust, and anger, reflecting the emotional distribution of each single piece of published information per day at each stage of the target public event.
[0119] Step 3. If the average value in the eighth information is included in the emotion intensity information of the emotion category, for each emotion category, under each content category, calculate the average of the average values in each stage of the eighth information as the baseline emotion, determine the deviation of the average value in each stage of the eighth information from the baseline emotion, and generate a line chart based on each deviation.
[0120] As Figure 5a and Figure 5b shown in the line charts corresponding to the eight emotion categories one by one. In the line chart corresponding to each emotion category, with the eight content categories of release, traffic, domestic, documentary, policy, international, emotion, and diagnosis and treatment, it can reflect the change of the emotion induced by each single piece of released information in each content category every day relative to the baseline emotion. In the state shown in the figure, taking the endpoints of all the line charts on the recovery side as a reference, in the order from top to bottom, explain the content categories corresponding to each line chart in turn. Among them, for the line chart corresponding to anticipation, the content categories corresponding to each line in turn are domestic, international, traffic, diagnosis and treatment, emotion, documentary, policy, and release; for the line chart corresponding to happiness, the content categories corresponding to each line in turn are emotion, traffic, policy, international, documentary, diagnosis and treatment, release, and domestic; for the line chart corresponding to surprise, the content categories corresponding to each line in turn are international, documentary, emotion, traffic, policy, diagnosis and treatment, domestic, and release; for the line chart corresponding to sadness, the content categories corresponding to each line in turn are emotion, traffic, release, international, policy, diagnosis and treatment, documentary, and domestic; for the line chart corresponding to trust, the content categories corresponding to each line in turn are emotion, traffic, diagnosis and treatment, release, documentary, international, domestic, and policy; for the line chart corresponding to fear, the content categories corresponding to each line in turn are international, traffic, diagnosis and treatment, domestic, documentary, emotion, policy, and release; for the line chart corresponding to disgust, the content categories corresponding to each line in turn are international, documentary, traffic, emotion, diagnosis and treatment, policy, domestic, and release; for the line chart corresponding to anger, the content categories corresponding to each line in turn are international, documentary, traffic, emotion, diagnosis and treatment, policy, domestic, and release.
[0121] It can be seen from the figure that for the same emotion category, different content categories will have different effects on the emotion, that is, they can regulate the emotion to different degrees. Thus, a regulatory effect is produced on the emotion margin.
[0122] For example, the content category of emotion raises public expectations during the outbreak period, but has no similar effect during the recovery period. At the same time, the content category of emotion has a significant effect on enhancing public trust in all three stages. The content category of transportation reduces public expectations and trust during the outbreak period, but increases public expectations and trust during the recovery period instead. The content category of diagnosis and treatment has a limited effect on enhancing public trust during the outbreak period, but has a significant effect during the recovery period. The content category of international affairs not only intensifies the public's emotions of fear, disgust, and anger, but also raises the public's expectations for the domestic content category.
[0123] Step 4: If the average value in the seventh information is included in the emotion intensity information of the emotion category, for each stage, based on the average value in the seventh information in the emotion intensity information of each emotion category, calculate the covariance matrix. Based on the average value in the seventh information in the emotion intensity information of each emotion category and the covariance matrix, generate an undirected graph according to the undirected graph model, and the undirected graph is used to represent the emotion correlation between each emotion category.
[0124] The undirected graph model therein can be the space method based on the nearest neighbor selection method, the interior point optimization method, etc., the clustering graph LASSO, the Bayesian adaptive graph LASSO, the joint graph LASSO, etc. based on the graph LASSO method. In the undirected graph model, when estimating the emotion correlation between each emotion category, the cross-validation method based on indicators such as the mean square error can be used to select the hyperparameter of the penalty factor.
[0125] A graph is composed of several vertices and edges connected to each other. An edge is connected by two vertices, and a graph without a direction is an undirected graph. An undirected graph can include a sparse graph and a dense graph. A sparse graph is a graph with few edges. A dense graph is a graph with many edges.
[0126] In this embodiment, the generated undirected graph has few edges. Therefore, a sparse graph is generated.
[0127] Such as Figure 6a 、 Figure 6b and Figure 6c As shown, visually display the emotion correlation network diagrams of each stage during the outbreak period (such as from January 20th to February 18th), the containment period (such as from February 19th to March 31st), and the recovery period (such as from April 1st to June 10th). Among them, the vertices represent the average daily emotion intensity (that is, the average value in the seventh information). Among them, the edges indicated by N1, N2, N3, and N4 represent negative correlations, and other edges represent positive correlations. The thickness of the edges represents the strength of the correlation between the two emotions.
[0128] In this embodiment, the generated box plot, line graph, and undirected graph can reflect emotions in more detail.
[0129] Based on any of the above embodiments, the above method may further include: fusing the generated first box plot, second box plot, line chart, and undirected graph. Specifically, the first box plot, second box plot, line chart, and undirected graph are placed in one graph. In this way, visual display can be performed together.
[0130] Based on any of the above embodiments, the above method may further include:
[0131] For each emotion category, based on the emotion intensity values of each content category in the emotion category, determine the content category with the maximum emotion intensity value and the content category with the minimum emotion intensity value in the emotion category. In this way, the interpretability of emotions is further increased.
[0132] In practical applications, the transparency, immediacy, and extensiveness of social media information dissemination provide timely development opportunities for risk communication based on information updates, and gradually become an important channel for risk communication. The information dissemination in the social media era has broken through the limitations of traditional time and space. By means of the transparent social environment of social media, various information, and the two-way communication platform, the public's ideas can be understood at the first time when a public emergency occurs, and rapid feedback and processing can be given in combination with the public's behavior. Only by adopting a rational discussion and communication method can the trust between risk managers and the public be better enhanced, and the efficiency of emergency management can be effectively improved. The changes in public behavior induced by social media information will significantly affect the development of public events, and the public's health decisions are also closely related to emotions. Therefore, in the context of public emergencies, especially when facing the social media environment, how to more deeply understand public emotions and optimize risk communication content according to public feedback is an urgent problem to be solved in emergency management coordination work.
[0133] Traditional emotion research mainly starts from psychological scales and sociological surveys to study the strategies of risk communication, public emotions, and their impacts on decision-making. There have been some modeling attempts, but the hypothesis basis mainly comes from the qualitative conclusions in empirical research.
[0134] In the embodiments of the present invention, a multi-dimensional emotion framework that simultaneously includes the expression of complex emotions and cognitive emotions (such as cognitive emotions like expectation and trust, etc.) is used. Based on an emotion dictionary of specific corpora, a regression analysis modeling method with strong interpretability is applied to understand and evaluate the characteristics and intensity differences of public emotion expressions under public emergencies, as well as the influence of the released information of each content category in each stage of the event's evolution over time on the modulation of emotion margins.
[0135] Specifically, the embodiments of the present invention mainly relate to the fields of decision analysis, natural language processing, and statistical analysis. Taking a large amount of text data such as the published information on social media and the public's comment information under sudden public events as input, through machine learning and statistical methods such as natural language processing, dimensionality reduction and clustering analysis, association analysis, and regression analysis, a content classification space for the information published on the platform is constructed, the category characteristics of the emotions expressed by the public in the comment text are qualitatively identified, the intensity differences of the public emotions are quantitatively evaluated, the influence effect of the content on the emotions is mined, and the above content is visually displayed using charts. On this basis, combined with each stage of the evolution of sudden public events over time, taking the average value of the public emotion intensity in each stage as the emotion baseline, the marginal adjustment effect of the emotions induced by various types of content relative to the emotion baseline is mined, and the visual display of the temporal evolution of the content and emotions is realized. Further, based on the undirected graph model estimation method, the temporal correlation of emotions in different disaster stages is described and visually displayed.
[0136] In addition, by automatically identifying the public emotions induced by the published information in social media, mining the emotion categories, emotion intensities, and emotion correlation relationships, and confirming the content categories and marginal adjustment degrees that affect the public emotions, it can help in the design of risk communication strategies.
[0137] Figure 7 An example of a schematic physical structure diagram of an electronic device is shown as Figure 7 As shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communication interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 can call the logical instructions in the memory 730 to execute the emotion mining method under public events, and the method includes: obtaining each piece of published information corresponding to the target public event and each piece of comment information of the published information; for each piece of the published information, determining the content category to which the published information belongs; for each content category, based on each piece of comment information of each piece of the published information under the content category, determining the emotion intensity information of the content category in each emotion category.
[0138] In addition, when the logical instructions in the above-mentioned memory 730 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0139] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the emotion mining method under the public events provided by the above-mentioned various methods. The method includes: obtaining each piece of release information corresponding to the target public event and each piece of comment information of the release information; for each piece of the release information, determining the content category to which the release information belongs; for each content category, based on each piece of the comment information of each piece of the release information under the content category, determining the emotion intensity information of the content category in each emotion category.
[0140] In yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the emotion mining method under the public events provided by the above-mentioned various methods. The method includes: obtaining each piece of release information corresponding to the target public event and each piece of comment information of the release information; for each piece of the release information, determining the content category to which the release information belongs; for each content category, based on each piece of the comment information of each piece of the release information under the content category, determining the emotion intensity information of the content category in each emotion category.
[0141] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0142] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for emotion mining under a public event, characterized in that, it includes: Obtain each piece of release information corresponding to the target public event and each piece of comment information of the release information; For each piece of the release information, determine the content category to which the release information belongs; For each content category, based on each piece of comment information of each piece of the release information under the content category, determine the emotion intensity information of the content category in each emotion category; The step of, for each content category, based on each piece of comment information of each piece of the release information under the content category, determining the emotion intensity information of the content category in each emotion category includes: Input all the comment information into a word segmentation tool to obtain a second word segmentation result output by the word segmentation tool; For each content category, under the content category, for each piece of comment information, obtain the words included in the comment information from the second word segmentation result, compare the words included in the comment information with the words in a pre-constructed emotion dictionary, and based on the comparison result, determine the words included in the comment information that exist in the emotion dictionary. For each emotion discrimination model corresponding to each emotion category, input the words included in the comment information that exist in the emotion dictionary into the emotion discrimination model corresponding to the emotion category to obtain the emotion intensity value of the comment information in the emotion category output by the emotion discrimination model corresponding to the emotion category. Based on the emotion intensity values of all the comment information under the content category in each emotion category, determine the emotion intensity information of the content category in each emotion category; wherein, the emotion discrimination model is trained based on the emotion dictionary and the emotion intensity value labels of each word in the emotion dictionary in the emotion category.
2. The method for emotion mining under a public event according to claim 1, characterized in that, the step of, for each piece of the release information, determining the content category to which the release information belongs includes: Input all the release information into a word segmentation tool to obtain a first word segmentation result output by the word segmentation tool; Obtain the word vector space obtained by converting the first word segmentation result; Obtain the probability that each word vector in the dimension-reduced word vector space belongs to each content category after soft clustering of the word vectors; For each piece of the release information, obtain the word vectors corresponding to the release information from the dimension-reduced word vector space, and based on the probability that the word vectors corresponding to the release information belong to each content category, determine the content category to which the release information belongs.
3. The method for emotion mining under a public event according to claim 2, characterized in that, the step of, based on the probability that the word vectors corresponding to the release information belong to each content category, determining the content category to which the release information belongs includes: For each content category, calculate the average value of the probabilities that the word vectors corresponding to the release information belong to the content category; determine the content category with the highest average value as the content category to which the release information belongs; Or, Cluster the word vectors corresponding to the release information. Based on the clustering category with the largest number of included word vectors, for each content category, calculate the average probability that each word vector in the clustering category belongs to the content category, and determine the content category to which the release information belongs as the content category with the highest average value.
4. The sentiment mining method under a public event according to claim 2, wherein, before obtaining the word vector space obtained by converting the first word segmentation result, it further includes: obtaining different values of hyperparameters used in the representation learning model; for each value of the hyperparameter, based on the representation learning model corresponding to the value, respectively convert the first word segmentation result to obtain a word vector space; based on the first input of the user, select the word vector space obtained by the representation learning model corresponding to the value that meets the requirements.
5. The sentiment mining method under a public event according to claim 2, wherein, before obtaining the probability that each word vector in the word vector space after dimensionality reduction belongs to each content category after soft clustering, it further includes: obtaining a preset number of clustering numbers; for each clustering number, perform soft clustering on each word vector in the word vector space after dimensionality reduction according to the clustering number to obtain the probability that each word vector belongs to each content category and the result evaluation index of the soft clustering; based on the second input of the user, select the probability that each word vector belongs to each content category corresponding to the clustering number that meets the requirements.
6. The sentiment mining method under a public event according to claim 5, wherein, before performing soft clustering on each word vector in the word vector space after dimensionality reduction according to the clustering number, it further includes: based on principal component analysis, perform principal component dimensionality reduction on each word vector in the word vector space, wherein the dimensionality after dimensionality reduction is determined based on the third input of the user.
7. The sentiment mining method under a public event according to claim 1, wherein, the sentiment dictionary is constructed in the following manner: obtain each comment information sample, where the comment information sample is marked with the sentiment intensity value of the comment information sample in each sentiment category, and use the sentiment intensity value of the comment information sample in each sentiment category as the sentiment intensity value of the word segmentation in the comment information sample in each sentiment category; under each sentiment category, based on the sentiment intensity value of the word segmentation in each comment information sample in the sentiment category, determine the sum of the sentiment intensity values of the same word segmentation in the sentiment category, and obtain the set of sentiment feature words for the sentiment category based on the preset percentage of the word segmentation with the largest sum of the sentiment intensity values of the sentiment category; obtain a set of high-frequency words based on the preset number of word segments with the highest word frequency in the word segments of each comment information sample; Based on the word segmentation of each of the comment information samples, at least one combined word is determined, and based on the at least one combined word, a combined word set is obtained, where the combined word includes at least two word segments with an associated relationship obtained based on an association analysis algorithm; Based on the emotion feature word sets of each of the emotion categories, the high-frequency word set, and the combined word set, the emotion dictionary is obtained.
8. The emotion mining method under a public event according to claim 7, characterized in that, the word segmentation dictionary of the word segmentation tool is supplemented with the proprietary new words corresponding to the target public event and / or the colloquial feature words in social media, so that the word segmentation tool can output the proprietary new words and / or the colloquial feature words.
9. The emotion mining method under a public event according to claim 7, characterized in that, the emotion intensity information of the content category in the emotion category includes at least one of the following information: The first information includes the emotion intensity value of the content category in the emotion category, and the emotion intensity value of the content category in the emotion category is the sum of the emotion intensity values of all the comment information under the content category in the emotion category; The second information includes the percentage of the emotion intensity value of the content category in the emotion category to the sum of the emotion intensity values of the content category in each of the emotion categories; The third information includes at least one of the average value, median, upper quartile, lower quartile, maximum value, and minimum value of the emotion intensity values of each of the release information of the content category in the emotion category; where the emotion intensity value of the release information in the emotion category is the sum of the emotion intensity values of all the comment information of the release information in the emotion category; The fourth information includes, in each stage of the evolution of the target public event over time, the emotion intensity value of the content category in the emotion category in the stage; where the emotion intensity value of the content category in the emotion category in the stage is the sum of the emotion intensity values of each of the release information of the content category in the emotion category in the stage; The fifth information includes at least one of the average value, median, upper quartile, lower quartile, maximum value, and minimum value of the emotion intensity values of each of the release information of the content category in the emotion category in each stage of the evolution of the target public event over time; The sixth information includes, in each stage of the evolution of the target public event over time, the emotion intensity value of the content category in each preset unit duration in the stage; the emotion intensity value of the content category in the preset unit duration in the stage is the sum of the emotion intensity values of each of the release information of the content category in the preset unit duration in the stage; The seventh information includes at least one of the average value, median, upper quartile, lower quartile, maximum value, and minimum value of the emotion intensity values of the content category in each preset unit duration in the emotion category in each stage of the evolution of the target public event over time; The eighth information includes, for each stage of the evolution of the target public event over time, at least one of the average, median, upper quartile, lower quartile, maximum, and minimum of the emotional intensity values of each single piece of the release information within the preset unit duration in the stage for the content category, and the emotional intensity value of each single piece of the release information within the preset unit duration in the stage for the content category is the average of the emotional intensity values of each piece of the release information within the preset unit duration in the stage for the content category.
10. The method for emotion mining under a public event according to claim 9, wherein, it further includes: generating a visual chart based on the emotional intensity information of the content category in each emotional category.
11. The method for emotion mining under a public event according to claim 10, wherein, the generating a visual chart based on the emotional intensity information of the content category in each emotional category includes: if the median, upper quartile, lower quartile, maximum, and minimum in the seventh information are included in the emotional intensity information of the emotional category, generating a first box plot based on the median, upper quartile, lower quartile, maximum, and minimum in the seventh information included in the emotional intensity information of the content category in each emotional category; if the median, upper quartile, lower quartile, maximum, and minimum in the eighth information are included in the emotional intensity information of the emotional category, generating a second box plot based on the median, upper quartile, lower quartile, maximum, and minimum in the eighth information included in the emotional intensity information of the content category in each emotional category; if the average in the eighth information is included in the emotional intensity information of the emotional category, for each emotional category, under each content category, averaging the averages in the eighth information for each stage as the baseline emotion, determining the deviation from the mean of the averages in the eighth information for each stage and the baseline emotion, and generating a line chart based on each deviation from the mean; if the average in the seventh information is included in the emotional intensity information of the emotional category, for each stage, calculating a covariance matrix based on the averages in the seventh information of the emotional intensity information of each emotional category, and generating an undirected graph based on the averages in the seventh information of the emotional intensity information of each emotional category and the covariance matrix according to an undirected graph model, where the undirected graph is used to represent the emotional correlation between each emotional category.
12. The method for emotion mining under a public event according to claim 9, wherein, it further includes: for each emotional category, determining the content category with the maximum emotional intensity value and the content category with the minimum emotional intensity value of the emotional category based on the emotional intensity values of the content category in the emotional category.
13. The method for emotion mining under a public event according to claim 1, wherein, Each of the emotion categories includes: anticipation, joy, trust, fear, surprise, sadness, disgust, and anger.
14. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the program, the steps of the emotion mining method under the public event described in any one of claims 1 to 13 are implemented.