Multi-level sentiment classification method and system for network public opinion
Through distributed message queues and multi-level emotion classification methods, social media data is obtained in real time, and multi-level emotion classification results are generated using Streaming K-means and Q-learning algorithms, which solves the problem of insufficient accuracy of emotion monitoring in the existing technology, and realizes real-time adaptation and accurate analysis of online public opinion.
Patent Information
- Application Number
- CN202510430252.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The existing online public opinion emotion classification methods mainly rely on a single data source and a simple classification model, which is difficult to adapt to the rapid migration of network hotspots and diversified emotional expression, resulting in insufficient accuracy of emotion monitoring.
Social media data is obtained in real time through distributed message queues, online clustering is used to use Streaming K-means algorithm to generate initial topic collections, and appropriate classification levels are selected based on topic state analysis model and Q-learning algorithm. Through dependency path analysis and information entropy value calculation, an information entropy value data set is constructed, and finally a hierarchical weight prediction model and emotion classification model are used to generate multi-level emotion classification results.
It achieves real-time adaptability and accuracy of online public opinion, can process multimodal data, provide more comprehensive and accurate sentiment analysis, and provides strong technical support for public opinion analysis and management.
Smart Images

Figure CN120354265A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network public opinion monitoring, and in particular, to a multi-level sentiment classification method and system for network public opinion. Background Art
[0002] The technology of network public opinion sentiment classification plays an important role in network public opinion analysis. With the popularization of social media, public sentiment and opinion have a profound impact on corporate brands, political decisions, and even social trends. Through natural language processing (NLP) technology, we can extract valuable information from a vast amount of network texts, understand public sentiment, and provide data support for decision-makers.
[0003] Existing network public opinion sentiment classification methods mainly rely on single data sources and simple classification models. For example, some methods only use text data for sentiment analysis, ignoring the rich information contained in multi-modal data such as images and videos. In addition, traditional classification models often rely on fixed label systems, making it difficult to adapt to the rapid migration of network hotspots and diverse sentiment expressions, resulting in a reduction in the accuracy of sentiment monitoring during the process of network public opinion analysis. Therefore, improvement is needed. Summary of the Invention
[0004] In order to improve the accuracy of sentiment monitoring during the process of network public opinion analysis, the present application provides a multi-level sentiment classification method and system for network public opinion.
[0005] In a first aspect, the above-mentioned invention object of the present application is achieved through the following technical solutions:
[0006] A multi-level sentiment classification method for network public opinion, the method comprising the steps of:
[0007] Capturing real-time data from social media platforms through a distributed message queue, filtering and cleaning non-text noise data, and forming a temporary corpus set for the real-time data of each time window of the social media platform according to a pre-set time window division mechanism;
[0008] A pre-set topic generation model performs online clustering on the temporary corpus set based on the Streaming K-means algorithm to generate an initial topic set;
[0009] A pre-set topic status analysis model analyzes the initial topic set to construct a topic status space, where the topic status space = (topic popularity Ht, sentiment dispersion Et, time span Δt));
[0010] The pre-set hierarchical depth decision model selects the action space a ∈ {three-level classification, four-level classification} through the Q-learning algorithm. The hierarchical depth decision model is provided with an accuracy reward function, which is used to improve the selection accuracy. The three-level classification structure is a three-layer logical relationship of theme, sub-theme, and sentiment, and the four-level classification structure is a four-layer logical relationship of event, theme, sub-theme, and sentiment;
[0011] Perform dependency path analysis on the current hierarchical structure, calculate the information entropy value of each node, and record the information entropy value of each node based on the pre-set timeline to construct an information entropy value dataset;
[0012] The pre-set hierarchical weight prediction model analyzes the information entropy value dataset based on the machine self-learning algorithm to generate the predicted weight coefficients of each level;
[0013] The pre-set sentiment classification model classifies the topics based on the predicted weight coefficients of each level to generate a multi-level sentiment classification result, which is used for online public opinion monitoring.
[0014] By adopting the above technical solutions, real-time acquisition and processing of social media data are carried out through a distributed message queue, online clustering is performed using the Streaming K-means algorithm to generate an initial topic set, a topic state space is constructed through a topic state analysis model, a suitable classification level is selected using the Q-learning algorithm, and an information entropy value dataset is constructed through dependency path analysis and information entropy value calculation. Finally, through the hierarchical weight prediction model and the sentiment classification model, a multi-level sentiment classification result is generated for online public opinion monitoring. This method can adapt to the changes in online public opinion in real time, improve the accuracy and timeliness of sentiment classification, and provide strong technical support for public opinion analysis and management.
[0015] In a preferred example of the present application, it can be further configured as follows: after the step of the pre-set topic generation model performing online clustering on the temporary corpus using the Streaming K-means algorithm to generate an initial topic set, the following steps are included:
[0016] Construct a dependency syntax tree for the text data within the same initial topic set, and extract semantic role labels as the first feature vector;
[0017] Use YOLOv7 to detect significant objects for the associated image data within the same initial topic set, and extract the HSV color space histogram as the second feature vector;
[0018] The pre-set feature fusion model calculates the fusion feature based on the cross-modal attention mechanism, and the fusion feature is used for the topic state analysis model to analyze the sentiment dispersion.
[0019] By adopting the above technical solution, multi-modal sentiment analysis can be performed on network public opinion data, and multi-level sentiment classification results can be generated. This method can not only process text data, but also combine image data to provide more comprehensive and accurate sentiment analysis, providing strong technical support for network public opinion monitoring.
[0020] In a preferred example of the present application, it can be further configured as follows: in the step where the pre-set hierarchical depth decision model selects the action space a ∈ {three-level classification, four-level classification} through the Q-learning algorithm, the steps include:
[0021] The hierarchical depth decision model extracts features from the topic state space, compares the topic heat with a pre-set topic heat standard value to generate a topic heat comparison result, and compares the sentiment dispersion with a pre-set sentiment dispersion standard value to generate a sentiment dispersion comparison result;
[0022] Based on the pre-set hierarchical classification rules, comprehensive analysis is performed on the topic heat comparison result and the sentiment dispersion comparison result to generate a selection result of the action space, and the selection result of the action space is used to determine the classification level.
[0023] By adopting the above technical solution, based on the preset rules, the model can automatically judge and select an appropriate classification level, improving the decision-making efficiency. The rules can be adjusted according to actual needs to adapt to different types of topics and application scenarios.
[0024] In a preferred example of the present application, it can be further configured as follows: after the step of performing comprehensive analysis on the topic heat comparison result and the sentiment dispersion comparison result based on the pre-set hierarchical classification rules to generate a selection result of the action space, and the selection result of the action space is used to determine the classification level, when the selection result of the action space is three-level classification, the steps include:
[0025] The pre-set LDA topic analysis model collects topics from the initial topic set to generate a topic word set T = {T1, T2,..., T k};
[0026] The pre-set sub-topic clustering analysis model performs sub-topic clustering on the topic word set based on the dynamic density clustering rule of the density clustering algorithm;
[0027] The pre-set BERT-CRF joint model analyzes the corresponding temporary corpus set in the sub-topic clustering to generate the sentiment polarity and intensity values corresponding to the sub-topics;
[0028] Associate the set of subject terms, the sub - topic clustering, and the sentiment polarity and intensity values to generate a sentiment classification result for three - level classification.
[0029] By adopting the above - mentioned technical solution, a sentiment classification result for three - level classification is generated, providing sentiment analysis from macro to micro, improving the detail level of classification, integrating the theme, sub - theme, and sentiment analysis results, and providing comprehensive data support for network public opinion monitoring.
[0030] In a preferred example, this application can be further configured as follows: After comprehensively analyzing the topic heat comparison result and the sentiment dispersion comparison result based on pre - set hierarchical classification rules to generate a selection result of the action space, and the selection result of the action space is used to determine the classification level. When the selection result of the action space is four - level classification, the following steps are included:
[0031] Extract event features from the initial topic set based on the BERT - CRF joint model to generate a set of keywords for emergency events;
[0032] The pre - set topic association model associates the topics corresponding to the keywords for emergency events based on the knowledge graph algorithm technology to construct an event - topic association matrix;
[0033] The pre - set event relationship reasoning model analyzes the event - topic association matrix based on the machine self - learning algorithm to generate the correlation degree of implicit event associations, and the correlation degree of implicit event associations is used to judge the correlation between two events and is used for subsequent prediction and judgment of the public opinion direction of events.
[0034] By adopting the above - mentioned technical solution, the implicit associations between different events can be discovered, providing a basis for public opinion prediction and judgment. By analyzing the correlation between events, the development direction of public opinion can be predicted in advance to assist in decision - making.
[0035] In a preferred example, this application can be further configured as follows: After the step where the pre - set event relationship reasoning model analyzes the event - topic association matrix based on the machine self - learning algorithm to generate the correlation degree of implicit event associations, the following steps are included:
[0036] Obtain the propagation node features of relevant events and associate the node features with the events to construct an event - node data set;
[0037] The pre - set node intervention model analyzes the event - node data set to identify the key nodes for event propagation, and the key nodes are used to improve the effectiveness of public opinion intervention.
[0038] By adopting the above technical solutions, multi-modal sentiment analysis can be performed on network public opinion data to generate multi-level sentiment classification results. Especially in the aspect of identifying key nodes in the event dissemination process, by constructing an event-node dataset and a node intervention model, the key nodes that have the greatest impact on event dissemination can be accurately identified, providing a scientific basis for public opinion intervention. This method can not only process text data but also combine image data to provide more comprehensive and accurate sentiment analysis, providing powerful technical support for network public opinion monitoring.
[0039] In a preferred example of the present application, it can be further configured as follows: after comprehensively analyzing the topic heat comparison result and the sentiment dispersion comparison result based on preset hierarchical classification rules to generate a selection result of the action space, and the selection result of the action space is used to determine the classification level. If it is detected that the subject has cross-platform dissemination characteristics, the four-level classification module is automatically activated.
[0040] If there are conflicting labels in the sentiment label set, the four-level classification is forced to be enabled.
[0041] In a second aspect, the above object of the present invention of the present application is achieved by the following technical solutions:
[0042] A multi-level sentiment classification device for network public opinion, the device includes: a temporary corpus set construction unit, configured to capture real-time data of social media platforms through a distributed message queue, filter and clean non-text noise data, and form a temporary corpus set for the real-time data of social media platforms corresponding to each time window according to a preset time window division mechanism;
[0043] An initial topic set generation unit, configured to preset a topic generation model to perform online clustering on the temporary corpus set based on the Streaming K-means algorithm to generate an initial topic set;
[0044] A topic state space construction unit, configured to preset a topic state analysis model to analyze the initial topic set to construct a topic state space, where the topic state space = (topic heat Ht, sentiment dispersion Et, time span Δt));
[0045] An action space selection unit, configured to preset a hierarchical depth decision model to select an action space a ∈ {three-level classification, four-level classification} through the Q-learning algorithm, where the hierarchical depth decision model is provided with an accuracy reward function, and the accuracy reward function is used to improve the selection accuracy. The three-level classification structure is a three-layer logical relationship of theme, sub-theme, and sentiment, and the four-level classification structure is a four-layer logical relationship of event, theme, sub-theme, and sentiment;
[0046] An information entropy value dataset construction unit, which is used to perform dependency path analysis on the current hierarchical structure, calculate the information entropy values of each node, and record the information entropy values of each node based on a preset time axis to construct an information entropy value dataset;
[0047] A prediction weight coefficient generation unit, which is used to preset a hierarchical weight prediction model to analyze the information entropy value dataset based on a machine self-learning algorithm to generate prediction weight coefficients for each level;
[0048] A sentiment classification generation unit, which is used to preset a sentiment classification model to classify topics based on the prediction weight coefficients of each level to generate multi-level sentiment classification results, and the sentiment classification results are used for network public opinion monitoring.
[0049] Thirdly, the above object of the present application is achieved by the following technical solutions:
[0050] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above multi-level sentiment classification method for network public opinion are implemented.
[0051] Fourthly, the above object of the present application is achieved by the following technical solutions:
[0052] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above multi-level sentiment classification method for network public opinion are implemented.
[0053] In summary, the present application includes at least one of the following beneficial technical effects:
[0054] 1. Real-time obtain and process social media data through a distributed message queue, perform online clustering using the Streaming K-means algorithm to generate an initial topic set. Construct a topic state space through a topic state analysis model, use the Q-learning algorithm to select an appropriate classification level, and construct an information entropy value dataset through dependency path analysis and information entropy value calculation. Finally, generate multi-level sentiment classification results through a hierarchical weight prediction model and a sentiment classification model for network public opinion monitoring. This method can adapt to the changes of network public opinion in real time, improve the accuracy and timeliness of sentiment classification, and provide strong technical support for public opinion analysis and management;
[0055] 2. It can perform multi-modal sentiment analysis on network public opinion data and generate multi-level sentiment classification results. This method can not only process text data, but also combine image data to provide more comprehensive and accurate sentiment analysis, providing strong technical support for network public opinion monitoring;
[0056] 3. It can discover the implicit associations between different events, providing a basis for public opinion prediction and judgment. By analyzing the correlation between events, it can predict the development direction of public opinion in advance and assist in decision-making.
[0057] 4. It can perform multi-modal sentiment analysis on online public opinion data and generate multi-level sentiment classification results. Especially in the identification of key nodes in event dissemination, by constructing an event-node dataset and a node intervention model, it can accurately identify the key nodes that have the greatest impact on event dissemination, providing a scientific basis for public opinion intervention. This method can not only process text data but also combine image data to provide more comprehensive and accurate sentiment analysis, providing strong technical support for online public opinion monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 is a flowchart of a multi-level sentiment classification method for online public opinion in an embodiment of the present application;
[0059] Figure 2 is a schematic block diagram of a multi-level sentiment classification system for online public opinion in an embodiment of the present application;
[0060] Figure 3 is a schematic diagram of an electronic device in an embodiment of the present application.
[0061] Reference Numerals in the Drawings:
[0062] 1. Temporary Corpus Construction Unit; 2. Initial Topic Set Generation Unit; 3. Topic State Space Construction Unit; 4. Action Space Selection Unit; 5. Information Entropy Value Dataset Construction Unit; 6. Prediction Weight Coefficient Generation Unit; 7. Sentiment Classification Generation Unit. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] The present application will be further described in detail below with reference to the accompanying drawings.
[0064] In one embodiment, as Figure 1 shown, the present application discloses a multi-level sentiment classification method for online public opinion, specifically including the following steps:
[0065] S10: Capture real-time data from social media platforms through a distributed message queue, filter and clean non-text noise data, and form a temporary corpus for the real-time data of each time window of the social media platform according to a pre-set time window division mechanism;
[0066] Specifically, in the embodiment of the present application for step S10, an example is:
[0067] Data acquisition: Use a distributed message queue (such as Apache Kafka) to obtain data in real time from social media platforms such as Weibo, WeChat, and Douyin. This data includes text, pictures, videos, etc. posted by users.
[0068] Data preprocessing: Clean and denoise the acquired data, removing irrelevant non-text data (such as the noisy parts in pictures and videos) and duplicate content.
[0069] Time window division: Divide the real-time data into multiple temporary corpus sets according to a preset time window (such as every 10 minutes as a window). For example, within the time window from 10:00 to 10:10, all relevant data collected constitutes a temporary corpus set.
[0070] According to step S10, social media data can be acquired and processed in real time, ensuring the timeliness of the analysis results. By cleaning and denoising, the data quality is improved, providing a reliable data basis for subsequent analysis. Dividing the data by time window facilitates subsequent clustering and analysis.
[0071] S20: A preset topic generation model performs online clustering on the temporary corpus set based on the Streaming K-means algorithm to generate an initial topic set;
[0072] Specifically, use the Streaming K-means algorithm to perform online clustering on the temporary corpus set within each time window. This algorithm can dynamically adjust the clustering center to adapt to the real-time changes of the data, extract features such as keywords and topic words from the text, combine multi-dimensional data such as user information and posting time to form a comprehensive feature vector, and identify new clustering clusters through the clustering results. These clusters represent emerging topics. For example, if a clustering cluster contains a large number of discussions about "vaccine side effects", then generate the initial topic "Concerns about vaccine side effects".
[0073] The Streaming K-means algorithm can process data streams in real time, dynamically adjust the clustering results to adapt to the rapid changes of online public opinion, can automatically discover emerging topics, providing a basis for subsequent sentiment analysis. The online clustering algorithm can operate efficiently on large-scale data to meet the needs of real-time analysis.
[0074] S30: A preset topic status analysis model analyzes the initial topic set to construct a topic status space, where the topic status space = (topic popularity Ht, emotional dispersion Et, time span Δt));
[0075] Specifically,
[0076] The topic popularity Ht is obtained by calculating the occurrence frequency of each topic within a specific time window and the user engagement (such as the number of likes and comments). For example, the topic "Concerns about vaccine side effects" appeared 1000 times within the time window from 10:00 to 10:10, and the user engagement was 500 times. The calculated Ht is 0.8. Sentiment analysis is performed on the text under each topic to calculate the dispersion degree Et of the sentiment distribution. For example, the sentiment distribution under the topic "Concerns about vaccine side effects" is: positive 20%, neutral 30%, negative 50%, and the calculated Et is 0.6. The duration Δt of the topic is recorded. For example, the topic "Concerns about vaccine side effects" started at 10:00 and lasted until 10:30, and Δt is 30 minutes.
[0077] By constructing a topic state space, comprehensively describe the popularity, sentiment distribution, and time span of the topic, providing multi-dimensional data support for subsequent analysis, being able to accurately capture the sentiment distribution and popularity changes of the topic, improving the accuracy of the analysis, and updating the topic state space in real time to adapt to the dynamic changes of the topic.
[0078] S40: The pre-set hierarchical depth decision model selects the action space a ∈ {three-level classification, four-level classification} through the Q-learning algorithm, where the hierarchical depth decision model is provided with an accuracy reward function, and the accuracy reward function is used to improve the selection accuracy;
[0079] In this application, the three-level classification structure is a three-layer logical relationship of theme, sub-theme, and sentiment, and the four-level classification structure is a four-layer logical relationship of event, theme, sub-theme, and sentiment;
[0080] Specifically, use the Q-learning algorithm to train the hierarchical depth decision model. The model selects the action space a ∈ {three-level classification, four-level classification} according to the current state of the topic (such as popularity, sentiment dispersion degree, time span), defines the reward function, and gives a positive reward when the classification level selected by the model matches the complexity of the actual data, otherwise gives a negative reward. For example, if a sudden public opinion event (such as "privacy leakage") requires more detailed classification, the model gets a positive reward for selecting the four-level classification. For a regular topic (such as "general discussions in daily life"), the model selects the three-level classification (theme → sub-theme → sentiment); for a sudden public opinion (such as "major social events"), the model selects the four-level classification (event → theme → sub-theme → sentiment).
[0081] The model can automatically select the appropriate classification level according to the complexity and importance of the topic, improving the flexibility and accuracy of classification. Through the Q-learning algorithm and the accuracy reward function, continuously optimize the selection strategy of the model, improve the accuracy of classification, and be able to adapt to different types of public opinion events, from regular topics to sudden public opinions, and provide appropriate classification structures.
[0082] S50: Perform dependency path analysis on the current hierarchical structure, calculate the information entropy value of each node, and record the information entropy value of each node based on a preset time axis to construct an information entropy value dataset;
[0083] Perform dependency path analysis on the current hierarchical structure (such as three-level classification or four-level classification) to determine the dependency relationships between nodes. For example, in four-level classification, the event "privacy leakage" depends on the topic "data security", the topic "data security" depends on the sub-topic "user privacy", and the sub-topic "user privacy" depends on the emotion "worry". Calculate the information entropy value of each node, which reflects the uncertainty of the node. For example, for the node "privacy leakage", calculate its probability distribution of occurrence in all relevant discussions, and the obtained information entropy value is 0.5. Based on a preset time axis (such as recording once every 10 minutes), record the information entropy value of each node to construct an information entropy value dataset.
[0084] Quantify the uncertainty of each node through the information entropy value, provide data support for subsequent weight prediction, and record the change of the information entropy value in real time to reflect the dynamic characteristics of the topic.
[0085] S60: A preset hierarchical weight prediction model analyzes the information entropy value dataset based on a machine learning algorithm to generate prediction weight coefficients for each level;
[0086] Specifically, use a machine learning algorithm (such as neural network, random forest, etc.) to analyze the information entropy value dataset. The model predicts the weight coefficient of each level according to the change trend of the information entropy value. For example, for the node "privacy leakage", predict that its weight coefficient at the event level is 0.7, at the topic level is 0.5, at the sub-topic level is 0.3, and at the emotion level is 0.2. Apply the predicted weight coefficients to subsequent sentiment classification to improve the accuracy of classification.
[0087] Dynamically predict the weight coefficient according to the information entropy value, reflect the importance of each level node, and improve the accuracy of sentiment classification through the adjustment of the weight coefficient. The model can continuously learn and optimize the weight prediction to adapt to the change of data.
[0088] S70: A preset sentiment classification model classifies the topic based on the predicted weight coefficients of each level to generate a multi-level sentiment classification result, and the sentiment classification result is used for network public opinion monitoring;
[0089] Specifically, a pre-trained sentiment classification model (such as BERT) is used in combination with the predicted weight coefficients to perform sentiment classification on the topics at each level. For example, for the four-level classification structure "Event → Topic → Sub-topic → Sentiment", the model first classifies the event "Privacy Leakage", then classifies the topic "Data Security", then classifies the sub-topic "User Privacy", and finally classifies the sentiment "Concern". The generated multi-level sentiment classification results are used for online public opinion monitoring to help relevant departments timely understand the public's sentiment tendency and focus of attention.
[0090] Provide multi-level sentiment classification from macro to micro to comprehensively reflect the structure and sentiment distribution of online public opinion. Through multi-level classification, accurately locate the key nodes and sentiment tendency of public opinion, improve the efficiency and effect of monitoring, provide a scientific basis for public opinion management, brand marketing, policy formulation, etc., and assist in decision-making.
[0091] In summary, for steps S10 - S70: Real-time obtain and process social media data through a distributed message queue, use the Streaming K-means algorithm for online clustering to generate an initial topic set. Construct a topic state space through a topic state analysis model, use the Q-learning algorithm to select an appropriate classification level, and construct an information entropy value data set through dependency path analysis and information entropy value calculation. Finally, through a hierarchical weight prediction model and a sentiment classification model, generate multi-level sentiment classification results for online public opinion monitoring. This method can adapt to the changes in online public opinion in real time, improve the accuracy and timeliness of sentiment classification, and provide strong technical support for public opinion analysis and management.
[0092] After step S20: The pre-set topic generation model performs online clustering on the temporary corpus based on the Streaming K-means algorithm to generate an initial topic set, the following steps are included:
[0093] S21: Construct a dependency syntax tree for the text data within the same initial topic set, and extract semantic role labels as the first feature vector;
[0094] Specifically, assume we have an initial topic set that contains the following text data:
[0095] "The company should strengthen data security protection to prevent privacy leakage."
[0096] "The introduction of data security regulations will help reduce privacy leakage incidents."
[0097] "Users express concern about the privacy leakage issue."
[0098] Extract semantic role labels from the dependency syntactic tree, such as subject (nsubj), object (dobj), adverbial (advmod), etc., to form the first feature vector. For example, for the above text, the first feature vector can be represented as:
[0099] [Company, strengthen, data security protection, prevent privacy leakage]
[0100] Through the dependency syntactic tree, the semantic structure of the text can be deeply understood, key semantic roles can be extracted, and the accuracy of sentiment classification can be improved. As a feature vector, semantic role labels can provide rich semantic information, enhancing the model's expressive ability and classification effect.
[0101] S22: Use YOLOv7 to detect significant objects in the associated image data within the same initial topic set, and extract the HSV color space histogram as the second feature vector;
[0102] Specifically, assume that in the same initial topic set, there is associated image data, such as pictures posted by users. We use the YOLOv7 detection algorithm to detect significant objects in these images. For example, it is detected that the images contain significant objects such as "group meeting", "crowd", "employee", etc.
[0103] Then, extract the HSV color space histogram for the detected significant object regions. For example, use the OpenCV library to calculate the HSV histogram to obtain the distribution of each color channel. Take these histogram data as the second feature vector.
[0104] S23: The pre-set feature fusion model calculates the fusion feature based on the cross-modal attention mechanism, and the fusion feature is used for the topic state analysis model to analyze the sentiment dispersion;
[0105] Specifically, combined with the image data, through significant object detection and color space analysis, the feature vector is enriched, and the comprehensiveness and accuracy of sentiment classification are improved. By detecting significant objects in the images, the user's emotional expression can be better understood, especially in the case of insufficient or unclear text information.
[0106] Suppose we have extracted the first feature vector (semantic role labels) from the text data and the second feature vector (HSV color space histogram) from the image data. The feature fusion model fuses these two feature vectors based on the cross-modal attention mechanism.
[0107] The cross-modal attention mechanism dynamically adjusts the contributions of both by calculating the attention weights between the text and image features. For example, if the text feature is more important in a certain topic, the model will give it a greater weight; vice versa. The fused feature vector can be represented as:
[0108] Fusion feature = α * first feature vector + β * second feature vector
[0109] Where α and β are weight coefficients calculated through the attention mechanism.
[0110] The fused feature vector is used by the topic state analysis model to analyze the emotional dispersion. For example, the emotional dispersion Et is obtained by calculating the emotional distribution of the fusion feature.
[0111] In summary, for steps S21 - S23, multi-modal sentiment analysis can be performed on network public opinion data to generate multi-level sentiment classification results. This method can not only process text data but also combine image data to provide more comprehensive and accurate sentiment analysis, providing strong technical support for network public opinion monitoring.
[0112] In step S40: In the step where the pre-set hierarchical depth decision model selects the action space a ∈ {three-level classification, four-level classification} through the Q-learning algorithm, the following steps are included:
[0113] S41: The hierarchical depth decision model extracts features from the topic state space, compares the topic heat with the pre-set topic heat standard value to generate a topic heat comparison result, and compares the emotional dispersion with the pre-set emotional dispersion standard value to generate an emotional dispersion comparison result;
[0114] Specifically, assume that the pre-set topic heat standard value is 0.7. For a certain topic, the calculated topic heat Ht is 0.8. Comparing Ht with the standard value, the generated topic heat comparison result is "higher than the standard value".
[0115] Assume that the pre-set emotional dispersion standard value is 0.5. For the same topic, the calculated emotional dispersion Et is 0.6. Comparing Et with the standard value, the generated emotional dispersion comparison result is "higher than the standard value".
[0116] S42: Based on the pre-set hierarchical classification rules, comprehensive analysis is performed on the topic heat comparison result and the emotional dispersion comparison result to generate a selection result for the action space, and the selection result of the action space is used to determine the classification level.
[0117] Specifically, the hierarchical classification rules: The pre-set hierarchical classification rules are as follows:
[0118] If the topic heat is higher than the standard value and the emotional dispersion is higher than the standard value, then select four-level classification.
[0119] Otherwise, select three-level classification.
[0120] For the topic in this application, since both the topic popularity and the emotional dispersion degree are higher than the standard values, according to the rules, the selection result of the action space is a four-level classification.
[0121] For steps S41 - S42, based on preset rules, the model can automatically judge and select an appropriate classification level, improving the decision-making efficiency. The rules can be adjusted according to actual needs to adapt to different types of topics and application scenarios.
[0122] After step S42: Based on the preset hierarchical classification rules, comprehensively analyze the topic popularity comparison result and the emotional dispersion degree comparison result to generate the selection result of the action space, and the selection result of the action space is used to determine the classification level. When the selection result of the action space is a three-level classification, the following steps are included:
[0123] S421: The preset LDA topic analysis model performs topic collection on the initial topic set to generate a topic word set T = {T1, T2,..., T k};
[0124] Specifically, for example: Suppose we have an initial topic set that contains the following text data:
[0125] "The company should strengthen data security protection to prevent privacy leakage."
[0126] "The introduction of data security regulations will help reduce privacy leakage incidents."
[0127] "Users are concerned about the privacy leakage problem."
[0128] Use the LDA topic analysis model to perform topic collection on these texts. First, preprocess the texts, including operations such as word segmentation and stop word removal. Then, convert the processed texts into a bag-of-words model. Next, run the LDA model. Suppose we set the number of topics to 3. After model training, the following topic word set is obtained:
[0129] Topic 1: Data security, protection, privacy leakage;
[0130] Topic 2: Rules, introduction, reduction, incidents;
[0131] Topic 3: Users, concern, problem.
[0132] The LDA model can automatically discover the hidden topics in the texts, summarize a large amount of text data into several topics, improve the analysis efficiency, and the generated topic word set can clearly reflect the core content of each topic, providing a basis for subsequent sub-topic clustering and sentiment analysis.
[0133] S422: The pre-set sub-topic clustering analysis model performs sub-topic clustering on the set of topic words based on the dynamic density clustering rule of the density clustering algorithm;
[0134] For the set of topic words generated above, use the density clustering algorithm (such as DBSCAN) for sub-topic clustering. Assume that the parameter settings of the density clustering algorithm are: eps = 0.5, min_samples = 2. Cluster the set of topic words to obtain the following sub-topic clustering results:
[0135] Sub-topic 1: Data security, protection, privacy leakage;
[0136] Sub-topic 2: Regulations, introduction, reduction, events;
[0137] Sub-topic 3: Users, concerns, problems.
[0138] The density clustering algorithm can automatically divide sub-topics according to the density distribution of data, adapt to data distributions of different shapes and densities, and the dynamic density clustering rule can adjust the clustering results according to the real-time changes of data, improving the adaptability and accuracy of the model.
[0139] S423: The pre-set BERT-CRF joint model analyzes the corresponding temporary corpus set in the sub-topic clustering to generate the sentiment polarity and intensity values corresponding to the sub-topics;
[0140] Specifically, for the above sub-topic clustering results, use the BERT-CRF joint model to perform sentiment analysis on the corresponding temporary corpus set. For example, for sub-topic 1 (data security, protection, privacy leakage), the sentiment polarity and intensity values obtained by the model analysis are as follows:
[0141] Text 1: "The company should strengthen data security protection to prevent privacy leakage." Sentiment polarity: Positive, intensity value: 0.8;
[0142] Text 2: "The introduction of data security rules will help reduce privacy leakage events." Sentiment polarity: Positive, intensity value: 0.7;
[0143] Text 3: "Users are concerned about privacy leakage problems." Sentiment polarity: Negative, intensity value: 0.6.
[0144] The BERT-CRF joint model can accurately identify the sentiment polarity and intensity in the text, provide detailed sentiment analysis results, generate corresponding sentiment polarity and intensity values for each sub-topic, and help to understand the sentiment distribution under different sub-topics more deeply.
[0145] S424: Associate the set of topic words, the sub-topic clustering, and the sentiment polarity and intensity values to generate a sentiment classification result with three-level classification;
[0146] Specifically, the above-mentioned set of subject terms, sub-topic clustering, and sentiment polarity and intensity values are associated to generate the following sentiment classification results for the three-level classification:
[0147] Theme 1: Data security, protection, privacy leakage;
[0148] Sub-theme 1: Data security, protection, privacy leakage;
[0149] Sentiment polarity: Positive, intensity value: 0.8;
[0150] Sentiment polarity: Positive, intensity value: 0.7;
[0151] Sentiment polarity: Negative, intensity value: 0.6;
[0152] Theme 2: Regulations, introduction, reduction, events;
[0153] Sub-theme 2: Regulations, introduction, reduction, events;
[0154] Sentiment polarity: Positive, intensity value: 0.7;
[0155] Theme 3: Users, concerns, problems;
[0156] Sub-theme 3: Users, concerns, problems;
[0157] Sentiment polarity: Negative, intensity value: 0.6.
[0158] Generate the sentiment classification results for the three-level classification, provide sentiment analysis from macro to micro, improve the classification detail, and integrate the theme, sub-theme and sentiment analysis results to provide comprehensive data support for online public opinion monitoring.
[0159] It should be noted that when the cross-platform propagation feature of the said theme is detected, the four-level classification module is automatically activated;
[0160] When there are conflicting labels in the sentiment label set, the four-level classification is forced to be enabled.
[0161] When the selection result of the action space is four-level classification, the steps are as follows:
[0162] S425: Extract event features from the initial topic set based on the BERT-CRF joint model to generate a set of keywords for emergencies;
[0163] Specifically, it can automatically identify the keywords of emergencies from a large amount of text, quickly respond to the changes in public opinion, and the generated set of keywords for emergencies provides a basis for subsequent event analysis and association.
[0164] S426: The pre-set topic association model associates the topics corresponding to the keywords of the emergency event based on the knowledge graph algorithm technology to construct an event-topic association matrix;
[0165] For the above-mentioned set of emergency event keywords, use the knowledge graph algorithm technology to construct an event-topic association matrix. For example, the constructed association matrix is as follows:
[0166] Privacy leakage → Topic 1 (data security, protection, privacy leakage);
[0167] Data security regulations → Topic 2 (regulations, introduction, reduction, events).
[0168] Beneficial effects: Through the knowledge graph algorithm technology, the association relationship between the emergency event and the topic can be clearly displayed.
[0169] Structured representation: The event-topic association matrix provides structured data support for subsequent event relationship reasoning.
[0170] S427: The pre-set event relationship reasoning model analyzes the event-topic association matrix based on the machine self-learning algorithm to generate the correlation degree of implicit event associations, and the correlation degree of implicit event associations is used to judge the correlation between two events and is used to predict and judge the public opinion direction of events in the future;
[0171] Specifically, it can discover the implicit associations between different events, provide a basis for public opinion prediction and judgment, and by analyzing the correlation between events, it can predict the development direction of public opinion in advance and assist in decision-making.
[0172] In summary, the present application can perform multi-modal sentiment analysis on network public opinion data and generate multi-level sentiment classification results. This method can not only process text data, but also combine image data to provide more comprehensive and accurate sentiment analysis, providing strong technical support for network public opinion monitoring.
[0173] After the step of S427: The pre-set event relationship reasoning model analyzes the event-topic association matrix based on the machine self-learning algorithm to generate the correlation degree of implicit event associations, the following steps are included:
[0174] S4271: Obtain the propagation node characteristics of relevant events and associate the node characteristics with the events to construct an event-node data set;
[0175] Specifically, the node characteristics include text content, publishing platform, publishing time, user interaction situation, etc. Associate these node characteristics with the events to construct an event-node data set.
[0176] S4272: The pre-set node intervention model analyzes the event-node data set to identify the key nodes for event propagation, and the key nodes are used to improve the effectiveness during public opinion intervention;
[0177] Specifically, use the pre-set node intervention model to analyze the above-mentioned event-node data set. The model identifies key nodes based on the following steps:
[0178] Node feature extraction: Extract the features of each node, such as the degree (number of connections), betweenness (number of times passing through the shortest path), PageRank value, etc. of the node.
[0179] Node importance evaluation: Evaluate the importance of each node according to the extracted features. For example, the degree of node A is 100, the betweenness is 50, and the PageRank value is 0.8; the degree of node B is 80, the betweenness is 30, and the PageRank value is 0.6; the degree of node C is 120, the betweenness is 60, and the PageRank value is 0.7.
[0180] Key node identification: According to the evaluation results, identify the key nodes that have the greatest impact on event propagation. For example, node C has the highest degree and betweenness, and also has a relatively high PageRank value, so it is identified as a key node.
[0181] Beneficial effects: By identifying key nodes, public opinion intervention can be carried out accurately, the effectiveness of intervention can be improved, resources can be concentrated on intervening key nodes, waste of resources can be avoided, and the intervention efficiency can be improved. The node intervention model can dynamically adjust the identification of key nodes according to the real-time propagation situation of events to adapt to changes in public opinion.
[0182] In summary, it is possible to perform multi-modal sentiment analysis on network public opinion data and generate multi-level sentiment classification results. Especially in the aspect of identifying key nodes in event propagation, by constructing an event-node data set and a node intervention model, the key nodes that have the greatest impact on event propagation can be accurately identified, providing a scientific basis for public opinion intervention. This method can not only process text data, but also combine image data to provide more comprehensive and accurate sentiment analysis, providing strong technical support for network public opinion monitoring.
[0183] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0184] In one embodiment, a multi-level sentiment classification device for network public opinion is provided, and the multi-level sentiment classification device for network public opinion corresponds one-to-one with the multi-level sentiment classification method for network public opinion in the above embodiment. As Figure 2As shown in the figure, the multi-level sentiment classification device for online public opinion includes a temporary corpus set construction unit 1, which is used to capture real-time data of social media platforms through a distributed message queue, filter and clean non-text noise data, and form a temporary corpus set for the real-time data of social media platforms corresponding to each time window according to a preset time window division mechanism;
[0185] An initial topic set generation unit 2, which is used to preset a topic generation model to perform online clustering on the temporary corpus set based on the Streaming K-means algorithm to generate an initial topic set;
[0186] A topic state space construction unit 3, which is used to preset a topic state analysis model to analyze the initial topic set to construct a topic state space, where the topic state space = (topic popularity Ht, sentiment dispersion Et, time span Δt));
[0187] An action space selection unit 4, which is used to preset a hierarchical depth decision model to select an action space a {three-level classification, four-level classification} through the Q-learning algorithm, where the hierarchical depth decision model is set with an accuracy reward function, and the accuracy reward function is used to improve the selection accuracy. The three-level classification structure is a three-layer logical relationship of theme, sub-theme, and sentiment, and the four-level classification structure is a four-layer logical relationship of event, theme, sub-theme, and sentiment;
[0188] An information entropy value data set construction unit 5, which is used to perform a dependency path analysis on the current hierarchical structure, calculate the information entropy values of each node, and record the information entropy values of each node based on a preset time axis to construct an information entropy value data set;
[0189] A prediction weight coefficient generation unit 6, which is used to preset a hierarchical weight prediction model to analyze the information entropy value data set based on a machine self-learning algorithm to generate prediction weight coefficients for each level;
[0190] A sentiment classification generation unit 7, which is used to preset a sentiment classification model to classify topics based on the prediction weight coefficients of each level to generate a multi-level sentiment classification result, and the sentiment classification result is used for online public opinion monitoring.
[0191] For the specific limitations of the multi-level sentiment classification device for online public opinion, reference can be made to the limitations of the multi-level sentiment classification method for online public opinion in the above text, which will not be elaborated here. Each module in the above multi-level sentiment classification device for online public opinion can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of the processor in the electronic device in the form of hardware, or stored in the memory of the electronic device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0192] In one embodiment, an electronic device is provided. The electronic device may be a server, and its internal structure diagram may be as shown in Figure 3 . The electronic device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the electronic device is used to store the database. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a multi-level sentiment classification method for network public opinion.
[0193] In one embodiment, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0194] Capture real-time data of the social media platform through a distributed message queue, filter and clean non-text noise data, and form a temporary corpus set for the real-time data of the social media platform corresponding to each time window according to a preset time window division mechanism;
[0195] The preset topic generation model performs online clustering on the temporary corpus set based on the Streaming K-means algorithm to generate an initial topic set;
[0196] The preset topic status analysis model analyzes the initial topic set to construct a topic status space, where the topic status space = (topic heat Ht, sentiment dispersion Et, time span Δt));
[0197] The preset hierarchical depth decision model selects an action space a ∈ {three-level classification, four-level classification} through the Q-learning algorithm. The hierarchical depth decision model is provided with an accuracy reward function, and the accuracy reward function is used to improve the selection accuracy. The three-level classification structure is a three-layer logical relationship of theme, sub-theme, and sentiment, and the four-level classification structure is a four-layer logical relationship of event, theme, sub-theme, and sentiment;
[0198] Perform dependency path analysis on the current hierarchical structure, calculate the information entropy values of each node, and record the information entropy values of each node based on a preset time axis to construct an information entropy value data set;
[0199] The pre-set hierarchical weight prediction model analyzes the information entropy value data set based on a machine self-learning algorithm to generate prediction weight coefficients for each level;
[0200] The pre-set sentiment classification model classifies topics based on the prediction weight coefficients of each level to generate a multi-level sentiment classification result, and the sentiment classification result is used for network public opinion monitoring.
[0201] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0202] Capture real-time data of the social media platform through a distributed message queue, filter and clean non-text noise data, and form a temporary corpus set for the real-time data of the social media platform corresponding to each time window according to a pre-set time window division mechanism;
[0203] The pre-set topic generation model performs online clustering on the temporary corpus set based on the Streaming K-means algorithm to generate an initial topic set;
[0204] The pre-set topic status analysis model analyzes the initial topic set to construct a topic status space, where the topic status space = (topic popularity Ht, sentiment dispersion Et, time span Δt));
[0205] The pre-set hierarchical depth decision model selects an action space a ∈ {three-level classification, four-level classification} through the Q-learning algorithm, where the hierarchical depth decision model is provided with an accuracy reward function, and the accuracy reward function is used to improve the selection accuracy. The three-level classification structure is a three-layer logical relationship of theme, sub-theme, and sentiment, and the four-level classification structure is a four-layer logical relationship of event, theme, sub-theme, and sentiment;
[0206] Perform dependency path analysis on the current hierarchical structure, calculate the information entropy value of each node, and record the information entropy value of each node based on a pre-set time axis to construct an information entropy value data set;
[0207] The pre-set hierarchical weight prediction model analyzes the information entropy value data set based on a machine self-learning algorithm to generate prediction weight coefficients for each level;
[0208] The pre-set sentiment classification model classifies topics based on the prediction weight coefficients of each level to generate a multi-level sentiment classification result, and the sentiment classification result is used for network public opinion monitoring.
[0209] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0210] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0211] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A multi-level sentiment classification method for online public opinion, characterized in that, The method includes the steps of: capturing real-time data of a social media platform through a distributed message queue, filtering and cleaning non-text noise data, and forming a temporary corpus set for the real-time data of the social media platform corresponding to each time window according to a preset time window division mechanism; The preset topic generation model performs online clustering on the temporary corpus set based on the Streaming K-means algorithm to generate an initial topic set; The preset topic status analysis model analyzes the initial topic set to construct a topic status space, where the topic status space = (topic popularity Ht, emotional dispersion Et, time span Δt)); The preset hierarchical depth decision model selects an action space a ∈ {three-level classification, four-level classification} through the Q-learning algorithm, where the hierarchical depth decision model is provided with an accuracy reward function for improving the selection accuracy. The three-level classification structure is a three-layer logical relationship of theme, sub-theme, and emotion, and the four-level classification structure is a four-layer logical relationship of event, theme, sub-theme, and emotion; Perform dependency path analysis on the current hierarchical structure, calculate the information entropy value of each node, and record the information entropy value of each node based on a preset time axis to construct an information entropy value data set; The preset hierarchical weight prediction model analyzes the information entropy value data set based on a machine self-learning algorithm to generate the predicted weight coefficients of each level; The preset sentiment classification model classifies topics based on the predicted weight coefficients of each level to generate a multi-level sentiment classification result, and the sentiment classification result is used for network public opinion monitoring.
2. A multi-level sentiment classification method for online public opinion according to claim 1, characterized in that, After the step where the preset topic generation model performs online clustering on the temporary corpus set based on the Streaming K-means algorithm to generate an initial topic set, the following steps are included: Construct a dependency syntax tree for the text data in the same initial topic set, and extract semantic role labels as the first feature vector; Use YOLOv7 to detect significant objects for the associated image data in the same initial topic set, and extract the HSV color space histogram as the second feature vector; The preset feature fusion model calculates the fusion feature based on the cross-modal attention mechanism, and the fusion feature is used by the topic status analysis model to analyze the emotional dispersion.
3. A multi-level sentiment classification method for online public opinion according to claim 2, characterized in that In the step where the preset hierarchical depth decision model selects an action space a ∈ {three-level classification, four-level classification} through the Q-learning algorithm, the following steps are included: The hierarchical depth decision model extracts features from the topic status space, compares the topic popularity with a preset topic popularity standard value to generate a topic popularity comparison result, and compares the emotional dispersion with a preset emotional dispersion standard value to generate an emotional dispersion comparison result; Based on a preset hierarchical classification rule, comprehensively analyze the topic popularity comparison result and the emotional dispersion comparison result to generate a selection result of the action space, and the selection result of the action space is used to determine the classification level.
4. A multi-level sentiment classification method for online public opinion according to claim 3, characterized in that, After comprehensively analyzing the topic popularity comparison result and the sentiment dispersion comparison result based on a preset hierarchical classification rule to generate a selection result of the action space, where the selection result of the action space is used to determine the level of classification, when the selection result of the action space is a third-level classification, the following steps are included: The pre-set LDA topic analysis model performs topic collection on the initial topic set to generate a set of topic words T = {T1, T2,..., T k}; The preset sub-topic clustering analysis model performs sub-topic clustering on the set of topic words based on the dynamic density clustering rule of the density clustering algorithm; The preset BERT-CRF joint model analyzes the corresponding temporary corpus set in the sub-topic clustering to generate the sentiment polarity and intensity values corresponding to the sub-topics; Associate the set of topic words, the sub-topic clustering, and the sentiment polarity and intensity values to generate a sentiment classification result for the third-level classification.
5. A multi-level sentiment classification method for online public opinion according to claim 4, characterized in that, After comprehensively analyzing the topic popularity comparison result and the sentiment dispersion comparison result based on a preset hierarchical classification rule to generate a selection result of the action space, where the selection result of the action space is used to determine the level of classification, when the selection result of the action space is a fourth-level classification, the following steps are included: Extract event features from the initial topic set based on the BERT-CRF joint model to generate a set of keywords for unexpected events; The preset topic association model associates the topics corresponding to the keywords for unexpected events based on the knowledge graph algorithm technology to construct an event-topic association matrix; The preset event relationship reasoning model analyzes the event-topic association matrix based on the machine self-learning algorithm to generate the correlation degree of implicit event associations, and the correlation degree of implicit event associations is used to judge the correlation between two events and is used for subsequent prediction and judgment of the public opinion direction of events.
6. A multi-level sentiment classification method for online public opinion according to claim 5, characterized in that, After the step of the preset event relationship reasoning model analyzing the event-topic association matrix based on the machine self-learning algorithm to generate the correlation degree of implicit event associations, the following steps are included: Obtain the propagation node features of relevant events and associate the node features with the events to construct an event-node data set; The preset node intervention model analyzes the event-node data set to identify the key nodes for event propagation, and the key nodes are used to improve the effectiveness during public opinion intervention.
7. A multi-level sentiment classification method for online public opinion according to claim 5, characterized in that, After comprehensively analyzing the topic popularity comparison result and the sentiment dispersion comparison result based on a preset hierarchical classification rule to generate a selection result of the action space, where the selection result of the action space is used to determine the level of classification, if it is detected that the topic has cross-platform propagation characteristics, the fourth-level classification module is automatically activated; If conflict labels appear in the sentiment label set, the fourth-level classification is forcibly enabled.
8. A multi-level sentiment classification device for online public opinion, which is applied to the multi-level sentiment classification method for online public opinion according to any one of claims 1-7, and is characterized in that, The device includes: A temporary corpus set construction unit (1) for capturing real-time data of social media platforms through a distributed message queue, filtering and cleaning non-text noise data, and forming a temporary corpus set for the real-time data of social media platforms corresponding to each time window according to a preset time window division mechanism; An initial topic set generation unit (2) is configured to preset a topic generation model to perform online clustering on a temporary corpus based on the Streaming K-means algorithm to generate an initial topic set; A topic state space construction unit (3) is configured to preset a topic state analysis model to analyze the initial topic set to construct a topic state space, where the topic state space = (topic popularity Ht, emotional dispersion Et, time span Δt)); An action space selection unit (4) is configured to preset a hierarchical depth decision model to select an action space a {three-level classification, four-level classification} through the Q-learning algorithm, where the hierarchical depth decision model is provided with an accuracy reward function, and the accuracy reward function is used to improve the selection accuracy. The three-level classification structure is a three-layer logical relationship of theme, sub-theme, and emotion, and the four-level classification structure is a four-layer logical relationship of event, theme, sub-theme, and emotion; An information entropy value data set construction unit (5) is configured to perform dependency path analysis on the current hierarchical structure, calculate the information entropy value of each node, and record the information entropy value of each node based on a preset time axis to construct an information entropy value data set; A prediction weight coefficient generation unit (6) is configured to preset a hierarchical weight prediction model to analyze the information entropy value data set based on a machine self-learning algorithm to generate prediction weight coefficients for each level; An emotion classification generation unit (7) is configured to preset an emotion classification model to classify topics based on the prediction weight coefficients of each level to generate a multi-level emotion classification result, and the emotion classification result is used for network public opinion monitoring.
9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of a multi-level emotion classification method for network public opinion according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of a multi-level emotion classification method for network public opinion according to any one of claims 1 to 7.
Citation Information
Patent Citations
Coarse-grained emotion analysis method based on hierarchical BERT neural network
CN110147452A
Public opinion analysis method and system based on social network platform
CN116522013A
Public opinion evaluation method, device and system, electronic equipment and storage medium
CN117371442A
Public opinion analysis and prediction method and system based on propagation big data analysis
CN118395301A
Automatic generation of statement-response sets from conversational text using natural language processing
US20200279075A1
Cited By
Client guarantee demand analysis and insurance product recommendation method based on data elements
CN121120271A
A Data-Based Approach to Customer Protection Needs Analysis and Insurance Product Recommendation
CN121120271B