Network public opinion intelligent classification and emergency decision-making system based on multi-modal fusion and dynamic evolution

Through multimodal fusion and dynamic evolution of the network public opinion intelligent classification and emergency decision-making system, the problem of insufficient cross-modal data alignment and conflict quantification in network public opinion analysis is solved, the collaborative analysis and dynamic decision-making of text semantics and visual features are realized, and the real-time and accuracy of emergency response are improved.

CN120611220APending Publication Date: 2025-09-09CHONGQING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510736608.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing technologies in online public opinion analysis lack the ability to align cross-modal data in real time and quantify conflicts, the dynamic communication model has a low coupling degree with stakeholder behavior prediction, and there are breakpoints in the emergency decision-making chain due to manual intervention, making it difficult to achieve real-time perception, dynamic deduction and intelligent decision-making.

Method used

The intelligent classification and emergency decision-making system for online public opinion adopts multimodal fusion and dynamic evolution, integrates natural language processing, computer vision, complex network analysis and system dynamics simulation technology, and realizes real-time analysis and conflict quantification of multimodal data such as text, images and videos through multi-source data collection and preprocessing, multi-dimensional classification engine, event graph construction and anomaly detection, dynamic risk assessment of stakeholders and intelligent decision-making and emergency response modules.

Benefits of technology

It realizes the collaborative analysis of text semantics and visual features, accurately quantifies online public opinion conflicts, supports local anomaly detection to global risk prediction, provides hierarchical response strategies and iteratively updates emergency handling rules, improving the real-time and accuracy of emergency response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611220A_ABST
    Figure CN120611220A_ABST
Patent Text Reader

Abstract

The invention relates to a network public opinion intelligent classification and emergency decision-making system based on multi-modal fusion and dynamic evolution, and belongs to the field of network public opinion monitoring and big data analysis and artificial intelligence. The system comprises a multi-source data acquisition and preprocessing module used for crawling multi-modal data, constructing a propagation path map after preprocessing, and identifying key propagation nodes; the multi-dimensional classification engine module is used for carrying out conflict intensity quantification on public opinion events and dynamically updating a rule word bank to keep the adaptability of a conflict intensity quantification model; the event graph construction and anomaly detection module is used for constructing a public opinion propagation path and public opinion event generality logic chain mode, monitoring public opinion propagation speed and giving an alarm; the stakeholder dynamic risk assessment module is used for finely classifying network public opinion participants, providing a basis for differential propagation intervention and simulating public opinion evolution to carry out risk simulation; and the intelligent decision-making and emergency response module executes different levels of emergency measures based on the risk index according to the hierarchical response strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of network public opinion monitoring, big data analysis and artificial intelligence, and relates to an intelligent classification and emergency decision-making system for network public opinion based on multimodal fusion and dynamic evolution. Background Art

[0002] Online public opinion analysis aims to use intelligent technology to conduct real-time monitoring, risk warning, and decision-making support for massive amounts of information on the internet. However, its technological development still faces multiple challenges. Traditional methods often adopt a single-dimensional analysis model, such as text classification based on sentiment polarity (positive, negative, neutral) or using the LDA topic model to predict topic popularity. Although these basic methods can initially capture the overall tendency of public opinion, they expose significant systemic shortcomings in complex real-world scenarios. Taking multimodal data processing as an example, current mainstream technologies are still highly focused on text analysis, and there is a clear gap in the ability to parse unstructured data such as images and videos. For example, in public health event monitoring, traditional models often rely solely on keywords in text reports (such as "cluster infection"), but ignore the critical value of visual signals. Video analysis can identify risk indicators such as whether people at gatherings are wearing masks and whether the crowd density exceeds the standard. The technical bottleneck is further reflected in real-time requirements: parsing short videos and live streaming data requires simultaneous processing of high-resolution images (such as 1080P) and multiple concurrent data streams. Due to limited computing resources, traditional stand-alone architectures often have processing delays exceeding 5 seconds. However, elastic computing solutions based on distributed clusters can compress latency to less than 200 milliseconds through dynamic resource scheduling, which is crucial for the timeliness of emergency response. A typical case of a protest in a certain place in 2023 highlights this problem. Because the monitoring system failed to interpret the content of the slogan "Oppose the traffic restriction policy" in the on-site video, it mistakenly classified the high-risk event as an ordinary rally, which directly delayed the emergency response by 3 hours and exposed the fatal flaw of single-modal analysis.

[0003] When it comes to modeling the dynamic evolution of public opinion, most traditional models rely on static indicator systems, making them incapable of addressing the complex dynamics of frequent emergencies and nonlinear transmission paths. For example, in government public opinion incidents, there is a strong correlation between official response time and online outbursts: when official statements are delayed for more than two hours, the spread of negative sentiment increases exponentially. However, static models, unable to capture this dynamic causal relationship, often lag in risk prediction. A company's product crisis offers a more profound lesson. Due to the lack of real-time tracking of the evolving communication paths of key opinion leaders (KOLs), negative information spread rapidly through social media, with related topics reaching over 500 million views within three hours, directly causing the company's market value to plummet by over 1 billion yuan in a single day. Experimental data comparisons show that predictions based on static rule engines have a prediction error rate as high as 42%. However, a dynamic evolution model using an LSTM (Long Short-Term Memory) network combined with an attention mechanism significantly reduces this error rate to 12% by capturing the nonlinear characteristics of time series. This demonstrates the necessity of dynamic modeling.

[0004] The decision-making support link of existing technologies also has intelligent breakpoints. Most systems still rely on manual experience to formulate disposal strategies and lack an automated closed-loop mechanism for multi-model collaboration. For example, in the content control of social platforms, traditional rule engines use a fixed keyword library (such as directly blocking sensitive words such as "rights protection"), but it is difficult to identify the semantic implicit risks of emerging network terms - the word "lying flat" is neutral on the surface, but may imply negative social emotions in a specific context; homophonic puns such as "Barbie Q" may evade conventional detection rules. Test data from a certain platform shows that such semantic traps cause the traditional rule engine to have a misjudgment rate of 35%, and manual review requires checking each item one by one, which takes more than 2 hours, completely missing the golden 30-minute window for public opinion handling. Compared to existing technologies, the patent application with publication number CN112685621A constructs a multi-dimensional monitoring framework, but its knowledge graph update cycle is as long as 24 hours, making it impossible to achieve real-time pattern matching between emergencies and similar historical events (such as environmental protests in different years). While the graph neural network model proposed in the academic field achieves an F1 score of 0.88 in predicting the spread of structured social networks, its accuracy in detecting violent images and abnormal behavior in short videos is less than 60%, and it does not incorporate analysis of the game relationships between stakeholders (such as government agencies, media, and online groups). Commercial tools such as Qingbo Big Data can provide sentiment polarity distribution and topic popularity rankings, but lack quantitative indicators of conflict intensity (such as the group emotional polarization index and the probability of offline action conversion). This led to a local government misjudging the conflict level 40% of the time during policy disputes, triggering subsequent offline gatherings.

[0005] Overall, the current technology system suffers from three core gaps: First, the real-time alignment and conflict quantification capabilities of cross-modal data are insufficient, making it difficult to achieve collaborative analysis of text semantics and visual features. Second, the dynamic communication model is poorly coupled with stakeholder behavior prediction, lacking a closed-loop deduction mechanism for "event evolution-subject behavior-environmental feedback." Finally, the emergency decision-making chain still suffers from manual intervention breakpoints, necessitating the establishment of an autonomous optimization closed loop from risk identification to strategy generation to effectiveness evaluation. These shortcomings severely restrict the practical effectiveness of public opinion systems in major public events. There is an urgent need to build a next-generation public opinion analysis system capable of real-time perception, dynamic deduction, and intelligent decision-making through the integration and innovation of multimodal deep learning, dynamic knowledge graphs, and reinforcement learning. Summary of the Invention

[0006] In view of this, the purpose of the present invention is to provide an intelligent classification and emergency decision-making system for online public opinion based on multimodal fusion and dynamic evolution. By integrating natural language processing (NLP), computer vision (CV), complex network analysis and system dynamics simulation technology, it can realize real-time analysis and conflict quantification of multimodal data such as text, images, and videos. It is suitable for public opinion monitoring, risk assessment and automated emergency response of governments, enterprises and public institutions.

[0007] In order to achieve the above object, the present invention provides the following technical solutions:

[0008] An intelligent classification and emergency decision-making system for online public opinion based on multimodal fusion and dynamic evolution, which includes a multi-source data acquisition and preprocessing module, a multidimensional classification engine module, a causal graph construction and anomaly detection module, a stakeholder dynamic risk assessment module and an intelligent decision-making and emergency response module.

[0009] The multi-source data acquisition and preprocessing module is used to crawl multimodal data from the network platform, pre-process the crawled data, perform spatiotemporal labeling, and construct a propagation path map to identify key propagation nodes;

[0010] The multi-dimensional classification engine module quantifies the conflict intensity of public opinion events based on the data output by the multi-source data acquisition and preprocessing module according to three attributes: sentiment polarity, event classification, and stakeholder relevance. At the same time, it adjusts the attribute threshold through a dynamic optimization mechanism and dynamically updates the rule vocabulary to incorporate emerging network terms, thereby maintaining the long-term adaptability of the conflict intensity quantification model.

[0011] The event graph construction and anomaly detection module is used to construct the public opinion propagation path from a micro level and the common logic chain model of public opinion events from a macro level, while monitoring the speed of public opinion propagation and issuing alarms;

[0012] The stakeholder dynamic risk assessment module refines the classification of online public opinion participants by constructing differentiated judgment rules and quantitative models, and divides the influence levels of the participants, thereby providing a basis for differentiated communication intervention; the stakeholder dynamic risk assessment module also simulates the evolution of public opinion through system dynamics and conducts risk simulation;

[0013] The intelligent decision-making and emergency response module executes different levels of emergency measures based on the risk index according to the hierarchical response strategy.

[0014] Furthermore, the multi-source data acquisition and preprocessing module includes a data acquisition submodule, a preprocessing submodule and a spatiotemporal label generation submodule.

[0015] Among them, the data collection submodule crawls text, images, videos and live streaming data from the network platform through a distributed crawler cluster architecture;

[0016] The preprocessing submodule is used to process the crawled multimodal data. For text data, non-Chinese characters, advertising information, and sensitive words are first filtered out. Then, key entities including names, organizations, places, and times are annotated. Finally, a hybrid sentiment analysis of the text content is performed using the SnowNLP framework and a custom dictionary. For image and video data, key visual features including crowd density and slogan text are first extracted. Then, a classification model is used to detect high-risk sensitive content such as violent scenes, rallies, and political slogans. At the same time, multi-label classification technology is used to perform multi-label semantic analysis on complex scenes to generate semantic labels.

[0017] The spatiotemporal tag generation submodule adds millisecond-level timestamps and city-level geographic tags to the processed data, and constructs a propagation path map with node influence weights based on social network forwarding relationships.

[0018] Furthermore, the multi-dimensional classification engine module includes a conflict strength quantification submodule and a dynamic optimization submodule.

[0019] Among them, the conflict intensity quantification submodule quantifies the conflict intensity of abstract opposing relationships in public opinion events based on emotional polarity, event classification and entity correlation of stakeholders; the dynamic optimization submodule improves the adaptability of the conflict intensity quantification submodule through fuzzy logic and rule iteration, wherein the fuzzy logic uses trapezoidal membership function to dynamically adjust the classification threshold to improve the sensitivity of conflict warning, and the rule iteration updates the vocabulary mapping rules by mining emerging network terms and combining the semantic vector space model to ensure that the classification rules are synchronized with the evolution of network public opinion.

[0020] Furthermore, the event graph construction and anomaly detection module includes a graph construction submodule and an anomaly detection submodule.

[0021] Among them, the graph construction submodule constructs a graph of events from the micro and macro levels; from the micro level, it identifies the subjects of public opinion events and the interactive actions between subjects, constructs a chain-type propagation relationship triple containing the subjects and interactive actions, and defines the logical edge type at the same time, and stores the relationship network in the Neo4j graph database; from the macro level, it performs thematic clustering on the historical case library, abstracts the common logical chain pattern of "emergency → official delayed response → negative emotion outbreak → secondary propagation", calculates the similarity of the element set between the new public opinion event and the common logical chain pattern through the Jaccard coefficient, triggers the historical pattern association when the Jaccard coefficient reaches a certain level, and generates an event evolution trend forecast report;

[0022] The anomaly detection submodule detects the node propagation speed and sentiment change rate through the STGCN model, and constructs a propagation speed prediction curve through the LDA topic model; when the node propagation speed suddenly increases by more than 500% or the sentiment change rate deviates from the historical mean by ±2 times the standard deviation, the corresponding node is marked as an abnormal account and a preliminary warning is triggered; when the actual propagation speed deviates from the predicted propagation speed by more than ±20%, a real-time alarm is triggered.

[0023] Furthermore, the stakeholder dynamic risk assessment module includes a stakeholder classification submodule and a risk assessment submodule.

[0024] The stakeholder classification submodule categorizes the participants of public opinion events and divides the influence levels of the participants through the communication power index; the participants of public opinion events are divided into five categories: direct stakeholders, indirect stakeholders, key opinion leaders, public interest parties and marginal groups. The communication power index I is expressed as:

[0025] I = 0.3 × number of followers + 0.4 × forwarding rate + 0.2 × historical influence rate + 0.1 × government relevance

[0026] Participants with I ≥ 0.8 were classified into the high-influence level, participants with 0.5 ≤ I < 0.8 were classified into the medium-influence level, and participants with I < 0.5 were classified into the low-influence level.

[0027] The risk assessment submodule simulates the evolution of public opinion through system dynamics and conducts risk simulation, including: defining core variables, namely government response speed VR, event sensitivity ES and KOL intervention intensity KI, and constructing the differential equation between the negative emotion change rate ΔN and the core variables ξ is a random disturbance term that obeys the normal distribution; the system dynamics simulation output includes a risk indicator time series prediction curve and a sensitivity analysis report. The risk indicator time series prediction curve shows the predicted values ​​of the negative emotion change rate and the speed of the spread range in the next 24 hours, with a time resolution of 15 minutes. The sensitivity analysis report shows the impact of adjusting the core variables on the evolution of public opinion.

[0028] Furthermore, the intelligent decision-making and emergency response module includes an emergency response submodule and a closed-loop feedback submodule.

[0029] Among them, the emergency response submodule executes different levels of emergency measures according to the risk index RI; when RI ≥ 0.8, a first-level response is triggered, the whole network emergency linkage mechanism is activated, real-time alerts are pushed to the head of the emergency management department through the SMS gateway, and the social platform collaboration interface is simultaneously called to automatically issue an authoritative rumor-refuting statement; in addition, based on the OAuth 2.0 protocol, real-time linkage with the platform is carried out to perform permission restriction operations on abnormal accounts, and high-risk IP clusters are locked through distributed log tracking; when 0.6 ≤ RI < 0.8, a second-level response is triggered, and geo-fencing technology is used to delineate high-risk areas. Authoritative information is pushed to the target user terminal through the CDN node, and the BERT model and rule engine dual-channel review mechanism is activated to ensure that the accuracy rate of sensitive content recognition exceeds 95% and the review delay is less than 5 seconds; when 0.4 ≤ RI < 0.6, a third-level response is triggered, and a multi-dimensional report including a conflict heat map and a propagation tree map is automatically generated, prompting manual intervention for review, and a dynamic monitoring mode is turned on, updating the risk scoring model parameters every 30 minutes;

[0030] The closed-loop feedback submodule quantitatively evaluates the disposal effect through multi-dimensional indicators, iteratively updates the emergency disposal strategy and disposal rules, and verifies the updated rule set in a simulation environment of typical public opinion scenarios.

[0031] The closed-loop feedback submodule quantitatively evaluates the treatment effect through multi-dimensional indicators, including evaluating the treatment effect through the speed of public opinion cooling down, the blocking rate of the transmission path and the public satisfaction; the speed of public opinion cooling down is the decrease in the proportion of negative emotions per hour; the blocking rate of the transmission path is the ratio of the difference in the forwarding volume of the abnormal account before and after the flow is restricted to the forwarding volume before the flow is restricted, and when the ratio is not less than 0.6, it is judged to be effectively blocked; the public satisfaction is the proportion of positive evaluations on the government's response speed and information transparency in a sampling survey with a sample size greater than or equal to 1,000, and it is qualified when the proportion of positive evaluations is greater than or equal to 60%.

[0032] The risk index RI is expressed as:

[0033] RI=ω1·CI+ω2·AR+ω3·SI

[0034] Where CI represents conflict intensity, AR represents the comprehensive anomaly risk value, SI represents the stakeholder influence index, and ω1, ω2, and ω3 are dynamic weights calibrated through historical data regression analysis. The corresponding emergency response thresholds are calibrated using the DBSCAN algorithm for clustering based on a historical case database. Their effectiveness is evaluated monthly using a confusion matrix. When the missed detection rate for Level 1 response events exceeds 5%, threshold fine-tuning is triggered.

[0035] The beneficial effects of the present invention are:

[0036] 1) The present invention uses a distributed crawler cluster architecture to crawl multimodal data such as text, images, videos, and live streaming data from the network platform, and through the spatiotemporal alignment of multimodal heterogeneous data, realizes precise association of cross-modal data through four-dimensional spatiotemporal coordinates (timestamp, longitude, latitude, and propagation level), solves the problem of temporal misalignment of images and text in traditional methods, and realizes collaborative analysis of text semantics and visual features.

[0037] 2) This paper proposes a fuzzy quantification model for conflict intensity. Based on trapezoidal membership functions and conflict heat maps, it transforms abstract oppositional relationships such as "government statements vs. netizens' doubts" into a computable conflict index. By mining emerging network terms and combining them with semantic vector space models to update vocabulary mapping rules, it ensures that classification rules are synchronized with the evolution of online public opinion, enabling the conflict quantification model to accurately quantify online public opinion conflicts.

[0038] 3) The present invention adopts a two-layer event graph architecture to identify the subjects of public opinion events and the interactive actions between subjects at the micro level, and construct a chain-type propagation relationship including the subjects and interactive actions. At the macro level, the historical case library is clustered by subject, and the common logical chain pattern of "emergency → official delayed response → negative emotional outbreak → secondary propagation" is abstracted. The STGCN model is used to detect the node propagation speed and sentiment change rate, and the LDA topic model is used to construct a propagation speed prediction curve, supporting multi-scale analysis from local anomaly detection to global risk prediction.

[0039] 4) The present invention adopts a graded response strategy based on risk index, adopts response measures of different urgency for different levels of public opinion risks, and quantitatively evaluates the disposal effect through multi-dimensional indicators. At the same time, it iteratively updates the emergency disposal strategy and disposal rules, and verifies the updated rule set in a simulation environment of typical public opinion scenarios to ensure that the response mechanism continues to adapt to the evolution of the public opinion ecosystem.

[0040] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:

[0042] Figure 1 This is the overall system architecture diagram;

[0043] Figure 2 It is a logical topology diagram between the functional modules of the system;

[0044] Figure 3 This is a schematic diagram of the complete processing flow of the system from data collection to emergency response;

[0045] Figure 4 This is the system dynamics simulation flow chart;

[0046] Figure 5 Schematic diagram of closed-loop feedback optimization system;

[0047] Figure 6 Schematic diagram of data flow between modules of the system. DETAILED DESCRIPTION

[0048] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0049] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.

[0050] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0051] An embodiment of the present invention provides an intelligent classification and emergency decision-making system for network public opinion based on multimodal fusion and dynamic evolution, such as Figure 1 As shown in the figure, the system includes multi-source data acquisition and preprocessing module, multi-dimensional classification engine module, event graph construction and anomaly detection module, stakeholder dynamic risk assessment module and intelligent decision-making and emergency response module. The data flow between the modules is as follows: Figure 2 shown.

[0052] 1. Multi-source data acquisition and preprocessing module

[0053] This module crawls multimodal data from the network platform, including text, images, videos, and live streaming data, preprocesses these data, and then performs multimodal annotation.

[0054] (1) The data collection phase uses a distributed crawler cluster architecture to capture multimodal data from various online platforms such as Weibo, Douyin, and Twitter. The details are as follows:

[0055] 1) System Architecture Construction: A distributed crawler cluster is built based on the Scrapy-Redis framework. The master node distributes tasks through Redis, and the slave nodes perform crawling. It supports the collection of text, images, videos, and live streaming data from platforms such as Weibo, Douyin, and Twitter.

[0056] Each node is configured with a dynamic IP proxy pool, which changes the IP every 100 requests through a polling mechanism. The proxy pool contains no less than 500 available IPs to avoid anti-crawl.

[0057] 2) Performance Optimization: We use asynchronous I / O technology (Tornado framework) to process network requests, combined with an 8GB Redis memory cache, to achieve a throughput of 10,000 data items per second on Node A, with latency controlled within 50 milliseconds.

[0058] Deploy a cluster monitoring module to automatically expand capacity when the node CPU usage exceeds 80% or the memory usage exceeds 70%, ensuring 24 / 7 operation.

[0059] 3) Anti-crawl strategy: Simulates browser fingerprinting technology, has 100 built-in User-Agents that rotate randomly every 20 requests, and dynamically generates cookies through Selenium with a 30-minute validity period.

[0060] The request frequency is controlled within 0.5 to 2 seconds (mean 1.2 seconds, standard deviation 0.3 seconds) through a normal distribution algorithm, reducing the risk of platform detection.

[0061] 4) API Data Integration: Access government public data, news media databases, and third-party public opinion monitoring platforms through RESTful interfaces to supplement structured data such as event timelines and geographic locations. For example, the government microblog API uses OAuth 2.0 authentication for daily synchronization, the Xinhuanet API uses HTTP Basic Auth to obtain real-time geotagged data, and the Eagle Eye Speed ​​Reading Platform accesses data via API keys.

[0062] Use Kafka message queues to build a streaming processing pipeline, set up three consumer groups for parallel processing, ensure the consistency of multi-source data time series based on message timestamps, and retain messages for 7 days.

[0063] (2) In the data preprocessing stage, different preprocessing methods are adopted for data of different modalities.

[0064] 1) For text data, data denoising is achieved through the BERT pre-trained model, effectively filtering non-Chinese characters, advertising information, and repetitive content. This process includes: first, filtering non-Chinese characters and advertising links with regular expressions, and then combining it with a preset library of 500 sensitive words for semantic filtering, with a denoising accuracy exceeding 95%. Then, the entity recognition process based on the BiLSTM-CRF model (embedding layer 300 dimensions, hidden layer 200 dimensions) is initiated. Based on 800,000 mixed corpus training, key entities such as names, organizations, locations, and time are annotated (F1-score reaches 0.91), supporting single batch processing of 100,000 items, significantly improving the efficiency of structured data extraction. Finally, sentiment analysis combines SnowNLP with a custom dictionary of more than 2,000 emerging vocabulary words, breaking through the limitations of binary classification and supporting the graded quantitative expression of mixed emotions of "60% negative emotions + 40% neutral attitudes", with processing delays of less than 100 milliseconds.

[0065] Finally, the output is the clean text after BERT denoising, the entity list of names / institutions / places / times extracted by BiLSTM-CRF (F1-score 0.91), the mixed sentiment quantization vector generated by SnowNLP (such as [negative: 0.6, neutral: 0.4]), and the sensitive word filtering results (custom blocking rules), providing fine-grained labels for in-depth interpretation of public opinion.

[0066] 2) For image and video data, we use the FFmpeg tool to extract one key frame per second, and combine it with OpenCV to perform scene segmentation and object detection, accurately extracting key visual features such as crowd density and slogan text. Specifically:

[0067] FFmpeg is used to extract one key frame per second. After OpenCV preprocessing (Gaussian blur 5×5 kernel, GrabCut scene segmentation), 20 types of targets (crowds, slogans, etc.) are detected through the YOLOv5 model, and visual features such as crowd density and slogan text are accurately extracted.

[0068] Sensitive content identification uses the ResNet50 model (ImageNet pre-training), with a detection accuracy of 92% for 10 types of high-risk elements (rallies, violent scenes, political slogans, etc.). It supports batch processing of 500 images at a time, reducing manual review costs.

[0069] The multi-label classification technology based on the CLIP model (BERT text encoder + ViT image encoder) supports semantic parsing and cross-modal retrieval functions for complex scenarios such as "protest slogans + not wearing masks", generating no less than 5 semantic labels, and the cross-modal retrieval response time is less than 500 milliseconds.

[0070] Finally, the output is high-risk visual feature labels such as "violent scenes" and "political slogans" detected by ResNet50 (with an accuracy of 92%), and cross-modal semantic vectors generated by the CLIP model (supporting the analysis of the combined features of "protest slogans + not wearing masks").

[0071] (3) In terms of labeling, time annotation records millisecond timestamps, synchronizes cluster time through the NTP server (error <1ms), and supports the reconstruction of the time trajectory of the event propagation chain with a resolution of 10ms. Geographic location resolution is based on a database containing 2 million IP segments, achieving city-level administrative district accuracy (error <50km). When there is no IP data, text entity annotation is used to supplement positioning, providing support for spatial dimension monitoring. The propagation path traceability constructs a directed graph of the social network, uses the PageRank algorithm to calculate the node influence (damping coefficient 0.85, iteration 100 times), outputs a weighted network topology, and automatically identifies the top 5% key propagation nodes.

[0072] Ultimately, the system outputs millisecond-level timestamps, city-level administrative district-level IP location codes (with controllable errors), and weighted social network forwarding relationship chains (recording direct / secondary forwarding attenuation coefficients), providing a data basis for the analysis of communication patterns.

[0073] In summary, in the multi-source data acquisition and preprocessing module, data acquisition nodes acquire raw data through dynamic IP proxies and anti-crawl strategies, and transmit it to the cleaning module via a Kafka queue. Text data undergoes denoising, entity labeling, and sentiment analysis, while image and video data undergo keyframe extraction, object detection, and multi-label classification. The processed data enters the spatiotemporal label generation module, where millisecond-level timestamps and city-level geographic tags are added, while a propagation path map is constructed. The final structured data (including text labels, visual features, spatiotemporal information, and propagation weights) is stored in an HBase distributed database (10GB per RegionServer shard) for subsequent public opinion analysis.

[0074] 2. Multidimensional classification engine module

[0075] This module achieves a multi-dimensional characterization of conflict intensity by defining core attributes such as sentiment polarity, event classification, and stakeholder correlation, while ensuring the synchronous optimization of classification rules and public opinion evolution through a dynamic optimization mechanism.

[0076] (1) Quantification of conflict intensity

[0077] Sentiment polarity quantification adopts a fuzzy membership quantification mechanism and uses the Sigmoid function to calculate the negative sentiment membership. The formula is:

[0078]

[0079] Where x represents the frequency of negative keywords in the text, k and x0 are dynamically adjusted according to the event type (such as social events, political events) to achieve differentiated measurement of sentiment intensity in different scenarios.

[0080] Event classification is based on the TF-IDF and LDA topic models, supporting multi-label joint annotation for social, political, economic, and entertainment topics. Combining TF-IDF text feature extraction with the LDA topic model, it supports multi-label joint annotation for social, political, economic, and entertainment topics. Using Dirichlet allocation to model topic probability distribution, it achieves hierarchical analysis of event semantics.

[0081] Stakeholder identification is based on the calculation of entity relevance. Taking government relevance as an example, we first count the frequency of occurrence of official agency keywords such as "State Council" and "local government" in the text. The higher the frequency, the higher the initial relevance. Secondly, we combine entity linking technology (such as DBpedia entity library) to confirm the authority and relevance of the agency corresponding to the keyword, and give high-authority agencies such as "State Council" a higher weight, followed by local agencies. Finally, we construct a relevance matrix and use the formula β(S 政府)=∑(keyword frequency×institutional weight), combining keyword frequency and institutional weight to calculate the semantic association strength between the subject and the government agency. The larger the value, the higher the association between the subject and the government, and the greater the contribution to CI in the conflict intensity calculation.

[0082] The application of rough set theory further optimizes the conflict quantification model. Specifically, attribute simplification and conflict feature screening are achieved based on the ROSETTA tool. The importance of attributes is calculated through information entropy, and redundant features such as high-frequency function words are eliminated to generate a streamlined decision table containing core decision attributes. The conflict matrix is ​​constructed based on statistical analysis of the confrontational relationship between different subjects, such as the combined conflict samples of "government statements" and "netizen doubts". These samples, together with non-confrontational samples such as "policy interpretation and media reprints" and "activity publicity and netizen interaction", constitute the total sample, covering all interaction instances of subjects such as government, netizens, enterprises, and media (including conflict and non-conflict scenarios), multi-dimensional attribute annotation data (emotional polarity, event classification labels, stakeholder correlation, etc.) and all interaction records. A two-dimensional conflict matrix is ​​constructed by calculating the proportion of combined conflict samples in the total sample, and the conflict intensity distribution is visualized with a heat map. The conflict index calculation formula is:

[0083]

[0084] The numerator is calculated by summing the sentiment polarity strength μ of each conflict sample. neg , event classification adjustment coefficient α(T), the product of stakeholder correlation β(S), comprehensively reflects the contribution of multi-dimensional attributes to conflict intensity, the denominator N total is the total sample size, and normalization is performed to make the conflict intensity of events of different scales comparable.

[0085] (2) Dynamic optimization mechanism

[0086] The dynamic optimization mechanism improves model adaptability through fuzzy logic and rule iteration.

[0087] Fuzzy logic uses trapezoidal membership functions to dynamically adjust the classification threshold. For specific scenarios such as public health events, the negative emotion judgment threshold is lowered from 50% in conventional scenarios to 30%, increasing the conflict warning sensitivity by 40% and effectively identifying hidden conflicts in sudden public opinion.

[0088] The optimization of the rule base lies in iterative updates based on historical data every month. The Apriori algorithm is used to mine frequent item sets of emerging Internet buzzwords such as "amazing" and "YYDS", and the lexical mapping rules are updated in combination with semantic vector space models (such as Word2Vec) to ensure that the classification rules are synchronized with the evolution of online public opinion. During the rule iteration process, the performance of the model is evaluated through a confusion matrix. When the F1-score is detected to be lower than 0.85 for three consecutive times, a forced update mechanism is triggered to ensure the long-term adaptability of the conflict quantification model.

[0089] The core feature data output by the multi-dimensional classification engine module directly drives the subsequent modules: key attribute sets such as high-frequency entity words screened by rough set theory and negative sentiment membership degrees calculated by the Sigmoid function, as well as the normalized conflict index (a continuous value from 0 to 1, reflecting the intensity of subject confrontation) and the heat map of the conflict matrix recording the proportion of subject confrontation relationship samples, constitute the conflict quantification result; the multi-label theme vectors generated by TF-IDF + LDA (such as [Society: 0.7, Politics: 0.3]) and the stakeholder association degree matrix calculated based on the keyword frequencies of official agencies form the event classification result; the sentiment determination threshold adjusted by fuzzy logic (such as the threshold drops to 30% in public health events) and the semantic features of emerging words (adapted to Internet buzzwords) are used as dynamic optimization parameters.

[0090] 3. Event事理图谱Construction and Anomaly Detection Module

[0091] (1) Event事理图谱Construction

[0092] The construction of the graph is divided into a dual structure of a micro layer and a macro layer, and the deep analysis of public opinion events is achieved through multi-level modeling. The micro layer constructs the "organization -网民" confrontation relationship edge based on entity association degree, uses the conflict intensity index as the edge weight, and the event theme label as the node attribute. When the macro layer clusters historical cases through the DBSCAN algorithm, the cosine similarity (>0.85) is calculated by introducing the conflict matrix and the theme vector, and common logical chains such as "emergency event → official delayed response → negative outbreak" are abstracted, as follows:

[0093] It should be noted that "事理图谱" seems to be a specific term in Chinese that might not have a direct and common English equivalent. Here I've just transliterated it as "事理图谱" for the sake of literal translation. You may need to adjust it according to the actual meaning it represents in your specific context.The micro-layer uses the Stanford CoreNLP tool to extract entities, accurately identify event subjects (including entity types such as people and institutions) and interactive actions (behavioral labels such as forwarding, commenting, and liking), and construct chain-propagation relationship triples in the form of "A forwards B→C comments D". Causal modeling defines logical edge types such as "questioning→diffusion" and "refuting→cooling down", and stores the relationship network in the Neo4j graph database. In the Neo4j graph database, node attributes include heat value (dynamically calculated based on the forwarding volume and time decay coefficient) and sentiment intensity (quantified value of the proportion of negative emotions), and edge attributes record the propagation time difference and path weight (such as the direct forwarding weight is 1, and the secondary forwarding weight decays to 0.8). Specifically, the node attribute configuration is: the heat value is dynamically calculated based on the forwarding volume and the time decay coefficient (α=0.95 / hour), and the calculation formula is Among them F i is the amount of forwarding for the i-th time, tt i is the difference in hours from the current time; the sentiment intensity is quantified by the proportion of negative sentiment, with a value range of [0, 1]; the edge attribute records the propagation time difference (accurate to minutes) and the path weight (direct forwarding weight is 1.0, and secondary and above forwarding weight is exponentially decayed by 0.8).

[0094] The macro level focuses on mining the evolution patterns of events, uses the DBSCAN algorithm to perform topic clustering on the historical case library, uses a cosine similarity > 0.85 as the threshold for determining topic similarity, abstracts common logical chain patterns such as "emergency → official delayed response → outbreak of negative emotions → secondary dissemination", and calculates the similarity of event element sets using the Jaccard coefficient. When the Jaccard coefficient is > 0.7, it triggers historical pattern association and generates an event evolution trend forecast report.

[0095] (2) Anomaly Detection

[0096] Anomaly detection is achieved through the collaboration of spatiotemporal graph convolutional network (STGCN) and LDA topic model.

[0097] The STGCN model takes as input features the node's propagation speed (number of forwards per unit time, measured in pieces per minute) and the rate of change in sentiment (the absolute difference in the proportion of negative sentiment over the minute). It constructs a three-layer spatiotemporal convolutional architecture: a temporal convolution kernel of size 3 (to capture the temporal features of the first, middle, and last three minutes) and a spatial convolution kernel of size 2 (to aggregate features of adjacent nodes). These three layers of spatiotemporal convolution output a node anomaly probability value (with a threshold set at 0.8). If a node's propagation speed suddenly increases by more than 500% or its sentiment intensity deviates from the historical mean by ±2 standard deviations, it is automatically flagged as an abnormal account and triggers a preliminary warning, accurately capturing micro-level anomalous nodes.

[0098] The LDA topic model focuses on the analysis of macro-propagation trends, combines the topic vector sequence with the conflict index time series to predict the inflection point of propagation heat, constructs the propagation speed prediction curve by fitting historical data, and dynamically learns the time series law of public opinion propagation. First, the propagation speed deviation rate D output by LDA is calculated. LDA (Percentage deviation between actual value and predicted value) is normalized and dimension effects are eliminated by extreme value mapping:

[0099]

[0100] Where ε is the minimum value (take 10 -6 Avoid denominator to be zero), min(|D LDA |) and max(|D LDA |) is the minimum and maximum absolute value of the deviation in the current calculation cycle. When the actual propagation speed deviates from the predicted value by more than ±20%, the normalized D LDA_norm Trigger macro trend deviation alarms to effectively identify inflection point events such as sudden increases or decreases in transmission volume.

[0101] The two models form a collaborative detection capability through functional complementarity and result fusion: STGCN focuses on abnormal behavior detection of key nodes at the micro level, while LDA is responsible for deviation analysis of propagation trends at the macro level. During the specific fusion calculation, a fixed weighted model is used to generate a comprehensive abnormal risk value AR:

[0102] AR=ω STGCN ·P STGCN +ω LDA ·D LDA_norm

[0103] Among them, P STGCN is the node abnormality probability output by STGCN (threshold 0.8), ω STGCN With ω LDA The default weights of 0.6 and 0.4, respectively, reflect a differentiated focus on micro-anomalies (60%) and macro-trends (40%). This approach not only preserves the details of individual node anomalies (such as sudden, high-frequency forwarding by high-impact accounts) but also captures deviations in platform-wide dissemination trends (such as a sudden drop in dissemination after an official debunking of a rumor). Through a parallel computing architecture and real-time data processing pipeline, the latency for anomaly detection is controlled to less than 5 minutes, providing timely and accurate risk warning support for a tiered response strategy.

[0104] The synergistic mechanism of the two models effectively compensates for the limitations of a single model. The STGCN model focuses on the microscopic level. While it can accurately identify abnormal characteristics such as the spread speed and sentiment changes of individual nodes, it may lead to misjudgments due to local fluctuations. The LDA topic model focuses on capturing spread trends at a macroscopic level, but it tends to overlook the triggering effects of key microscopic nodes. Through this collaboration, when STGCN detects an anomaly such as a sudden increase in the spread speed of a node, LDA can simultaneously verify whether the current spread trend deviates from the prediction, avoiding misjudgments or omissions caused by a single model's reliance on local or one-sided information. Secondly, the synergistic mechanism significantly improves the timeliness and accuracy of detection. By fusing detection results from the spatiotemporal dimensions (with a fusion weight of 0.6:0.4), dual verification of microscopic anomalies and macroscopic trends is achieved. For example, after STGCN issues an initial warning, LDA can quickly confirm whether the macroscopic spread speed has simultaneously deviated from expectations. If both trigger an anomaly, risk assessment can be performed in a much shorter time. This model controls the delay in abnormal event detection to less than 5 minutes. Compared with a single model operating independently, it can not only reduce false alarms, but also identify risks more promptly and accurately, thereby significantly improving the overall efficiency of public opinion risk identification.

[0105] 4. Stakeholder Dynamic Risk Assessment Module

[0106] This module mainly classifies and quantifies stakeholders, and simulates the evolution of public opinion through system dynamics to conduct risk simulation.

[0107] (1) Stakeholder classification and quantification

[0108] By building differentiated judgment rules and quantitative models, we can achieve a refined classification of online public opinion participants. Public opinion participants include five categories: direct stakeholders, indirect stakeholders, key opinion leaders, public interest parties, and marginal groups.

[0109] in:

[0110] 1) Direct stakeholders are identified based on both entity relevance and event mention frequency. When β(S) is greater than 0.6 and the event mention frequency is ≥ 3 times, the role is marked.

[0111] 2) Indirect stakeholders rely on the preset industry knowledge graph (including 500+ industry categories and related relationships) and automatically deduce the indirect relationship between the subject and the event through a path search algorithm (such as breadth-first search). When there are ≥2 indirect relationship paths in the graph, this type of subject is included.

[0112] 3) The selection of key opinion leaders (KOLs) incorporates multiple indicators:

[0113] Node heat value H NIt is used to measure the popularity of a single piece of content, based on the dynamic forwarding volume and time decay effect. The calculation formula is: where R i is the amount of transmission of the i-th forwarding, ω i is the forwarding level weight (direct forwarding ω1=1.0, secondary forwarding ω2=0.8, and so on), λ is the time decay coefficient (default value is 0.1), t i The hour difference between the forwarding time and the current time;

[0114] Propagation path weight W P Characterize the breadth and depth of content diffusion, calculated by forwarding level and key node influence, the formula is where ω j is the forwarding weight of the jth layer (the deeper the layer, the lower the weight), F j is the number of fans of the j-th layer forwarding node;

[0115] The historical influence score (HI) combines content interaction and sentiment guidance capabilities. The formula is HI = 0.4R + 0.3C + 0.3|E|, where R is the average number of reposts over the past six months (normalized to [0, 1]), C is the average number of comments (normalized to [0, 1]), and E is the sentiment guidance coefficient (ranging from [-1, 1], with the sign retained to distinguish between positive and negative guidance capabilities).

[0116] The basic threshold for the fan size F is F ≥ 10 4 .

[0117] The following conditions must be met at the same time when determining KOL: N ≥0.6 and W P ≥50 (node ​​heat value reflects the short-term explosive power of a single piece of content, and the transmission path weight measures the structural efficiency of content diffusion); historical comprehensive ability HI ≥ 0.7 and E ≥ 0.2 (requires that the historical influence score meets the standard and the emotional guidance coefficient is positive to ensure positive guidance ability); and fan size F ≥ 10 4 .

[0118] 4) Public interest parties are directly identified through account authentication information, including official authentication labels such as government agency authentication, media organization authentication, and public welfare group authentication. Authentication information is stored in a whitelist database (updated in real time).

[0119] 5) The marginal group is defined as ordinary netizens with less than 100 followers and no certification labels.

[0120] The above five types of participants were divided into high (I ≥ 0.8), medium (0.5 ≤ I < 0.8), and low (I < 0.5) influence levels according to the communication index using the K-means clustering algorithm (the initial number of cluster centers was set to 3). The communication index I was calculated as follows:

[0121] I = 0.3 × number of followers + 0.4 × forwarding rate + 0.2 × historical influence rate + 0.1 × government relevance

[0122] This hierarchical division directly provides a differentiated basis for system dynamics simulation and graded response strategies: the communication behavior of high-influence entities is given a higher weight in the system dynamics model, while the graded response strategy implements precise intervention according to the level (such as the first-level response gives priority to reaching high-influence KOLs).

[0123] (2) System dynamics simulation of public opinion evolution

[0124] System dynamics (SD) simulation models the dynamic interaction relationship between multiple agents by defining core variables and differential equations (such as Figure 4 Core variables include:

[0125] Government response speed (VR): the time interval between the occurrence of an incident and the first official response (unit: minutes), which is log-normalized and included in the model;

[0126] Event sensitivity (ES): a comprehensive indicator calculated based on the emotional polarity and spread of the event (value [0,1]);

[0127] KOL Involvement Intensity (KI): The weighted value of the frequency of KOL publishing relevant content and emotional tendency (positive content weight +1, negative content weight -1).

[0128] A differential equation is constructed to describe the dynamic relationship between the negative emotion change rate (ΔN) and the core variables. The basic model is:

[0129]

[0130] Where ξ is a random disturbance term (obeying normal distribution N(0,0.1)). The model parameters are calibrated through historical data regression analysis, and the least squares method is used to fit the equation coefficients. The goodness of fit R 2 The parameter validity was confirmed when ≥0.85.

[0131] The simulation output consists of two parts: 1) Risk indicator time series prediction curve: outputs the predicted values ​​of indicators such as the rate of change of negative emotions and the speed of expansion of the spread range in the next 24 hours, with a time resolution of 15 minutes; 2) Sensitivity analysis report: By adjusting core variables (such as shortening the government response speed by 50%), the impact on the evolution of public opinion is calculated. Quantitative indicators include peak delay time, fluctuation amplitude change rate, etc., providing decision makers with an evaluation of the effectiveness of different intervention strategies.

[0132] The system dynamics simulation module supports dynamic parameter configuration and can adjust variable weights and equation structures according to specific event types (such as social events and public health events). The simulation time step is set to 5 minutes, and the time required for a single complete simulation is controlled within 30 seconds, which can meet real-time decision support needs.

[0133] The results of stakeholder classification form a close underlying data support relationship with system dynamics simulation. Specifically, the classification results directly serve as core input variables for the simulation model. Attributes such as the KOL's follower size and historical influence score directly determine the weighted calculation logic for the KOL intervention intensity (KI) in the system dynamics model. The number of industry-linked paths for indirect stakeholders influences the calculation of the spread parameter in the event sensitivity (ES). The authentication information of public stakeholders is used to calibrate the weight distribution of government connections in the spread index model, ensuring that model parameters align with the characteristics of actual stakeholders. Furthermore, the influence hierarchy classification drives the differentiated modeling of the simulation model. The forwarding and commenting behaviors of high-influence entities (such as KOLs) are given higher weights in the differential equations (for example, their contribution to the spread rate is set to five times that of low-influence entities). The behavioral patterns of marginal groups are defaulted to the basic spread coefficient. This differentiated design enables the simulation results to accurately reflect the network characteristics of "key nodes leading spread," improving the model's ability to fit complex public opinion dissemination scenarios.

[0134] 5. Intelligent decision-making and emergency response module

[0135] First, the risk index (RI) is defined as the core basis for tiered response. A comprehensive quantitative system is formed by integrating core indicators output by multiple modules. Its basic indicator inputs integrate the conflict intensity index (CI) from the multidimensional classification engine module, the comprehensive anomaly risk value (AR) from the event-based graph construction and anomaly detection module, and the stakeholder influence index (SI) from the stakeholder dynamic risk assessment module. (SI is constructed through expert evaluation and historical data analysis to construct a weight matrix for the communication index hierarchy (high, medium, and low). The weights of each hierarchy are calculated using the AHP method, and the final weighted summation generates a comprehensive index, providing a quantitative basis for risk grading and emergency response.) On this basis, the RI is calculated using the multi-factor linear weighting formula RI = ω1·CI + ω2·AR + ω3·SI (ω1, ω2, and ω3 are dynamic weights calibrated through historical data regression analysis, with a default ratio of 0.4:0.3:0.3). This model adheres to the principles of conflict coreness (sensitive scenarios can increase the CI weight to 0.5), anomaly sensitivity adjustment (AR weight increases by 20% in the event of sudden anomalies), and subject influence correction (negative communication by high-influence subjects amplifies SI contribution). The RI thresholds (level 1 RI ≥ 0.8, level 2 0.6 ≤ RI ≤ 0.8, level 3 0.4 ≤ RI ≤ 0.6) are calibrated through clustering using the DBSCAN algorithm through the historical case library, and the effects are evaluated monthly through the confusion matrix. When the missed rate of level 1 events is greater than 5%, the threshold fine-tuning (step size 0.05) is triggered to ensure that the risk level division is in line with the actual evolution of public opinion, providing a scientific and dynamic quantitative basis for differentiated emergency response.

[0136] Differentiated emergency response is achieved through risk index (RI) quantification, and a three-level response strategy system is established:

[0137] Level 1 response (RI ≥ 0.8): Triggering the network-wide emergency linkage mechanism, the system pushes alert information containing real-time risk indexes and core transmission nodes to the head of the emergency management department through an SMS gateway that supports the SMPP protocol, with a delay of less than 10 seconds. Synchronously call the social platform collaboration interface based on OAuth 2.0 authentication, such as automatically publishing an authoritative rumor-refuting statement that has been verified by the BERT model semantic consistency (similarity threshold ≥ 0.95) through the Weibo Gold V account API. At the technical management level, the system interacts with the platform in real time based on the OAuth 2.0 protocol. For abnormal propagation accounts marked with an abnormal probability ≥ 0.8 by the STGCN model, the system implements a 15-minute permission restriction prohibiting forwarding and commenting through the platform's open API interface. It also starts a distributed log tracking system that stores logs in an Elasticsearch cluster (retrieval delay < 200ms). Based on the city-level precision IP geolocation database, it locks down high-risk IP clusters that have detected abnormal propagation behavior for five consecutive times.

[0138] Level 2 Response (0.6≤RI<0.8): Focusing on regional risk control, we utilize geo-fencing technology based on the AutoNavi Map API to delineate polygonal areas, identifying high-risk administrative areas with district- and county-level precision. We then push authoritative information through CDN nodes covering over 80% of users in the target area, with a push latency of ≤3 seconds. We also activate a dual-channel content review mechanism using the BERT model and rule engine. A pre-trained BERT model in Chinese performs semantic risk assessment (with a negative sentiment threshold set at 0.6). The rule engine matches a pre-set sensitive vocabulary containing over 5,000 dynamically updated words. The results of the two are combined using a voting mechanism with a BERT weight of 0.6 and a rule engine weight of 0.4, ensuring a ≥95% accuracy rate in identifying sensitive content and a single piece of content review latency of <5 seconds.

[0139] Level 3 Response (0.4 ≤ RI < 0.6): Focused on decision support, when the risk index (RI) falls within this range, the system automatically generates multi-dimensional analysis reports, including conflict heatmaps based on geolocation data density visualization and propagation dendrograms derived from Neo4j graphs. These reports are pushed to the public opinion monitoring console via a web interface, prompting manual intervention for review. Dynamic monitoring mode is also enabled, collecting ≥1,000 new data points every 30 minutes. The risk scoring model parameters are updated using a stochastic gradient descent algorithm with a learning rate of 0.01 to optimize the accuracy of RI predictions. Operators can manually adjust the intervention threshold in steps of 0.05 based on the real-time risk situation. After adjustment, the system automatically verifies the effectiveness of the new threshold using a KS test with a test efficiency of ≥0.75 to ensure that it accurately distinguishes different risk levels. However, since updates to the risk scoring model parameters can indirectly affect the distribution of RI calculations (e.g., model optimization may cause an overall RI shift), threshold recalibration is required, either manually or through the system, to accommodate dynamic risk grading requirements.

[0140] In addition, the intelligent decision-making and emergency response module is also equipped with a closed-loop feedback optimization mechanism. The closed-loop feedback optimization system improves the handling effect through multi-dimensional indicator evaluation and strategy iteration:

[0141] Quantitative evaluation of treatment effects (such as Figure 5 shown):

[0142] 1) Public opinion cooling speed: The core indicator is the hourly decrease in the proportion of negative sentiment. This decrease is calculated in real time by dividing the difference between the current and previous moments' negative sentiment content ratios by the previous moment's ratio. This is calculated in real time by the sentiment analysis module that integrates Snow NLP and a custom dictionary to reflect the effectiveness of short-term interventions within a 24-hour evaluation cycle.

[0143] 2) Transmission path blocking rate: This is calculated by comparing the forwarding volume of abnormal accounts 30 minutes before and 30 minutes after the restriction. Specifically, it is calculated as the ratio of the difference in forwarding volume before and after the restriction to the forwarding volume before the restriction. When this ratio is ≥ 0.6, it is considered to be effectively blocked. Path changes are calculated using the Neo4j graph traversal algorithm.

[0144] 3) Public Satisfaction: Based on sampling survey data with a sample size of ≥1,000, the proportion of positive evaluations of government responsiveness and information transparency is calculated (a threshold of ≥60% is considered acceptable). The survey data is unstructured through sentiment analysis combined with natural language processing technology for topic extraction.

[0145] Policy iteration and rule updating:

[0146] 1) A / B testing framework: Set up control groups for different treatment combinations (such as "flow control strategy" and "flow control + KOL guidance"), use a 72-hour test period, and conduct a two-sample T-test with a significance level of α = 0.05 on the core indicators of a single group of ≥ 500 communication records (hourly decrease in negative emotions, communication path blocking rate) to compare the differences in indicators.

[0147] 2) Historical Case Analysis: A random forest algorithm with 100 trees and a feature importance threshold of 0.1 was used to analyze historical disposal data with a capacity of ≥ 100,000 (such as product quality crises and public health event case libraries), and dynamically adjust rule weights (used to characterize the priority or influence of specific variables in the disposal strategy. The higher the weight, the more critical the role of the variable in rule matching and strategy selection. For example, the influence weight of government response speed in the first-level response ranges from 0.3 to 0.5, indicating that in the decision-making rules of the first-level response strategy, the impact of government response speed on the disposal effect accounts for 30%-50%).

[0148] 3) Update of the entire rule base: Word vector changes are detected monthly through the Word2Vec model (new words are judged as words with a similarity < 0.7). At the same time, sensitive words that are missed during manual review (such as emerging variants) are added to the BERT custom filtering rule base to improve the filtering accuracy of non-Chinese characters and advertising information, integrate emerging network terms and cross-platform collaborative disposal strategies (such as the TikTok-Weibo linkage flow control rules), and verify the updated rule set in a simulation environment containing 20 typical public opinion scenarios to ensure that the rule coverage rate is ≥ 90% and the conflict rate is < 5%. The error compensation parameters of the IP territorial resolution model are corrected through the analysis of the transmission path blocking rate, further narrowing the error range of municipal administrative district identification, strengthening the temporal consistency of timestamps and transmission relationship chains, and ensuring that the response mechanism continues to adapt to the evolution of the public opinion ecosystem.

[0149] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.

Claims

1. An intelligent classification and emergency decision-making system for online public opinion based on multimodal fusion and dynamic evolution, characterized by: The system includes a multi-source data acquisition and preprocessing module, a multi-dimensional classification engine module, a fact map construction and anomaly detection module, a stakeholder dynamic risk assessment module, and an intelligent decision-making and emergency response module; The multi-source data acquisition and preprocessing module is used to crawl multimodal data from the network platform, pre-process the crawled data, perform spatiotemporal labeling, and construct a propagation path map to identify key propagation nodes; The multi-dimensional classification engine module quantifies the conflict intensity of public opinion events based on the data output by the multi-source data acquisition and preprocessing module according to three attributes: sentiment polarity, event classification, and stakeholder relevance. At the same time, it adjusts the attribute threshold through a dynamic optimization mechanism and dynamically updates the rule vocabulary to incorporate emerging network terms, thereby maintaining the long-term adaptability of the conflict intensity quantification model. The event graph construction and anomaly detection module is used to construct the public opinion propagation path from a micro level and the common logic chain model of public opinion events from a macro level, while monitoring the speed of public opinion propagation and issuing alarms; The stakeholder dynamic risk assessment module refines the classification of online public opinion participants by constructing differentiated judgment rules and quantitative models, and divides the influence levels of the participants, thereby providing a basis for differentiated communication intervention; the stakeholder dynamic risk assessment module also simulates the evolution of public opinion through system dynamics and conducts risk simulation; The intelligent decision-making and emergency response module executes different levels of emergency measures based on the risk index according to the hierarchical response strategy.

2. The system according to claim 1, wherein: The multi-source data acquisition and preprocessing module includes a data acquisition submodule, a preprocessing submodule and a spatiotemporal label generation submodule; The data collection submodule uses a distributed crawler cluster architecture to crawl text, images, videos, and live streaming data from the network platform. It configures a dynamic IP proxy pool and simulates browser fingerprint technology to circumvent the anti-crawling mechanism of the network platform. It accesses government public data, news media databases, and third-party public opinion monitoring platforms through a RESTful interface to supplement structured data such as event timelines and geographic locations. At the same time, it uses Kafka message queues to implement streaming processing to avoid data gaps and maintain temporal consistency. The preprocessing submodule is used to process the crawled multimodal data. For text data, non-Chinese characters, advertising information, and sensitive words are first filtered out. Then, key entities including names, organizations, places, and times are annotated. Finally, a hybrid sentiment analysis of the text content is performed using the SnowNLP framework and a custom dictionary. For image and video data, key visual features including crowd density and slogan text are first extracted. Then, a classification model is used to detect high-risk sensitive content such as violent scenes, rallies, and political slogans. At the same time, multi-label classification technology is used to perform multi-label semantic analysis on complex scenes to generate semantic labels. The spatiotemporal tag generation submodule adds millisecond-level timestamps and city-level geographic tags to the processed data, and constructs a propagation path map with node influence weights based on social network forwarding relationships.

3. The system according to claim 1, wherein: The multidimensional classification engine module includes a conflict intensity quantification submodule and a dynamic optimization submodule; the conflict intensity quantification submodule quantifies the conflict intensity of abstract opposing relationships in public opinion events based on sentiment polarity, event classification, and entity relevance of stakeholders; The dynamic optimization submodule improves the adaptability of the conflict intensity quantification submodule through fuzzy logic and rule iteration, wherein the fuzzy logic uses a trapezoidal membership function to dynamically adjust the classification threshold to improve the sensitivity of conflict warning; the rule iteration updates the vocabulary mapping rules by mining emerging network terms and combining them with a semantic vector space model to ensure that the classification rules are synchronized with the evolution of online public opinion.

4. The system according to claim 3, characterized in that The sentiment polarity is quantified by the following formula: In the formula, x represents the frequency of negative keywords in the text, k and x0 are dynamically adjusted according to the event type to achieve differentiated measurement of sentiment intensity in different scenarios; The event classification is based on TF-IDF text feature extraction and LDA topic model, supports multi-label joint annotation, and models topic probability distribution through Dirichlet distribution to achieve hierarchical analysis of event semantics; The entity relevance is obtained by weighted summation based on the frequency of occurrence of the institution keyword in the text, β(S) = ∑(keyword frequency × institution weight); The conflict intensity is represented by the ratio of conflict samples to total samples, where the total samples include conflict samples and non-conflict samples. The conflict intensity is represented by: In the formula, the numerator is calculated by summing the sentiment polarity strength μ of each conflict sample. neg , event classification adjustment coefficient α(T) and entity relevance β(S), which comprehensively reflect the contribution of multi-dimensional attributes to conflict intensity, N total is the total sample size, and the conflict intensity comparison of events of different scales is achieved through normalization processing.

5. The system according to claim 1, wherein: The said event graph construction and anomaly detection module includes a graph construction submodule and an anomaly detection submodule; The graph construction submodule constructs a graph of events from both micro and macro levels. At the micro level, it identifies the subjects of public opinion events and the interactions between them, constructs a chain-type propagation relationship triple containing the subjects and the interactions, defines the logical edge types, and stores the relationship network in the Neo4j graph database. At the macro level, it performs thematic clustering on the historical case library, abstracting the common logical chain pattern of "emergency → official delayed response → negative emotional outbreak → secondary propagation". The similarity between the element set of the new public opinion event and the common logical chain pattern is calculated using the Jaccard coefficient. When the Jaccard coefficient reaches a certain level, the historical pattern association is triggered, and a forecast report on the event evolution trend is generated. The anomaly detection submodule detects the node propagation speed and sentiment change rate through the STGCN model, and constructs a propagation speed prediction curve through the LDA topic model; when the node propagation speed suddenly increases by more than 500% or the sentiment change rate deviates from the historical mean by ±2 times the standard deviation, the corresponding node is marked as an abnormal account and a preliminary warning is triggered; when the actual propagation speed deviates from the predicted propagation speed by more than ±20%, a real-time alarm is triggered.

6. The system according to claim 5, characterized in that In the anomaly detection submodule, the comprehensive anomaly risk value AR is obtained by weighted fusion of the node propagation speed and the normalized propagation speed prediction deviation rate to achieve collaborative early warning; wherein, the comprehensive anomaly risk value AR is expressed as: AR=ω STGCN ·P STGCN +ω LDA ·D LDA_norm Where, P STGCN is the node propagation speed, ω STGCN With ω LDA All are fixed weights; min(|D LDA |) and max(|D LDA |) is the minimum and maximum absolute value of the deviation in the current calculation cycle, and ε is a very small constant.

7. The system according to claim 1, wherein: The stakeholder dynamic risk assessment module includes a stakeholder classification submodule and a risk assessment submodule; The stakeholder classification submodule categorizes the participants of public opinion events and divides the influence levels of the participants of public opinion events through the communication power index; the participants of public opinion events are divided into five categories: direct stakeholders, indirect stakeholders, key opinion leaders, public interest parties and marginal groups; the communication power index I is expressed as: I = 0.3 × number of followers + 0.4 × forwarding rate + 0.2 × historical influence rate + 0.1 × government relevance Participants with I ≥ 0.8 were classified as high-influence level, participants with 0.5 ≤ I < 0.8 were classified as medium-influence level, and participants with I < 0.5 were classified as low-influence level; The risk assessment submodule simulates the evolution of public opinion through system dynamics and conducts risk simulation, including: defining core variables, namely government response speed VR, event sensitivity ES and KOL intervention intensity KI, and constructing the differential equation between the negative emotion change rate ΔN and the core variables ξ is a random disturbance term that obeys the normal distribution; the system dynamics simulation output includes a risk indicator time series prediction curve and a sensitivity analysis report. The risk indicator time series prediction curve shows the predicted values ​​of the negative emotion change rate and the speed of the spread range in the next 24 hours, with a time resolution of 15 minutes. The sensitivity analysis report shows the impact of adjusting the core variables on the evolution of public opinion.

8. The system according to claim 1, wherein: The intelligent decision-making and emergency response module includes an emergency response submodule and a closed-loop feedback submodule; The emergency response submodule implements different levels of emergency measures based on the risk index (RI). When the RI is ≥ 0.8, a level 1 response is triggered, activating the network-wide emergency linkage mechanism. A real-time alert is sent to the head of the emergency management department via the SMS gateway, and the social platform collaboration interface is simultaneously invoked to automatically issue an authoritative rumor-refuting statement. Furthermore, based on the OAuth 2.0 protocol, the submodule interacts with the platform in real time to restrict permissions for accounts that spread abnormal information, and locks down high-risk IP clusters through distributed log tracking. When 0.6≤RI<0.8, a Level 2 response is triggered. Geofencing technology is used to demarcate high-risk areas, authoritative information is pushed to target user terminals through CDN nodes, and a dual-channel review mechanism using the BERT model and rule engine is activated to ensure an accuracy rate of over 95% in identifying sensitive content and a review delay of less than 5 seconds. When 0.4≤RI<0.6, a Level 3 response is triggered. A multi-dimensional report including a conflict heat map and a propagation tree diagram is automatically generated, prompting manual intervention for review. A dynamic monitoring mode is also activated, and the risk scoring model parameters are updated every 30 minutes. The closed-loop feedback submodule quantitatively evaluates the disposal effect through multi-dimensional indicators, iteratively updates the emergency disposal strategy and disposal rules, and verifies the updated rule set in a simulation environment of typical public opinion scenarios.

9. The system according to claim 8, characterized in that The closed-loop feedback submodule quantitatively evaluates the treatment effect through multi-dimensional indicators, including evaluating the treatment effect through the speed of public opinion cooling down, the blocking rate of the transmission path and the public satisfaction; the speed of public opinion cooling down is the decrease in the proportion of negative emotions per hour; the blocking rate of the transmission path is the ratio of the difference in the forwarding volume of the abnormal account before and after the flow is restricted to the forwarding volume before the flow is restricted, and when the ratio is not less than 0.6, it is judged to be effectively blocked; the public satisfaction is the proportion of positive evaluations on the government's response speed and information transparency in a sampling survey with a sample size greater than or equal to 1,000, and it is qualified when the proportion of positive evaluations is greater than or equal to 60%.

10. The system according to claim 8, wherein: The risk index RI is expressed as: RI=ω1·CI+ω2·AR+ω3·SI Where CI represents conflict intensity, AR represents comprehensive abnormal risk value, SI represents stakeholder influence index, ω1, ω2, and ω3 are dynamic weights calibrated through historical data regression analysis; The level thresholds corresponding to emergency measures are calibrated through clustering using the DBSCAN algorithm through the historical case library, and the effects are evaluated monthly through the confusion matrix. When the missed detection rate of first-level response events is greater than 5%, threshold fine-tuning is triggered.

Citation Information

Patent Citations

  • Online public opinion detection system and method integrating public opinion wind direction tracking and civil condition prediction

    CN112685621A

Cited By

  • Intelligent decision support system based on big data and artificial intelligence

    CN120851666A

  • Intelligent decision support system based on big data and artificial intelligence

    CN120851666B

  • Traffic public opinion risk label classification early warning method based on large model

    CN121071617A

  • Rail transit data analysis method and system based on cloud computing

    CN121167221A

  • Quality report generation method and system based on multi-dimensional data

    CN121187911A