Cross-platform information acquisition and public opinion processing optimization method and system
By building a cross-platform public opinion processing system, intelligent collaborative collection and standardized integration of multi-source heterogeneous data have been achieved. In-depth analysis is carried out using multi-dimensional analysis models, which solves the problems of incomplete information collection and single analysis dimensions in existing technologies, and improves the intelligence and response efficiency of public opinion monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ALPHA RISK CONTROL TECH CO LTD
- Filing Date
- 2026-01-20
- Publication Date
- 2026-05-08
AI Technical Summary
Existing online public opinion monitoring systems lack dynamic adaptability, cross-platform data collaboration, multimodal analysis, and in-depth source tracing capabilities, resulting in incomplete and untimely information collection, a single analysis dimension, serious resource waste, and difficulty in providing accurate intelligence support.
A cross-platform information collection and public opinion processing system is constructed. Through unified data format, multi-dimensional analysis model and intelligent scheduling, it realizes intelligent collaborative collection and standardized integration of multi-source heterogeneous data. Combined with dissemination heat calculation, risk assessment and anomaly detection, it generates public opinion analysis results and triggers early warning.
It enables dynamic and optimal allocation of cross-platform resources, enhances the depth and breadth of public opinion analysis, improves the system's intelligence level and response efficiency, and ensures timely information collection and accurate early warning.
Smart Images

Figure CN121998632A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of Internet information processing technology, specifically to a cross-platform information collection and public opinion processing optimization method and system. Background Technology
[0002] With the rapid development of internet technology, social media, news portals, forums, and blogs have become the main channels for the public to express opinions and disseminate information, resulting in an explosive growth of online public opinion. Online public opinion is not only a barometer of social sentiment but also has a profound impact on public management, business operations, and social stability. Therefore, timely, accurate, and comprehensive monitoring, analysis, and early warning of online public opinion have become crucial issues that government agencies and enterprises urgently need to address.
[0003] Currently, numerous online public opinion monitoring systems exist in the market. However, these systems generally suffer from the following shortcomings in practical applications: First, they largely rely on manually preset keywords and data sources, making it difficult to dynamically adapt to changes in online information, resulting in incomplete and untimely data collection with significant lag. Second, their analysis methods often focus on the sentiment or thematic analysis of single texts, lacking the ability to gain macro-level, real-time insights into the overall online public opinion trend, dissemination dynamics, and cross-platform dissemination networks. Third, systems from different service providers or departments operate independently, with varying technical architectures, data standards, and algorithm models, leading to significant duplication of effort and resource waste in basic data collection, storage, cleaning, and model development, creating information silos. Finally, existing systems lack the ability to identify and trace the source of false information and abnormal dissemination patterns, making it difficult to provide decision-makers with accurate and in-depth intelligence support in complex online environments.
[0004] While some patented technologies have proposed solutions to address the aforementioned issues, such as coordinating resources through cloud platforms or introducing artificial intelligence for real-time monitoring, significant shortcomings remain in areas such as intelligent collaborative data collection across platforms, integrated application of multimodal deep analysis models, and dynamic resource optimization and in-depth tracing based on analysis results. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a cross-platform information collection and public opinion processing optimization method and system, aiming to solve the problems of low collection efficiency, single analysis dimension, system redundancy and insufficient in-depth analysis capability in the prior art.
[0006] In a first aspect, embodiments of the present invention provide a cross-platform information collection and public opinion processing optimization method, comprising: collecting online public opinion data from multiple independent public opinion monitoring subsystems; performing standardized preprocessing on the collected online public opinion data from different public opinion monitoring subsystems to generate standard public opinion data with a unified data format; performing public opinion analysis on the standard public opinion data based on a preset multi-dimensional analysis model to generate public opinion analysis results, wherein the public opinion analysis includes at least the calculation of public opinion dissemination heat and public opinion risk assessment; and generating a public opinion monitoring report and early warning information based on the public opinion analysis results.
[0007] Secondly, embodiments of the present invention provide a cross-platform information collection and public opinion processing optimization system, comprising: a data collection module for collecting online public opinion data from multiple independent public opinion monitoring subsystems; a data preprocessing module for standardizing the collected online public opinion data from different public opinion monitoring subsystems to generate standard public opinion data with a unified data format; a data analysis module for performing public opinion analysis on the standard public opinion data based on a preset multi-dimensional analysis model to generate public opinion analysis results, wherein the public opinion analysis includes at least the calculation of public opinion dissemination heat and public opinion risk assessment; and a report and early warning module for generating public opinion monitoring reports and early warning information based on the public opinion analysis results.
[0008] Thirdly, embodiments of the present invention provide an electronic device, including: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to perform the method as described in the first aspect.
[0009] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in the first aspect.
[0010] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: First, by building a unified cross-platform processing framework, intelligent collaborative collection and standardized integration of multi-source heterogeneous public opinion data are realized. This not only effectively breaks down information silos, but also achieves dynamic optimal allocation of collection resources across subsystems through intelligent scheduling from a global perspective, thus avoiding duplication of construction and waste of resources from the source.
[0011] Secondly, by introducing a multi-dimensional analysis model integrating propagation analysis, network analysis, semantic sentiment analysis, and comprehensive risk assessment, the system can comprehensively understand public opinion trends from macro to micro and from static to dynamic perspectives, significantly enhancing the depth and breadth of analysis. Furthermore, by constructing a real-time feedback loop of "analysis-early warning-resource scheduling" based on the analysis results, the in-depth understanding of public opinion trends is directly transformed into dynamic allocation instructions for system data collection and computing resources, as well as precise early warning actions. This achieves a paradigm shift from passive monitoring and manual handling to proactive perception and intelligent control, greatly improving the overall system's intelligence level and response efficiency. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart of a cross-platform information collection and public opinion processing optimization method provided in an embodiment of the present invention.
[0014] Figure 2 This is a schematic diagram of the architecture of a cross-platform information collection and public opinion processing optimization system provided in an embodiment of the present invention.
[0015] Figure 3 This is a detailed flowchart of the data acquisition and preprocessing steps in one embodiment of the present invention.
[0016] Figure 4 This is a detailed flowchart of the multi-dimensional public opinion analysis steps in one embodiment of the present invention.
[0017] Figure 5 This is a detailed flowchart of the report generation and resource optimization steps in one embodiment of the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the embodiments of this invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0019] It should be noted that, unless otherwise specified, the embodiments of the present invention and the features thereof can be combined with each other.
[0020] In this invention, "cross-platform" has a dual meaning: first, it refers to the diversity of data sources, including different types of online platforms such as social media, news portals, forums, and blogs; second, it refers to the compatibility of the system architecture, which can connect and integrate public opinion monitoring subsystems independently built by different service providers and departments to achieve unified processing of multi-source heterogeneous data. Example 1
[0021] See Figure 1 This illustrates the overall flow 100 of the cross-platform information collection and public opinion processing optimization method provided by an embodiment of the present invention. This method can be executed by a server, server cluster, or dedicated computing device integrated with relevant software, and includes the following steps: Step S110: Collect online public opinion data from multiple independent public opinion monitoring subsystems.
[0022] This step aims to integrate the data collection capabilities of multiple independently operating public opinion monitoring subsystems (e.g., systems from different service providers or business units). The implementing entity connects to these subsystems through standardized application programming interfaces (APIs) or data buses. The collected data includes, but is not limited to, text content such as news releases, social media posts, forum comments, and blog articles, as well as related metadata (such as publication time, publisher, source URL, number of reposts, number of comments, number of likes, etc.).
[0023] In an optional implementation, this step also includes an intelligent data acquisition and scheduling sub-step, the process of which can be found in [reference needed]. Figure 3 The upper part. Specifically, it includes: S111, obtain the network collection sites (such as specific news website URLs, social media account lists) and collection frequency currently configured for each access public opinion monitoring subsystem.
[0024] S112 performs a global statistical analysis of the configuration information of all subsystems. For example, it counts the total number of data collection tasks configured by different subsystems for each target website within a preset time window (e.g., the past 24 hours). .
[0025] S113, Calculate the average number of configurations across all websites. By setting threshold parameters α and β (e.g., α=2.0, β=0.5), a high configuration threshold is achieved. Low configuration threshold .
[0026] S114, Identification and Adjustment: For any website, if , it is determined as a "redundant collection website". The system sends an instruction to the global collection scheduler to reduce the collection frequency of this site (for example, from once per minute to once every five minutes) to save network bandwidth and computing resources.
[0027] If , it is determined as a "low - attention website". The system increases its collection frequency (for example, from once every ten minutes to once per minute), or recommends it to relevant subsystems for supplementary configuration to prevent information omission.
[0028] This intelligent scheduling mechanism effectively optimizes the collection resources across platforms and reduces redundant overhead.
[0029] Step S120, perform standardized pre - processing on the collected online public opinion data from different public opinion monitoring subsystems to generate standard public opinion data with a unified data format.
[0030] Since the data formats of the accessed subsystems are diverse (possibly XML, JSON, HTML, or private binary formats), the core of this step is data cleaning and format standardization. See Figure 3 the lower part of S121, data cleaning: Process the original text data. Use regular expressions and stop - word lists to remove meaningless stop words such as "de", "le", "zai", etc. Calculate text fingerprints using algorithms based on SimHash or MinHash, and remove duplicate information with content similarity exceeding a preset threshold (such as 95%).
[0031] S122, format conversion: Convert the cleaned data through a pre - written format adapter (Adapter) into the standard data format agreed upon within the system. Preferably, use the JSON format and define a set of core fields, for example: json { "id": "Unique identifier", "title": "Article title", "content": "Main text content", "publish_time": "Publication time in ISO 8601 format", "source_url": "Source link", "author": "Author / publisher", "media_type": "Media type (news, Weibo, etc.)", "meta_data": { "repost_count": Number of reposts, "comment_count": Number of comments "like_count": Number of likes } } The conversion process includes field name mapping, data type casting (such as converting a string date to a timestamp), and data validity validation.
[0032] S123, Feature Extraction: Extract features from standardized text data using natural language processing techniques.
[0033] Semantic feature extraction: Use pre-trained word vector models (such as Word2Vec, BERT) to convert text into a sequence of word vectors; or use TF-IDF (term frequency-inverse document frequency) to extract keywords and their weights.
[0034] Statistical feature extraction: Extract numerical features from metadata, such as the number of reposts, the number of comments, and the number of followers of the publisher (if available).
[0035] Ultimately, each piece of public opinion data is represented as a structured feature vector, namely "standard public opinion data".
[0036] Step S130: Based on the preset multi-dimensional analysis model, perform public opinion analysis on the standard public opinion data to generate public opinion analysis results.
[0037] This step is the core analysis phase, utilizing an integrated multi-dimensional analysis model to deeply mine standard public opinion data. This model consists of multiple sub-models, including propagation dynamics analysis, social network analysis, semantic sentiment analysis, and comprehensive risk assessment. Each sub-model processes data in parallel, and the output results (such as propagation indicators, network nodes, and sentiment scores) are aggregated into the risk assessment engine for fusion calculation, ultimately generating a unified, multi-dimensional public opinion analysis result.
[0038] See Figure 4 Specifically, it includes the following sub-steps: S131, Popularity Analysis: Calculate the dissemination index for each public opinion topic or key article.
[0039] Number of transmissions Total readership or exposure during the statistical period.
[0040] propagation rate : Calculate the increment of propagation quantity per unit time (e.g., per hour). .
[0041] Fan Spread Index This is an indicator used to measure the spread of public opinion among key groups. One calculation formula is:
[0042] in, Let be the weight of the i-th dissemination behavior, used to distinguish the basic influence differences between different behavior types (e.g., a forwarding weight of 1.0 and a comment weight of 0.8). n represents the total number of dissemination behaviors for this public opinion within the statistical time period. The influence score for the user initiating this action is a composite score that considers the user's attribute weights (such as social identity, whether they are verified media outlets, domain experts, credibility of past behavior, account activity, etc.) and their number of followers. N represents the total number of dissemination actions within the statistical period (i.e., N = n, used here as a normalized denominator to ensure comparability of indices for public opinion events of different scales). Weights User Influence Score Together, they determine the overall impact contribution of a single communication action.
[0043] This index more accurately reflects the penetrating influence of public opinion. The calculated fan diffusion index... This will serve as a key influence indicator for subsequent risk assessment and early warning. For example, the number of times a public opinion event spreads... Not high, but fan spread index When the risk level is abnormally high, it indicates that the information may be spreading rapidly among key populations. The system will mark such events as "high concealment risk" and appropriately increase their risk level. .
[0044] The attribute weights of the users who initiate the behavior can be derived from a comprehensive scoring model built on multiple dimensions, including their social identity (such as whether they are certified media or domain experts), the credibility of their historical behavior, and account activity, and are then compared with the number of followers. Together they are used to calculate the impact contribution of their dissemination behavior.
[0045] S132, Propagation Network Analysis: Construct a directed graph of public opinion propagation based on forwarding, @, and citation relationships. Nodes represent users or media, and edges represent propagation behaviors. Community detection algorithms (such as the Louvain algorithm) are used to identify propagation communities, and centrality algorithms (such as betweenness centrality and PageRank) are used to identify key propagation nodes (hubs).
[0046] S133, Risk Assessment and Classification: Topic classification: Input the text feature vector into a pre-trained text classification model (such as TextCNN, BERT classifier) to perform multi-level topic classification of public opinion (such as "technology-artificial intelligence-ethics").
[0047] Sentiment Analysis: A sentiment analysis model based on a pre-trained BERT model and fine-tuned on public opinion domain data is used to determine the sentiment polarity (positive, negative, neutral) and intensity of the text. The intensity of negative sentiment is taken from the negative category probability value output by the model, ranging from [0,1].
[0048] Comprehensive Risk Assessment: Constructing a Risk Assessment Model
[0049] in, The value of the logarithm of the propagation rate. For negative emotional intensity, The source authority level is 0-1. For the risk attribute values of the key propagation nodes associated with the network, - These are the weights obtained through training on historical data. According to... The risk levels are divided into high, medium, and low.
[0050] The authority of the information source is determined based on historical data and a multi-dimensional scoring model. Specifically, this includes: For news media or official accounts, an authority baseline is constructed based on the authenticity of their historical content, their record of debunking rumors, and their social credibility rating. For self-media accounts, the following formula is used to calculate the following: This is based on a combination of factors including their verification status, number of followers, quality of historical content, and number of reports received.
[0051] in, Let k be the evaluation metric (such as historical authenticity rate, certification level, number of followers, etc.). For the corresponding weights; The system supports model training based on manually labeled samples and optimizes weight allocation through supervised learning.
[0052] The sentiment analysis model is based on a pre-trained BERT model and fine-tuned on labeled data in the public opinion domain. The model outputs the probability distribution of sentiment polarity (positive, negative, neutral), where the negative sentiment intensity Sentiment_Neg takes the probability value of the negative category, ranging from [0,1].
[0053] Risk attribute values of key propagation nodes This method measures the potential risk of key nodes in a communication network by calculating the correlation between historical node behavior and current public opinion. It is based on historical node behavior (the proportion of nodes participating in high-risk events in the past 30 days). ), node's own credibility (Calculation method similar to source authority), and centrality indicators in the current propagation network. (Such as the normalized result of PageRank values) Comprehensive calculation: The specific calculation formula is as follows:
[0054] in, The proportion of historically high-risk events. For node credibility, The centrality index is α, β, and γ, which are configurable empirical weighting coefficients. The system default values are 0.4, 0.4, and 0.2.
[0055] Risk level classification thresholds are based on historical public opinion events. The distribution and manually labeled levels (high / medium / low risk) are set. The system uses the percentile method to determine the threshold. Collect risk scoring data for public opinion events over the past year; according to Sort the data and use the top 10% as the high-risk threshold (e.g., 0.75) and the top 30% as the medium-risk threshold (e.g., 0.5). The system allows administrators to dynamically adjust thresholds based on actual early warning effects.
[0056] Anomaly detection threshold The settings are based on the DTW distance distribution between historical typical propagation pattern sequences and normal sequences. Specific steps include: Collect a large number of known types of public opinion dissemination sequences (such as natural fermentation, marketing promotion, and rumor outbreak); Calculate the DTW distance between each type of sequence and the typical pattern sequence, and statistically analyze the distance distribution; The 90th percentile of the distance between the rumor outbreak sequence and the rumor pattern is used as the initial threshold. ; The system supports dynamic adjustment of thresholds based on false alarm rate and false negative rate.
[0057] In addition, the fan spread index It can also be introduced as a correction factor in risk assessment. For example, the risk assessment model can be extended to: ,in The corresponding weights are used to capture the spread risk based on user influence.
[0058] Furthermore, while calculating the propagation index in S131, the system introduces an independent anomaly detection module. This module samples the propagation volume of public opinion at fixed time intervals (e.g., 5 minutes) to form a time series. ,Will Compared with various typical propagation pattern sequences stored in the historical database (such as the "natural fermentation pattern") "Marketing and promotion model" "Rumor Outbreak Mode" A comparison was made using the Dynamic Time Warping (DTW) algorithm. The DTW distance between the sequence and each typical sequence. A similarity threshold is set. (For example, DTW distance is less than) Then it is considered that the shapes are similar). If and The DTW distance is less than And at the same time significantly smaller than its and If the distance is specified, the system determines that the current public opinion's propagation pattern is abnormal and matches the characteristics of a rumor, thus generating a high-level anomaly warning. This method can effectively identify abnormal propagation behaviors that are difficult to detect using conventional thresholds.
[0059] In a preferred embodiment, the topic classification step further integrates semantic understanding and tag generation technology based on a Large Language Model (LLM) to improve classification accuracy and semantic coherence. Specifically, this includes: Fine-grained semantic tag generation: The system uses an internally integrated LLM service client to call the API of a locally deployed large language model (such as ChatGLM-6B) that has been fine-tuned for text commands in the public opinion domain. The client constructs structured prompts, for example: "You are a public opinion analysis expert. Please generate three key semantic tags for the following text, in the format: Level 1 tag; Level 2 tag; Level 3 tag. Text: [] The preprocessed text is filled with prompts and sent to the LLM. After the model returns the results, the client extracts the labels at each level by parsing the semicolon ";" separator. For example, for a public opinion report about "a new energy vehicle battery catching fire", the model returns the text "new energy vehicle; battery safety; safety accident; social concern", which will be parsed to obtain four levels of labels, thus transcending the fixed category system of traditional classification models.
[0060] Confidence calibration and manual review queue: The large language model outputs a confidence score for each generated label. The system sets a confidence threshold (e.g., 0.85). For data where all labels have a confidence score higher than the threshold, the label is automatically adopted to complete the classification; for data with low-confidence labels or where the model outputs "uncertain," it is automatically added to the "low-confidence review queue" and pushed to the manual review interface. This, combined with label suggestions generated by the model, significantly improves the efficiency of manual review.
[0061] The tagging system evolves dynamically: the system periodically (e.g., weekly) analyzes newly emerging, high-frequency model-generated tags. When a new tag appears more than a set threshold (e.g., 100 times) and its semantic similarity to any tag in the existing tag library is less than 0.4 (calculated using word vector cosine similarity), the system will prominently display a notification in the background, suggesting that the administrator include it in the formal classification system. This enables the classification model and tagging system to evolve collaboratively, adapting to the rapid changes in online public opinion.
[0062] By introducing the deep semantic understanding capabilities of large language models, more accurate, flexible, and interpretable topic identification of public opinion content has been achieved. This solves the problem of insufficient generalization ability of traditional text classification models when facing new topics and complex expressions. Furthermore, it forms an efficient collaborative mechanism of "automatic machine classification as the main method and precise manual review as a supplement," which significantly improves the automation level and classification quality of large-scale public opinion processing.
[0063] Step S140: Based on the public opinion analysis results, generate a public opinion monitoring report and early warning information.
[0064] See Figure 5 This step transforms the analysis results into decision support information and forms a closed-loop processing mechanism.
[0065] S141, Report Generation: The system generates visual public opinion monitoring reports periodically (e.g., hourly) or in real-time. Report content includes: a ranking of trending topics (sorted by popularity), a chart showing the spread of key public opinion events, a list of risky public opinion events, a dissemination network graph, and a sentiment distribution pie chart. Reports can be presented in web dashboards, PDF documents, and other formats.
[0066] S142, Warning Triggered: The system background continuously monitors for S133. and in S131 When any public opinion Exceeding the "high-risk threshold", or When the number of cases surges beyond the "outbreak threshold" within a short period (e.g., 5 minutes), the alert engine is triggered. The alert information is pushed to designated alarm channels via message queues, such as monitoring dashboards, SMS, email, DingTalk / WeChat Work robots, etc.
[0067] In a preferred embodiment, the early warning triggering and push mechanism further includes multi-mode intelligent routing and closed-loop feedback functions to ensure that the early warning information can accurately reach the relevant responsible persons and form a closed-loop handling mechanism.
[0068] Warning Level - Responsible Person Mapping Matrix: The system has built-in configurable "Warning Level - Responsible Person" mapping rules. For example, warnings can be divided into three levels: "Red (high risk), Orange (medium risk), and Yellow (low risk)," and associated with different responsible departments or personnel groups (e.g., a red warning is simultaneously pushed to the public relations department, legal department, and senior management; a yellow warning is only pushed to the operations staff on duty).
[0069] Multi-channel adaptive push and confirmation: The system maintains a user status table and obtains users' real-time status by polling or subscribing to the online status interface of office software (such as WeChat Work). Simultaneously, the system records the historical average confirmation response time for various alert channels (such as instant messages, SMS, and phone calls) as "historical feedback preferences." The alert engine adaptively selects the push channel based on the recipient's current online status and preferences. For example, for currently online users, instant messages are prioritized via office software; if unread within 180 seconds, an SMS notification is superimposed; for alerts marked as red and unconfirmed for a long period (e.g., 300 seconds), a phone call is automatically triggered. Each alert message requires the recipient to click "Received" or "Processing," and the system records the confirmation time and the person making the confirmation.
[0070] Closed-loop feedback mechanism and early warning effectiveness evaluation: The system provides a convenient feedback portal for early warning response. Responsible personnel can directly submit brief response measures (e.g., "Verified as false information; a clarification statement is being drafted") in the attachment to the early warning message. The system automatically collects this feedback and correlates it with subsequent changes in the dissemination data of the public opinion (e.g., the rate of decline in popularity, shift in sentiment, etc.), regularly generating an "Early Warning Response Effectiveness Report" to evaluate the actual effectiveness of different early warning strategies and provide data support for the subsequent optimization of early warning thresholds and rules.
[0071] By constructing a complete early warning closed loop of "precise push, mandatory confirmation, convenient feedback, and effect evaluation", the pain points of traditional public opinion system early warnings that "only send out warnings, do not deliver them, and do not work" have been solved. This ensures that key early warning information can be responded to and handled in a timely and effective manner, and the handling experience is precipitated into system optimization knowledge, thereby improving the intelligence and refinement of the entire public opinion emergency management.
[0072] S143, Dynamic Resource Optimization (Extended Closed Loop): The system dynamically adjusts system resources based on the analysis results (heat and risk level) of all current active public opinion. For example, specific strategies include: the resource scheduler (e.g., a Kubernetes-based Horizontal Pod Autoscaler) monitors the load of the data analysis module in real time. When the task queue for processing high-heat, high-risk public opinion accumulates, it automatically increases the CPU and memory quotas for the computing Pod of the data analysis module 230; simultaneously, it sends instructions to the intelligent scheduler of the data collection module 210 to increase the concurrency of collection threads or the collection frequency for relevant high-risk information source websites. Conversely, for low-heat public opinion, resource allocation is appropriately reduced. This dynamic resource scheduling mechanism based on real-time feedback ensures optimal configuration of system resources at the global level.
[0073] It should be noted that the aforementioned adjustment of collection frequency based on configuration statistics (step S114) and adjustment of collection resources based on public opinion risk level (step S143) work together on the global collection strategy. When the two conflict, the system prioritizes the monitoring needs of high-risk public opinion sources, that is, temporarily overriding the frequency reduction strategy for redundant sites to ensure the timely acquisition of key information. The system maintains a dynamic 'list of key monitoring sites', which is updated in real time driven by the analysis results of high-risk public opinion.
[0074] Furthermore, for public opinion events marked as high-risk or exhibiting abnormal spread, the system initiates a deep source tracing analysis module. This module performs the following operations: Constructing an event knowledge graph: Using current public opinion events as core nodes, and leveraging entity recognition and relationship extraction technologies, an automatic knowledge graph is constructed that includes entities such as "people, institutions, locations, events, and claims" and their relationships (such as "publishing", "questioning", and "involving").
[0075] Knowledge Fusion and Contradiction Detection: Align the event graph with the system-maintained "Trusted Fact Knowledge Base" (constructed from structured information such as authoritative media and official reports). Utilize graph reasoning rules to detect whether there are logical contradictions or timeline conflicts between the "claims" nodes in the event graph and those in the Trusted Fact Knowledge Base.
[0076] Enhanced source propagation chain analysis: Combining the propagation network obtained from S132, we not only analyze the information diffusion path, but also focus on the credibility assessment of early sources upstream of the propagation chain (based on the verification records of their historical published content).
[0077] Generate a source tracing report: Output a structured report, such as: "Regarding the rumors about 'Event X,' its core claim A contradicts fact B released by authoritative source Y on date Z; this information originated from account C, which has a credibility rating of 'low.' Based on this assessment, the information is highly likely to be false and we recommend thorough investigation." This function greatly enhances the system's ability to identify and analyze information in complex environments.
[0078] As an optional extended implementation scheme, to address the issues of data privacy and collaborative modeling among different public opinion monitoring subsystems, this invention can integrate a federated learning framework. Under this scheme, each subsystem acts as a client of the federated learning, training a shared global model (such as a sentiment analysis model) locally using its own data, and only encrypting the gradient update values of the model parameters before uploading them to the central server (i.e., the system of this invention). The server aggregates updates from multiple clients using a secure aggregation algorithm (such as Secure Aggregation), generating an optimized global model, which is then distributed to each client. This iterative process enables the fusion of cross-platform data value and the collaborative improvement of model performance without sharing the original data.
[0079] Example 2 To further illustrate the specific implementation process and technical effects of the present invention, this embodiment takes a fictional "new energy vehicle battery fire" public opinion event as an example to demonstrate the complete execution process of the method in a real-world scenario.
[0080] Assume this system is connected to three independent public opinion monitoring subsystems: Subsystem A (monitoring news portals), Subsystem B (monitoring social media), and Subsystem C (monitoring industry forums). At 9:00 AM one morning, information about "the battery of a new energy vehicle of brand XX catching fire while driving" began appearing on multiple platforms.
[0081] Cross-platform data collection: The system receives data streams from three subsystems in real time via standard APIs. Simultaneously, the intelligent data acquisition and scheduling submodule (corresponding to...) Figure 3 (Upper part) Initiate global statistical analysis. Statistics show that a certain car forum was repeatedly configured with a data collection task 5 times within 24 hours across three subsystems, while the average number of times all websites were configured was 2. Based on the set threshold parameters (α=2.0, β=0.5), the high threshold was calculated. 4. Low configuration threshold Due to the number of configurations on a certain car forum The system marked it as a redundant data collection website and automatically reduced its global data collection frequency from once per minute to once every 5 minutes. Meanwhile, another "tech blog" was also flagged for having fewer data collection attempts than [presumably a specific website name]. The website was marked as low-interest, and the system increased its data collection frequency from once every 10 minutes to once per minute. This scheduling mechanism optimized data collection resources early in the event, reducing redundant network requests by approximately 15%.
[0082] Data standardization preprocessing: The collected raw data varied in format (including XML, JSON, and HTML). The system first cleaned the text, removing stop words and deduplicating it using the SimHash algorithm to remove duplicate posts with a content similarity higher than 95%, achieving a deduplication rate of 22%. Subsequently, a format adapter was used to uniformly convert the data into the system's internal standard JSON format (fields such as id, title, content, publish_time, meta_data, etc., as defined in Example 1). Finally, the BERT word vector model was used to extract semantic features from the text, and statistical features such as the number of reposts and comments were extracted from the metadata to form structured, standard public opinion data.
[0083] Multi-dimensional public opinion analysis: Popularity calculation: The system calculates the cumulative exposure of this event within 1 hour. 10,000 times. Within the first 30 minutes, exposure surged from 50,000 to 500,000, calculating the propagation rate. 10,000 / hour. Fan Spread Index According to the formula The calculation, which combines the weights of behaviors such as forwarding, commenting, and liking, as well as the user's influence score, yields the following result: .
[0084] Propagation network analysis: Based on the forwarding and @ relationship, a directed propagation graph was constructed to identify two key propagation nodes (automotive industry KOLs and authoritative media accounts), and the Louvain algorithm was used to discover two main propagation clusters: "car owner community" and "industry media community".
[0085] Risk assessment and classification: Topic classification: Deep semantic understanding of the text is performed through an integrated large language model (such as a finely tuned ChatGLM-6B) to generate multi-level tags: "New energy vehicles > Battery safety > Safety accidents > Social concerns".
[0086] Sentiment Analysis: Determining the Intensity of Negative Sentiment in a Text .
[0087] Comprehensive risk score: Substituted into the risk assessment model Among them, the authority of the information source Key node risk value Calculated Based on the risk level classification rules, this event is marked as high risk.
[0088] Anomaly Detection: The Dynamic Time Warping (DTW) module compares the time series of the event's propagation volume with historical "rumor outbreak patterns" sequences, and measures a DTW distance of 12.3 (less than the threshold). This triggers an abnormal propagation warning.
[0089] Report generation and early warning push: Report generation: The system automatically generates visual public opinion monitoring reports, including hot topic rankings, dissemination trend curves, dissemination network maps, sentiment distribution maps, etc., which are pushed to administrators in real time through a web dashboard.
[0090] Warning Triggering and Closed-Loop Push: A red warning is triggered when the risk score exceeds the high-risk threshold (0.8). Based on the preset "Warning Level - Responsible Person" mapping matrix, the warning information is simultaneously pushed to the Public Relations Department (WeChat Work), Legal Department (SMS + Email), and senior management (DingTalk robot). The system adaptively selects the push channel based on the recipient's online status and historical feedback preferences, requiring the recipient to confirm within 180 seconds. The warning message includes an event summary, source tracing suggestions, and a feedback entry point for handling the situation.
[0091] Dynamic resource optimization and in-depth source tracing: Based on the high popularity and high risk of the event, the resource scheduler automatically increases computing resources for the data analysis module (the number of CPU cores increases from 4 to 8) and increases the collection frequency of key information sources.
[0092] The deep tracing module was activated simultaneously, constructing an event knowledge graph. It identified the earliest posting account as a low-credibility self-media outlet and compared it with a trusted fact database, finding that the "vehicle model" mentioned did not match official records. The system included the tracing conclusion in its warning report: "Information is questionable; verification is recommended."
[0093] As can be seen from this embodiment, the method of the present invention can: achieve intelligent scheduling of cross-platform resources and reduce redundant collection while ensuring coverage in real public opinion events; complete the efficient conversion from multi-source heterogeneous data to standardized and structured features; conduct in-depth analysis through multi-dimensional models, accurately calculate the spread popularity, identify key nodes, and complete risk assessment and anomaly detection within 30 minutes; generate visual reports and trigger hierarchical early warnings, and realize the fully automated response of "precise push - confirmation feedback - closed-loop handling - resource optimization".
[0094] Example 3 This embodiment provides an architecture for a cross-platform information collection and public opinion processing optimization system 200, such as... Figure 2 As shown. The system 200 includes: Data acquisition module 210: responsible for executing step S110, including multiple data acquisition adapters for connecting to different public opinion monitoring subsystems; and an intelligent scheduler for dynamically adjusting the acquisition frequency.
[0095] Data preprocessing module 220: responsible for executing step S120, including data cleaning unit, format conversion unit and feature extraction unit.
[0096] Data analysis module 230: Responsible for executing step S130, it is the intelligent core of the system. It includes: Model Library 231: Stores multi-dimensional analysis models, such as propagation analysis models, network analysis models, classification models, risk assessment models, etc.
[0097] Computing Engine 232: Based on big data computing frameworks such as Spark and Flink, it performs distributed computing on massive amounts of standard public opinion data.
[0098] Anomaly detection unit 233: used to perform the aforementioned DTW anomaly detection.
[0099] Knowledge graph construction and tracing unit 234: used to support the aforementioned deep tracing analysis, and to perform entity recognition, relation extraction, knowledge fusion and contradiction detection.
[0100] Reporting and alerting module 240: responsible for executing step S140, including a visual report generator, a real-time alerting engine, and a resource scheduler.
[0101] Storage and Database 250: Used to store raw data, standard data, analysis results, model parameters, and knowledge graphs.
[0102] The system's 200 modules work together to achieve full-process automation and optimization, from multi-source data access, intelligent processing, in-depth analysis to result output.
[0103] Example 4 This embodiment provides an electronic device including one or more processors and a storage device. One or more programs are stored on the storage device. When the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to perform the method described in Embodiment 1. The electronic device may be a cloud server, a local server, or a high-performance workstation.
[0104] Example 5 This embodiment provides a computer-readable storage medium, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk. The storage medium stores a computer program, which, when executed by a processor, implements the method described in Embodiment 1.
[0105] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this invention, and these modifications or substitutions should all be covered within the scope of protection of this invention. Therefore, the scope of protection of this invention should be determined by the scope of the claims.
Claims
1. A cross-platform information collection and public opinion processing optimization method, characterized in that, include: Collect online public opinion data from multiple independent public opinion monitoring subsystems; The collected online public opinion data from different public opinion monitoring subsystems are subjected to standardized preprocessing to generate standard public opinion data with a unified data format; Based on a pre-set multi-dimensional analysis model, the standard public opinion data is analyzed to generate public opinion analysis results, wherein the public opinion analysis includes at least the calculation of public opinion dissemination heat and public opinion risk assessment. Based on the public opinion analysis results, a public opinion monitoring report and early warning information are generated.
2. The method according to claim 1, characterized in that, The collection of online public opinion data from multiple independent public opinion monitoring subsystems includes: Obtain information on the network data collection sites configured in each of the aforementioned public opinion monitoring subsystems; Based on statistical analysis of the network data collection site information configured in each of the aforementioned public opinion monitoring subsystems, redundant data collection sites that have been configured repeatedly are identified. Based on the identification results, the global acquisition frequency of the redundant acquisition sites is dynamically adjusted.
3. The method according to claim 2, characterized in that, The step of dynamically adjusting the global acquisition frequency of the redundant acquisition sites based on the identification results includes: The number of times each target website was subjected to data collection tasks configured by different public opinion monitoring subsystems within a preset time period was counted. Calculate the average number of configurations across all websites, and set high and low configuration thresholds; Websites with a configuration frequency exceeding the high configuration threshold are identified as redundant collection websites, and their global collection frequency is reduced. Websites with fewer configuration attempts than the low configuration threshold are identified as low-interest websites, and their global collection frequency is increased.
4. The method according to claim 1, characterized in that, The standardization preprocessing of the collected online public opinion data from different public opinion monitoring subsystems to generate standard public opinion data with a unified data format includes: The collected online public opinion data is cleaned to remove stop words and duplicate data; The cleaned data is converted into a predefined platform standard data format, wherein the conversion includes field mapping, data type conversion and data format validation; Feature extraction is performed on the converted data to generate standard public opinion data containing semantic and statistical features.
5. The method according to claim 1, characterized in that, The method of performing public opinion analysis on the standard public opinion data based on a preset multi-dimensional analysis model includes: Calculate the amount of public opinion dissemination, the dissemination rate, and the fan diffusion index; Identify the propagation path and key nodes of public opinion; Based on at least one of the following factors: dissemination volume, dissemination rate, fan diffusion index, dissemination path, and key dissemination nodes, a risk assessment and classification of public opinion is conducted.
6. The method according to claim 5, characterized in that, The calculation of the fan spread index takes into account the weights of forwarding, commenting, liking behaviors, and the attributes of the users who initiated the behaviors.
7. The method according to claim 1, characterized in that, The step of generating a public opinion monitoring report and early warning information based on the public opinion analysis results includes: Based on the aforementioned public opinion analysis results, a visualized public opinion monitoring report is generated, which includes trending topics, dissemination trends, and risk distribution. When the indicators in the public opinion analysis results exceed the preset threshold, an early warning message is automatically generated and pushed.
8. The method according to claim 1, characterized in that, The method further includes: Based on the public opinion heat and risk level in the aforementioned public opinion analysis results, the allocation of computing resources in the data collection and data analysis stages is dynamically adjusted.
9. A cross-platform information collection and public opinion processing optimization system, characterized in that, include: The data acquisition module is used to collect online public opinion data from multiple independent public opinion monitoring subsystems; The data preprocessing module is used to perform standardized preprocessing on the collected online public opinion data from different public opinion monitoring subsystems to generate standard public opinion data with a unified data format. The data analysis module is used to perform public opinion analysis on the standard public opinion data based on a preset multi-dimensional analysis model to generate public opinion analysis results. The public opinion analysis includes at least the calculation of public opinion dissemination heat and public opinion risk assessment. The reporting and early warning module is used to generate public opinion monitoring reports and early warning information based on the public opinion analysis results.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 8.