Cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception

Through a hybrid crawler architecture and improved deep learning model, the data missing, model accuracy and early warning lag of the comment public opinion monitoring system are solved, and efficient and accurate real-time monitoring of public opinion is achieved.

CN120371898AInactive Publication Date: 2025-07-25HUNAN UNIV OF SCI & TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510854641.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing comments and public opinion monitoring system has many shortcomings in data collection and characterization, model architecture and performance, early warning mechanism and response, resulting in high data missing rate, low model accuracy, high warning false alarm rate, and lagging response, which cannot meet the real-time monitoring needs of emergencies.

Method used

The hybrid crawler architecture is used to dynamically capture multi-platform comment data, perform quality evaluation and self-correction, and use an improved two-way long and short-term memory network and multi-head attention mechanism for deep learning, combining temperature scaling and dual-threshold warning mechanism to achieve emotional classification and real-time warning.

Benefits of technology

It has achieved the improvement of data capture rate, improved model accuracy, reduced early warning false alarm rate and shortened response time, meeting the real-time monitoring needs of emergencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371898A_ABST
    Figure CN120371898A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception. The method comprises the following steps: dynamically capturing multi-platform comment public opinion data based on a hybrid crawler architecture; performing dynamic quality evaluation and self-correction on the comment public opinion data, performing integrity verification, noise filtering and multi-modal alignment, and outputting a standardized multi-modal tensor data set; training the data set; a temperature scaling adjustment model is adopted to output probability distribution calibration confidence; outputting an emotion classification result and a confidence score, and storing the result and the confidence score in a time sequence database; double-threshold dynamic hierarchical real-time early warning is adopted, and a real-time early warning signal is output; and updating the keyword library and the model parameters, and generating disposal suggestions and a visual thermodynamic diagram in a management background. According to the method, dynamic crawlers and multiple modes are fused, limitation of a single data source is broken through, conjoint analysis of short video bullet screens and image-text comments is achieved for the first time, and public opinion monitoring dimensions are expanded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and particularly to a cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception. Background Art

[0002] With the further development of AI technology, comment public opinion analysis will be more accurate and efficient, providing stronger support for decision-making. The existing technology has the following three core defects, which restrict the practical application of the comment public opinion monitoring system, specifically:

[0003] Defect 1: Data collection and representation bottleneck; specifically manifested as:

[0004] First, the lack of dynamic content capture; traditional crawlers based on Requests or BeautifulSoup are difficult to parse JavaScript dynamic rendering pages such as the waterfall flow of Weibo comment areas and Douyin bullet screens, resulting in too high a data loss rate;

[0005] Second, the fragmentation of multi-modal data; existing systems usually process text, emojis (such as Emoji, kaomoji), and rating data separately, lacking a unified semantic mapping framework. For example, the correlation between a five-star rating and the comment text "very satisfied" has not been effectively mined;

[0006] Third, the lack of data real-time performance; adopting a timed batch crawling strategy, it is impossible to capture the instantaneous propagation characteristics of sudden public opinions. For example, within 1 hour after a certain hot event breaks out, the comment volume surges by 500% but the system fails to respond in time;

[0007] Defect 2: Model architecture and performance limitations; specifically manifested as:

[0008] 1. Defects in long-range dependence modeling; due to the vanishing gradient problem of the standard long short-term memory network (LSTM), the accuracy of judging the sentiment tendency of long texts with more than 50 basic processing units (tokens) of text data drops by 12.7%;

[0009] 2. Rigid feature weight allocation; the traditional attention mechanism uses a single attention head and is difficult to capture the complex interaction relationships between sentiment words (such as "angry"), degree adverbs (such as "extremely"), and negative words (such as "not");

[0010] 3. Vulnerability to adversarial samples; existing models have poor robustness to adversarial attacks such as containing typos (such as using "fuwucha" instead of "service is poor") and inserting irrelevant characters;

[0011] Defect 3: Early warning mechanism and response lag; specifically manifested as:

[0012] 1) The single threshold strategy fails. Relying solely on the sentiment polarity threshold (e.g., the proportion of negative comments > 20%), it is prone to false alarms. For example, during a promotional event, "the price is too high" may be normal user feedback rather than a crisis event;

[0013] 2) The lack of spatio-temporal dimension analysis fails to conduct multi-dimensional research and judgment by combining the geographical distribution of comments and time aggregation;

[0014] 3) Insufficient human-machine collaboration: After the early warning is triggered, it relies on manual review, with a long average response time, unable to meet the requirements for handling emergencies. Summary of the Invention

[0015] To solve the above technical problems, the present invention provides a cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception, which has a simple algorithm and high working efficiency.

[0016] The technical solution of the present invention to solve the above technical problems is: A cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception, including the following steps:

[0017] S1: Heterogeneous data collection, dynamically crawling multi-platform comment public opinion data based on a hybrid crawler architecture;

[0018] S2: Dynamically evaluate and self-correct the quality of the comment public opinion data, perform integrity verification, noise filtering, and multi-modal alignment, and output a standardized multi-modal tensor data set, which is stored in a distributed database;

[0019] S3: Use an improved model based on a bidirectional long short-term memory network and an attention mechanism to perform deep learning training on the data set, and deeply mine and understand the semantic information of the comment public opinion;

[0020] S4: Adjust the model output probability distribution using temperature scaling to calibrate the confidence level;

[0021] S5: Output the sentiment classification result and the confidence score, and store them in a time series database;

[0022] S6: Adopt a dual-threshold dynamic hierarchical real-time early warning, output a real-time early warning signal, and realize the dynamic perception of the comment public opinion;

[0023] S7: Update the keyword library and model parameters, and generate disposal suggestions and a visualization heat map in the management background, and dynamically generate a comment public opinion decision-making method.

[0024] For the above cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception, the specific process of the step S1 is as follows:

[0025] S11, Uniform Resource Locator (URL) Seed Injection: The initial URL list is injected into the Redis database queue through the start_urls field of the distributed crawler framework Scrapy-Redis; the Bloom Filter is used for URL deduplication;

[0026] S12, Dynamic Task Allocation: Each Selenium Worker node pulls the URLs to be crawled from the Redis queue of the in-memory key-value storage system through the distributed crawler framework;

[0027] S13, Headless Chrome Configuration: Each Worker node starts multiple Chrome instances;

[0028] S14, Page Loading and Rendering: The basic loading loads the page through driver.get and sets the timeout;

[0029] S15, Structured Data Extraction: Use the XPath of the Extensible Markup Language (XML) or the Cascading Style Sheets (CSS) selector to locate elements; perform multi-modal data processing. For text data, directly extract and clean it, removing HyperText Markup Language (HTML) tags and special characters; for emoji data, parse Unicode or image links and map them to semantic tags; perform normalization processing on the rating data;

[0030] S16, Data Storage: Includes real-time writing and deduplication mechanisms;

[0031] Real-time Writing: Batch write the data into the sharded cluster of the non-relational database system MongoDB through asynchronous Input / Output (I / O);

[0032] Deduplication Mechanism: Generate a unique key based on the user identifier user_id + timestamp + the hash value of the content content_hash to avoid duplicate storage.

[0033] In the above cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception, in step S2, the process of noise filtering is as follows: Apply the feature selection algorithm based on information entropy to automatically eliminate duplicate comments and content with poor information volume.

[0034] In the above cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception, in step S2, the process of multi-modal alignment is as follows: Design a joint embedding space based on contrastive learning, map text, emojis, and ratings to the same high-dimensional space, and calculate the cross-modal correlation weights through cosine similarity.

[0035] In the above cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception, in step S3, the improved model includes a bidirectional long short-term memory network layer, a multi-head attention layer, and an adversarial training layer;

[0036] In the bidirectional LSTM layer, a gated recurrent unit (GRU) is adopted, with 128 hidden units set to capture forward and backward temporal dependencies;

[0037] In the multi-head attention layer, 8 independent attention heads are deployed, focusing on sentiment words, degree modifiers, negation words, and entity nouns respectively, and the multi-head outputs are weighted and fused through a learnable matrix;

[0038] In the adversarial training layer, character-level perturbations are injected during the training phase. The character-level perturbations include random substitution, insertion, and deletion to enhance the deep understanding of semantics, and the projected gradient descent (PGD) attack is used to generate adversarial samples to improve the robustness of the model.

[0039] In the above cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception, in step S4, the temperature scaling formula is used to scale the original output logits vector of the model, and by adjusting the temperature, the smoothness of the probability distribution is controlled;

[0040]

[0041] Among them, is the temperature parameter scalar; is the class probability after the i-th calibration; is the i-th original output logits vector of the model; is the j-th original output logits vector of the model; K is the number of model outputs, equal to the total number of classes;

[0042] The optimization objective is: to minimize the negative log-likelihood loss on the calibration set and find the optimal temperature;

[0043] First, 20% of the training data is reserved as the calibration set, and then the negative log-likelihood loss is minimized:

[0044]

[0045] Among them is the optimal temperature obtained through optimization, is the one-hot encoding indicating whether the m-th sample of the true label belongs to the n-th class, is the predicted probability that the calibrated model predicts the m-th sample belongs to the n-th class, and N is the total number of samples.

[0046] For the above cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception, in step S5, the time series database is selected as InfluxDB; the data retention policy is: the original data is retained for 30 days, and the downsampled aggregated data is retained for 1 year; the storage optimization policy is: batch writing, and batch writing is triggered every time 1000 records are accumulated or every 1 second; the index design is: an inverted index is established for the platform and sentiment fields; data compression enables Snappy algorithm compression.

[0047] For the above cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception, in step S6, set the sentiment intensity threshold α and the keyword density threshold β;

[0048] For the sentiment intensity threshold α, dynamically calibrate based on historical data, and use the Exponentially Weighted Moving Average (EWMA) algorithm to update the sentiment intensity threshold α;

[0049] For the keyword density threshold β, construct an emergency event word library, and use the Term Frequency-Inverse Document Frequency (TF-IDF) weighted statistical density. The keyword density threshold β is divided according to the event level. When the keyword density is 0.1, it is defined as a first-level event, and when the keyword density is 0.05, it is defined as a second-level event.

[0050] For the above cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception, in step S6, the process of real-time warning is as follows: First, calculate the sentiment intensity, and judge whether the sentiment intensity reaches the sentiment intensity threshold α. If it does not reach, no warning is issued; if it reaches, continue to calculate the keyword density, and then judge whether the keyword density reaches the keyword density threshold β. If it does not reach, a first-level warning is issued and automatically pushed to the mobile terminal of the on-duty personnel; if it reaches, a second-level warning is issued, linking the emergency management system to trigger SMS notifications and the start of the public relations plan.

[0051] For the above cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception, in step S7, use the loss function α-balanced Focal Loss and a dynamic learning rate to update the parameters of the model to optimize the performance of the model; and add spatio-temporal correlation analysis to integrate time and space information;

[0052] Loss function design: Use the α-balanced Focal Loss function, where the class weight factor is z, and z = [0.2, 0.3, 0.5] corresponds to positive, neutral, and negative classes, and set the focusing parameter γ to focus on difficult-to-classify samples;

[0053] The dynamic learning rate applies the OneCycle learning rate strategy, and the initial learning rate is 5e -4 , and it warms up to 1e after 20 cycles -3Post-exponential decay;

[0054] Spatio-temporal correlation analysis includes geographical clustering and time series prediction;

[0055] Geographical clustering: Combining IP addresses with user registration information, using the DBSCAN algorithm to identify comment aggregation areas to mark potential risk hotspots;

[0056] Time series prediction: Based on the Prophet model to fit the comment growth curve, if the actual value exceeds the 95th percentile of the prediction interval, the enhanced monitoring mode is triggered.

[0057] The beneficial effects of the present invention are as follows:

[0058] 1. The present invention combines dynamic crawling with multi-modal fusion, breaks through the limitation of a single data source, and for the first time realizes the joint analysis of short video bullet screens and text and picture comments, expanding the dimension of public opinion monitoring.

[0059] 2. The present invention first extracts fine-grained features by independent attention heads, and then dynamically fuses them through a gating mechanism, with the F1-score of the index improved by 3.8% compared with the traditional multi-head attention.

[0060] 3. The present invention provides visualization of feature contribution degrees, marks the keywords triggering early warnings and their emotional intensities, assists in manual review decisions, and the measured efficiency is improved by 60%.

[0061] 4. At the data level, the dynamic content capture rate of the present invention is ≥95%, and the multi-modal data alignment error is ≤0.15; at the model level, the F1-score of sentiment classification is 92.1%, and the inference delay is <200 ms / item; at the application level, the false alarm rate of early warning drops to 8.7%, and the emergency response time is shortened to within 8 minutes. Description of the Drawings

[0062] Figure 1 It is the overall flowchart of the present invention.

[0063] Figure 2 It is the structural diagram of the improved model of the present invention.

[0064] Figure 3 It is the flowchart of the early warning threshold decision of the present invention. Detailed Embodiments

[0065] The following further describes the present invention in conjunction with the drawings and embodiments.

[0066] As Figure 1 shown, a cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception includes the following steps:

[0067] S11, Uniform Resource Locator (URL) Seed Injection: The initial URL list is injected into the Redis queue through the start_urls of the distributed crawler framework Scrapy-Redis; the Bloom Filter is used for URL deduplication;

[0068] S12, Dynamic Task Allocation: Each Selenium Worker node pulls the URLs to be crawled from the Redis queue of the in-memory key-value storage system through the distributed crawler framework;

[0069] S13, Headless Chrome Configuration: Each Worker node starts multiple Chrome instances;

[0070] S14, Page Loading and Rendering: The basic loading loads the page through driver.get and sets the timeout;

[0071] S15, Structured Data Extraction: Use the path language XPath of Extensible Markup Language (XML) or Cascading Style Sheets CSS selectors to locate elements; perform multi-modal data processing. For text data, directly extract and clean it, removing HyperText Markup Language HTML tags and special characters; for emoji data, parse Unicode or image links and map them to semantic tags; perform normalization processing on the scoring data;

[0072] S16, Data Storage: Includes real-time writing and deduplication mechanism;

[0073] Real-time writing: Batch write the data into the sharded cluster of the non-relational database system MongoDB through asynchronous input / output (I / O);

[0074] Deduplication mechanism: Generate a unique key based on the user identifier user_id + timestamp + content hash value content_hash to avoid duplicate storage.

[0075] S2: Dynamically evaluate and self-correct the quality of comment public opinion data, perform integrity verification, noise filtering, and multi-modal alignment, and output a standardized multi-modal tensor data set, which is stored in the distributed database.

[0076] The process of noise filtering is as follows: Apply the feature selection algorithm based on information entropy to automatically eliminate duplicate comments and low-quality information content (such as "666", "passing by").

[0077] The process of multi-modal alignment is as follows: Design a joint embedding space based on contrastive learning, map text, emojis, and scores to the same high-dimensional space, and calculate the cross-modal correlation weights through cosine similarity.

[0078] S3: Use an improved model based on a bidirectional long short-term memory network and an attention mechanism to train the dataset, and deeply mine and understand the semantic information of comment public opinion.

[0079] As Figure 2 shown, the improved model includes a bidirectional long short-term memory network layer, a multi-head attention layer, and an adversarial training layer;

[0080] In the bidirectional LSTM layer, the gated recurrent unit GRU is adopted, and 128 hidden units are set to capture forward and backward temporal dependencies;

[0081] In the multi-head attention layer, 8 independent attention heads are deployed, which respectively focus on sentiment words, degree modifiers, negation words, and entity nouns, and the multi-head outputs are weighted and fused through a learnable matrix;

[0082] In the adversarial training layer, character-level perturbations are injected during the training phase. The character-level perturbations include random replacement, insertion, and deletion, which enhance the in-depth understanding of semantics, and the projected gradient descent PGD attack is used to generate adversarial samples to improve the robustness of the model.

[0083] S4: Use temperature scaling to adjust the model output probability distribution to calibrate the confidence.

[0084] Use the temperature scaling formula to perform temperature scaling on the model's original output logits vector, and control the smoothness of the probability distribution by adjusting the temperature;

[0085]

[0086] Among them, is the temperature parameter scalar; is the class probability after the i-th calibration; is the i-th original output logits vector of the model; is the j-th original output logits vector of the model; K is the number of model outputs, which is equal to the total number of classes;

[0087] The optimization objective is: minimize the negative log-likelihood loss on the calibration set to find the optimal temperature;

[0088] First, reserve 20% of the training data as the calibration set, and then minimize the negative log-likelihood loss:

[0089]

[0090] Among them is the optimal temperature obtained through optimization, is the one-hot encoding indicating whether the m-th sample of the true label belongs to the n-th class, is the predicted probability that the calibrated model assigns the m-th sample to the n-th class, and N is the total number of samples.

[0091] S5: Output the sentiment classification result and confidence score, and store them in the time series database.

[0092] The time series database selected is InfluxDB, which supports high-throughput writing (>100,000 records / second), automatic time partitioning, and downsampling query optimization.

[0093] The data retention policy is as follows: the original data is retained for 30 days, and the downsampled aggregated data (such as the hourly average) is retained for 1 year.

[0094] The storage optimization strategy is as follows: write in batches, trigger a batch write every 1000 records accumulated or every 1 second, reducing the input / output (I / O) overhead.

[0095] The index design is as follows: create an inverted index for the platform and sentiment fields to accelerate multi-dimensional queries; data compression, enable the Snappy algorithm for compression, saving 60% of the storage space.

[0096] S6: Adopt a dual-threshold dynamic hierarchical real-time warning, output real-time warning signals, and realize the dynamic perception of comment public opinion.

[0097] Set the sentiment intensity threshold α and the keyword density threshold β.

[0098] For the sentiment intensity threshold α, dynamically calibrate it based on historical data, and use the Exponentially Weighted Moving Average (EWMA) algorithm to update the sentiment intensity threshold α; for example, in normal scenarios, α = 0.7, and during holidays, α = 0.65, as users express stronger emotions during holidays.

[0099] For the keyword density threshold β, construct an emergency event word library containing more than 2000 keywords such as "fire" and "fraud". Adopt the Term Frequency-Inverse Document Frequency (TF-IDF) weighted statistical density. The keyword density threshold β is divided according to the event level. When the keyword density is 0.1, it is defined as a first-level event, and when the keyword density is 0.05, it is defined as a second-level event.

[0100] As Figure 3 shown, the process of real-time warning is as follows: First, calculate the sentiment intensity, and judge whether the sentiment intensity reaches the sentiment intensity threshold α. If it does not reach, no warning is issued; if it reaches, continue to calculate the keyword density, and then judge whether the keyword density reaches the keyword density threshold β. If it does not reach, a first-level warning is issued and automatically pushed to the mobile terminal of the duty personnel; if it reaches, a second-level warning is issued, linking with the emergency management system to trigger SMS notifications and the start of the public relations plan.

[0101] S7: Update the keyword library and model parameters, generate disposal suggestions and visualization heatmaps in the management background, and dynamically generate comment public opinion decision-making methods. Use the loss function α-balanced Focal Loss and a dynamic learning rate to update the model parameters to optimize the model performance; and add spatio-temporal correlation analysis to integrate time and space information to reduce monitoring deviation.

[0102] Loss function design: Use the α-balanced Focal Loss function, where the class weight factor z = [0.2, 0.3, 0.5] corresponds to positive, neutral, and negative categories, and set the focusing parameter γ to focus on difficult-to-classify samples.

[0103] Dynamic learning rate: Apply the OneCycle learning rate strategy, with an initial learning rate of 5e -4 , which warms up to 1e after 20 epochs -3 and then decays exponentially. Compared with the traditional Adam optimizer, the convergence speed is increased by 36%.

[0104] Spatio-temporal correlation analysis:

[0105] Geographical clustering: Combine the IP address and user registration information, and use the DBSCAN algorithm to identify comment aggregation areas to mark potential risk hotspots.

[0106] Time series prediction: Based on the Prophet model, fit the comment growth curve. If the actual value exceeds the 95% quantile of the prediction interval, trigger the enhanced monitoring mode.

[0107] Example 1: Verification of dynamic crawler performance.

[0108] Test platforms: Weibo, Douyin, Bilibili;

[0109] Results: Compared with the traditional Requests library, the integrity of dynamic content capture is increased from 67% to 93%;

[0110] The process is as follows:

[0111] Step 1: Configure the test environment. The hardware environment is a server: 4-core CPU, 32GB memory, 1Gbps bandwidth;

[0112] The software environment is an operating system: Ubuntu 20.04 LTS, Python version: 3.8.10; the dependent libraries are Scrapy 2.6.0, Selenium 4.1.0, ChromeDriver 95.0.4638.69; browser simulation: Headless Chrome (headless mode).

[0113] Step 2: Selection of the testing platform: The target testing platforms are Weibo, Douyin, and Bilibili.

[0114] Step 3: Determination of test cases: Select 10 popular topic pages on Weibo; select 50 popular videos on Douyin; select 20 long-form graphic columns on Bilibili.

[0115] Step 4: Data capture. After the crawler is started, the crawling tasks are allocated according to the platform priority; then for each page, first obtain the basic HTML through Scrapy, and then start Selenium to render the dynamic content; finally, extract the structured fields.

[0116] Step 5: Setting of integrity benchmarks. It includes manual sampling verification and API comparison verification. Manual sampling verification is to randomly select 100 comments from each platform and manually visit the page to check the data integrity; for platforms that support open APIs, API comparison verification is to call the official API to obtain the benchmark data and compare the results captured by the crawler.

[0117] Step 6: Design of a comparative experiment. The hybrid crawler and the control group run simultaneously to crawl the same target pages. Record the completion time, data volume, and exception logs of each batch of tasks. After data cleaning, calculate the capture completeness respectively.

[0118] Step 7: Record and analyze the test results.

[0119] Step 8: Draw key conclusions. The hybrid crawler significantly improves the data capture completeness through dynamic rendering and anti-crawling strategies.

[0120] Example 2: Model comparative experiment.

[0121] Dataset: Self-built Chinese comment corpus (100,000 pieces, including 5 types of sentiment labels);

[0122] Comparative models: TextCNN, standard LSTM, BERT-base;

[0123] Result: The F1-score of the model of the present invention reaches 92.1%, and the inference speed is 4.3 times faster than that of BERT.

Claims

1. A cross-modal review public opinion decision-making method based on deep semantic understanding and dynamic perception, characterized in that It includes the following steps: S1: Heterogeneous data collection, dynamically crawling multi-platform review and public opinion data based on a hybrid crawler architecture; S2: Dynamically evaluate and self-correct the quality of review and public opinion data, perform integrity verification, noise filtering, and multi-modal alignment, output a standardized multi-modal tensor data set, and store it in a distributed database; S3: Use an improved model based on a bidirectional long short-term memory network and an attention mechanism to perform deep learning training on the data set, and deeply mine and understand the semantic information of review and public opinion; S4: Adjust the model output probability distribution using temperature scaling to calibrate the confidence; S5: Output the sentiment classification result and confidence score, and store it in a time series database; S6: Adopt a dual-threshold dynamic hierarchical real-time warning, output a real-time warning signal, and realize the dynamic perception of review and public opinion; S7: Update the keyword library and model parameters, and generate disposal suggestions and a visualization heat map in the management background, and dynamically generate a decision-making method for review and public opinion.

2. The cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception according to claim 1, wherein The specific process of step S1 is as follows: S11, Uniform Resource Locator URL seed injection: The initial URL list is injected into the Redis database queue through the start_urls field of the distributed crawler framework Scrapy-Redis; Use a Bloom Filter to remove duplicate URLs; S12, Dynamic task allocation: Each Selenium Worker node pulls the URL to be crawled from the Redis queue of the in-memory key-value storage system through the distributed crawler framework; S13, Headless Chrome configuration: Each Worker node starts multiple Chrome instances; S14, Page loading and rendering: The basic loading loads the page through driver.get and sets the timeout; S15, Structured data extraction: Use the path language XPath of Extensible Markup Language XML or Cascading Style Sheets CSS selector to locate elements; Perform multi-modal data processing, directly extract and clean text data, remove HyperText Markup Language HTML tags and special characters; For emoji data, parse Unicode or image links and map them to semantic tags; Normalize the rating data; S16, Data storage: Includes real-time writing and deduplication mechanisms; Real-time writing: Batch write data into the non-relational database system MongoDB sharding cluster through asynchronous input / output I / O; Deduplication mechanism: Generate a unique key based on the user identifier user_id + timestamp + the hash value of the content content_hash to avoid duplicate storage.

3. The cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception according to claim 1, wherein In step S2, the process of noise filtering is: Apply a feature selection algorithm based on information entropy to automatically eliminate duplicate comments and content with poor information volume.

4. The cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception according to claim 1, characterized in that, In step S2, the process of multi-modal alignment is: Design a joint embedding space based on contrast learning, map text, emojis, and ratings to the same high-dimensional space, and calculate the cross-modal correlation weight through cosine similarity.

5. The cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception according to claim 1, characterized in that, In step S3, the improved model includes a bidirectional long short-term memory network layer, a multi-head attention layer, and an adversarial training layer; In the bidirectional long short-term memory network layer, the gated recurrent unit (GRU) is adopted, with 128 hidden units set to capture forward and backward temporal dependencies. In the multi-head attention layer, 8 independent attention heads are deployed, which respectively focus on sentiment words, degree modifiers, negation words, and entity nouns, and the outputs of multiple heads are weighted and fused through a learnable matrix. In the adversarial training layer, character-level perturbations are injected during the training phase. The character-level perturbations include random replacement, insertion, and deletion to enhance the in-depth understanding of semantics, and the projected gradient descent (PGD) attack is used to generate adversarial samples to improve the robustness of the model.

6. The cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception according to claim 1, characterized in that In step S4, the temperature scaling formula is used to perform temperature scaling on the original output logits vector of the model, and by adjusting the temperature, the smoothness of the probability distribution is controlled. ; Among them, is the temperature parameter scalar; is the class probability after the i-th calibration; is the i-th original output logits vector of the model; is the j-th original output logits vector of the model; K is the number of times the model outputs, equal to the total number of classes; The optimization objective is to minimize the negative log-likelihood loss on the calibration set and find the optimal temperature. First, 20% of the training data is reserved as the calibration set, and then the negative log-likelihood loss is minimized. ; where is the optimal temperature obtained by optimization, is the one - hot encoding indicating whether the m - th sample of the true label belongs to the n - th class, is the predicted probability that the calibrated model assigns the m - th sample to the n - th class, and N is the total number of samples.

7. The cross-modal review public opinion decision-making method based on deep semantic understanding and dynamic perception according to claim 6, characterized in that In step S5, the time series database selected is InfluxDB; the data retention policy is as follows: the original data is retained for 30 days, and the downsampled and aggregated data is retained for 1 year; the storage optimization strategy is: batch writing, and batch writing is triggered every time 1000 records are accumulated or every 1 second; the index design is: an inverted index is established for the platform and emotion fields; data compression enables the Snappy algorithm for compression.

8. The cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception according to claim 1, characterized in that In step S6, the sentiment intensity threshold α and the keyword density threshold β are set. Regarding the sentiment intensity threshold α, it is dynamically calibrated based on historical data, and the exponential weighted moving average (EWMA) algorithm is used to update the sentiment intensity threshold α. Regarding the keyword density threshold β, an emergency event word library is constructed, and the term frequency-inverse document frequency (TF-IDF) is used to weightedly statistically calculate the density. The keyword density threshold β is divided according to the event level. When the keyword density is 0.1, it is defined as a first-level event, and when the keyword density is 0.05, it is defined as a second-level event.

9. The cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception according to claim 8, characterized in that, In step S6, the process of real-time warning is as follows: First, the sentiment intensity is calculated to determine whether the sentiment intensity reaches the sentiment intensity threshold α. If not, no warning is issued; if it reaches, the keyword density is calculated continuously, and then it is determined whether the keyword density reaches the keyword density threshold β. If not, a first-level warning is issued and automatically pushed to the mobile device of the on-duty personnel; if it reaches, a second-level warning is issued, linking with the emergency management system to trigger the SMS notification and the start of the public relations plan.

10. The cross-modal comment public opinion decision-making method based on deep semantic understanding and dynamic perception according to claim 8, wherein In step S7, the loss function α-balanced Focal Loss and a dynamic learning rate are adopted to update the parameters of the model to optimize the performance of the model; and spatio-temporal correlation analysis is added to integrate time and space information. Loss function design: The α-balanced Focal Loss function is adopted, where the class weight factor is z, and z = [0.2, 0.3, 0.5] corresponds to positive, neutral, and negative classes, and the focusing parameter γ is set to focus on difficult-to-classify samples. The dynamic learning rate applies the OneCycle learning rate strategy, with an initial learning rate of 5e -4 , which warms up to 1e after 20 epochs -3 and then decays exponentially; Spatio-temporal correlation analysis includes geographical clustering and time series prediction. Geographical clustering: Combine IP addresses with user registration information, and use the DBSCAN algorithm to identify comment aggregation areas to mark potential risk hotspots; Time series prediction: Based on the Prophet model to fit the comment growth curve, if the actual value exceeds the 95th percentile of the prediction interval, the enhanced monitoring mode will be triggered.

Citation Information

Patent Citations

  • Scenic spot public opinion monitoring and analyzing system and method

    CN111461553A

  • Entity relation joint extraction method

    CN117875416A

  • Public opinion prediction system, method and equipment based on large model and storage medium

    CN118861396A

  • Model training method and device, equipment and medium

    CN118861693A

  • An intelligent prediction system for social public opinion risk

    CN119760361A