Network public opinion risk identification and monitoring early warning method based on intelligent crawler technology

By building an intelligent crawler model and a comprehensive identification and early warning model, and automatically crawling and analyzing Internet public opinion information, the problem of inefficient public opinion management in the existing technology is solved, and more accurate and efficient public opinion monitoring and management is achieved.

CN120030212APending Publication Date: 2025-05-23SUZHOU LIANRUIKE ELECTRONIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510105829.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In the prior art, the public opinion management method has the problem of inefficient screening, which is difficult to effectively deal with the massive information processing needs on the Internet. It is also susceptible to subjective judgments, resulting in inaccuracy and inconsistency of the analysis results.

Method used

The network public opinion risk identification and monitoring and early warning method is adopted based on intelligent crawling technology. By building an intelligent crawler model and a comprehensive identification and early warning model, the Internet public opinion information is automatically captured and analyzed, and the warning information is generated and pushed to relevant personnel. At the same time, processing suggestions are automatically generated in combination with historical case libraries, expert systems or machine learning algorithms.

Benefits of technology

It improves the efficiency of capturing and analyzing public opinion information, avoids the influence of subjective judgment, ensures the accuracy and consistency of analysis results, shortens the response time of public opinion management, and improves management efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030212A_ABST
    Figure CN120030212A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a network public opinion risk identification and monitoring early warning method based on an intelligent crawler technology, and the method comprises the steps: firstly constructing an intelligent crawler model and a comprehensive identification early warning model; based on an intelligent crawler model, information related to public opinions is obtained based on the intelligent crawler model, and the captured information is encrypted and transmitted to a preprocessing module for preprocessing; and analyzing and identifying the preprocessed data on the basis of a comprehensive identification early warning model, automatically generating targeted processing suggestions on the basis of early warning information in combination with a historical case library, an expert system or a machine learning algorithm, and pushing the processing suggestions to related personnel, thereby solving the problem of low screening efficiency of a public opinion management mode in the prior art. The method solves the technical problems that mass information processing requirements on the Internet are difficult to effectively meet, and analysis results are inaccurate and inconsistent due to the fact that the method is easily influenced by subjective judgment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a network public opinion risk identification and monitoring and early warning method based on intelligent crawler technology. Background Art

[0002] With the rapid development of the Internet and information technology, online public opinion has gradually become a key component of social public opinion. The rapid spread and wide impact of online public opinion have made governments, enterprises and other organizations pay more and more attention to public opinion management. Effective public opinion management can help these organizations quickly capture the dynamics of public opinion, predict possible risks, and make timely and accurate decisions.

[0003] However, the traditional public opinion management method mainly relies on manual search, screening and analysis, which has many shortcomings. First, manual search and screening are inefficient and difficult to cope with the processing needs of massive amounts of information on the Internet; second, manual analysis is easily affected by factors such as personal experience and subjective judgment, resulting in inaccurate and inconsistent analysis results, thus affecting the scientific nature of decision-making.

[0004] In summary, the existing public opinion management methods have the problem of low screening efficiency, which makes it difficult to effectively cope with the massive information processing needs on the Internet. At the same time, they are easily affected by subjective judgment, resulting in inaccurate and inconsistent analysis results. Summary of the invention

[0005] The purpose of the present invention is to provide a method for identifying and monitoring network public opinion risks based on intelligent crawler technology, aiming to solve the problem of low screening efficiency in the public opinion management method in the prior art, which is difficult to effectively cope with the massive information processing needs on the Internet, and is easily affected by subjective judgment, resulting in inaccurate and inconsistent analysis results.

[0006] To achieve the above purpose, the present invention adopts a network public opinion risk identification and monitoring and early warning method based on intelligent crawler technology, which includes the following steps:

[0007] Step 1: First, build an intelligent crawler model and a comprehensive identification and early warning model;

[0008] Step 2: Based on the intelligent crawler model, start the crawler program, widely crawl information related to public opinion on the Internet according to preset keywords, website lists or access paths generated by algorithms, and encrypt and transmit the crawled information to the preprocessing module for preprocessing;

[0009] Step 3: Based on the comprehensive recognition and early warning model, analyze and identify the pre-processed data, and generate early warning information based on the analysis and identification results and push it to relevant personnel;

[0010] Step 4: Based on the warning information, combined with the historical case library, expert system or machine learning algorithm, targeted processing suggestions are automatically generated and pushed to relevant personnel.

[0011] Among them, the specific method of building an intelligent crawler model is as follows:

[0012] First, a module is designed to collect information on web page structure, content format, and update frequency on the Internet to form an initial training dataset;

[0013] Extract the URL features, content features, and link features of web pages from the initial training data set, and select the features that have a greater impact on crawler efficiency and accuracy as training features;

[0014] The training features are trained using machine learning algorithms (such as random forest, support vector machine or neural network) to form an intelligent crawler model.

[0015] After the intelligent crawler model is built, the model is tuned according to the test data set to improve the crawler's crawling efficiency and accuracy.

[0016] Among them, the specific methods of constructing a comprehensive identification and early warning model are as follows:

[0017] First, feature vectors are extracted from historical data;

[0018] Choose a machine learning algorithm (such as SVM, random forest, neural network, etc.) or a deep learning model (such as LSTM, BERT);

[0019] The extracted feature vectors are input into the model for training to learn the feature representation and classification rules of public opinion data;

[0020] Use the test data set to evaluate the model and optimize the model based on the evaluation results, such as adjusting model parameters and adding features;

[0021] According to the needs of public opinion monitoring, set warning rules (such as sensitive word triggers, sentiment tendency thresholds, topic clustering results, etc.);

[0022] Integrate warning rules into the model to automatically trigger warnings during real-time monitoring.

[0023] Among them, after the comprehensive identification and early warning model is constructed, the model is optimized through cross-validation and grid search technology to improve its recognition accuracy and generalization ability.

[0024] Among them, the method of encrypting and transmitting data is:

[0025] First, the captured information is encapsulated into a specific data packet format to facilitate subsequent encrypted transmission;

[0026] Before data packets are transmitted, symmetric encryption algorithms, asymmetric encryption algorithms or hash algorithms are used to encrypt data packets;

[0027] Use secure transmission protocols (such as HTTPS, SFTP, etc.) to transmit the encrypted data packets to the pre-processing module;

[0028] After receiving the encrypted data packet, the preprocessing module uses the corresponding decryption algorithm to decrypt the data packet.

[0029] Among them, the preprocessing module preprocesses the data in the following way:

[0030] Remove noise, redundancy, and outliers from data;

[0031] Convert data into a format suitable for subsequent analysis and recognition, such as converting text data into structured data;

[0032] The data is scaled or normalized to eliminate differences in dimensions and numerical ranges between different features.

[0033] Among them, the specific methods of automatically generating targeted processing suggestions based on early warning information, combined with historical case libraries, expert systems or machine learning algorithms are as follows:

[0034] First, receive the warning information from the comprehensive recognition warning model and extract the key elements;

[0035] Using similarity calculation and cluster analysis algorithms, the current warning information is matched with cases in the historical case library to find the most similar or relevant cases;

[0036] Based on the matched cases, the handling measures are extracted, and adjusted and modified in combination with the specific circumstances of the current warning information to form handling suggestions.

[0037] When pushing information to relevant personnel, SMS, email or APP is used. During the pushing process, the accuracy and timeliness of the information should be ensured.

[0038] Among them, in the process of risk identification and monitoring and early warning, log information of key operations (such as data capture time, preprocessing results, warning information generation time, etc.) is recorded to facilitate subsequent problem tracking and performance analysis.

[0039] The present invention discloses a network public opinion risk identification and monitoring and early warning method based on intelligent crawler technology. First, an intelligent crawler model and a comprehensive identification and early warning model are constructed. Based on the intelligent crawler model, a crawler program is started to widely crawl information related to public opinion on the Internet according to preset keywords, website lists or access paths generated by algorithms, and the crawled information is encrypted and transmitted to a preprocessing module for preprocessing. Based on the comprehensive identification and early warning model, the preprocessed data is analyzed and identified, and early warning information is generated according to the analysis and identification results and pushed to relevant personnel. Based on the early warning information, in combination with a historical case library, an expert system or a machine learning algorithm, targeted processing suggestions are automatically generated and pushed to relevant personnel. In this way, the problem of low screening efficiency in the public opinion management method in the prior art is solved, and it is difficult to effectively cope with the massive information processing needs on the Internet. At the same time, it is easily affected by subjective judgment, resulting in inaccurate and inconsistent analysis results. Technical problems.

[0040] Through the application of intelligent crawlers and comprehensive identification and early warning models, the system can more accurately capture and analyze online public opinion information, avoiding the limitations of keyword search-based systems in dealing with complex and changing information.

[0041] At the same time, it can automatically generate early warning information and handling suggestions, and quickly push them to relevant personnel through effective channels, thereby greatly shortening the response time of public opinion management and improving management efficiency.

[0042] The early warning information and treatment suggestions provided are based on in-depth analysis and professional knowledge, which can provide decision makers with more scientific and accurate basis and help them make more effective decisions.

[0043] Both intelligent crawlers and comprehensive identification and early warning models are highly adaptable and can be adjusted and optimized as the network environment and public opinion information changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0045] Figure 1 It is a flow chart of the steps of the network public opinion risk identification and monitoring and early warning method based on intelligent crawler technology of the present invention. DETAILED DESCRIPTION

[0046] Embodiments of the present invention are described in detail below. Examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, but should not be construed as limiting the present invention.

[0047] See also Figure 1 , Figure 1 It is a flow chart of the steps of the network public opinion risk identification and monitoring and early warning method based on intelligent crawler technology of the present invention.

[0048] The present invention provides a network public opinion risk identification and monitoring and early warning method based on intelligent crawler technology, comprising the following steps:

[0049] S1. First, build an intelligent crawler model and a comprehensive identification and early warning model;

[0050] For this specific implementation, the specific method of constructing the intelligent crawler model is as follows:

[0051] First, a module is designed to collect information on web page structure, content format, and update frequency on the Internet to form an initial training dataset;

[0052] Extract the URL features, content features, and link features of web pages from the initial training data set, and select the features that have a greater impact on crawler efficiency and accuracy as training features;

[0053] The training features are trained using machine learning algorithms (such as random forest, support vector machine or neural network) to form an intelligent crawler model.

[0054] After the intelligent crawler model is built, the model is tuned according to the test data set to improve the crawler's crawling efficiency and accuracy.

[0055] The specific methods of constructing a comprehensive identification and early warning model are as follows:

[0056] First, feature vectors are extracted from historical data;

[0057] Choose a machine learning algorithm (such as SVM, random forest, neural network, etc.) or a deep learning model (such as LSTM, BERT);

[0058] The extracted feature vectors are input into the model for training to learn the feature representation and classification rules of public opinion data;

[0059] Use the test data set to evaluate the model and optimize the model based on the evaluation results, such as adjusting model parameters and adding features;

[0060] According to the needs of public opinion monitoring, set warning rules (such as sensitive word triggers, sentiment tendency thresholds, topic clustering results, etc.);

[0061] Integrate warning rules into the model to automatically trigger warnings during real-time monitoring.

[0062] After the comprehensive identification and early warning model is constructed, the model is optimized through cross-validation and grid search technology to improve its recognition accuracy and generalization ability.

[0063] S2. Based on the intelligent crawler model, the crawler program is started to widely crawl information related to public opinion on the Internet according to preset keywords, website lists or access paths generated by algorithms, and the crawled information is encrypted and transmitted to the preprocessing module for preprocessing;

[0064] For this specific implementation, the method of encrypting and transmitting data is as follows:

[0065] First, the captured information is encapsulated into a specific data packet format to facilitate subsequent encrypted transmission;

[0066] Before data packets are transmitted, symmetric encryption algorithms, asymmetric encryption algorithms or hash algorithms are used to encrypt data packets;

[0067] The encryption algorithm is as follows:

[0068] Symmetric encryption algorithm:

[0069] Such as AES (Advanced Encryption Standard), which has the advantages of fast encryption and decryption speed and relatively simple key management.

[0070] In encrypted information transmission, the sender and receiver use the same key for encryption and decryption.

[0071] Asymmetric encryption algorithm:

[0072] For example, RSA (Rivest-Shamir-Adleman algorithm) provides higher security based on the mathematical problem of factorizing large numbers.

[0073] In encrypted information transmission, the sender uses the receiver's public key to encrypt, and the receiver uses his or her own private key to decrypt. This method is often used in scenarios such as key exchange and digital signatures.

[0074] Hash Algorithm:

[0075] Such as MD5, SHA-256, etc., which are used to generate data summaries or hash values ​​and have the characteristics of irreversibility and anti-collision.

[0076] In information encryption transmission, hash algorithms can be used to verify the integrity and authenticity of data. For example, the sender hashes the data before encryption and transmits the hash value together with the encrypted data; the receiver recalculates the hash value of the data after decryption and compares it with the received hash value to verify whether the data has been tampered with.

[0077] Use secure transmission protocols (such as HTTPS, SFTP, etc.) to transmit the encrypted data packets to the pre-processing module;

[0078] Among them, HTTPS is an encrypted HTTP protocol based on SSL / TLS protocol, which is widely used in Web communication and data transmission; it provides security features such as confidentiality, integrity and authentication of data during transmission.

[0079] SFTP is a file transfer protocol based on the SSH protocol. It allows users to securely transfer files through an encrypted channel and provides confidentiality and integrity protection for data during transmission.

[0080] After receiving the encrypted data packet, the preprocessing module uses the corresponding decryption algorithm to decrypt the data packet.

[0081] Among them, the preprocessing module preprocesses the data in the following way:

[0082] Remove noise, redundancy, and outliers from data;

[0083] Convert data into a format suitable for subsequent analysis and recognition, such as converting text data into structured data;

[0084] The data is scaled or normalized to eliminate differences in dimensions and numerical ranges between different features.

[0085] S3. Based on the comprehensive recognition and early warning model, the pre-processed data is analyzed and identified, and early warning information is generated and pushed to relevant personnel based on the analysis and identification results;

[0086] S4. Based on the early warning information, combined with the historical case library, expert system or machine learning algorithm, targeted processing suggestions are automatically generated and pushed to relevant personnel.

[0087] For this specific implementation, firstly, the warning information transmitted by the comprehensive identification warning model is received, and the key elements are extracted;

[0088] Using similarity calculation and cluster analysis algorithms, the current warning information is matched with cases in the historical case library to find the most similar or relevant cases;

[0089] Based on the matched cases, the handling measures are extracted, and adjusted and modified in combination with the specific circumstances of the current warning information to form handling suggestions.

[0090] When pushing information to relevant personnel, use SMS, email or APP. During the push process, the accuracy and timeliness of the information should be ensured.

[0091] During the risk identification and monitoring and early warning process, log information of key operations (such as data capture time, preprocessing results, warning information generation time, etc.) is recorded to facilitate subsequent problem tracking and performance analysis.

[0092] A network public opinion risk identification and monitoring and early warning method based on intelligent crawler technology of the present embodiment is used. First, an intelligent crawler model and a comprehensive identification and early warning model are constructed. Based on the intelligent crawler model, the crawler program is started to widely crawl public opinion-related information on the Internet according to preset keywords, website lists or access paths generated by algorithms, and the crawled information is encrypted and transmitted to a preprocessing module for preprocessing. Based on the comprehensive identification and early warning model, the preprocessed data is analyzed and identified, and early warning information is generated based on the analysis and identification results and pushed to relevant personnel. Based on the early warning information, in combination with a historical case library, an expert system or a machine learning algorithm, targeted processing suggestions are automatically generated and pushed to relevant personnel. In this way, the problem of low screening efficiency in the public opinion management method in the prior art is solved, and it is difficult to effectively cope with the massive information processing needs on the Internet. At the same time, it is easily affected by subjective judgment, resulting in inaccurate and inconsistent analysis results. Technical problems.

[0093] Through the application of intelligent crawlers and comprehensive identification and early warning models, the system can more accurately capture and analyze online public opinion information, avoiding the limitations of keyword search-based systems in dealing with complex and changing information.

[0094] At the same time, it can automatically generate early warning information and handling suggestions, and quickly push them to relevant personnel through effective channels, thereby greatly shortening the response time of public opinion management and improving management efficiency.

[0095] The early warning information and treatment suggestions provided are based on in-depth analysis and professional knowledge, which can provide decision makers with more scientific and accurate basis and help them make more effective decisions.

[0096] Both intelligent crawlers and comprehensive identification and early warning models are highly adaptable and can be adjusted and optimized as the network environment and public opinion information changes.

[0097] The present invention also has the following beneficial effects:

[0098] 1. Improve information capture efficiency:

[0099] The intelligent crawler model can automatically and efficiently capture information related to public opinion on the Internet, greatly reducing the time and cost of manual information collection.

[0100] By optimizing crawler algorithms and performance, the speed and accuracy of information capture can be further improved, ensuring the real-time and accuracy of public opinion information.

[0101] 2. Enhance information processing capabilities:

[0102] The intelligent crawler model not only has powerful information crawling capabilities, but also has data processing and analysis capabilities.

[0103] The captured public opinion information can be preprocessed, classified, summarized and other operations can be performed to provide strong support for subsequent analysis and decision-making.

[0104] 3. Improve the intelligence level of public opinion monitoring:

[0105] By introducing advanced technologies such as natural language processing and machine learning, the intelligent crawler model can understand and analyze public opinion information more deeply.

[0106] It can automatically identify key features of public opinion information such as emotional tendencies and dissemination trends, providing more comprehensive and in-depth insights for public opinion monitoring.

[0107] 4. Promote scientific and standardized public opinion management:

[0108] The public opinion information obtained through the intelligent crawler model can provide a scientific basis for public opinion management.

[0109] Based on this information, more scientific and reasonable public opinion management strategies can be formulated to improve the efficiency and effectiveness of public opinion management.

[0110] 5. Ensure information security and privacy protection:

[0111] When building an intelligent crawler model, you can focus on the design of information security and privacy protection.

[0112] Through encrypted transmission technology, the security of captured public opinion information during transmission and storage is ensured to avoid information leakage and abuse.

[0113] What is disclosed above is only a preferred embodiment of the present invention, and it certainly cannot be used to limit the scope of rights of the present invention. Ordinary technicians in this field can understand that all or part of the processes of the above embodiment and equivalent changes made according to the claims of the present invention still fall within the scope of the invention.

Claims

1. A network public opinion risk identification and monitoring and early warning method based on intelligent crawler technology, characterized in that: The steps include: First, build an intelligent crawler model and a comprehensive identification and early warning model; Based on the intelligent crawler model, the crawler program is started to widely crawl information related to public opinion on the Internet according to preset keywords, website lists or access paths generated by algorithms, and the crawled information is encrypted and transmitted to the preprocessing module for preprocessing; Based on the comprehensive recognition and early warning model, the pre-processed data is analyzed and identified, and early warning information is generated and pushed to relevant personnel based on the analysis and identification results; Based on the early warning information, combined with the historical case library, expert system or machine learning algorithm, targeted processing suggestions are automatically generated and pushed to relevant personnel.

2. The network public opinion risk identification and monitoring and early warning method based on intelligent crawler technology according to claim 1 is characterized in that: The specific method of building an intelligent crawler model is as follows: First, a module is designed to collect information on web page structure, content format, and update frequency on the Internet to form an initial training dataset; Extract the URL features, content features, and link features of web pages from the initial training data set, and select the features that have a greater impact on crawler efficiency and accuracy as training features; The training features are trained using machine learning algorithms to form an intelligent crawler model.

3. The network public opinion risk identification and monitoring and early warning method based on intelligent crawler technology as claimed in claim 2 is characterized in that: After the intelligent crawler model is built, the model is tuned according to the test data set.

4. The network public opinion risk identification and monitoring and early warning method based on intelligent crawler technology as claimed in claim 3 is characterized in that: The specific methods of constructing a comprehensive identification and early warning model are as follows: First, feature vectors are extracted from historical data; Choose a machine learning algorithm or deep learning model; The extracted feature vectors are input into the model for training to learn the feature representation and classification rules of public opinion data; Use the test data set to evaluate the model and optimize the model based on the evaluation results; Set early warning rules according to the needs of public opinion monitoring; Integrate warning rules into the model to automatically trigger warnings during real-time monitoring.

5. The network public opinion risk identification and monitoring and early warning method based on intelligent crawler technology as claimed in claim 4 is characterized in that: After the comprehensive identification and early warning model is constructed, the model is optimized through cross-validation and grid search technology.

6. The network public opinion risk identification and monitoring and early warning method based on intelligent crawler technology as claimed in claim 5 is characterized in that: Methods for encrypting data transmission: First, the captured information is encapsulated into a specific data packet format to facilitate subsequent encrypted transmission; Before data packets are transmitted, symmetric encryption algorithms, asymmetric encryption algorithms or hash algorithms are used to encrypt data packets; Using a secure transmission protocol to transmit the encrypted data packets to the pre-processing module; After receiving the encrypted data packet, the preprocessing module uses the corresponding decryption algorithm to decrypt the data packet.

7. The network public opinion risk identification and monitoring and early warning method based on intelligent crawler technology according to claim 6 is characterized in that: The preprocessing module preprocesses the data in the following way: Remove noise, redundancy, and outliers from data; Convert data into a format suitable for subsequent analysis and recognition, such as converting text data into structured data; The data is scaled or normalized to eliminate differences in dimensions and numerical ranges between different features.

8. The network public opinion risk identification and monitoring and early warning method based on intelligent crawler technology as claimed in claim 7 is characterized in that: Based on the early warning information, combined with the historical case library, expert system or machine learning algorithm, the specific way to automatically generate targeted processing suggestions is as follows: First, receive the warning information from the comprehensive recognition warning model and extract the key elements; Using similarity calculation and cluster analysis algorithms, the current warning information is matched with cases in the historical case library to find the most similar or relevant cases; Based on the matched cases, the handling measures are extracted, and adjusted and modified in combination with the specific circumstances of the current warning information to form handling suggestions.

9. The network public opinion risk identification and monitoring and early warning method based on intelligent crawler technology as claimed in claim 8, characterized in that: When pushing information to relevant personnel, use SMS, email or APP. During the push process, the accuracy and timeliness of the information should be ensured.

10. The network public opinion risk identification and monitoring and early warning method based on intelligent crawler technology according to claim 9, characterized in that: During the risk identification and monitoring and early warning process, log information of key operations is recorded.