Tobacco public opinion monitoring method
Through a multi-dimensional encoder and deep learning model, combined with an external knowledge base and prompt learner, a tobacco public opinion monitoring model is built, which solves the misjudgment problem of complex contexts and malicious behavior recognition in the existing technology, and achieves a more efficient and reliable tobacco public opinion analysis.
Patent Information
- Application Number
- CN202510390474.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-22
AI Technical Summary
The existing tobacco public opinion monitoring technology has problems of misjudgment and inefficiency in dealing with complex contexts and malicious behavior recognition, and lacks a systematic evaluation mechanism, making it difficult to quantify the effectiveness of the strategy.
A multi-dimensional encoder and deep learning model are adopted, combined with an external knowledge base and prompt learner, tobacco public opinion monitoring model is constructed through sentiment analysis, topic classification, decision-making disposal and false public opinion recognition, and the model performance is optimized using the self-attention layer and loss function.
It significantly improves the accuracy and reliability of tobacco public opinion analysis, can better capture complex contexts and identify malicious behaviors, and provide scientific decision-making support.
Smart Images

Figure CN120354199A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of public opinion monitoring, and more specifically, to a method for monitoring tobacco public opinion. Background Art
[0002] In the tobacco industry, public opinion monitoring is particularly important because tobacco-related public opinion not only involves consumers' health awareness and social opinions, but also is closely related to national policies, industry supervision, and market competition.
[0003] With the rapid development of artificial intelligence technology, especially the rise of large-scale pre-trained models (such as GPT, BERT, etc.), public opinion analysis technology has entered a new stage of development.
[0004] However, traditional large models mainly rely on keyword matching and simple machine learning algorithms, and have limited ability to understand complex contexts. For example, when dealing with ironic, metaphorical, or context-related emotional expressions, these models are prone to misjudgment and cannot accurately capture the true attitudes of the public. This results in a possible deviation between the public opinion analysis results and the actual situation, affecting the effectiveness of decision-making.
[0005] Moreover, some malicious behaviors are often highly concealed and usually disguise themselves by imitating the activity patterns of real users, including but not limited to posting frequency, content style, interaction behaviors, etc., making it difficult for traditional detection methods based on rules or simple machine learning models to identify their abnormal features. Especially when malicious users can dynamically adjust their strategies to adapt to changes in the detection mechanism, this disguise becomes even more difficult to detect.
[0006] In addition, existing public opinion handling strategies often rely on personal experience and intuition, which are not only inefficient, but also easily affected by subjective factors. And due to the lack of a systematic evaluation mechanism, it is difficult to quantify the effects of different strategies, and thus optimize future response plans.
[0007] In summary, although the existing technology has achieved certain achievements in some aspects, there are still obvious deficiencies in terms of accuracy, ability to handle complex situations, and scientific decision-making support. Summary of the Invention
[0008] In view of this, the present invention provides a method for monitoring tobacco public opinion, which effectively makes up for these deficiencies by introducing advanced deep learning models, multi-dimensional verification systems, and intelligent public opinion handling strategies, and provides a more intelligent, efficient, and reliable network public opinion monitoring solution for tobacco enterprises.
[0009] To achieve the above object, the present invention adopts the following technical solutions:
[0010] A method for monitoring tobacco public opinion, comprising
[0011] Obtain tobacco public opinion monitoring data and conduct tobacco public opinion monitoring based on the tobacco public opinion monitoring model;
[0012] The tobacco public opinion monitoring model includes a multi-dimensional encoder and a deep learning model,
[0013] The multi-dimensional encoder includes a sentiment analysis encoder, a topic classification encoder, a decision-making and handling encoder, an alarm level encoder, and a false public opinion identification encoder, which are used to encode the tobacco public opinion monitoring data for different tasks;
[0014] The deep learning model is used to conduct tobacco public opinion monitoring according to the encoding vectors of each encoder.
[0015] Preferably, a prompt learner is used to generate a task prompt embedding vector according to the user task prompt text, and the task prompt embedding vector is respectively added to the encoding sequences of different tasks.
[0016] Preferably, an external knowledge base is constructed, relevant content in the external knowledge base is obtained through keywords in the public opinion monitoring data, a knowledge embedding vector is generated by the prompt learner according to the relevant content, and the knowledge embedding vector is respectively added to the sentiment analysis encoding sequence, the subject classification encoding sequence, and the decision-making and handling encoding sequence.
[0017] Preferably, the external knowledge base is a knowledge graph constructed based on tobacco-related materials, Elasticsearch is used for full-text search according to the keywords of the public opinion monitoring data, and the Jaccard coefficient is used to measure the similarity between the input keywords and the retrieved relevant content.
[0018] Tobacco-related materials include tobacco-related regulatory documents, tobacco industry research reports / latest reports, and tobacco-related scientific research papers / technical reports.
[0019] Preferably, the sentiment analysis encoding sequence, the subject classification encoding sequence, and the decision-making and handling encoding sequence are preferentially format-converted through a projection layer to align with the task prompt embedding vector or the knowledge embedding vector.
[0020] Preferably, before the encoding vectors obtained by each encoder are input into the deep learning model, the weights are dynamically adjusted through a self-attention layer to better capture long-distance dependence relationships.
[0021] Preferably, the multi-dimensional encoder further includes a word vector encoder, which is used to encode the user's question and obtain a question-and-answer vector through the deep learning model based on the encoding vector.
[0022] Preferably, the tobacco public opinion monitoring model is pre-trained using a first loss function, and the first loss function is the weighted cross-entropy of sentiment classification, topic classification, and warning level. The expression is as follows:
[0023]
[0024] In the formula, represents the first loss function, and respectively represent the cross-entropy of sentiment classification, topic classification, and warning level, and α, β, and γ are the respective corresponding weight coefficients;
[0025] and supervised fine-tuning is performed using a second loss function. The expression of the second loss function is as follows:
[0026]
[0027] In the formula, represents the second loss function, represents the cross-entropy loss of the tobacco public opinion monitoring model, α represents the regularization coefficient, ||·||F represents the Frobenius norm of the matrix, λ represents the regularization coefficient, and U and V represent the low-rank matrices obtained by decomposing the weight matrix of the tobacco public opinion monitoring model.
[0028] Preferably, the public opinion monitoring data is obtained through multi-platform crawling. The steps include:
[0029] Setting crawling rules and creating an index. The crawling rules include crawler ID, crawler name, keywords, and crawling period;
[0030] Performing deduplication processing, format unification, and invalid information filtering on the crawled data.
[0031] Preferably, an opinion content table, a false public opinion information table, and a decision-making suggestion record table are automatically generated based on the monitoring results.
[0032] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a tobacco public opinion monitoring method, aiming to significantly improve the accuracy and reliability of tobacco public opinion analysis through an external retrieval enhancement strategy, and at the same time realize comprehensive tobacco public opinion monitoring based on a deep learning large model, including sentiment analysis, subject classification, warning level, false public opinion identification, and decision-making and disposal. Description of the Drawings
[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on the provided drawings.
[0034] Figure 1 Schematic diagram of the tobacco public opinion monitoring model of the present invention;
[0035] Figure 2 Flowchart of tobacco public opinion monitoring based on the tobacco public opinion monitoring model of the present invention. Detailed implementation manners
[0036] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0037] The embodiments of the present invention disclose a tobacco public opinion monitoring method, which significantly improves the accuracy and reliability of false information identification through a multi-dimensional monitoring system and in-depth analysis of user behavior patterns.
[0038] Among them, for the multi-dimensional monitoring system, the present application can not only evaluate the authenticity and credibility from the text content itself, but also assist in enhancing the judgment by cross-referencing relevant information; for the in-depth analysis of user behavior, the present application constructs a monitoring framework, which can not only capture conventional abnormal behavior characteristics and identify changes in behavior patterns that are difficult for humans to detect, such as posting a large number of similar contents or comments in a short period of time, but also realize the decision-making and disposal of public opinion.
[0039] In one implementation manner, the tobacco public opinion monitoring method includes:
[0040] Obtain tobacco public opinion monitoring data and conduct tobacco public opinion monitoring based on the tobacco public opinion monitoring model;
[0041] The tobacco public opinion monitoring model includes a multi-dimensional encoder and a deep learning model.
[0042] The multi-dimensional encoder includes a sentiment analysis encoder, a topic classification encoder, a decision-making and disposal encoder, an alarm level encoder, and a false public opinion identification encoder, which are used to encode the tobacco public opinion monitoring data for different tasks.
[0043] The deep learning model is used to conduct tobacco public opinion monitoring according to the encoded vectors of each encoder.
[0044] The following is illustrated by specific examples.
[0045] Example 1
[0046] This application aims to provide accurate sentiment analysis, topic classification, alarm level assessment, and virtual public opinion detection functions, and generate scientific and reasonable decision-making suggestions for public opinion management.
[0047] Furthermore, the Retrieval-Augmented Generation (RAG) technology is incorporated to retrieve relevant information from a wide range of government policies and tobacco knowledge databases, so as to improve the accuracy and relevance of the content, and ensure the practicality and credibility in the intelligent decision-making process. Finally, a large tobacco public opinion analysis model is formed, specifically referring to Figure 1 , Figure 1 as the structural diagram of the tobacco public opinion monitoring model;
[0048] 1. First, construct an external knowledge base.
[0049] The external knowledge base is a knowledge graph constructed based on tobacco-related materials, and Elasticsearch is used for full-text search according to the keywords of public opinion monitoring data, and the Jaccard coefficient is utilized to measure the similarity between the input keywords and the retrieved relevant content.
[0050] Specifically, the data sources of the external database come from: tobacco-related policies, laws, and regulations issued by national and local governments at all levels; integrating authoritative materials such as market research reports and health impact research reports in the tobacco industry. Gathering reports on the tobacco industry from mainstream news media, covering the latest developments, event reviews, etc. Screening tobacco-related scientific research papers and technical reports from academic databases (such as CNKI, Web of Science).
[0051] Construction of the knowledge graph of the external knowledge base: Apply named entity recognition (NER) technology to extract key entities in the documents, such as person names, place names, and organizational names. Through dependency syntax analysis and pattern matching methods, establish the semantic relationships between entities to form a knowledge graph. Adopt the triple storage method of a graph database (Neo4j) to save the constructed knowledge graph and support efficient query.
[0052] Use dependency syntax analysis and pattern matching methods to determine the semantic relationships between entities. Suppose we have a set of sentences S, and through a dependency syntax analysis tool, we can obtain the subject-predicate-object structure or more complex grammatical relationships in each sentence. For example, for the sentence "Zhang San likes to smoke Furongwang", we can obtain the relationship R (Zhang San, likes, Furongwang). Let R be the set of all extracted relationships, then: R = {r1, r2, r3,..., rn}, where each represents a specific semantic relationship.
[0053] The external knowledge base selects the Elasticsearch full-text search engine, and configures the index to optimize the search speed. According to the keywords in the input text (extracted by named entity recognition NER technology), quickly locate the relevant content in the knowledge graph to ensure the relevance and accuracy of the retrieval results. And introduce the Jaccard coefficient to measure the similarity between the input text and the retrieved content.
[0054] II. Introduce a prompt learner
[0055] Use the prompt learner to convert the retrieved information into prompt embeddings and pass them as part of the input to the deep learning model to assist it in better understanding the text content and generating accurate content; in this embodiment, the deep learning model selects Tongyi Qianwen large model. Tongyi Qianwen large model can understand more complex semantic structures, especially when dealing with texts involving technical terms, metaphorical expressions, or specific cultural backgrounds, and provide more accurate sentiment analysis results.
[0056] In this embodiment, the prompt learner consists of learnable basic external knowledge prompt embeddings and the task modality encoder E enc E enc converts the input text for prompts into task prompt embeddings (which can be multiple) and then adds them to the encoding sequences of different tasks respectively. At the same time, E base and E dec together form a total prompt embedding and are combined with the sentiment analysis encoding sequence, the subject classification encoding sequence, and the decision-making and handling encoding sequence and then jointly input into the tobacco public opinion analysis large model.
[0057] Specifically, in the public opinion sentiment analysis modality encoder, through retrieval-enhanced generation technology, relevant information is obtained from a wide range of tobacco industry knowledge bases and converted into prompt embeddings. Specifically, the prompt learner first identifies the external knowledge fragments most relevant to the input public opinion text, and then converts these knowledge fragments into prompt embeddings. The prompt embeddings and the original public opinion text encoding are jointly input into the Tongyi Qianwen large model, and the self-attention mechanism is used to achieve effective fusion between the two;
[0058] When it is necessary to classify the theme of the public opinion text, the relevant knowledge graph nodes are automatically retrieved and added to the input sequence as prompt embeddings. Let P represent the set of prompt embeddings generated by the prompt learner, and T represent the public opinion text encoding. Then the final representation form input to the model can be expressed as X = [T; P], where; represents the vector concatenation operation;
[0059] The sentiment decision-making and handling modal encoder combines professional suggestions and response strategies in a pre-constructed knowledge base of the tobacco industry, enabling the system to provide customized decision-making support for managers.
[0060] Through the retrieval-augmented generation technology, the model can access external knowledge bases during the training process, thereby enhancing its information acquisition and understanding capabilities in multi-modal data processing.
[0061] To further optimize the above technical solution, the sentiment analysis coding sequence, the subject classification coding sequence, and the decision-making and handling coding sequence of this application are preferentially format-converted through a projection layer to align with the task prompt embedding vector or the knowledge embedding vector.
[0062] Meanwhile, before the encoded vectors obtained by each encoder are input into the deep learning model, they pass through a self-attention layer for dynamic weight adjustment to better capture long-range dependencies.
[0063] In some embodiments, the multi-dimensional encoders share the same internal structure, that is, each encoder first converts the input text data into a low-dimensional vector representation through an embedding layer. For the input public opinion text T and the relevant information P retrieved from the external knowledge base, the two first pass through their respective embedding layers to obtain the corresponding vector representations ET and EP. Then, under a common self-attention mechanism, these two vectors are fused together to form a new representation X = [ET; EP], where; represents the vector concatenation operation.
[0064] The self-attention mechanism enables the model to dynamically adjust the weights between different words, thereby better capturing long-range dependencies.
[0065] Specifically, let Q, K, and V be the Query, Key, and Value matrices respectively, then the output of the self-attention layer can be expressed as: where dk is the dimension of the key vector. By stacking multiple such self-attention layers, the model can refine the input information layer by layer and gradually construct more abstract and high-level feature representations. During this process, each layer adjusts its own parameters according to the output of the previous layer to optimize the final prediction result.
[0066] Finally, the high-level feature representations after multi-layer attention processing and the prompts provided by the external knowledge base will be passed to the tobacco public opinion large model for performing specific tasks such as sentiment analysis, topic classification, or decision-making and handling. These modal encoders can dynamically adjust the activation function through the identifier F = {f0, f1, f2, f3...}.
[0067] In one embodiment, the multidimensional encoder also includes a word vector encoder for encoding user questions and obtaining a question-answer vector based on the encoded vector through a deep learning model. In this embodiment, the large model can customize personalized communication strategies according to the characteristics of different groups by learning a large amount of historical data and social dynamics; that is, for different audiences, such as young consumers, policymakers, public health advocates, etc., the most suitable language style and expression can be used to more effectively reach the target audience and improve communication effectiveness. For example, for young people, the system will use more easy-to-understand and life-like language; in professional fields, more rigorous and authoritative expressions are used to ensure the professionalism and accuracy of information communication.
[0068] Embodiment 2
[0069] In order to further optimize the above technical solution, the tobacco public opinion monitoring model of this application is trained and fine-tuned in the following ways:
[0070] 1. Data preprocessing, including:
[0071] a. De-duplication: The unique identifier of each record is calculated through the MD5 algorithm to remove duplicates and ensure the uniqueness of the data set. The MD5 algorithm converts tobacco review text of any length into a fixed-length 128-bit hash value and removes duplicates by comparing the hash values between different reviews.
[0072] b. Format unification: Convert data from different sources into a unified structured JSON format for subsequent processing.
[0073] c. Invalid information filtering: remove HTML tags, special characters, and obviously irrelevant noise data, such as advertisements, meaningless emoticons, etc.
[0074] 2. Public Opinion Comment Data Labeling
[0075] Using an auxiliary annotation model based on the LLAMA2 large language model, each tobacco public opinion comment is annotated with positive, negative or neutral sentiment, topic and warning, providing a basis for subsequent fine-tuning.
[0076] a. Emotional label generation: Automatically add emotional labels (positive / negative / neutral) based on the review text content and perform manual review.
[0077] b. Topic tag generation: The TF-IDF method is used to extract keywords, and combined with the K-means clustering algorithm, the comments are preliminarily classified by topic, such as tobacco control policy, product evaluation, health impact, etc. Then, the language model based on LLAMA2 is used to automatically assign one or more topic tags (tobacco control policy, product evaluation, etc.) to each comment text, and manual audit is performed.
[0078] c. Generation of public opinion warning level labels: Automatically assign one or more public opinion warning labels (very severe / severe / medium / slight / none) to each comment text and conduct manual verification.
[0079] III. Data augmentation
[0080] a. Establish a thesaurus related to the tobacco industry, covering professional terms and their variants. And on the premise of maintaining the original semantics, randomly replace some words with synonyms to increase data diversity.
[0081] b. Use the LLAMA2 model to generate contexts similar to but not exactly the same as the existing comments to enrich the dataset. Integrate the Elasticsearch engine API for searching to retrieve relevant information from a wide range of document databases to assist in generating more diverse background materials.
[0082] IV. Data feature engineering
[0083] a. Use the Jieba tokenizer to tokenize Chinese text, considering the proper nouns and abbreviations in the tobacco industry. For English content, apply the Lancaster Stemmer for stemming; for Chinese, implement named entity recognition and part-of-speech tagging through the HMM+CRF model.
[0084] b. Calculate the TF-IDF value to reflect the importance of words in the document, and introduce the VADER (Valence Aware Dictionary and sEntiment Reasoner) scoring system to measure the sentiment intensity of each comment. VADER is a sentiment analysis tool specifically designed for social media text and is particularly suitable for short texts with emotional colors.
[0085] V. Dataset preparation
[0086] a. Randomly divide the training set, validation set, and test set in a ratio of 8:1:1 to ensure sufficient differences and representativeness among the sets. Adopt the K-fold cross-validation strategy to further verify the stability and generalization ability of the model performance.
[0087] b. Form a tobacco public opinion analysis dataset with 60,398 pieces of data, among which 42,098 pieces of data can be used for public opinion sentiment classification, 41,293 pieces of data for public opinion topic classification, and 4,912 pieces of data for public opinion warning level prediction. And it includes 11,000 pieces of public opinion comments generated by LLAMA, including 5,000 pieces of malicious attack comments.
[0088] VI. Pre-training of the tobacco public opinion monitoring model
[0089] In this embodiment, Tongyi Qianwen's foundation large model is adopted for large-scale pre-training on the Chinese tobacco public opinion analysis dataset. The cross-entropy loss function is used to optimize the model parameters, and the learning rate and other hyperparameters are set to supervised guide the model to distinguish positive, negative, and neutral comments, identify different types of public opinion topics (such as tobacco control policies, health impacts, market dynamics, etc.), and identify the public opinion warning levels.
[0090] Specifically, the first loss function is used for pre-training. The first loss function is the weighted cross-entropy of sentiment classification, topic classification, and warning level, and the expression is:
[0091]
[0092] In the formula, represents the first loss function, and respectively represent the cross-entropy of sentiment classification, topic classification, and warning level, and α, β, and γ are the respective corresponding weight coefficients.
[0093] VII. Supervised Fine-tuning of the Tobacco Public Opinion Monitoring Model
[0094] The objective function of supervised fine-tuning can be expressed as: Among them, is the cross-entropy loss, which is used to measure the difference between the model prediction and the true label; is the regularization term, where ||·||F represents the Frobenius norm of the matrix (the square root of the sum of the squares of the elements), and α is the regularization coefficient. It is used to constrain the model parameters to prevent overfitting; λ is the regularization coefficient to balance the weights of the two terms.
[0095] In the supervised fine-tuning stage, the pre-trained large model is further fine-tuned on a specific dataset, so as to perform dissimilar fine-tuning on the five specific functions of public opinion sentiment analysis, public opinion topic classification, false public opinion analysis, public opinion disposal decision-making, and public opinion warning.
[0096] The difference between it and pre-training is that: the goal of supervised fine-tuning (SFT) is to make the model adapt to the task requirements of the vertical domain through fine-tuning, only fine-tuning some parameters or adding an adaptation layer, so as to quickly complete the adaptation of professional data.
[0097] In the specific implementation of low-rank adaptation, the weight matrix of the model is decomposed into the product of two low-rank matrices, that is, W = UV, where U and V are low-rank matrices.
[0098] In this way, the number of model parameters is significantly reduced while maintaining its key feature representation in multimodal data processing. During the fine-tuning process, the optimization objective includes not only traditional loss functions (such as cross-entropy loss) but also a regularization term to constrain the rank of the low-rank matrix. This ensures that the model can still effectively process complex multimodal data while reducing the number of training parameters.
[0099] Preferably, as new tobacco public opinion data continuously flows in, the training dataset is updated regularly, and the model is retuned to maintain its accuracy under the latest public opinion changes. This application continuously optimizes its deep learning prediction model according to public opinion feedback, adjusts communication strategies, and ensures that the expected public opinion management goal is finally achieved. In addition, new training samples are introduced, and only the new part is fine-tuned to reduce the time cost of retraining the entire model. In addition, the model performance is evaluated regularly to ensure its stability and reliability in practical applications.
[0100] By introducing a data-driven method empowered by large models and a continuous evaluation mechanism, this application can provide a more scientific and reasonable public opinion disposal plan, which not only improves the decision-making efficiency but also enhances the effectiveness and pertinence of countermeasures, and can maintain a high degree of flexibility and adaptability in a rapidly changing network environment.
[0101] The present invention uses a large tobacco public opinion analysis model as the core technology. Through domain adaptation optimization, it can understand more complex semantic structures, especially when dealing with texts involving professional terms, metaphorical expressions, or specific cultural backgrounds, and provide more accurate sentiment analysis results.
[0102] Embodiment III
[0103] The monitoring process based on the tobacco public opinion monitoring model is as Figure 2 shown; including:
[0104] 1. Specify crawling rules and use a crawler to crawl data from specified sources;
[0105] In this embodiment, a crawler module is constructed to crawl tobacco public opinion data from various different platforms to ensure the comprehensiveness, accuracy, and timeliness of the data. Specifically, it includes:
[0106] 1.1 Crawler initialization
[0107] Clarify the target platform and its characteristics, and determine the required data types. For tobacco network public opinion monitoring, this application needs to cover multiple platforms such as Weibo, Tieba, Xiaohongshu, forums, official announcements, and government policies. The integration of multi-source data not only increases the breadth of public opinion monitoring but also improves its depth, enabling the system to quickly evaluate the social response after the release of tobacco-related policies and provide a basis for subsequent adjustments. At the same time, using this cross-platform data, the system can more accurately identify potential public opinion hotspots and development trends, improving the overall accuracy and precision of public opinion monitoring.
[0108] Each platform has its unique structure and access rules. In tobacco network public opinion monitoring, platforms such as Weibo and Xiaohongshu provide official API interfaces, allowing legal and efficient data acquisition; while forums such as Tieba and Tianya Community, as well as official announcement websites, use Web Scraping technology to parse web page content and extract information. For some platforms such as Zhihu and Douyin, both API and Web Scraping are used to supplement data.
[0109] For platforms that support API interfaces, this application preferentially uses the API because this method is more stable and compliant with platform specifications, reducing the risk of being blocked. For those platforms that do not provide APIs or have limited API functions, Web Scraping technology is adopted.
[0110] Furthermore, the Python library BeautifulSoup is used to effectively parse HTML pages, locate and extract the required elements. At the same time, to improve efficiency and reduce risks, this method sets up a proxy IP pool to avoid IPs being blocked due to frequent requests, and configures the User-Agent to simulate browser behavior, reducing the probability of being detected as an automated tool. In addition, for content that requires user login to access, it is very necessary to implement the simulated login process (using Selenium) to ensure that a complete dataset can be obtained.
[0111] In addition to the above-mentioned proxy IP pool and User-Agent settings, this application simulates the login mechanism for specific platforms to ensure smooth access to restricted content. When writing the specific crawler logic, the key lies in parsing the page structure, accurately locating and extracting the target elements. An error retry mechanism is added to ensure that the entire process will not be interrupted in case of network fluctuations or other unexpected situations. The captured data should be immediately saved to the local file system or database for subsequent processing. For unstructured text data (such as HTML fragments), it can be converted into a structured JSON object or CSV table form to simplify subsequent data processing steps.
[0112] 1.2 Crawler Rule Management
[0113] A. Establish an indexing structure
[0114] First, an effective indexing structure needs to be established for the crawler rules. This is achieved through indexes in the database. In a relational database, indexes are created for the key fields of the crawler rules (crawler ID, crawler name, keywords, etc.) to accelerate query operations. Since the system requires higher query speed, common rules are loaded into memory at startup and an efficient search structure is used for indexing.
[0115] B. Design a query interface
[0116] Next, design a user-friendly query interface that allows administrators to find specific crawler rules based on different conditions (crawler name, keywords, execution period, etc.). This interface can be through a web interface. This interface can define the parameters required for the query, including the crawler name `name`, keywords `keywords`, execution period range `period_range`, etc. Moreover, this system supports the fuzzy query function, allowing partial matching of the crawler name to find relevant rules more flexibly.
[0117] C. Build query logic
[0118] Based on the query conditions provided by the user, write an SQL query statement or a search algorithm in memory to retrieve the data that meets the conditions from the set of stored crawler rules. For complex queries, multiple conditions need to be combined and efficiency issues need to be considered. For a relational database, construct an SQL query statement and use the WHERE clause to filter out the records that meet the conditions. If it is an in-memory index, the hash table can be directly traversed and the corresponding filtering conditions can be applied.
[0119] D. Result display and pagination
[0120] Finally, display the query results to the user in tabular form and provide a pagination function to handle a large amount of data. Ensure that each result shows key information, such as crawler ID, name, status, etc., for the user to further operate.
[0121] Use an HTML table to display the query results, including columns such as crawler ID, name, keywords, last execution time, etc. When the query results are too many, implement a pagination mechanism to limit the number of results returned each time and provide navigation links (first page, previous page, next page, last page).
[0122]
[0123] Through the above steps, an efficient and user-friendly crawler rule query system is implemented to help administrators quickly locate the required rules and perform necessary management and adjustments. The developed user interface allows administrators to add, modify, and delete crawler rules, including specifying the crawled website (selected from a predefined list), keywords (supporting combined queries for multiple keywords), crawling period, and other parameters. Its functions are as follows:
[0124] Add crawler rules: Define a new crawler rule by setting the crawler ID, crawler name, execution period, last execution time, and the number of crawled texts;
[0125] Modify crawler rules: Update any one or more pieces of information in the existing crawler rules;
[0126] Delete crawler rules: Remove the unnecessary crawler rules to ensure they are no longer executed;
[0127] Query crawler rules: Implement the query function for the set rules to quickly locate the required rules.
[0128] E. Task Scheduling and Timed Execution
[0129] Set timed tasks to trigger the running of the crawler script regularly to ensure the timeliness and persistence of data. This system is implemented using the Python library APScheduler, which supports multiple scheduling strategies, including one-time, periodic, and CRON expression-based scheduling methods. Such a design not only improves the flexibility of the system but also ensures the regularity and stability of data scraping.
[0130] Finally, establish a sound monitoring and maintenance mechanism. Log records can not only help track historical operations but also assist in troubleshooting. In terms of performance optimization, adjust parameters such as the concurrency number and request frequency according to the actual running situation to improve efficiency while avoiding overloading the platform server.
[0131] 2. Data preprocessing and integration to ensure that data from different sources can be comprehensively analyzed on the same platform; (The processing steps are the same as before and will not be elaborated here)
[0132] 3. Fine-tune and train the tobacco public opinion monitoring model, and use the adjusted model combined with the tobacco public opinion professional knowledge base for public opinion analysis, malicious comment monitoring, and public opinion handling decision-making.
[0133] 4. Data post-processing
[0134] 5. Result visualization and data management
[0135] In this embodiment, the steps to achieve visualization include:
[0136] A. Environment Configuration and Project Initialization
[0137] First, ensure that the necessary tools and dependencies are installed in the development environment, and initialize a new front-end project. This includes installing Node.js and npm, creating a new project based on Vue.js, and installing ECharts and its Vue plugin.
[0138] B. Build the basic page structure
[0139] Build the basic page layout in the Vue project, including the navigation bar, main content area, etc. Set up the routing to support different view pages, such as the line chart view of public opinion sentiment changing over time and the bar chart comparison view of the attention levels of different topics.
[0140] C. Integrate ECharts for data visualization
[0141] Add ECharts charts to each view page to display different types of analysis results. Specifically: The line chart shows the change of sentiment over time: Use a line chart to represent the changing trends of positive, negative, and neutral emotions over time. Users can intuitively see the fluctuations of sentiment through this view. The bar chart compares the attention levels of different topics: Show the attention level comparison of each topic (tobacco control policies, product evaluations, brand promotions, etc.) through a bar chart to help users quickly understand which topics are more concerned.
[0142] D. Dynamically load data
[0143] Obtain real-time data from the backend API and dynamically update the chart content. Ensure that the latest public opinion data can be obtained every time the page is loaded or there is user interaction, so as to maintain the timeliness and accuracy of information. Improve the UI / UX design to ensure that the user interface is friendly and easy to operate. Define global style rules to make the entire application have a consistent visual style. At the same time, adopt the responsive design principle to ensure that the charts can adapt to different screen sizes and provide a better mobile experience.
[0144] E. Deployment and release
[0145] After the development is completed, deploy the project to the production environment to make it publicly accessible. Generate optimized static resource files and select a suitable hosting platform to upload these files. Configure domain name resolution so that users can access the tobacco network public opinion monitoring system through the Internet.
[0146] Furthermore, modular design is carried out accordingly, including a data collection module, which is used to collect public opinion information from the Internet through crawlers or other means and store this information in the "public opinion content table";
[0147] as well as an emotion analysis module, a theme classification module, a decision-making and handling module, a public opinion warning module, and a false public opinion monitoring module, which are used to monitor public opinions based on the tobacco public opinion monitoring model according to the text in the "public opinion content table" and display analysis charts;
[0148] This application finally automatically generates a public opinion content table, a false public opinion information table, and a decision-making suggestion record table based on the monitoring results.
[0149] The present invention provides a tobacco public opinion monitoring method based on a deep learning model, which has functions of public opinion emotion analysis, theme classification, public opinion warning, false detection, and decision-making and handling. It aims to overcome the limitations of the prior art in dealing with complex contexts, multi-source heterogeneous data fusion, and personalized emotion classification, and significantly improve the accuracy and reliability of tobacco public opinion analysis. This method can not only capture and interpret the trends and changes of public opinions more accurately, but also provide strong technical support for the public opinion management and decision-making in the tobacco industry.
[0150] This application not only overcomes the technical limitations of traditional methods in the face of false information under complex disguises, but also provides more accurate and comprehensive support for public opinion monitoring by introducing deeper content understanding and behavior pattern analysis. This mechanism is particularly suitable for identifying and curbing false public opinions caused by highly organized malicious behaviors, ensuring the authenticity and reliability of public opinion analysis results, and maintaining a healthy ecosystem in the cyberspace.
[0151] Compared with the prior art,
[0152] 1) The present invention adopts a fine-tuned large model for tobacco public opinion analysis, which is trained with a large number of Chinese texts, enabling the model to understand more complex semantic structures. Especially when dealing with texts involving professional terms, metaphorical expressions, or specific cultural backgrounds, it provides more accurate emotion analysis, helping to capture the attitude changes of the public towards tobacco products more accurately. At the same time, the system integrates data from social platforms such as Weibo, Tieba, and Xiaohongshu, ensuring that information from different sources can be comprehensively analyzed on the same platform, expanding the coverage of public opinion monitoring, and improving the comprehensiveness and accuracy of analysis results;
[0153] 2) Innovatively introduce a comprehensive evaluation system based on the large model for tobacco public opinion analysis, which not only evaluates the authenticity and credibility of the text content itself, but also enhances the judgment by cross-verifying relevant information sources. This method can automatically check the historical records, authority, and consistency of the message sources; at the same time, utilize the powerful correlation analysis ability of the large model to achieve cross-platform and cross-time information comparison, ensuring that the obtained data is the latest and most reliable.
[0154] Meanwhile, with the help of Tongyi Qianwen's large model's learning and understanding capabilities for massive data, an efficient abnormal behavior detection framework is constructed. This framework can not only capture conventional abnormal behavior characteristics but also identify changes in behavior patterns that are difficult for humans to detect, thus significantly improving the recognition rate of false information;
[0155] 3) A public opinion management strategy based on the combination of a large model for tobacco public opinion analysis and manual intervention is proposed. For common public opinion issues, the system can automatically generate response templates or recommended measures to ensure a response in the first place. At the same time, the system reserves room for manual intervention. Especially when dealing with sensitive topics or complex events, this hybrid approach can ensure both rapid response and flexibility and adaptability
[0156] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.
[0157] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A tobacco public opinion monitoring method, characterized in that obtain tobacco public opinion monitoring data and conduct tobacco public opinion monitoring based on a tobacco public opinion monitoring model; the tobacco public opinion monitoring model includes a multi-dimensional encoder and a deep learning model, the multi-dimensional encoder includes a sentiment analysis encoder, a topic classification encoder, a decision-making and handling encoder, an alarm level encoder, and a false public opinion identification encoder, which are used to encode the tobacco public opinion monitoring data for different tasks; the deep learning model is used to conduct tobacco public opinion monitoring according to the encoded vectors of each encoder.
2. The tobacco public opinion monitoring method according to claim 1, wherein Use a prompt learner to generate a task prompt embedding vector according to the user task prompt text, and add the task prompt embedding vector to the encoded sequences of different tasks respectively.
3. The tobacco public opinion monitoring method according to claim 1, wherein Construct an external knowledge base, obtain relevant content in the external knowledge base through keywords in the public opinion monitoring data, use a prompt learner to generate a knowledge embedding vector according to the relevant content, and add the knowledge embedding vector to the sentiment analysis encoding sequence, the subject classification encoding sequence, and the decision-making and handling encoding sequence respectively.
4. The tobacco public opinion monitoring method according to claim 3, characterized in that The external knowledge base is a knowledge graph constructed based on tobacco-related materials, and Elasticsearch is used for full-text search according to the keywords of the public opinion monitoring data, and the Jaccard coefficient is used to measure the similarity between the input keywords and the retrieved relevant content.
5. The tobacco public opinion monitoring method according to claim 2 or 3, characterized in that, The sentiment analysis encoding sequence, the subject classification encoding sequence, and the decision-making and handling encoding sequence are preferentially format-converted through a projection layer to align with the task prompt embedding vector or the knowledge embedding vector.
6. The tobacco public opinion monitoring method according to claim 1, wherein Before the encoded vectors obtained by each encoder are input into the deep learning model, the weights are dynamically adjusted through a self-attention layer.
7. The tobacco public opinion monitoring method according to claim 1, characterized in that The multi-dimensional encoder also includes a word vector encoder, which is used to encode the user's question and obtain a question-and-answer vector through the deep learning model based on the encoded vector.
8. The tobacco public opinion monitoring method according to claim 1, characterized in that, The tobacco public opinion monitoring model is pre-trained using a first loss function, and the first loss function is the weighted cross-entropy of sentiment classification, topic classification, and alarm level, and the expression is: In the formula, represents the first loss function, and respectively represent the cross-entropies of sentiment classification, topic classification, and alert level, and α, β, and γ are the respective corresponding weight coefficients; and supervised fine-tuning is performed using a second loss function, and the expression of the second loss function is: In the formula, represents the second loss function, represents the cross-entropy loss of the tobacco public opinion monitoring model, α represents the regularization coefficient, ||·||F represents the Frobenius norm of the matrix, λ represents the regularization coefficient, and U and V represent the low-rank matrices obtained by decomposing the weight matrix of the tobacco public opinion monitoring model.
9. The tobacco public opinion monitoring method according to claim 1, characterized in that The public opinion monitoring data is obtained through multi-platform crawlers, and the steps include: set the crawler rules and create an index, and the crawler rules include crawler ID, crawler name, keyword, and crawler period; perform deduplication processing, format unification, and invalid information filtering on the crawler data.
10. The tobacco public opinion monitoring method according to claim 1, characterized in that Automatically generate an opinion content table, a false public opinion information table, and a decision-making suggestion record table based on the monitoring results.