Intelligent tobacco public opinion monitoring and analyzing method based on knowledge distillation

Through the method based on knowledge distillation and graph convolution network, the problem that traditional tobacco public opinion monitoring methods cannot cope with massive complex information is solved, and comprehensive, accurate and timely monitoring of tobacco public opinion is achieved, cost is reduced, analysis and traceability accuracy is improved, and effective early warning and decision-making is supported.

CN120216744APending Publication Date: 2025-06-27GUANGDONG TOBACCO JIEYANG CITY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510288629.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Traditional tobacco public opinion monitoring and analysis methods cannot effectively deal with massive and complex information, and it is difficult to accurately grasp the unique semantic and emotional tendencies of the tobacco field, and lack effective traceability and early warning technologies.

Method used

The intelligent monitoring and analysis method of tobacco public opinion based on knowledge distillation is adopted, multi-platform data is obtained through crawling technology, student models are used to perform emotional prediction and knowledge distillation, and knowledge graphs are built in combination with graph convolution networks and multi-relational graph mechanisms to achieve accurate traceability and early warning of public opinion.

Benefits of technology

It realizes comprehensive, accurate and timely monitoring and analysis of tobacco public opinion, reduces the cost of deployment reasoning, improves the accuracy of sentiment analysis and traceability accuracy, and supports effective early warning and decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216744A_ABST
    Figure CN120216744A_ABST
Patent Text Reader

Abstract

The invention discloses a tobacco public opinion intelligent monitoring analysis method based on knowledge distillation, which comprises the following steps: determining a crawler website and a crawler rule, and obtaining tobacco public opinion data; performing word segmentation on the tobacco public opinion data; according to the word segmentation result, emotion prediction is carried out through a student model, the student model carries out knowledge distillation on tobacco public opinion text analysis tasks through a teacher model, and prediction differences are measured by introducing KL divergence distillation loss of temperature parameters. The method can break through the limitation of a single platform, and achieves the comprehensive collection of tobacco data. And meanwhile, the reasoning cost on deployment can be effectively reduced, and specific semantic and emotional tendencies in the tobacco field can be accurately grasped, so that the recognition accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tobacco public opinion monitoring, and more specifically, to an intelligent monitoring and analysis method for tobacco public opinion based on knowledge distillation. Background Art

[0002] In today's digital age, the tobacco industry is facing unprecedented challenges in the public opinion environment. The booming development of social media, news websites, and various professional forums has led to a flood of tobacco-related public opinion information, which spreads extremely fast and has a wide range of influence. Tobacco companies, regulatory authorities, and related research institutions urgently need to comprehensively and timely understand the public's views and attitudes towards tobacco products, policies, and industry trends.

[0003] However, traditional public opinion monitoring and analysis methods are often limited to a single platform or simple data collection, and are unable to cope with such a large amount of complex information.

[0004] On the other hand, existing public opinion analysis methods have deficiencies in accuracy and depth. Traditional sentiment analysis models are difficult to accurately grasp the semantics and sentiment tendencies unique to the tobacco field, have limited text interpretation capabilities in complex contexts, and large models have high inference costs in deployment. At the same time, in terms of public opinion tracing and early warning, due to the lack of effective technical means, traditional methods are difficult to clarify the dissemination context of information and cannot detect potential crises in advance. In addition, the dissemination of tobacco public opinion information involves many entities and complex relationships, and traditional analysis methods are unable to effectively model and analyze these entities and relationships, resulting in difficulties in tracking the source and dissemination path of public opinion.

[0005] Therefore, in the face of scattered multi-platform and multi-type data sources, how to overcome the above problems and achieve accurate and intelligent monitoring of tobacco public opinion is an urgent problem for those skilled in the art to solve. Summary of the Invention

[0006] In view of this, the present invention provides an intelligent monitoring and analysis method for tobacco public opinion based on knowledge distillation, aiming to provide a comprehensive and innovative solution for tobacco public opinion management.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] An intelligent monitoring and analysis method for tobacco public opinion based on knowledge distillation, comprising:

[0009] Determine the crawler websites and crawler rules to obtain tobacco public opinion data;

[0010] Segment the tobacco public opinion data;

[0011] Perform sentiment prediction using the student model according to the segmentation results;

[0012] The student model performs knowledge distillation on the tobacco public opinion text analysis task through the teacher model, and measures the prediction difference through the following loss function:

[0013]

[0014] In the formula, L KD represents the KL divergence distillation loss, y T is the output probability distribution of the teacher model, y S is the output probability distribution of the student model, T represents the temperature parameter, and i represents the sample index.

[0015] Preferably, the crawler websites include Weibo platform, WeChat official accounts, news website platforms, and Tieba; the crawler rules include crawler keywords and crawler periods.

[0016] Furthermore, the tobacco public opinion data is preferably preprocessed, including: converting to lowercase, removing punctuation, and removing redundancy through hash values.

[0017] Furthermore, the tobacco public opinion data is segmented, including performing dictionary matching after segmenting the tobacco public opinion data, and determining the vocabulary boundaries of unmatched word strings using the hidden Markov model according to the transition probabilities between characters.

[0018] Furthermore, considering the cross-entropy loss to measure the prediction difference, the expression of the total loss function is:

[0019] L total = λL KD + (1 - λ)L CD

[0020] In the formula, λ represents the balance parameter, L CD represents the cross-entropy loss, and the expression is:

[0021]

[0022] In the formula, y true (i) represents the true label of the i-th sample of the student model, y pred (i) is the predicted label of the i-th sample of the student model.

[0023] Furthermore, performing sentiment prediction using the student model based on the segmentation results includes:

[0024] Inputting each segment into the student model for sentiment prediction, and counting the number of sentiment tendencies in different categories;

[0025] And inputting the original text corresponding to the segment into the student model for sentiment prediction, and counting the number of sentiment tendencies in different categories.

[0026] Further, perform traceability analysis through the knowledge graph based on the prediction results; the knowledge graph is constructed based on tobacco public opinion data.

[0027] Further, the knowledge graph predicts entity categories through the following steps:

[0028] Obtain entity nodes in the tobacco public opinion data, and propagate them layer by layer through the graph convolutional network to obtain node embedding representations; during propagation, the node update formula is:

[0029]

[0030] In the formula, represents the node feature vector of the (l + 1)-th layer, σ represents the activation function, N(i) represents the set of neighbor nodes of node i, c i,j represents the normalization coefficient, W(l) represents the feature transformation matrix, represents the weight matrix;

[0031] Input the node embedding representation into the linear classifier to obtain the entity category.

[0032] Further, the knowledge graph uses a bilinear decoder to predict entity relationships, and the prediction formula is:

[0033]

[0034] In the formula, P represents and have a relationship r ij the probability, σ represents the activation function, R r represents the bilinear matrix corresponding to the relationship r, represents the embedding representation of node i, represents the embedding representation of node j.

[0035] Further, the bilinear decoder uses a contrastive loss function based on negative sampling:

[0036]

[0037] In the formula, β represents the hyperparameter for controlling the balance between positive and negative samples.

[0038] Further, generate a public opinion report based on the sentiment prediction results, and perform early warning and visual display.

[0039] The present invention discloses a tobacco public opinion intelligent monitoring and analysis method based on knowledge distillation. Compared with the prior art,

[0040] This application can break through the limitations of a single platform and achieve comprehensive collection of tobacco data;

[0041] This application trains a student model based on knowledge distillation, which greatly reduces the inference cost in deployment. At the same time, the KL divergence distillation loss with temperature parameter is introduced to measure the difference in probability distributions between the teacher model and the student model outputs, ensuring the recognition accuracy of the student model.

[0042] This application introduces a graph convolutional network and a multi-relationship graph mechanism to construct a knowledge graph for accurate tracing of public opinion on tobacco, thus providing strong support for early warning. Brief Description of the Drawings

[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0044] Figure 1 It is a flowchart example of the intelligent monitoring and analysis method for tobacco public opinion based on knowledge distillation of the present invention;

[0045] Figure 2 It is a display diagram of customized crawler programs for multiple social platforms of the present invention;

[0046] Figure 3 It is a flowchart of sentiment analysis and tracing of the present invention;

[0047] Figure 4 It is a comprehensive visual display of public opinion analysis data of the present invention. Detailed Embodiments

[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0049] The embodiments of the present invention disclose an intelligent monitoring and analysis method for tobacco public opinion based on knowledge distillation.

[0050] In one embodiment, the steps refer to Figure 1 , Figure 1 It is a flowchart example of a method for intelligent monitoring and analysis of tobacco public opinion of the present invention.

[0051] 1. In one embodiment, the collection and storage of tobacco public opinion data are realized through crawler technology, including:

[0052] 1.1 Develop customized backend crawler programs for multiple social platforms

[0053] Develop customized multi-attribute (main posts, replies, and comments from various information sources) web crawler programs for social media platforms (such as Weibo, WeChat official accounts, etc.), news websites, and professional forums. Considering the significant differences in page architectures and data acquisition specifications among different platforms, custom crawler programs are innovatively developed for the following four platforms, referring to Figure 2 ;

[0054] 1.1.1 Weibo Platform

[0055] The Weibo platform has a complex page architecture, extensively uses JavaScript dynamic loading technology, and data is transmitted in JSON format, containing rich meta-information such as user profiles and interaction data. At the same time, its anti-crawler mechanism is strict, restricting the IP access frequency and requiring login verification. In terms of technical implementation, the Scrapy framework is used to build the crawler main body, and the custom Spider class is used to plan the crawling logic. The Selenium library combined with ChromeDriver is used to simulate user login, and the obtained Cookie is passed to Scrapy to break through the login restriction.

[0056] In the data extraction stage, XPath expressions are used to accurately locate data nodes. For example, / / div[@class='weibo-content'] is used to locate Weibo content, / / a[@class='name'] is used to extract the publisher information, and / / span[@class='ct'] is used to extract the publication time.

[0057] For some content with special formats in Weibo data, such as topic tags and @ mentions, regular expressions #([^#]+)#|@([^@\s]+) are designed for recognition and processing. In terms of data storage, according to the previously designed unified SQL table format, after organizing the crawled data into a suitable format, SQL insert statements are executed through the pymysql library to insert the data into the initial data public_opinion_data table.

[0058] 1.1.2 WeChat Official Account Platform

[0059] The acquisition channels of WeChat official account articles are relatively special. They need to be accessed within the WeChat client environment, and data is mainly obtained through specific interfaces. Complex parameters and signature information need to be carried during requests. At the same time, WeChat's anti-crawler mechanism is extremely complex, and it strictly guards against data crawling from non-official channels. To achieve data crawling, first, deeply analyze the interface request logic of WeChat official account articles. Use the sha1 algorithm in the hashlib library to perform specific hash calculations on multiple key parameters (timestamp, random number, article link, etc.) to generate a valid signature. During the actual crawling process, use the Scrapy framework to send HTTP requests with correct parameters. For some articles that require login permissions to access, use Selenium to simulate the WeChat login process. In the data parsing section, use the BeautifulSoup library to parse the obtained page content and extract key information such as article title, body, publisher, and publication time. For special symbols, emojis, etc. that may appear in WeChat official account articles, design a regular expression [\u2600-\u27BF]|\uD83C[\uDF00-\uDFFF]|\uD83D[\uDC00-\uDE4F] to clean and standardize the text content. When storing data, organize the data into the corresponding format according to the unified SQL table structure, and use the pymysql library to execute the insert operation to insert the data into the initial data public_opinion_data table.

[0060] 1.1.3 News website platform

[0061] Although the page structures of news websites are relatively standardized, there are still significant differences between different sites, and the data presentation form is mainly static HTML. To prevent malicious data collection, some news websites have set up simple anti-crawler mechanisms such as restricting access frequencies. When building a crawler based on the Scrapy framework, by carefully analyzing the page structure of the target news website, accurate XPath paths are determined for data extraction. For example, / / h1[@class='article-title'] is used to extract the article title, / / div[@class='article-body'] is used to extract the body content, / / span[@class='author-name'] is used to extract the author information, and / / time[@class='publish-time'] is used to extract the publication time. For the possible image links, hyperlinks, etc. in news articles, a regular expression (https?|ftp|file): / / [-A-Za-z0-9+&@# / %?=~_|!:,.;]+[-A-Za-z0-9+&@# / %=~_|] is designed for identification and processing. In terms of data storage, according to the unified SQL table format, the organized data is inserted into the initial data public_opinion_data table using the pymysql library.

[0062] 1.1.4 Forum platforms such as Tieba

[0063] The page structures of forums such as Tieba are relatively simple and intuitive, but there is some dynamically loaded content, and the data formats are relatively diverse. Its anti-crawler measures mainly include restricting the IP access frequency and popping up verification codes in specific situations. When designing the crawler, the Scrapy framework is used to build the main framework, and the Selenium library is combined to handle the dynamically loaded part. The post title is located through the XPath expression / / div[@class='threadlist_title'], the replies and comments are extracted from / / div[@class='d_post_content'], the publisher information is obtained from / / span[@class='tb_icon_author'], and the publication time is extracted from / / span[@class='date']. For the common emoticons, special formats, etc. in Tieba content, a regular expression \[\w+\] is designed for normalization processing. When storing data, following the unified SQL table design, after organizing the crawled data, the pymysql library is used to insert the organized data into the initial data public_opinion_data table.

[0064] As a preferred implementation, a user interaction interface is designed at the front end of the crawler in this embodiment; specifically,

[0065] 1.2.1 Designed the front-end interaction interface

[0066] In the design of the system's front-end interface, a very convenient website crawling selection module was constructed and presented to users in the form of a dropdown box, which comprehensively covered various common social media platforms, news websites, and professional forums where tobacco public opinion might appear. At the same time, a powerful keyword input box was set up, supporting users to input multiple keywords for combined queries. For example, users can input targeted keyword combinations such as "tobacco price policy" to accurately locate the required public opinion information. In addition, a flexible crawling cycle selection item was equipped, allowing users to independently select diverse crawling cycles such as hourly, daily, and weekly according to the urgency of public opinion monitoring and the data update frequency of different platforms. When users complete the above settings and click the "New" button, the system will quickly call the customized crawler program according to the rules set by users to crawl data and save it to the database, and immediately start the corresponding crawler task to ensure the timeliness and accuracy of data collection.

[0067] 1.2.2 Designed the crawler rule management and maintenance interface

[0068] To greatly improve the management efficiency and convenience of crawler rules for users, a fully functional crawler maintenance interface was specially created. In this interface, users can easily query, modify, and delete crawler rules. In terms of the query function, users only need to input relevant keywords or rule names in the search box, and the system will quickly search in the database and display the matching results in a clear list form, facilitating users to quickly locate the required rules. For the rule modification operation, after users select the rule to be modified, the system will pop up a detailed editing window where users can flexibly adjust key information such as the crawled website, keywords, and crawling cycle. After the modification is completed and the "Save" button is clicked, the system will automatically update the rule information in the database and ensure that the new rule takes effect immediately, so as to ensure that the crawler task can be executed efficiently according to the new settings. When deleting a rule, users only need to select the rule to be deleted and click the "Delete" button, and the system will pop up a confirmation dialog box to prevent misoperation. Once the user confirms the deletion, the system will immediately remove the rule from the database and stop executing the corresponding crawler task, effectively releasing system resources and improving the overall operation efficiency of the system. Through the implementation of the above series of functions, the efficient management and maintenance of crawler rules have been successfully achieved, greatly enhancing the flexibility and pertinence of public opinion data collection and laying a solid foundation for subsequent public opinion analysis work.

[0069] 2. In one embodiment, the crawled tobacco public opinion information is preferentially preprocessed for data;

[0070] In the process of cleaning tobacco public opinion information, a series of efficient and accurate technical means are adopted in this embodiment. First of all, for text data, a comprehensive normalization process is carried out on it. Specifically, using Python's string processing functions, the text content is uniformly converted to lowercase. For example, using the lower() function, if there is text "Tobacco isharmful!But some people still smoke.", after executing text.lower(), it becomes "tobacco isharmful!but some people still smoke". Then, punctuation marks are removed through regular expressions. Using Python's re module, executing re.sub(r'[^\w\s]', "", text) removes all punctuation marks in the text, resulting in "tobacco is harmfulbut some people still smoke".

[0071] At this time, the MD5 (Message-Digest Algorithm 5) hash algorithm is used to calculate the hash value of each data record. If the hash values of two texts are the same after normalization processing, they will be determined as duplicate data, and the system only retains one of the records. In this way, data redundancy is effectively reduced.

[0072] After that, construct the regular expression r'[^\x00-\x7F]+' to accurately match the garbled text. In Python, the findall function of the re module can be used to find all matching strings. For example, for the text “%^&#@!This is normaltext”, executing re.findall(r'[^\x00-\x7F]+', text) will return “%^&#@!”, indicating that this string is garbled and should be removed from the dataset. Here, [^\x00-\x7F] means to match all non-ASCII characters (the ASCII character range is from \x00 to \x7F), and + means to match one or more such characters. For a combination of pure numbers, use the regular expression r'^\d+$'. In Python, to determine whether a text is pure numbers, the re.match function can be used. For example, if re.match(r'^\d+$', '123456'), the returned result is a match object, indicating that this text is pure numbers and should be filtered and deleted. For a combination of pure letters, use the regular expression r'^[a-zA-Z]+$'. For example, if re.match(r'^[a-zA-Z]+$', 'abcdef'), if the returned result is a match object, it means that this text is a combination of pure letters and has no actual semantics and can be deleted from the dataset. Through these regular expression patterns for effective screening and deletion, ensure that each record in the dataset has certain semantic value and analysis significance, laying a solid foundation for subsequent data analysis and processing.

[0073] 3. In one embodiment, further perform sentiment analysis on the preprocessed public opinion data through the Xu student model and perform traceability analysis through the knowledge graph; specifically refer to Figure 3 ;

[0074] 3.1 First, pre-train and perform knowledge distillation on the student model;

[0075] Traditional sentiment analysis methods often have difficulty accurately grasping semantics and sentiment tendencies when dealing with tobacco public opinion texts in complex contexts, and there is still much room for improvement in their accuracy and depth. Although large language models have excellent sentiment understanding capabilities, when directly applied to tobacco public opinion sentiment analysis, the high inference cost becomes a bottleneck hindering their widespread use. Small language models have low inference costs but lack in sentiment analysis capabilities.

[0076] To solve these problems, this application proposes to optimize the sentiment analysis ability by using model pre-training and knowledge distillation. Specifically, the student model is pre-trained by collecting general and tobacco-related texts, enabling the student model to accumulate language understanding ability and domain knowledge. Then, with the help of knowledge distillation, the powerful sentiment analysis ability of the teacher model (such as Llama-2) on the tobacco public opinion dataset is transferred to the student model (T5-base) to further improve its sentiment analysis performance in the tobacco domain. This method can improve accuracy while reducing costs, meeting the actual needs of tobacco public opinion monitoring.

[0077] Specifically, first, collect data from different general domains, covering various information sources such as news, novels, and blog articles. These general domain data can provide the model with rich language expression paradigms and semantic understanding bases, helping to improve the model's adaptability to different language styles and contexts. On the other hand, focus on collecting public text datasets related to tobacco public opinion, which should contain rich emotional expressions and diverse tobacco-related topics, such as text content on tobacco product evaluations, tobacco control policy discussions, and awareness of the harms of smoking.

[0078] Then load the student model (such as T5-base) and perform pre-training.

[0079] During the pre-training process of this embodiment, the general domain dataset and the tobacco public opinion-related dataset are used, and the sequence-to-sequence (Seq2Seq) training objective is adopted, aiming to convert the input text data into another more concise form of expression. By pre-training the student model, the model's text processing ability can be improved, providing a basic ability for accurately analyzing the sentiment tendency of tobacco public opinion texts. The loss function used in pre-training is:

[0080]

[0081] where x i is the input text, u is the target text, and u t is the t-th token of the target text. This loss function is used to measure the model's performance in generating the target text, prompting the model to learn the text conversion and generation ability.

[0082] During the pre-training process, the Adam optimizer is used to update the model parameters, with the initial learning rate set to lr pre = 3e -4 and the number of training epochs set to 80. In each training epoch, according to the calculated loss function value, the model parameters are updated through the backpropagation algorithm, and the backpropagation formula is:

[0083]

[0084] where L is the loss function, w is the model's parameter, and zi is the output of the intermediate layer.

[0085] To make the model training more stable, a linear decay learning rate adjustment strategy can be adopted. The learning rate adjustment formula is:

[0086] l r = lr pre ·decay(t),

[0087] where decay(t) is a function that adjusts the learning rate according to the epoch t.

[0088] After completing the pre-training of the student model (T5-base), to further improve its performance on the tobacco public opinion sentiment analysis task, we select Llama-2 as the teacher model. Through pre-training on tobacco public opinion data, the teacher model (Llama-2) can learn specific language patterns, sentiment tendencies, and related semantics in the tobacco field, so as to better understand and process tobacco public opinion texts and meet the task requirements of tobacco public opinion analysis.

[0089] After that, the knowledge distillation method is used to transfer the output of the fine-tuned Llama-2 as the knowledge source to the student model (T5-base). During the distillation process, to guide the teacher model to output more targeted information, a specific prompt is designed: "Please analyze the sentiment tendency and key information expressed in the following tobacco public opinion text, and give the basis for evaluation", so that the student model can learn the ideas and methods of the teacher model for processing tobacco public opinion texts.

[0090] During the knowledge distillation process, in order to let the student model learn the output mode of the teacher model and enable the student model to imitate the decision-making logic of the teacher model, this method uses a distillation loss function based on KL divergence to measure the difference in the output distributions of the teacher model and the student model. It quantifies the difference between models in the form of probability distribution. Compared with traditional methods, it can more essentially reflect the similarity of model outputs and provides an efficient measurement method for knowledge transfer. The distillation loss function L based on KL divergence KD can be expressed as:

[0091]

[0092] where y T is the output probability distribution of the teacher model (Llama-2), and y S is the output probability distribution of the student model (T5-base).

[0093] However, directly using the above formula may lead to an overly sharp probability distribution, making it difficult for the student model to learn effectively. To address this issue, a temperature parameter T is introduced to make the probability distribution smoother. The generated soft labels carry more information, which helps the student model better learn the distribution of the teacher model. The adjusted formula is:

[0094]

[0095] where T is set to 2 to make the soft labels smoother. The soft labels contain more information about intermediate states, enabling the student model to not only focus on the final classification results during learning but also learn more details of the teacher model's decision-making process, enhancing the effect of knowledge transfer.

[0096] Meanwhile, to balance the distillation loss and the performance of the student model on the original task, ensuring that while learning the knowledge of the teacher model, the student model does not lose its basic capabilities, a balance parameter λ = 0.5 is introduced and combined with the original cross-entropy loss L CE =-∑ i y true (i) log y pred (i). The total distillation loss function is:

[0097] L total =λL KD +(1 - λ)L CD .

[0098] It comprehensively considers both aspects of the model learning the teacher's knowledge and maintaining its own capabilities, breaking the limitations of a single loss function. Its role is to provide more comprehensive guidance for the training of the student model. By adjusting the balance parameter λ, the weight between the student model learning the teacher model's knowledge and completing the original task can be flexibly controlled, achieving a balanced development of both, effectively preventing the student model from deviating too much from the original task when learning the teacher model's knowledge, and ensuring an improvement in its comprehensive performance on the tobacco public opinion sentiment analysis task.

[0099] During the distillation training process, the Adam optimizer is used to update the parameters of the student model, with the initial learning rate set to 5e -4 and the number of training epochs set to 60. For each sample, it is input into the teacher model and the student model. The loss is calculated according to the distillation loss function, and the parameters of the student model are updated through backpropagation, making the output distribution of the student model gradually approach the output distribution of the teacher model while taking into account its own performance on the original task, achieving effective knowledge transfer.

[0100] 3.2 Tobacco Public Opinion Sentiment Analysis with the Optimized Student Model

[0101] In a preferred solution, first, the preprocessed tobacco public opinion text data is read from the database with the help of the Jieba library in Python to perform word segmentation. The word segmentation mechanism of the Jieba library is based on efficient dictionary matching and statistical machine learning algorithms. Its core lies in constructing a complete dictionary system through a large-scale corpus, which contains rich vocabulary information and corresponding frequency data. During the word segmentation process, it matches the input text with the dictionary entries according to the maximum probability principle, and separates the corresponding vocabulary from the text when the match is successful.

[0102] For word strings that do not appear in the dictionary, the hidden Markov model (HMM) statistical technique will be used to infer the possible word boundaries based on the transition probabilities between characters, so as to achieve accurate word segmentation of continuous text.

[0103] Furthermore, for the segmented words, they are input into the T5-base model after pre-training and distillation for sentiment analysis.

[0104] In a preferred embodiment, for each word, the model determines its sentiment tendency E(w ij ) = softmax(MLP(Encoder(x))), where Encoder converts the input text or word into a hidden representation, MLP maps the hidden representation to the sentiment category dimension, and the softmax function converts the output into a probability distribution. The formula is Here C = 3 (indicating 3 sentiment categories: positive, negative, neutral).

[0105] To count the total number of word segments with different sentiment tendencies in all texts, three counters are defined respectively for counting the number of word segments with positive, negative, and neutral sentiment tendencies. Their update formulas are as follows:

[0106]

[0107] At the same time, for each text T i , the student model (T5-base) model is called to perform overall sentiment analysis on it, and the sentiment tendency E(w i ) of text T ij is obtained, which also belongs to the set {positive, negative, neutral}. To count the total number of different sentiment tendencies of each text in the database, three counters are defined respectively for recording the number of texts in the database that are judged to have positive, negative, and neutral sentiment tendencies. Their update formulas are:

[0108]

[0109] After completing the statistics of the above sentiment tendencies, store these six statistical results into the corresponding fields of the database respectively for subsequent evaluation and visual display.

[0110] 3.3 Traceability Analysis Using Knowledge Graph

[0111] In the analysis and traceability of tobacco public opinion data, traditional methods are difficult to effectively capture and analyze such complex network information, resulting in low accuracy and efficiency in entity extraction and relationship reasoning. Therefore, this application introduces a graph convolutional network (GCN) and a multi-relational graph mechanism, aiming to more effectively process the complex relationships in tobacco public opinion data, improve the accuracy of entity extraction and relationship reasoning, and thus provide more powerful support for the traceability analysis of public opinion data.

[0112] First, read the stored tobacco public opinion data from the database. Use a graph convolutional network (GCN) and a multi-relational graph mechanism for entity extraction and relationship reasoning. For entity extraction, the initial feature vector of each entity node is propagated layer by layer through a graph neural network. In tobacco public opinion data, the influence degrees of different relationships on entities are different. Assume a given knowledge graph G = (V, ε), where V is the set of entity nodes and ε is the set of relationships.

[0113] node v i has an initial feature vector of At each layer of graph convolution, its feature representation needs to be updated. represents the feature vector of node v i at the (l + 1)-th layer. The specific formula for updating the node representation is as follows:

[0114]

[0115] where N(i) is the set of neighbor nodes of node i. For each neighbor node j, its feature vector at the l-th layer is , first multiply by the weight matrix W(l) for feature transformation, then divide by the normalization coefficient c i,j to balance the contributions of different neighbor nodes, and finally sum them up. is then combined with the feature vector of node i itself at the l-th layer multiplied by the weight matrix Finally, perform a non-linear transformation on the above summation result through the ReLU activation function σ to obtain the updated feature vector W(l) is used to transform the feature vectors of neighbor nodes and determines the importance degree and transformation method of neighbor node information when updating the feature of the current node; Then, it controls the role of the features of the current node itself during the update process, and the two jointly affect the direction and degree of node feature update.

[0116] This method continuously updates the feature representation of nodes by layer-by-layer propagation and fusion of the information of neighbor nodes and itself, provides richer and more accurate node features for entity extraction, and enables the model to better distinguish different entities.

[0117] This method takes into account the different contributions of neighbor nodes to the feature update of the current node under different relationships, and flexibly adjusts the information fusion method. The neighbor nodes of different relationships have different weights for node feature update. Thus, it more accurately reflects the influence of different relationships on node features, enables node features to incorporate multiple relationship information, and provides richer semantic information for more accurate entity extraction and subsequent sentiment analysis and traceability analysis. After L-layer graph convolution, the final embedding representation of node i is obtained.

[0118] It is input into a linear classifier. First, it is multiplied by the weight matrix W. Then, the bias b is added. Finally, it is processed by the softmax function. The softmax function will convert the input values into a probability distribution, thereby obtaining the probability corresponding to each category, and the category with the highest probability is the predicted entity category. c It is multiplied by the weight matrix W. c Then, the bias b is added.

[0119] Among them, W. c is the weight matrix of the classification layer, which determines the mapping relationship between the final embedding representation. and each entity category. b. c is the bias term of the classification layer, which increases the flexibility of the model, helps the model better fit the data, and adjusts the classification boundary. This classifier classifies based on the rich node features obtained from the graph convolutional network. Compared with traditional classification methods, it can better utilize the structural information in the knowledge graph, improve the accuracy of entity classification, and more accurately identify different types of entities such as tobacco brands, related figures, and events.

[0120] In terms of relationship reasoning, a bilinear decoder based on graph convolution is used to predict the relationship between entities. For two entities v. i and v. j Their embedding representations after L-layer graph convolution are respectively. and. These two vectors are processed. First, is transposed and multiplied by the bilinear matrix R. r Then, it is multiplied by Multiply them, and finally map them to the interval [0, 1] through the sigmoid function σ to obtain the entity v i and v j The probability that there is a relationship r between them. The prediction formula for the relationship r using a bilinear decoder is:

[0121]

[0122] where R r is the bilinear matrix corresponding to the relationship r. It captures the complex semantic information of the relationships between different entities and determines the possibility of a specific relationship existing between entity pairs. σ is the sigmoid function that maps the calculation result to the interval [0, 1], giving it a probabilistic meaning and facilitating the judgment of the likelihood of the existence of a relationship.

[0123] In this method, the relationship features between entity embedding representations are mined through the bilinear matrix R r Compared with traditional methods, it can more accurately predict the relationships in complex knowledge graphs. In tobacco public opinion analysis, it can more precisely determine whether the relationship between tobacco companies and retailers is a "sales" relationship or other relationships, improving the accuracy of relationship reasoning.

[0124] Meanwhile, a contrastive loss function L based on negative sampling is introduced:

[0125]

[0126] where the first half represents the entity pair (v i , v j ) and its relationship r that truly exist in the knowledge graph. is the probability that the model predicts the existence of the relationship r for this entity pair. Taking the logarithm of this probability, summing them up, and adding a negative sign means that the model should maximize the prediction probability of the truly existing relationship during training, that is, make the model's prediction of the true relationship as accurate as possible. The second half represents the entity pairs and their relationships that do not exist in the knowledge graph (false samples obtained through negative sampling). β is a hyperparameter that controls the balance between positive and negative samples. is the probability that the model predicts the non-existence of a relationship. Similarly, taking the logarithm and summing them up, this part aims to make the model minimize the prediction probability of false relationships.

[0127] Furthermore, introducing a contrastive loss function based on negative sampling enables the model to effectively distinguish between correct and incorrect entity pairs and their relationships. In an actual knowledge graph, the number of positive samples (true relationships) is often much smaller than the number of possible negative samples (false relationships). If trained only based on positive samples, the model is prone to overfitting and difficult to perform well on unseen data. By introducing the incorrect samples obtained through negative sampling and using this contrastive loss function, the model can learn the patterns of false relationships and avoid misjudging false relationships as true relationships.

[0128] This application breaks the traditional mode that only relies on positive sample training. By leveraging the ideas of negative sampling and contrastive learning, it enriches the training information of the model, enabling the model to not only learn the characteristics of true relationships but also learn how to identify false relationships, thereby enhancing the accuracy and reliability of entity extraction and relationship reasoning in a complex knowledge graph environment and better coping with the challenges of data sparsity and complex relationships.

[0129] To achieve efficient storage and management of complex entities and their relationships, this application selects Neo4j as the graph database to construct a complete tobacco public opinion knowledge graph. Neo4j is based on a graph data model, representing entities as nodes and relationships between entities as edges. This data structure is naturally suitable for processing data with complex association relationships. In Neo4j, various tobacco public opinion-related entities, such as tobacco brands, relevant figures, events, etc., are stored in the form of independent nodes respectively. Each node has a unique identifier and is accompanied by key-value pairs describing the attributes of the entity. For example, a tobacco brand node may contain attributes such as brand name, founding time, and affiliated company. The tobacco public opinion data processed by a graph neural network or extracted by other means is imported according to the data format requirements of Neo4j. The batch import tool of Neo4j or a custom import script can be used to ensure that the data can be accurately and quickly stored in the database. At the same time, as new tobacco public opinion data is generated, the information in the graph database needs to be updated in a timely manner to ensure the timeliness and integrity of the knowledge graph.

[0130] The relationships between entities are connected in the form of edges. In order to fully express the rich semantic information of the relationship, the attributes of the edge are refined. The relationship type clearly records the nature of the association between entities, such as "production", "sales", "supervision", etc., providing clear guidance for understanding the business logic in public opinion. Strength information quantifies the closeness or influence of the relationship in a numerical way, for example, using a floating point number between 0 and 1. The closer the value is to 1, the closer the relationship is. In the relationship of "a tobacco company establishes long-term cooperation with a large retailer", if the two parties cooperate frequently and the business volume is large, the strength value can be set to 0.8; if the cooperation is relatively less and the business correlation is low, the strength value is set to 0.3. The time attribute accurately records the time point when the event occurs or the relationship is generated, and is stored in a standard time format (such as YYYY-MM-DDHH:MM:SS), providing key time clues for subsequent public opinion analysis and traceability based on time series. For example, when recording the relationship "A tobacco company released a new product on May 1, 2023", the time attribute is accurately recorded as "2023-05-01 00:00:00".

[0131] 4. In one embodiment, public opinion report generation and real-time early warning analysis are realized;

[0132] 4.1 For report generation, this application designs a structured public opinion report template in HTML format. In the public opinion overview section, the activity level of public opinion is measured by counting the total amount of public opinion data within a certain time span. For example, by counting the total number of public opinion information released in the past month, if the data volume reaches N pieces, it indicates that the tobacco public opinion is in a relatively active state during this period. At the same time, combined with the proportion data of sentiment analysis, the main sentiment tendency distribution is elaborated. In the sentiment analysis details section, visualization charts are generated with the help of the Matplotlib library for auxiliary explanation. When generating a bar chart comparing the number of comments of different sentiments, the sentiment type is used as the abscissa and the number of comments as the ordinate, and the height of each bar corresponds to the number of comments of the corresponding sentiment. Different sentiment categories are distinguished by different colors to visually present the quantity difference. For the change of sentiment tendency over time, a line chart is drawn with time as the abscissa and the sentiment proportion as the ordinate. Through the ups and downs of the line, the fluctuations of each sentiment tendency in the time series are clearly shown. In the analysis of the dissemination path, with the help of the visualization results of the Neo4j knowledge graph, the dissemination trajectory and key nodes of public opinion are shown. The nodes in the graph represent the main bodies participating in the dissemination of public opinion, such as social media accounts, news media organizations, opinion leaders, etc.; the edges represent the dissemination relationships between the main bodies. The visualized picture of the graph is embedded in the report, and important dissemination nodes are marked and explained. In the part of the association of key entities, the mutual influence relationships between the main tobacco brands, relevant figures and events are clearly presented in tabular form. The nature of each relationship is described in detail. For example, there may be an endorsement relationship between a brand and a person, and there may be a causal relationship between a brand and an event. For example, an event of a brand's product quality problem leads to damage to the brand image. For the degree of influence, it is quantitatively explained through the association strength index. If the association strength is 0.8 (with a full score of 1), it indicates that the relationship between the two is close and the event has a relatively significant impact on the brand;

[0133] Preferably, the above text description is called on the LLaMA-2 model to summarize and polish the text description in the report. At the same time, combined with prompt engineering technology, appropriate prompts are designed, such as "Please optimize the following text description about tobacco public opinion to make it more logical and readable: [original text content]", taking the text paragraphs in the report as input, obtaining the optimized output result, and storing the final report in the database.

[0134] 4.2 Public Opinion Early Warning and Prompt

[0135] This application first extracts key features from the public opinion text, such as the keyword occurrence frequency, theme category (predetermined using text classification algorithms), etc. For example, count the occurrence frequencies of keywords such as "hazards of smoking", "tobacco control policies", "tobacco prices" in the text at different time points as important elements of the feature vector. Let the number of occurrences of keyword i at time t be n i,tDetermine the topic category using text classification algorithms. Taking SVM as an example, map the text feature vector x to the topic category set C through f(x). Construct features based on time series and calculate the change rate of sentiment tendency. Growth rate of topic popularity Expansion rate of dissemination scope Where S t , H t , R t Are the sentiment tendency value, topic popularity value, and dissemination scope index at time t respectively. Divide the prepared dataset into a training set, a validation set, and a test set according to a certain ratio. For example, 70% is used for training, 20% for validation, and 10% for testing.

[0136] After that, construct a long short-term memory network (LSTM) model. Its network structure includes an input layer, a hidden layer, and an output layer. The input layer receives the time series data processed by feature engineering. The hidden layer consists of multiple LSTM units, which can effectively handle the long-term dependencies in the time series. In each LSTM unit, the sigmoid function is used as the activation function. The expression of the sigmoid function is It maps the input value to between 0 and 1 to control the flow of information and the gating mechanism. The output layer is set according to the prediction target of the model. If it is to predict the probability of an opinion crisis occurring, the output layer is a neuron, using the sigmoid activation function to map the output value to between 0 and 1, representing the probability value.

[0137] Divide the time series data at a fixed time step T. Use the data of the past T time points as the input sequence X to predict the opinion index y at the next time point. t+1 . During training, use the Adam optimization algorithm to update the weights W and biases b, and calculate and update the parameter θ according to m t , v t etc. Use the mean squared error loss as the objective function, and adopt the early stopping method to prevent overfitting. Stop training when the performance of the validation set does not improve for multiple consecutive epochs.

[0138] Set the threshold for opinion crisis early warning according to the performance of the trained time series prediction model on the test set and historical experience. For example, if the predicted probability of an opinion crisis occurring by the model exceeds 0.7 or the sentiment tendency changes sharply to the negative (such as the growth rate of the negative sentiment tendency exceeds a certain threshold), it is considered that a potential opinion crisis may occur.

[0139] The system retrieves the latest tobacco public opinion data in real time and processes it according to the previous feature engineering and data processing procedures. The processed feature data is input into the trained time series prediction model for prediction. The prediction result is compared with the set warning threshold. If it exceeds the threshold, the warning mechanism is immediately triggered. A prominent warning message is displayed on the system interface, with a pop-up window prompting "Potential public opinion crisis, please pay attention in time!", and at the same time, the display color of relevant data is changed (such as marking the data related to the warning public opinion as red) so that users can quickly identify and pay attention to the warning situation. When the public opinion warning is triggered, the system uses the smtplib library and email library in Python to construct the email content. The email subject is set to "Tobacco Public Opinion Warning Notice", and the email body is written in HTML format, including detailed information about the warning, such as an overview of the public opinion event that triggered the warning (generating a concise description using a text template and relevant data), the current prediction probability (clearly displayed in numerical form), and key analysis data (such as the sentiment tendency distribution shown in a chart or text form, a list of hot topics). The smtplib library is used to connect to the email server, and after login verification, the email is sent to the administrator's email address to ensure that the administrator can receive the warning information in time and take corresponding measures. At the same time, after the email is sent successfully or failed, the system records the sending result in a log file for subsequent viewing and troubleshooting.

[0140] 5. In one embodiment, the public opinion analysis data is comprehensively visualized and displayed.

[0141] In this application, as Figure 4 , the visual display includes three parts:

[0142] 5.1 Visual display of sentiment analysis statistics

[0143] First, read the sentiment tendency statistical results stored in the previous sentiment analysis from the database. For the number of word segments with different sentiment tendencies, use SQL query statements in MySQL to extract the required information, using SELECT N_word_positive,N_word_negative,N_word_neutral FROM sentiment_word_statistics; obtain relevant data from the table storing the word segment statistical data For the number of texts with different sentiment tendencies, use a similar query statement to extract from the corresponding table Such as SELECT N_text_positive,N_text_negative,N_text_neutral FROM sentiment_text_statistics.

[0144] The statistically analyzed public opinion data is visually presented in various chart forms using the Matplotlib library in Python. A bar chart is used to show the quantity comparison of comments in different sentiment categories. The height of the bars represents the number of comments, and different colors distinguish positive, neutral, and negative sentiments, enabling users to intuitively understand the sentiment tendency distribution of the public opinion. A pie chart is used to display the proportion of different topics in the public opinion, visually presenting the relative importance of each topic. It shows the sentiment tendency distribution, development trend, and topic importance of the tobacco public opinion from different dimensions, providing users with comprehensive and intuitive public opinion information and helping to deeply analyze the characteristics and development trend of the tobacco public opinion.

[0145] 5.2 Visualization Display of the Traceability of Public Opinion Data

[0146] In the traceability analysis stage of the tobacco public opinion data, based on the constructed knowledge graph, the traceability graph is displayed with the help of the powerful visualization interface of Neo4j. Users have great operational convenience in this process. They can quickly filter out nodes with positive, neutral, or negative evaluations according to the sentiment tendency attributes of the nodes through simple filtering operations. This filtering function enables users to focus on information with specific sentiment tendencies, providing convenience for further analysis. For each filtered node, using the efficient query and backtracking functions provided by Neo4j, users can dig deeper into the source information by tracing its associated edges. Taking a negative evaluation node as an example, assuming that this node was posted by a certain consumer on social media, when the user clicks on this node, they can perform a backtracking query along a relationship edge such as "publisher". Through this query method, users can quickly locate the account information of this consumer and the original text content they posted. This process not only helps users efficiently trace the source of negative evaluations but also enables them to comprehensively, deeply, and clearly understand the origin and dissemination context of the public opinion. Through this traceability visualization display, users can deeply understand the complex relationship network of the tobacco public opinion from multiple dimensions, including entity attributes, relationship types, relationship strengths, and time sequences, providing strong support for accurate public opinion response and management. It provides users with an intuitive and powerful tool to help them sort out a clear context from complex tobacco public opinion information, thereby better grasping the dynamic changes and internal connections of the public opinion and providing a scientific basis and guidance for subsequent decision-making and management work.

[0147] 5.3 Visualization Display of Public Opinion Early Warning

[0148] The system timely obtains the current early warning probability value and related early warning keywords from the backend program and displays these key information on the front-end interface. For the early warning probability value, it reflects the likelihood of the current public opinion crisis occurring; while the early warning keywords help users quickly locate the key factors triggering the public opinion. When the prediction result exceeds the set threshold, the early warning mechanism is immediately triggered.

[0149] On the system interface, to enable users to quickly detect potential public opinion crises, prominent warning messages will be displayed, prompting "Potential public opinion crisis, please pay attention in time!" in the form of a pop-up window. This pop-up window has a high visual salience to ensure that it can attract the user's attention at the first time. At the same time, the system will change the display color of relevant data. For example, the data related to the early warning public opinion will be marked in red. Through this intuitive color change, users can quickly identify and pay attention to the early warning situation. This visual design not only enhances the information transmission efficiency but also enables users to make a quick response.

[0150] The advantages of this application compared with the prior art include:

[0151] 1. Create a customized public opinion data collection system. For different platforms such as Weibo, WeChat, and news websites, use the Scrapy framework and Selenium library of Python to develop crawler programs that can adapt to the characteristics of each platform, and accurately capture information such as main posts, replies, and comments. The front-end interface design is extremely user-friendly. Users can select the target website through a drop-down box, input multiple keyword combinations, and flexibly set the crawling period. At the same time, the crawler maintenance interface supports querying, modifying, and deleting rules, providing a rich and high-quality data basis for subsequent analysis.

[0152] Traditional public opinion analysis methods are simple and have limitations. The sentiment analysis method based on large models will incur high inference costs in deployment. This method innovatively uses student model pre-training and knowledge distillation to optimize the sentiment analysis ability. The final small model can greatly reduce the inference cost in deployment. This method collects multi-source texts to pre-train the T5-base model, accumulates language and tobacco knowledge, uses Llama-2 as the teacher model, sets prompt words, and uses a unique function based on KL divergence and cross-entropy loss to balance learning and its own ability. After testing, the accuracy rate of public opinion sentiment recognition reaches 89.6%.

[0153] 2. Traditional distillation loss functions are insufficient in knowledge transfer and performance balance. This patent designs a loss function that combines the distillation loss function based on KL divergence and the cross-entropy loss. The distillation loss function based on KL divergence can measure the difference in the output probability distributions of the teacher model (Llama-2) and the student model (T5-base), prompting the student model to imitate the decision-making logic of the teacher model. Introduce a temperature parameter to smooth the probability distribution, enabling the student model to better learn the distribution of the teacher model. At the same time, combine the cross-entropy loss and a balance parameter to ensure that the student model does not lose its own ability in the original task while learning the knowledge of the teacher model.

[0154] 3. Traditional public opinion traceability analysis methods are difficult to handle the complex entity relationships and structural information in tobacco public opinion. This patent conducts traceability analysis on public opinion data based on graph convolutional networks. By introducing graph convolutional networks (GCN) and multi-relational graph mechanisms, the characteristics of entity nodes are propagated and updated layer by layer. Different relationships have different weights for node feature update, which can more accurately reflect the complex relationships between entities. A bilinear decoder based on graph convolution is used to predict entity relationships, and a contrastive loss function based on negative sampling is combined to enhance the generalization ability of the model. At the same time, a knowledge graph is constructed with Neo4j to accurately capture dependency relationships and mine propagation paths. Combined with time series prediction and early warning, the actual early warning success rate is 98.6%, providing strong support for the decision-making of relevant departments.

[0155] The various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.

[0156] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A tobacco public opinion intelligent monitoring and analysis method based on knowledge distillation, characterized in that: Determine the crawler website and crawler rules to obtain tobacco public opinion data; Segment the tobacco public opinion data; Use the student model to predict sentiment based on the word segmentation results; The student model performs knowledge distillation on the tobacco public opinion text analysis task through the teacher model, and measures the prediction difference through the following loss function: Where, L KD represents the KL divergence distillation loss, y T is the output probability distribution of the teacher model, y S is the output probability distribution of the student model, T represents the temperature parameter, and i represents the sample index.

2. The tobacco public opinion intelligent monitoring and analysis method according to claim 1 is characterized in that: Priority should be given to preprocessing tobacco public opinion data, including: converting to lowercase, removing punctuation, and removing redundancy through hash values.

3. The tobacco public opinion intelligent monitoring and analysis method according to claim 1 is characterized in that: The tobacco public opinion data is segmented, including dictionary matching after segmenting the tobacco public opinion data, and determining the vocabulary boundaries of unmatched word strings based on the transition probability between characters using a hidden Markov model.

4. The tobacco public opinion intelligent monitoring and analysis method according to claim 1 is characterized in that: Considering the cross entropy loss to measure the prediction difference, the total loss function expression is: THE total =λL KD +(1-λ)L CD In the formula, λ represents the equilibrium parameter, L CD Represents the cross entropy loss, expressed as: In the formula, y true (i) represents the true label of the i-th sample of the student model, y pred (i) The predicted label of the i-th sample of the student model.

5. The tobacco public opinion intelligent monitoring and analysis method according to claim 1 is characterized in that: Using the student model to predict sentiment based on the word segmentation results includes: Input each word segment into the student model for sentiment prediction and count the number of sentiment tendencies in different categories; The original text corresponding to the word segmentation is input into the student model for sentiment prediction, and the number of sentiment tendencies in different categories is counted.

6. The tobacco public opinion intelligent monitoring and analysis method according to claim 1 is characterized in that: According to the prediction results, source tracing analysis is performed through a knowledge graph; the knowledge graph is constructed based on tobacco public opinion data.

7. The tobacco public opinion intelligent monitoring and analysis method according to claim 6 is characterized in that: The knowledge graph predicts entity categories through the following steps: Obtain the entity nodes in the tobacco public opinion data, propagate them layer by layer through the graph convolutional network, and obtain the node embedding representation; during the propagation, the node update formula is: In the formula, represents the node feature vector of the (l+1)th layer, σ represents the activation function, N(i) represents the set of neighbor nodes of node i, c i,j represents the normalization coefficient, W(l) represents the feature transformation matrix, represents the weight matrix; The node embedding representation is input into the linear classifier to obtain the entity category.

8. The tobacco public opinion intelligent monitoring and analysis method according to claim 6 is characterized in that: The knowledge graph uses a bilinear decoder to predict entity relationships, and the prediction formula is: In the formula, P represents and There is a relationship ij The probability of, σ represents the activation function, R r Represents the bilinear matrix corresponding to the relation r, represents the embedding representation of node i, represents the embedding representation of node j.

9. The tobacco public opinion intelligent monitoring and analysis method according to claim 8 is characterized in that: The bilinear decoder adopts a contrastive loss function based on negative sampling: In the formula, β represents the hyperparameter that controls the balance between positive and negative samples.

10. The tobacco public opinion intelligent monitoring and analysis method according to claim 1 is characterized in that: Generate public opinion reports based on sentiment prediction results, and provide early warnings and visual displays.

Citation Information

Cited By

  • Emotion analysis method based on emergency scene

    CN120804331A