WeChat official account sensitive information analysis method and system
Through the WeChat public account sensitive information analysis methods, including data collection, cleaning, analysis and visual processing, the machine learning algorithm is used to realize data automation and intelligent processing, and the existing system has been solved, and more intuitive intelligent suggestions and optimization strategies are provided, which improves analysis efficiency and accuracy.
Patent Information
- Application Number
- CN202411965736.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-27
AI Technical Summary
The existing WeChat public account analysis system has poor visualization, which is difficult to provide companies with intuitive intelligent suggestions and optimization strategies, and has a low degree of automation.
Through the sensitive information analysis methods of WeChat public accounts, including data collection, cleaning, analysis and visual processing, machine learning algorithms are used to realize data automation and intelligent processing, and intuitive graph display and intelligent suggestions are provided.
It improves the efficiency and accuracy of data analysis, provides enterprises with more intuitive intelligent suggestions and optimization strategies, enhances the degree of automation, reduces manual intervention, and improves the scientificity and accuracy of decision-making.
Smart Images

Figure CN120045691A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information extraction, and more particularly, to a method and system for analyzing sensitive information of WeChat official accounts. Background Art
[0002] With the continuous progress of Internet technology, the WeChat official account, as a new type of social media platform, has become an important channel for people to obtain information and communicate. The amount of information on WeChat official accounts is huge, and it has the characteristics of real-time and interactivity. Analyzing and researching WeChat official accounts can help better understand user needs and market trends.
[0003] With the advent of the big data era, data has become an important resource for the development of enterprises and society. WeChat official account analysis systems need to process a large amount of data, including articles and keywords, and generally use big data processing technologies such as distributed computing and data mining to improve the efficiency and accuracy of data processing and analysis.
[0004] In recent years, scholars generally believe that the analysis of WeChat official account data is an important means to understand user needs and market trends. By collecting, cleaning, and analyzing WeChat official account data, useful information can be extracted to provide decision-making support for relevant personnel. In addition, data mining of WeChat official accounts can help enterprises better understand user behavior and needs and discover potential market opportunities. By mining and analyzing WeChat official account data, features such as user portraits, content types, and hot topics can be extracted to provide strong support for enterprises to formulate marketing strategies. Scholars emphasize that the accuracy of WeChat official account information extraction is a key factor affecting the analysis results. Therefore, advanced natural language processing technologies are needed to accurately extract information such as the theme, keywords, and sentiment tendency of articles to provide more in-depth understanding and analysis for relevant personnel.
[0005] Regarding the application of this technology to the research of WeChat official account analysis systems today, scholars such as Wang Xiaolong proposed a text classification model for WeChat official accounts based on deep learning. This model uses deep learning technologies such as convolutional neural networks and long short-term memory networks to classify and analyze WeChat official account texts, improving the accuracy and efficiency of text classification.
[0006] Scholars such as Zhao Yong proposed a sentiment analysis model for WeChat official accounts based on transfer learning. This model uses an existing sentiment analysis dataset for pre-training and then transfers the pre-trained model to the WeChat official account sentiment analysis task, improving the accuracy and efficiency of sentiment analysis.
[0007] Scholars such as Zhang Yong proposed a topic mining model for WeChat official accounts based on topic models. This model uses topic model technologies such as Latent Dirichlet Allocation to perform topic modeling and topic mining on the text of WeChat official accounts, helping enterprises discover potential market opportunities and user needs.
[0008] However, the existing WeChat official account analysis systems mainly focus on the classification and analysis of text, with relatively poor visualization, making it difficult to provide intuitive intelligent suggestions and optimization strategies for enterprises, and there is also the problem of low automation. Summary of the Invention
[0009] The technical problem to be solved by the present invention is:
[0010] To solve the problems that the existing WeChat official account analysis systems have relatively poor visualization, making it difficult to provide intuitive intelligent suggestions and optimization strategies for enterprises, and there is also the problem of low automation.
[0011] The technical solution adopted by the present invention to solve the above technical problems:
[0012] The present invention provides a method for analyzing sensitive information of WeChat official accounts, including the following steps:
[0013] S100. Collection of official account data, directly obtaining data through the WeChat official account background or using third-party tools for data collection. The official account data includes data details of multiple articles;
[0014] S200. Cleaning and preprocessing the official account data collected in step S100, including removing outliers and duplicate values, unifying data formats, standardizing field names, identifying and processing missing values, removing stop words, and performing text conversion;
[0015] S300. Analyzing the data cleaned and processed in step S200, including content label extraction, article quantity analysis, and article content analysis;
[0016] S400. Visualizing the data analyzed in step S300, and visually presenting the analysis results to users in the form of charts;
[0017] S500. Automatically and intelligently processing the official account data in real time through machine learning algorithms, for regularly automatically performing data analysis tasks and timely discovering abnormal situations.
[0018] Furthermore, in step S100, when directly obtaining data through the WeChat official account background, relevant data can be directly obtained through the API interface or data export function provided by the WeChat official account background. The specific operation steps are as follows:
[0019] S110. Log in to the WeChat official account background and enter the management interface;
[0020] S120. Locate the entrance of data export or API interface, and obtain the relevant API key or login information;
[0021] S130. Use the API key or login information to write corresponding code through a programming language to call the API interface or data export function of the WeChat official account background;
[0022] S140. Extract the required data according to the return result of the API interface or data export function.
[0023] Further, in step S100, when using a third-party tool for data collection, the specific operation steps are as follows:
[0024] (1). Select a data collection software or crawler program;
[0025] (2). Write a collection script or rule according to the data structure and pattern of the WeChat official account;
[0026] (3). Regularly or real-time collect data from the WeChat official account and save the collected data to local or cloud storage.
[0027] Further, in step S200, when performing data cleaning and processing, identify and remove outliers beyond the normal range by setting thresholds or using statistical methods, including
[0028] Identify and delete duplicate data. Use a data cleaning tool or write a script to identify duplicate entries in the dataset and delete such data, only retaining unique data records;
[0029] Unify the data format, making data from different sources or formats follow a unified format standard;
[0030] Standardize field naming, standardize words with the same meaning;
[0031] Set thresholds and handle outliers. For numerical fields, set a reasonable threshold range, consider values beyond this range as outliers, and delete, replace with the average or median for the identified outliers;
[0032] Identify and handle missing values. Check whether there are missing values in the dataset. If there are missing values, delete the records containing missing values, or use interpolation, mean filling methods to estimate and fill these missing values;
[0033] Remove stop words, delete the stop words in the article content;
[0034] Text conversion, including case conversion and traditional / simplified Chinese conversion.
[0035] Furthermore, in step S200, corresponding code is written in Python during text preprocessing, specifically including:
[0036] S210: Extract Chinese, English, and numbers, use regular expression re.compile(u"[^a-zA-Z0-9\u4e00-lu9fa5]") to remove emojis and special symbols, and only keep letters, numbers, and Chinese characters; where \u4e00-\u9fa5 is the Unicode Chinese character encoding range.
[0037] S220: Expand the WeChat exclusive word segmentation dictionary and stop words. First, use the official account nickname as part of the word segmentation dictionary. Second, add the words that often appear in the official account to the word segmentation dictionary. Finally, view and manually select the title, abstract, and short text body as the word segmentation dictionary and stop word dictionary.
[0038] S230: Load the word segmentation dictionary expanded in step S220, use jieba for word segmentation, use the accurate mode, and call jieba.enable_parallel() for parallel word segmentation; use the jieba.posseg module for part-of-speech tagging, and only keep nouns and verbs during the part-of-speech tagging stage.
[0039] S240: Remove stop words, including removing words with no clear meaning such as modal particles, adverbs, prepositions, and conjunctions.
[0040] Furthermore, in step S300, the content tags are extracted by text analysis of the article content to extract article tags; the article quantity analysis is to count and analyze the number of articles, including the change in the number of articles in different time periods and the distribution of the number of articles on different topics or keywords; the article content analysis is to perform text analysis on the user interaction data of the article, and the user interaction data includes comments, messages, or other interaction data after the article, and extract the user's emotional tendency and opinion feedback information.
[0041] Furthermore, in step S400, select a visualization tool according to the requirements during visualization processing, and use the visualization tool to create corresponding charts according to the analysis results and requirements. The charts include line charts, bar charts, pie charts, and scatter plots.
[0042] Furthermore, in step S500, the machine learning algorithms used are decision tree algorithm, Adaboost algorithm, Bagging algorithm, or support vector machine algorithm.
[0043] A WeChat official account sensitive information analysis system, which has program modules corresponding to the above steps and executes the steps in the above WeChat official account sensitive information analysis method when running.
[0044] A computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the WeChat official account sensitive information analysis method when called by a processor.
[0045] Compared with the prior art, the beneficial effects of the present invention are:
[0046] A WeChat official account sensitive information analysis method and system of the present invention uses a visualization tool to display the analysis results in the form of charts, and at the same time introduces automation tools and artificial intelligence technologies to realize the automation and intelligence of data collection, cleaning, and analysis, and provides intelligent suggestions and optimization strategies for the operation of the official account. Through the visualization display, users can more intuitively understand the data situation and trends of the WeChat official account, so as to better formulate marketing strategies and adjust product functions and other decisions. At the same time, automation and intelligent technologies can improve the efficiency and accuracy of data analysis and provide more accurate and comprehensive decision-making support for the operation of the official account. Among them, to improve decision-making efficiency and accuracy, through automated data collection, cleaning, and analysis, the system can quickly process a large amount of data and reduce manual intervention. This not only improves work efficiency but also reduces the incidence of human errors, making decision-making more scientific and accurate. Real-time monitoring and feedback, the system can achieve real-time monitoring of the operation data of the WeChat official account, discover problems in a timely manner and make adjustments. This rapid feedback mechanism enables the operation strategy to be dynamically adjusted according to the latest data changes, so as to better meet market demands. In-depth user analysis, using machine learning algorithms, can deeply analyze user behavior and identify potential user needs and preferences. This kind of analysis can help the operator of the official account formulate more targeted content and marketing strategies, and improve user stickiness and satisfaction. The automated and intelligent improvement algorithms support natural language processing (NLP) technology. The system can automatically analyze the content of articles and identify potential sensitive information and keywords. This automated processing not only improves the efficiency of sensitive information detection but also reduces the workload of manual review. Machine learning and predictive analysis use machine learning algorithms, and the system can identify patterns and trends in the data, so as to conduct predictive analysis. For example, it can predict the popularity of a certain type of content among specific user groups to help the operator optimize the content release strategy. The anomaly detection algorithm is to implement an anomaly detection algorithm, and the system can monitor abnormal situations in the data stream in real time, such as a sudden increase in user complaints or a decrease in interactions. This early warning mechanism enables the operator to take measures in a timely manner to avoid potential crises. Description of the Drawings
[0047] Figure 1It is a flowchart of a method for analyzing sensitive information of WeChat official accounts in an embodiment of the present invention;
[0048] Figure 2 It is a flowchart of text preprocessing using Python in an embodiment of the present invention;
[0049] Figure 3 It is a flowchart of the Bagging algorithm in an embodiment of the present invention;
[0050] Figure 4 It is the main data interface of the data collection software in an embodiment of the present invention;
[0051] Figure 5 It is a label statistical summary chart in an embodiment of the present invention. Detailed implementation manners
[0052] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the specific embodiments of the present invention in detail with reference to the accompanying drawings.
[0053] Specific implementation manner one: Combining Figure 1 and Figure 2 As shown, the present invention provides a method for analyzing sensitive information of WeChat official accounts, including the following steps:
[0054] S100. Collect official account data. The official account data includes data details of multiple articles, and the data details include article data, content data, abstract, title, author, and publication time;
[0055] When collecting official account data, it includes directly obtaining data through the WeChat official account background or using third-party tools for data collection. Among them,
[0056] When directly obtaining data through the WeChat official account background, relevant data can be directly obtained through the API interface or data export function provided by the WeChat official account background. The specific operation steps are as follows:
[0057] S110. Log in to the WeChat official account background and enter the management interface;
[0058] S120. Find the entry for data export or API interface, and obtain the relevant API key or login information for subsequent API calls;
[0059] S130. Use the API key or login information to write corresponding code through a programming language (such as Python, Java) to call the API interface or data export function of the WeChat official account background; all API calls require a valid AccessToken;
[0060] Write code in Python to call the API and extract the required data; the following is an example code for obtaining a list of articles published by a WeChat official account:
[0061]
[0062]
[0063] S140. Extract the required data according to the return result of the API interface or data export function;
[0064] If the WeChat official account background does not provide direct data export or API interface, third-party tools can be used for data collection. The specific operation steps are as follows:
[0065] (1) Select a suitable third-party tool, combined with Figure 4 as shown, such as a data collection software or a crawler program;
[0066] (2) According to the data structure and rules of the WeChat official account, write the corresponding collection script or rules, including:
[0067]
[0068] # Configure necessary parameters
[0069] COOKIES = "your_cookies_here" # Replace with the actual cookies
[0070] TOKEN = "your_token_here" # Replace with the actual token
[0071] FAKEID = "your_fakeid_here" # Replace with the fakeid of the official account
[0072] HEADERS = {
[0073] "Cookie": COOKIES,
[0074] "User-Agent": "Mozilla / 5.0 (Windows NT 10.0; Win64; x64) AppleWebKit / 537.36 (KHTML, like Gecko) Chrome / 91.0.4472.124 Safari / 537.36"
[0075] }
[0076] # Define the function to obtain the article list
[0077] def get_articles(offset):
[0078] url = f"https: / / mp.weixin.qq.com / cgi-bin / appmsg"
[0079] params = {
[0080] "action": "list_ex",
[0081] "begin": offset,
[0082] "count": 5, # Each time, get 5 articles, can be adjusted according to needs
[0083] "fakeid": FAKEID,
[0084] "type": 9,
[0085] "token": TOKEN,
[0086] "lang": "zh_CN",
[0087] "f": "json"
[0088] }
[0089] response = requests.get(url, headers = HEADERS, params = params)
[0090] data = response.json()
[0091] if data.get("base_resp", {}).get("ret") == 0:
[0092] return data.get("app_msg_list", [])
[0093] else:
[0094] print("Error fetching articles:", data.get("base_resp", {}).get("errmsg"))
[0095]
[0096] main()
[0097] (3) Collect data from WeChat official accounts regularly or in real time, and save the collected data to local or cloud storage;
[0098] S200. Clean and preprocess the official account data collected in step S100, remove outliers and duplicate values, to ensure the accuracy and reliability of the data, and thus improve the analysis efficiency for the next step of data analysis;
[0099] When performing data cleaning and processing, identify and remove outliers outside the normal range by setting thresholds or using statistical methods, to avoid negative impacts on the analysis results, thereby improving the accuracy and reliability of the data. Specifically, it includes:
[0100] Identify and delete duplicate data. Use data cleaning tools or write scripts to identify duplicate entries in the dataset according to certain rules. The rules can be the article title or the publication time. After identifying duplicate data, delete such data and only retain unique data records;
[0101] Unify the data format. For data from different sources or in different formats, such as dates and numbers, perform formatting processing to make it follow a unified format standard;
[0102] Standardize field naming. To ensure that the field naming in the dataset is unified and easy to understand, avoid using ambiguous or inconsistent terms, and standardize words with the same meaning; the unified naming rule is that all field names should follow a consistent naming rule. For example, use lowercase letters and underscores to separate words (such as article_title), and avoid using ambiguous or inconsistent terms; for fields that need code conversion, a specific prefix can be added before the original field name (such as c_ represents the standard code), while for extended added attribute fields, a suffix can be added after the original name (such as mc represents the name attribute); make the field names concise and clear, that is, the field names should be short but can accurately express their meanings, and avoid using overly complex or professional terms for downstream users to understand; type identification can be added to the field names. For example, date fields can start with date_, and numeric fields can start with num_, to quickly identify the data type;
[0103] Set thresholds and handle outliers. For some numeric fields, set a reasonable threshold range, consider values outside this range as outliers, and delete, replace with the average value, median, or perform other appropriate processing according to the specific situation for the identified outliers; the above criteria for setting thresholds can adopt the quartile method, and its basic steps are as follows:
[0104] Calculate the first quartile (Q1): the 25% quantile of the dataset;
[0105] Calculate the third quartile (Q3): the 75th percentile of the dataset;
[0106] Calculate the interquartile range (IQR): IQR = Q3 - Q1;
[0107] Set the threshold:
[0108] Lower limit = Q1 - 1.5 × IQR
[0109] Upper limit = Q3 + 1.5 × IQR
[0110] Any value outside this range is considered an outlier;
[0111] Specific example
[0112] Suppose there is a dataset containing the following values:
[0113] values = [10, 12, 15, 14, 100, 13, 11, 9, 200, 10]
[0114] By applying the quartile method, we can perform the following calculations:
[0115] Q1 = 12
[0116] Q3 = 15
[0117] IQR = Q3 - Q1 = 15 - 12 = 3
[0118] According to the above formula, the set threshold is:
[0119] Lower limit = Q1 - 1.5 × IQR = 12 - 1.5 × 3 = 3.5Q1 - 1.5 × IQR = 12 - 1.5 × 3 = 3.5
[0120] Upper limit = Q3 + 1.5 × IQR = 15 + 1.5 × 3 = 21.5Q3 + 1.5 × IQR = 15 + 1.5 × 3 = 21.5
[0121] Therefore, any value less than 3.5 or greater than 21.5 is considered an outlier; in this example, 100 and 200 are both outliers;
[0122] Identify and handle missing values, check if there are missing values in the dataset, that is, whether the values of some fields are empty or incomplete. If there are missing values, delete the records containing missing values, or use interpolation and mean filling methods to estimate and fill these missing values;
[0123] Remove stop words, delete the common stop words in the article content, such as "of" and "is", to reduce interference with subsequent analysis;
[0124] Text conversion, including case conversion and traditional / simplified Chinese conversion;
[0125] S300. Analyze the data after cleaning and processing in step S200, including content tag extraction, article quantity analysis, and article content analysis. Among them,
[0126] The content tag extraction is to extract article tags such as themes and keywords through text analysis of the article content for better classification and organization of data; natural language processing techniques such as TF-IDF, TextRank and other algorithms can be used for tag extraction;
[0127] The article quantity analysis is to count and analyze the quantity of articles, including the change in the quantity of articles in different time periods. The different time periods can be daily, weekly or monthly, and the distribution of the quantity of articles for different themes or keywords;
[0128] The article content analysis is to perform text analysis on the user interaction data of the article. The user interaction data includes comments and messages after the article or other interaction data, and extract the emotional tendency and opinion feedback information of the user for better understanding of the user's attitude and satisfaction towards the product; sentiment analysis algorithms or tools can be used for sentiment classification and sentiment scoring of user comments;
[0129] Among them, the common algorithms of sentiment analysis algorithms are: Naive Bayes is a probability-based classification algorithm, suitable for processing a large amount of text data, especially performing well in sentiment analysis; it judges the emotional tendency by calculating the conditional probability of each word in the text; Support Vector Machine (SVM) is a classification algorithm that maximizes the interval between different categories by constructing a hyperplane, suitable for high-dimensional data and commonly used in sentiment classification tasks; when dealing with text data, Convolutional Neural Network (CNN) extracts local features through convolutional layers, suitable for capturing context information; the improved CNN-LDA method combined with the LDA topic model can better mine the deep semantics of the text and improve the accuracy of sentiment analysis; BERT and its variants BERT are popular pre-trained models in recent years, which perform sentiment analysis through context understanding and have high accuracy;
[0130] In order to improve the accuracy and reliability of sentiment analysis, the following optimization directions can be taken:
[0131] Expand and update the dictionary. Sentiment analysis based on the dictionary depends on the quality of the sentiment vocabulary library; regularly update the dictionary to cover emerging words and expressions, which can significantly improve the recognition ability.
[0132] Introduce context information. When performing sentiment classification, considering context information can help to more accurately understand the user's intention; for example, deep learning models can be used to capture the semantic relationship between texts, thereby improving the classification effect;
[0133] Combined with multimodal data, in addition to text, multimodal data such as images and audio can also be utilized for comprehensive analysis, which will help to more comprehensively understand the user's emotional state;
[0134] Personalized support. By constructing a personalized model based on the user's historical data, the accuracy of emotion recognition for specific user groups can be improved. This method can be achieved by analyzing the user's behavior and preferences;
[0135] S400. Visualize the data analyzed in step S300, and visually display the analysis results to the user in the form of charts, helping the user to more intuitively understand the situation of the WeChat official account, quickly discover and analyze the laws and trends in the data, so as to better formulate marketing strategies and make decisions on adjusting product functions;
[0136] When performing visualization processing, select appropriate visualization tools according to requirements. These tools provide rich visualization charts and functions, which can help users more intuitively understand the data situation; use the visualization tools to create corresponding charts according to the analysis results and requirements; use a relational graph to display the relationships between enterprises; also integrate all the visualized charts into a report, adding necessary explanations and descriptions; Combine Figure 5 As shown, when the user clicks on the "Tag Statistics" option, the interface will display the article tag distribution of each official account, such as "housing price", "grain price", "border security"; the user can select a specific official account or date range for tag statistics and comparative analysis;
[0137] The visualization tools mentioned above include Tableau, Power BI, and ECharts;
[0138] The visualized charts include line charts, bar charts, pie charts, and scatter plots; for example, a line chart can be used to display the content release trend of the official account, and a bar chart can be used to display the number of articles in different categories;
[0139] After constructing the visualization chart, the attributes and styles of the chart can be configured; for example, the title, axis labels, colors, font attributes of the chart can be set, and the size and position styles of the chart can be adjusted; these configurations can be adjusted according to actual needs to make the chart more beautiful and easy to understand;
[0140] The constructed visualization charts can also be integrated into a report, that is, multiple charts are combined together to form a complete analysis report; necessary explanations and descriptions are added to the report to help users better understand the data situation and trends; at the same time, the report can be updated regularly to reflect the dynamic changes of the data;
[0141] S500. Automatically and intelligently process WeChat official account data in real time through machine learning algorithms, including automatic classification, clustering, and prediction operations, so as to automatically complete data analysis tasks. That is, by writing automated scripts or tools, regularly execute data analysis tasks automatically to reduce manual intervention, improve work efficiency, and promptly detect abnormal situations, providing rapid responses for enterprise decision-making, and helping the WeChat official account better meet user needs and achieve sustainable growth. For example, machine learning algorithms can be used to automatically classify article content and extract tags, automatically identify information such as the theme and keywords of articles, and can also improve the response speed and user experience of the WeChat official account analysis system.
[0142] Use machine learning algorithms to predict and analyze data, such as predicting user churn and content popularity; the accuracy and efficiency of prediction can be improved by training models.
[0143] Specific implementation plan two: Different from specific implementation plan one, corresponding code is written in Python during text preprocessing to expand the WeChat-specific word segmentation dictionary and WeChat-specific stop words for subsequent hot spot detection, specifically including:
[0144] S210. Extract Chinese, English, and numbers, use regular expressions, re.compile(u"[^a-zA-Z0-9\u4e00-lu9fa5]"), to remove emojis and special symbols, and only retain letters, numbers, and Chinese; where \u4e00-\u9fa5 is the Unicode Chinese character encoding range, with a total of 20,901 Chinese characters.
[0145] S220. Expand the WeChat-specific word segmentation dictionary and stop words. First, use 133 WeChat official account nicknames as part of the word segmentation dictionary. Secondly, add words that often appear in WeChat official accounts such as "Read the full text, Recommended reading, More news, Without permission" and words that often appear in news such as "Headlines, Information, Quick look, Breaking news, Express delivery" to the word segmentation dictionary. Finally, view the titles, abstracts, and short text bodies of 250,000 news items (conduct in-depth analysis and review of a certain theme or content), and manually select a part of the words such as "Share pictures, Quote of the day, Follow, Reminder", and a dictionary with a size of 415 is established to be used as the word segmentation dictionary and stop word dictionary.
[0146] S230. Load the extended word segmentation dictionary in step S220 and perform word segmentation using Jieba. Jieba supports three word segmentation modes, including the accurate mode, the full mode, and the search engine mode. The accurate mode attempts to cut the sentence most precisely and is suitable for text analysis. The full mode scans out all the words that can form words in the sentence, with very high speed, but it cannot resolve ambiguities. The search engine mode is based on the accurate mode and further segments long words to improve the recall rate, which is suitable for search engine word segmentation.
[0147] The present invention uses the accurate mode. To improve the word segmentation speed, Jieba.enable_parallel() is called for parallel word segmentation. It should be noted that the four core elements of news, namely time, place, person, and event, mainly come from five categories of words: nouns, verbs, non-predicate adjectives, tense words, and numerals. The present invention uses the Jieba.posseg module for part-of-speech tagging. In the comparative experiments of retaining nouns, verbs, adjectives and only retaining nouns and verbs, it is found that there is almost no difference in the topic detection effect between the two. Therefore, only nouns and verbs are retained in the part-of-speech tagging stage in this article.
[0148] S240. Remove stop words, including removing words with no clear meaning such as modal particles, adverbs, prepositions, and conjunctions, to reduce the impact on the topic extraction result. The stop word list selects the general stop word list on the Internet plus the extended WeChat exclusive stop word list.
[0149] Through the preprocessing process of the above steps, a relatively pure text set can be initially obtained.
[0150] Specific implementation plan three: Different from specific implementation plan two, when performing data analysis in step S300, the machine learning algorithms used include the decision tree algorithm, the Adaboost algorithm, the Bagging algorithm, or the support vector machine algorithm, where
[0151] (1) Decision tree algorithm. The decision tree algorithm is a process of classifying data through certain classification rules. Decision trees are divided into two types: classification trees and regression trees. Making a decision tree for discrete variables is called a classification tree, and making a decision tree for continuous variables is called a regression tree.
[0152] The process of generating a decision tree is as follows: First, a decision tree is generated from the training sample set. Then, a certain evaluation criterion is used to select a feature from the features of the training set data as the splitting criterion for the current node. According to the selected decision tree splitting criterion, child nodes are recursively generated from top to bottom until the growth stops when the branch stopping rule is satisfied. The generated decision tree is prone to overfitting. Generally, pruning of the decision tree is required, that is, the generated decision tree is tested and corrected. Mainly, the data in the test set is used to verify the preliminary rules generated during its generation process, and the branches that affect the prediction accuracy are pruned to reduce the scale of the decision tree structure and alleviate the overfitting phenomenon.
[0153] Decision tree algorithm data set division: The data set is divided into a training set and a test set, using a ratio of 80 / 20 or 70 / 30. The training set is used to construct the decision tree, and the test set is used to evaluate the model performance.
[0154] Evaluation criteria: The accuracy, precision, recall, and F1-score metrics are used to evaluate the classification effect of the model. The robustness of the model is further verified through cross-validation.
[0155] (2) Adaboost is an iterative ensemble algorithm that can be used for classification and regression. Its core idea is that in each training process, more attention is paid to the previous misclassified samples, and the weights of each sample in the sample set are re-assigned according to the sample distribution, so that the classifier is more accurate.
[0156] Adaboost algorithm data set division: The same division method of the training set and the test set is adopted. In each round of iteration, Adaboost adjusts the sample weights according to the misclassification situation of the previous round of classifier, so that the samples that are difficult to classify obtain higher weights in the next round.
[0157] Evaluation criteria: The classification accuracy is used to evaluate the performance of the final model. The performance of the model at different thresholds can be measured by plotting the ROC curve and calculating the AUC value.
[0158] For a given training data set:
[0159]
[0160] Among them, the variables X and Y represent the input factors and annotations respectively; R represents the real-valued output; n represents the dimension of the samples.
[0161] The initial weight vector of the training samples in the data set is:
[0162]
[0163] Among them, N is the number of dataset samples; w 1i is the rated weight of the training samples;
[0164] The specific training process is as follows:
[0165] (a) Use the training dataset with weight distribution to train and learn to obtain the weak classifier Gm(xi);
[0166] (b) Calculate the regression error rate during each iteration; for the k-th weak learner, calculate its maximum error on the training set, E m =max|y i -G m (x i ), i = 1, 2,... N. Calculate the relative error of the samples If using the squared error, then If using the exponential error, then From this, the regression error rate of the k-th weak learner G m (x) on the training dataset can be obtained: The smaller the error rate, the more accurate the prediction result of the base classifier for the data samples;
[0167] (c) Adopt the "weighted majority voting" method to combine multiple base classifiers; according to the calculated regression error rate e m , re-adjust the weights for the base classifiers; the formula for calculating the weight coefficient of the weak learner is: α m represents the importance of G m (x i ) in the final classifier;
[0168] Before generating the next base classifier, the Adaboost algorithm will adjust the sample weights. If the sample is misclassified, the weight will increase; if the sample is correctly classified, the weight of this sample will be reduced. For updating the sample weights D m+1 =(ω m+1,1 ,…,ω m+1,i ,…,ω m+1,N ), the weight coefficient of the sample set of the (k + 1)-th weak learner is ω m+1 ; where ∈ m represents the error rate of the weak classifier, the smaller the better; is a part when updating the sample weights, used to emphasize the samples misclassified by the weak classifier, so that they can obtain higher weights in the next iteration;
[0169] (d) Z m is the normalization factor,
[0170] In the Adaboost algorithm, the regression method is to take the weak learner corresponding to the median of the weights of the weighted weak learners as the strong learner. The final strong regression learner is:
[0171]
[0172] Among them, M is an important parameter that controls the complexity and performance of the AdaBoost model;
[0173] (3) Combining Figure 3 As shown, the Bagging algorithm uses the bootstrap method for sampling with replacement during the training and learning process. Assume that the total number of the sample data set is N, and n (n < N) samples are randomly drawn with replacement, and a decision tree is trained with this; because it is sampling with replacement each time, if it is executed p times repeatedly, p different sample sets can be taken out, and then a decision tree is constructed for each sample set; if the purpose of constructing the model is classification, then the final classification result is voted by the classification results of these decision trees; if the purpose is regression, then the predicted value of the explanatory variable is obtained by the simple average of the results of these decision trees;
[0174] Bagging algorithm data set division: The bootstrap method (Bootstrap Sampling) is adopted to randomly draw multiple subsets from the original data set, and each subset can be sampled repeatedly; each subset is used to train a basic learner, and finally the results of multiple learners are combined by voting or averaging
[0175] Evaluation criteria: The improvement effect of the Bagging model is evaluated by comparing the accuracy of the Bagging model and a single learner; cross-validation can also be used to ensure the consistency and reliability of the model
[0176] (4) The basic idea of support vector regression to solve problems is: First, through a non-linear mapping map the samples from the input space to a high-dimensional feature space, and then perform linear regression on the samples in the space H to find the optimal regression hyperplane, that is, fit the optimal regression function (w and b are parameters to be determined), where, represents the feature mapping function in the context of machine learning; finally, the optimal regression function is used to perform regression prediction on other samples; the expression of the standard support vector regression loss function is:
[0177]
[0178] Among them, ε represents a tolerance range, that is, if the error between the model predicted value f(x) and the true value y is within the range of ε, it will not be included in the loss;
[0179] Support Vector Machine (SVM) Algorithm Dataset Division: The dataset is divided into a training set and a test set with a ratio of 80 / 20. After feature selection, the features are processed by standardization or normalization to improve the classification performance of SVM.
[0180] Evaluation Criteria: Multiple metrics such as accuracy, precision, recall, and F1-score are used for comprehensive evaluation. The confusion matrix can be used to analyze the classification results in detail to understand the performance of the model in each category.
[0181] Among them, ε represents the maximum error allowed by the regression function, called the kernel width. Using the above support vector regression loss function can improve the generalization ability of the regression model. The support vector algorithm constructs the regression model based on the principle of structural risk minimization, that is, not only minimizing the empirical risk during training but also reducing the complexity of the model. The above problem of finding the optimal regression function can be transformed into an optimization problem.
[0182] Objective Function:
[0183]
[0184] Among them, w is the core parameter in SVM, used to define the optimal hyperplane, and its size and direction directly affect the performance and generalization ability of the classifier.
[0185] Its Constraint Conditions:
[0186] y i -f(x i )≤ε+ξ i
[0187]
[0188] Among them, C is the penalty coefficient, n is the number of samples, and ξ i , is the slack variable. According to the duality principle, when certain conditions are met, the solution of the dual problem is the solution of the above formula. Therefore, the optimal regression hyperplane is:
[0189]
[0190] a i , is the Lagrange multiplier, and 0≤a i , x i is the support vector that satisfies y i -f(x i )=ε or f(x i )-y i =ε, and b is a constant. Therefore, the support vector xi Find out as the calibration function.
[0191] Specific implementation plan four: The present invention provides a WeChat official account sensitive information analysis system, which has program modules corresponding to the above steps and executes the steps in the above WeChat official account sensitive information analysis method when running.
[0192] The other combinations and connection relationships of this implementation plan are the same as those of the specific implementation plan three.
[0193] Specific implementation plan five: The present invention provides a computer-readable storage medium, which stores a computer program configured to implement the steps of the WeChat official account sensitive information analysis method when called by a processor.
[0194] The other combinations and connection relationships of this implementation plan are the same as those of the specific implementation plan three.
[0195] The present invention has the following advantages:
[0196] 1. Multi-modal data fusion analysis, combining multi-modal data such as text, pictures, and videos for sensitive information analysis, rather than relying solely on text data.
[0197] Implementation method: Use image recognition technology to detect sensitive content in the article cover or illustrations; perform speech-to-text processing on video content and extract potential sensitive information in combination with a sentiment analysis model.
[0198] Advantage: Compared with single text analysis, it can capture sensitive information more comprehensively and is applicable to diverse content forms.
[0199] 2. Dynamic sensitive word library construction and update, introducing a dynamic sensitive word library update mechanism based on deep learning to learn newly emerging sensitive words in real time.
[0200] Implementation method: Use a pre-trained language model (such as BERT) combined with transfer learning to extract the latest sensitive words from social media and news; build a domain adaptation model to adjust the sensitive word weights according to the official account type (such as government affairs, education, entertainment).
[0201] Advantage: Solve the problem of low efficiency in manually maintaining the sensitive word library and improve the adaptability and accuracy of the system.
[0202] 3. Personalized risk assessment model, customize risk assessment indicators according to the official account type (personal version / enterprise version) and operation objectives (such as marketing promotion, article management).
[0203] Implementation method: For personal official accounts, focus on evaluating the risk of privacy leakage (such as the exposure of personal identity information); for enterprise official accounts, focus on the risk analysis of trade secret leakage and brand articles.
[0204] Advantages: Provide differentiated services for different user groups, improving the practicality of the system and user satisfaction.
[0205] 4. Intelligent visualization and decision support. Develop a visualization tool based on a knowledge graph to display sensitive information in association and provide optimization suggestions.
[0206] Implementation method: Build a knowledge graph to associate sensitive information with the context and visually present it through a graphical interface; provide automated decision-making suggestions, such as adjusting the push strategy or modifying article content.
[0207] Advantages: Improve users' ability to understand the analysis results while reducing the operation threshold.
[0208] 5. Real-time monitoring and early warning mechanism. Introduce a real-time monitoring function to instantly scan newly released content and prompt potential risks through a multi-level early warning mechanism.
[0209] Implementation method: Use streaming data processing technology to perform real-time parsing and analysis of newly released content; set different levels of early warning notifications according to the risk level (such as SMS reminders, highlighting marks).
[0210] Advantages: Achieve rapid response to sensitive information and reduce potential losses.
[0211] 6. Privacy protection and enhanced compliance. Integrate privacy protection technologies (such as differential privacy) to ensure the security of user data and comply with relevant laws and regulations (such as GDPR) at the same time.
[0212] Implementation method: Introduce a differential privacy algorithm in the data processing link to avoid exposing user privacy information; provide a compliance check function to automatically review whether the article content violates laws and regulations.
[0213] Advantages: Enhance the credibility of the system and meet the increasingly strict data protection requirements.
[0214] Although the present invention is disclosed as above, the protection scope of the present invention is not limited thereto. Those skilled in the art of the present invention can make various changes and modifications without departing from the spirit and scope of the present invention, and these changes and modifications will all fall within the protection scope of the present invention.
Claims
1. A method for analyzing sensitive information of WeChat public accounts, characterized in that: The following steps are involved: S100, public account data collection, using the WeChat public account backend to directly obtain data or using a third-party tool to collect data, the public account data includes data details of multiple articles; S200, cleaning and preprocessing the public account data collected in step S100, including removing abnormal values and duplicate values, unifying data formats, standardizing field naming, identifying and processing missing values, removing stop words, and performing text conversion; S300, analyzing the data cleaned and processed in step S200, including content tag extraction, article quantity analysis and article content analysis; S400, visualizing the data analyzed in step S300, and visually displaying the analysis results to the user in the form of charts; S500, uses machine learning algorithms to perform real-time automated and intelligent processing of public account data, which is used to automatically perform data analysis tasks regularly and detect abnormal situations in a timely manner.
2. A method for analyzing sensitive information of a WeChat public account according to claim 1, characterized in that: In step S100, when directly acquiring data using the WeChat official account backend, the relevant data can be directly acquired through the API interface or data export function provided by the WeChat official account backend; the specific operation steps are as follows: S110. Log in to the WeChat public account backend and enter the management interface; S120, find the entrance of data export or API interface, and obtain relevant API key or login information; S130, using the API key or login information, write corresponding code in a programming language to call the API interface or data export function of the WeChat public account background; S140: Extract required data according to the return result of the API interface or the data export function.
3. A method for analyzing sensitive information of a WeChat public account according to claim 1, characterized in that: In step S100, when a third-party tool is used to collect data, the specific operation steps are as follows: (1) Select data collection software or crawler program; (2) Write collection scripts or rules based on the data structure and rules of WeChat public accounts; (3) Collect data from WeChat public accounts on a regular or real-time basis and save the collected data to local or cloud storage.
4. A method for analyzing sensitive information of a WeChat public account according to claim 2 or 3, characterized in that: In step S200, when performing data cleaning and processing, outliers outside the normal range are identified and removed by setting thresholds or using statistical methods. include, Identify and delete duplicate data. Use data cleaning tools or write scripts to identify duplicate entries in the data set and delete such data to retain only unique data records. Unify data formats, so that data from different sources or formats follow a unified format standard; Standardize field naming and standardize words with the same meaning; Set thresholds and handle outliers. For numeric fields, set a reasonable threshold range and treat values outside this range as outliers. Delete identified outliers or replace them with the mean or median. Identify and handle missing values. Check whether there are missing values in the data set. If there are missing values, delete the records containing missing values, or use interpolation and mean filling methods to estimate and fill these missing values. Remove stop words and delete the stop words in the article content; Text conversion, including uppercase and lowercase conversion and traditional and simplified Chinese conversion.
5. A method for analyzing sensitive information of a WeChat public account according to claim 4, characterized in that: In step S200, the corresponding code is written in Python during text preprocessing, specifically including: S210, extract Chinese, English and numbers, use regular expression, re.compile(u"[^a-zA-Z0-9\u4e00-lu9fa5]"), remove emoticons and special symbols, and only keep letters, numbers and Chinese characters; where \u4e00-\u9fa5 is the Unicode Chinese character encoding range; S220, expanding the WeChat exclusive word segmentation dictionary and stop words, firstly taking the public account nickname as part of the word segmentation dictionary, secondly adding the words that frequently appear in the public account to the word segmentation dictionary, and finally checking and manually selecting the title, abstract, and short text body as the word segmentation dictionary and stop word dictionary; S230, load the word segmentation dictionary expanded in step S220, use jieba to perform word segmentation, use the precise mode, call jieba.enable_parallel0 to perform parallel word segmentation; use the jieba.posseg module to perform part-of-speech tagging, and only retain nouns and verbs in the part-of-speech tagging stage; S240, remove stop words, including removing modal particles, adverbs, prepositions, conjunctions and words without clear meaning.
6. A method for analyzing sensitive information of a WeChat public account according to claim 5, characterized in that: In step S300, the content tag extraction is to extract article tags by text analysis of the article content; the article quantity analysis is to count and analyze the number of articles, including the changes in the number of articles in different time periods and the distribution of the number of articles with different themes or keywords; the article content analysis is to perform text analysis on the user interaction data of the article, and the user interaction data includes comments and messages or other interaction data after the article, to extract the user's emotional tendencies and feedback information.
7. A method for analyzing sensitive information of a WeChat public account according to claim 6, characterized in that: In step S400, a visualization tool is selected according to requirements when performing visualization processing, and corresponding charts are created using the visualization tool according to the analysis results and requirements. The charts include line charts, bar charts, pie charts, and scatter plots.
8. A method for analyzing sensitive information of a WeChat public account according to claim 6, characterized in that: In step S500, the machine learning algorithm used is a decision tree algorithm, an Adaboost algorithm, a Bagging algorithm or a support vector machine algorithm.
9. A WeChat public account sensitive information analysis system, characterized by: The system has a program module corresponding to the steps of any one of claims 1 to 8 above, and executes the steps in the above-mentioned WeChat public account sensitive information analysis method when running.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the WeChat public account sensitive information analysis method described in any one of claims 1 to 8 when called by a processor.