Online public opinion positive and negative classification method
Through the positive and negative classification method of online public opinion, crawling technology and deep learning models are used to collect data, preprocess and feature extraction of network public opinion texts, solving the problem of low classification accuracy in the existing technology, and achieving efficient and accurate classification of public opinion sentiment.
Patent Information
- Application Number
- CN202510180451.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-30
AI Technical Summary
When the existing online public opinion classification method faces texts with vague semantics and insignificant emotional tendencies, the classification accuracy still needs to be improved, and the rules-based methods are inefficient, making it difficult to cover complex and changeable network languages.
A positive and negative classification method for network public opinion is adopted, and data collection, preprocessing, feature extraction, model construction and training is used, data is collected using crawler technology, tools such as Python NLTK are used to clean and participle, features are extracted using TfidfVectorizer and emotional dictionary, and convolutional neural network (CNN) model is built for training and prediction.
It realizes the emotional intensity identification and positive and negative classification of online public opinion texts, improves classification accuracy, can output detailed emotional grading results, and provides support for public opinion monitoring and emotional analysis.
Smart Images

Figure CN120067416A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of online public opinion, and specifically, to a method for classifying positive and negative online public opinion. Background Art
[0002] With the rapid development of the Internet, online public opinion data has grown explosively. Timely and accurately classifying positive and negative online public opinion is of great significance for various organizations such as governments and enterprises to understand public attitudes and make decisions. At present, common online public opinion classification methods mainly include rule-based methods and machine learning-based methods. Rule-based methods rely on a large number of manually written rules, with low efficiency and difficulty in covering complex and changing online languages; although machine learning-based methods have improved the classification efficiency to a certain extent, the classification accuracy still needs to be improved when facing texts with ambiguous semantics and unclear emotional tendencies. Therefore, there is an important practical need to develop an efficient and accurate method for classifying positive and negative online public opinion. Summary of the Invention
[0003] The present invention proposes a method for classifying positive and negative online public opinion, which solves the problems in the related art.
[0004] The technical solution of the present invention is as follows: A method for classifying positive and negative online public opinion, characterized by comprising: S1: Data collection: Using web crawler technology, collect online public opinion text data from major social media platforms, news websites, etc.; S2: Data preprocessing: Use tools such as Python NLTK to clean the data, perform word segmentation with Jieba and part-of-speech tagging with NLTK, and extract word stems with SnowballStemmer; S3: Feature extraction: Use TfidfVectorizer of scikit-learn to calculate TF-IDF values, initialize and set parameters, fit and transform the text data to obtain a feature matrix, and load a sentiment dictionary (such as HowNet) to extract sentiment word features in the text; S4: Model construction and training: Use TensorFlow / PyTorch to construct a CNN model. Input the feature vectors into the model, set hyperparameters, and then train with labeled public opinion data, use cross-entropy loss and Adam optimization, and incorporate sentiment grading labels; S5: Classification prediction: Preprocess and extract features of the online public opinion text to be classified according to the above steps, and then input it into the trained model to obtain positive and negative classification results including sentiment grading (-3, -2, -1, 0, 1, 2, 3).
[0005] 2. A method for classifying positive and negative online public opinions according to claim 1, characterized in that, in the step S1: Data collection: Using web crawler technology, the specific steps for collecting online public opinion text data from major social media platforms, news websites, etc. are as follows: Step S1: Data collection S1.1 Determine the data source: Target platform selection: According to the analysis requirements, determine the social media platforms to be crawled; Content type determination: Clearly define the content types to be collected; S1.2 Design the crawler strategy: Access frequency setting: According to the access restrictions and crawler etiquette of the target platform, reasonably set the access frequency to avoid causing excessive pressure on the platform server; Data scraping rules: Formulate data scraping rules, including the fields to be scraped; Countermeasure against anti-crawler mechanism: Study the anti-crawler mechanism of the target platform and prepare corresponding countermeasures; S1.3 Develop the crawler program: Select the programming language: According to the team's technology stack and crawler requirements, select a suitable programming language; Write the crawler code: Use libraries or frameworks such as requests, BeautifulSoup, Scrapy to write the crawler code to implement the data scraping function; Data storage design: Design a data storage scheme to store the scraped data; S1.4 Execute the crawler task: Deploy the crawler environment: Deploy the crawler environment on the server or local computer to ensure that the crawler program can run normally; Run the crawler program: Execute the crawler program to start scraping data; Depending on the data volume and scraping speed, it may need to run continuously for some time; Monitor the crawler status: Real-time monitor the running status of the crawler program and handle abnormal situations in a timely manner, such as network anomalies and changes in the target platform structure; S1.5 Data cleaning and preprocessing: Data deduplication: Remove duplicate scraped data to ensure the uniqueness of the data set; Data cleaning: Process the noise and irrelevant information in the data; Data formatting: Convert the data into a format suitable for subsequent analysis; S1.6 Data storage and backup: Store the cleaned data: Store the cleaned and preprocessed data in the database or file system for subsequent analysis; Data backup: Regularly back up the data to prevent data loss or damage.
[0006] 3. A method for classifying positive and negative online public opinions according to claim 1, characterized in that the data preprocessing in step S2: cleaning the data using tools such as Python NLTK, performing word segmentation using Jieba word segmentation, and performing part-of-speech tagging using NLTK. The specific steps of extracting word stems using SnowballStemmer are as follows: S2.1 Use the NLTK library or other text processing tools to clean the data: Install the NLTK library; Import the necessary modules; Remove HTML tags: Use regular expressions to remove HTML tags in the text; Remove special characters: Use regular expressions to remove non-alphanumeric characters; Remove stop words (optional): Download and load the stop word list of NLTK to remove stop words in the text; Apply the cleaning function: Apply the above cleaning function to the text data; S2.2 Use the Jieba word segmentation tool to perform word segmentation: Install the Jieba word segmentation tool; Import Jieba word segmentation; Perform word segmentation: Use Jieba word segmentation to segment the cleaned text; S2.3 Use the part-of-speech tagging tool of the NLTK library to perform part-of-speech tagging: Download the part-of-speech tagging resources (if not downloaded); Perform part-of-speech tagging: Use the part-of-speech tagger of NLTK to perform part-of-speech tagging on the segmented text; S2.4 Use tools such as SnowballStemmer to extract word stems: Install SnowballStemmer (usually installed with NLTK); Import SnowballStemmer; Extract word stems: Initialize SnowballStemmer and apply it to the segmented text (or the words after part-of-speech tagging).
[0007] 4. A method for classifying positive and negative online public opinions according to claim 1, characterized in that the step S3: feature extraction: calculating the TF-IDF value using TfidfVectorizer of scikit-learn, initializing and setting parameters, fitting and transforming the text data to obtain a feature matrix, and loading an emotion dictionary (such as HowNet), and the specific steps of extracting emotion vocabulary features in the text are as follows: S3.1 Calculate the TF-IDF values using TfidfVectorizer in scikit-learn: Initialize TfidfVectorizer: Initialize it using the TfidfVectorizer class in the scikit-learn library of Python; Set parameters: Set the parameters of TfidfVectorizer according to requirements; Fit and transform the text data: Input the preprocessed text data into TfidfVectorizer for fitting and transformation to obtain the TF-IDF feature matrix of the text; S3.2 Load the sentiment dictionary and extract sentiment vocabulary features: Load the sentiment dictionary: Select and load an appropriate sentiment dictionary; Extract sentiment vocabulary features: Extract the sentiment vocabulary features in the text according to the occurrence and sentiment intensity of sentiment words in the text; This can be achieved by traversing the words in the text and matching them with the sentiment dictionary; The matched sentiment words can be quantified according to their sentiment intensity and used as part of the features.
[0008] 5. According to a positive and negative classification method for online public opinion as described in claim 1, wherein the model construction and training in step S4: Construct a CNN model using TensorFlow / PyTorch. Input the feature vectors into the model, set the hyperparameters, and then train with the labeled public opinion data, using cross-entropy loss and Adam optimization, and integrating the sentiment grading labels. The specific steps are as follows: S4.1 Use TensorFlow or PyTorch to construct a convolutional neural network (CNN) model: Select a deep learning framework: According to requirements and familiarity, select TensorFlow and PyTorch as deep learning frameworks; Construct the CNN model: Use the construction tools of the selected framework to design and construct a convolutional neural network model; S4.2 Input the feature vectors into the model and set the hyperparameters: Input the feature vectors: Use the extracted feature vectors (such as the TF-IDF feature matrix and sentiment vocabulary features) as the input data of the model; Set the hyperparameters: Set appropriate hyperparameters according to the requirements of the model and the characteristics of the training data. These hyperparameters have an important impact on the performance and training effect of the model; S4.3 Use the labeled online public opinion data for training: Prepare training data: Use the labeled online public opinion data as the training data for the model; the data should include the text content and the corresponding sentiment classification labels; Train the model: Input the training data into the CNN model, and use the cross-entropy loss function to evaluate the difference between the predicted results of the model and the true labels; use the Adam optimizer to update the parameters of the model to minimize the loss function; Integrate sentiment classification labels: During the training process, integrate the sentiment classification labels into the training data so that the model can learn the features corresponding to different sentiment intensities. This helps the model to more accurately identify the sentiment intensity of the text in subsequent sentiment classification tasks.
[0009] 6. A method for classifying positive and negative online public opinion according to claim 1, characterized in that the classification prediction in step S5: Preprocess and extract features from the online public opinion text to be classified according to the above steps, and then input it into the trained model to obtain the positive and negative classification results including sentiment classification (-3, -2, -1, 0, 1, 2, 3). The specific steps are as follows: S5.1 Text preprocessing: Clean the online public opinion text to be classified; Use the Jieba word segmentation tool to segment the cleaned text; Perform part-of-speech tagging on the segmented text to further understand the structure and meaning of the text; Perform stemming to restore the vocabulary to its basic form for better feature matching; S5.2 Feature extraction: Use the TfidfVectorizer class in scikit-learn to calculate the TF-IDF values of the text to obtain the TF-IDF feature matrix of the text; Load the sentiment dictionary, such as the HowNet sentiment dictionary, and extract sentiment vocabulary features according to the occurrence and sentiment intensity of the sentiment vocabulary in the text; Combine the TF-IDF feature matrix and the sentiment vocabulary features into a complete feature vector; S5.3 Model input and prediction: Input the extracted feature vector into the trained convolutional neural network (CNN) model; The model processes the input feature vector, and extracts high-level features of the text through structures such as convolutional layers, pooling layers, and fully connected layers; The model outputs the positive and negative classification results including sentiment classification (-3, -2, -1, 0, 1, 2, 3); these sentiment classifications represent the sentiment intensity of the text, negative values indicate negative sentiment, positive values indicate positive sentiment, and 0 indicates neutral sentiment; S5.4 Result interpretation and application: Interpret the sentiment classification results of the text based on the output of the model; Apply the classification results to practical scenarios such as public opinion monitoring and sentiment analysis to assist in decision-making or provide insights.
[0010] The working principle and beneficial effects of the present invention are as follows: 1. The present invention systematically collects, preprocesses, and extracts the features of online public opinion text data. This method can accurately identify the sentiment intensity of the text. Specifically, web crawler technology is used to widely collect data, ensuring the richness and diversity of the data; data cleaning and word segmentation are performed using tools such as NLTK in Python, improving the quality of the data; TfidfVectorizer and sentiment dictionaries are used to extract features, enhancing the model's ability to understand text sentiment. A convolutional neural network model is constructed and trained. By optimizing hyperparameters and incorporating sentiment grading labels, the classification accuracy of the model is further improved. Finally, positive and negative classification results with detailed sentiment grading can be output, providing strong support for practical applications such as public opinion monitoring and sentiment analysis. This not only helps enterprises and governments better understand public opinions but also provides a scientific basis for decision-making, improving the efficiency and effectiveness of decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0012] Figure 1 It is a schematic diagram of the structural principle of the present invention. SPECIFIC EMBODIMENTS
[0013] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0014] Embodiment 1 As Figure 1 shown, this embodiment proposes a method for classifying positive and negative of online public opinion, including: S1: Data collection: Use web crawler technology to collect online public opinion text data from major social media platforms, news websites, etc.; S2: Data preprocessing: Clean the data using tools such as Python NLTK, perform word segmentation using Jieba, perform part-of-speech tagging using NLTK, and extract word stems using SnowballStemmer; S3: Feature extraction: Calculate the TF-IDF values using TfidfVectorizer in scikit-learn, initialize and set parameters, fit and transform the text data to obtain the feature matrix, and load the sentiment dictionary (such as HowNet) to extract the sentiment vocabulary features in the text; S4: Model construction and training: Build a CNN model using TensorFlow / PyTorch. Input the feature vectors into the model, set the hyperparameters, and then train with the labeled public opinion data, using cross-entropy loss and Adam optimization, and incorporate the sentiment grading labels; S5: Classification prediction: Preprocess and extract features from the network public opinion text to be classified according to the above steps, and then input it into the trained model to obtain the positive and negative classification results including sentiment grading (-3, -2, -1, 0, 1, 2, 3).
[0015] In this embodiment, in step S1: Data collection: Using web crawler technology, the specific steps for collecting network public opinion text data from major social media platforms, news websites, etc. are as follows: Step S1: Data collection S1.1 Determine the data source: Target platform selection: According to the analysis requirements, determine the social media platforms that need to be crawled; Content type determination: Clearly define the content types that need to be collected; S1.2 Design the crawler strategy: Access frequency setting: According to the access restrictions and crawler etiquette of the target platform, reasonably set the access frequency to avoid causing excessive pressure on the platform server; Data scraping rules: Formulate data scraping rules, including the fields that need to be scraped; Anti-crawler mechanism response: Study the anti-crawler mechanism of the target platform and prepare corresponding countermeasures; S1.3 Develop the crawler program: Select the programming language: According to the team's technology stack and crawler requirements, select a suitable programming language; Write the crawler code: Use libraries or frameworks such as requests, BeautifulSoup, Scrapy to write the crawler code to implement the data scraping function; Data storage design: Design a data storage scheme to store the scraped data; S1.4 Execute the crawler task: Deploy the crawler environment: Deploy the crawler environment on the server or local computer to ensure that the crawler program can run normally; Run the crawler program: Execute the crawler program to start scraping data; Depending on the size of the data volume and the scraping speed, it may be necessary to run continuously for a period of time; Monitor the crawler status: Real-time monitor the running status of the crawler program and promptly handle abnormal situations, such as network anomalies and changes in the target platform structure; S1.5 Data cleaning and preprocessing: Data deduplication: Remove duplicate crawled data to ensure the uniqueness of the dataset; Data cleaning: Process noise and irrelevant information in the data; Data formatting: Convert the data into a format suitable for subsequent analysis; S1.6 Data storage and backup: Store the cleaned data: Store the cleaned and preprocessed data in a database or file system for subsequent analysis; Data backup: Regularly back up the data to prevent data loss or corruption.
[0016] Specifically, the data preprocessing in step S2: Use tools such as Python NLTK to clean the data, perform word segmentation with Jieba, and conduct part-of-speech tagging with NLTK. The specific steps for stemming with SnowballStemmer are as follows: S2.1 Use the NLTK library or other text processing tools for data cleaning: Install the NLTK library; Import necessary modules; Remove HTML tags: Use regular expressions to remove HTML tags from the text; Remove special characters: Use regular expressions to remove non-alphanumeric characters; Remove stop words (optional): Download and load the NLTK stop word list to remove stop words from the text; Apply the cleaning function: Apply the above cleaning function to the text data; S2.2 Use the Jieba word segmentation tool for word segmentation: Install the Jieba word segmentation tool; Import Jieba for word segmentation; Word segmentation processing: Use Jieba to perform word segmentation on the cleaned text; S2.3 Use the part-of-speech tagging tool in the NLTK library for part-of-speech tagging: Download the part-of-speech tagging resources (if not downloaded); Part-of-speech tagging: Use the NLTK part-of-speech tagger to perform part-of-speech tagging on the segmented text; S2.4 Use tools such as SnowballStemmer for stemming: Install SnowballStemmer (usually installed with NLTK); Import SnowballStemmer; Stemming: Initialize SnowballStemmer and apply it to the tokenized text (or words after part-of-speech tagging).
[0017] In this embodiment, step S3: Feature extraction: Use TfidfVectorizer in scikit-learn to calculate TF-IDF values, initialize and set parameters, fit and transform the text data to obtain a feature matrix, and load a sentiment dictionary (such as HowNet). The specific steps for extracting sentiment word features from the text are as follows: S3.1 Calculate TF-IDF values using TfidfVectorizer in scikit-learn: Initialize TfidfVectorizer: Initialize using the TfidfVectorizer class in the scikit-learn library of Python; Set parameters: Set the parameters of TfidfVectorizer according to requirements; Fit and transform the text data: Input the preprocessed text data into TfidfVectorizer, perform fitting and transformation to obtain the TF-IDF feature matrix of the text; S3.2 Load the sentiment dictionary and extract sentiment word features: Load the sentiment dictionary: Select and load a suitable sentiment dictionary; Extract sentiment word features: According to the occurrence and sentiment intensity of sentiment words in the text, extract the sentiment word features in the text; This can be achieved by traversing the words in the text and matching them with the sentiment dictionary; The matched sentiment words can be quantified according to their sentiment intensity and used as part of the features.
[0018] Specifically, for model construction and training in step S4: Build a CNN model using TensorFlow / PyTorch. Input the feature vector into the model, set hyperparameters, and then train with the labeled public opinion data, using cross-entropy loss and Adam optimization, and integrating sentiment grading labels. The specific steps are as follows: S4.1 Build a convolutional neural network (CNN) model using TensorFlow or PyTorch: Select a deep learning framework: According to requirements and familiarity, select TensorFlow and PyTorch as deep learning frameworks; Build a CNN model: Use the construction tools of the selected framework to design and build a convolutional neural network model; S4.2 Input the feature vector into the model and set hyperparameters: Input feature vector: Use the extracted feature vectors (such as TF-IDF feature matrix and sentiment word features) as the input data of the model; Set hyperparameters: According to the requirements of the model and the characteristics of the training data, set appropriate hyperparameters, which have an important impact on the performance and training effect of the model; S4.3 Use the labeled online public opinion data for training: Prepare training data: Use the labeled online public opinion data as the training data of the model; the data should include text content and corresponding sentiment classification labels; Train the model: Input the training data into the CNN model, and use the cross-entropy loss function to evaluate the difference between the prediction result of the model and the true label; use the Adam optimizer to update the parameters of the model to minimize the loss function; Integrate sentiment classification labels: During the training process, integrate the sentiment classification labels into the training data so that the model can learn the features corresponding to different sentiment intensities. This helps the model to more accurately identify the sentiment intensity of the text in subsequent sentiment classification tasks.
[0019] In this embodiment, the classification prediction in step S5: Preprocess and extract features from the online public opinion text to be classified according to the above steps, and then input it into the trained model to obtain the positive and negative classification results including sentiment classification (-3, -2, -1, 0, 1, 2, 3). The specific steps are as follows: S5.1 Text preprocessing: Clean the online public opinion text to be classified; Use the Jieba word segmentation tool to perform word segmentation on the cleaned text; Perform part-of-speech tagging on the segmented text to further understand the structure and meaning of the text; Perform stemming to restore the vocabulary to its basic form for better feature matching; S5.2 Feature extraction: Use the TfidfVectorizer class in scikit-learn to calculate the TF-IDF values of the text to obtain the TF-IDF feature matrix of the text; Load the sentiment dictionary, such as the HowNet sentiment dictionary, and extract sentiment word features according to the occurrence and sentiment intensity of sentiment words in the text; Combine the TF-IDF feature matrix and sentiment word features into a complete feature vector; S5.3 Model input and prediction: Input the extracted feature vector into the trained convolutional neural network (CNN) model; The model processes the input feature vectors and extracts the high-level features of the text through structures such as convolutional layers, pooling layers, and fully connected layers; The model outputs positive and negative classification results that include sentiment ratings (-3, -2, -1, 0, 1, 2, 3); these sentiment ratings represent the intensity of the sentiment of the text, with negative values indicating negative sentiment, positive values indicating positive sentiment, and 0 indicating neutral sentiment; S5.4 Result Interpretation and Application: Interpret the sentiment classification results of the text based on the output of the model; Apply the classification results to actual scenarios, such as public opinion monitoring and sentiment analysis, to assist in decision-making or provide insights.
[0020] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for classifying positive and negative network public opinion, characterized in that: include: S1: Data collection: Using web crawler technology, collect online public opinion text data from major social media platforms, news websites, etc.; S2: Data preprocessing: using tools such as Python NLTK to clean the data, using Jieba to perform word segmentation and NLTK to perform part-of-speech tagging, and using SnowballStemmer to perform word stemming; S3: Feature extraction: Use scikit-learn's TfidfVectorizer to calculate the TF-IDF value, initialize and set parameters, fit the converted text data to get the feature matrix, and load the sentiment dictionary (such as HowNet) to extract the sentiment vocabulary features in the text; S4: Model construction and training: Use TensorFlow / PyTorch to build a CNN model. Input the feature vector into the model, set the hyperparameters, and then train it with annotated public opinion data, using cross entropy loss and Adam optimization, and incorporating sentiment classification labels; S5: Classification prediction: The online public opinion text to be classified is preprocessed and feature extracted according to the above steps, and then input into the trained model to obtain positive and negative classification results including sentiment rating (-3, -2, -1, 0, 1, 2, 3).
2. According to claim 1, a method for classifying positive and negative network public opinion is characterized in that: In step S1: Data collection: Using web crawler technology to collect online public opinion text data from major social media platforms, news websites, etc. The specific steps are as follows: Step S1: Data Collection S1.1 Determine the data source: Target platform selection: Determine the social media platforms that need to be crawled based on analysis requirements; Content type determination: Identify the type of content that needs to be collected; S1.2 Design crawler strategy: Access frequency setting: According to the access restrictions and crawler etiquette of the target platform, reasonably set the access frequency to avoid excessive pressure on the platform server; Data crawling rules: formulate data crawling rules, including the fields that need to be crawled; Anti-crawler mechanism response: Study the anti-crawler mechanism of the target platform and prepare corresponding response strategies; S1.3 Developing crawler programs: Choose a programming language: Choose a suitable programming language based on the team's technology stack and crawler requirements; Write crawler code: Use requests, BeautifulSoup, Scrapy library or framework to write crawler code and realize data crawling function; Data storage design: Design a data storage solution to store the captured data; S1.4 Execute crawler tasks: Deploy the crawler environment: deploy the crawler environment on the server or local computer to ensure that the crawler program can run normally; Run the crawler program: Execute the crawler program to start crawling data. Depending on the amount of data and the crawling speed, it may take some time to run continuously. Monitor crawler status: monitor the running status of the crawler program in real time and handle abnormal situations in a timely manner, such as network anomalies and changes in the target platform structure; S1.5 Data cleaning and preprocessing: Data deduplication: remove duplicate captured data to ensure the uniqueness of the data set; Data cleaning: dealing with noise and irrelevant information in the data; Data formatting: converting data into a format suitable for subsequent analysis; S1.6 Data storage and backup: Storing cleaned data: Storing cleaned and preprocessed data in a database or file system for subsequent analysis; Data backup: Back up data regularly to prevent data loss or damage.
3. According to claim 1, a method for classifying positive and negative network public opinion is characterized in that: The data preprocessing in step S2 includes: using tools such as Python NLTK to clean the data, using Jieba to segment the data, using NLTK to perform part-of-speech tagging, and using SnowballStemmer to extract the stems. The specific steps are as follows: S2.1 Use NLTK library or other text processing tools to clean the data: Install the NLTK library; Import necessary modules; Remove HTML tags: Use regular expressions to remove HTML tags from text; Remove special characters: Use regular expressions to remove non-alphanumeric characters; Remove stop words (optional): Download and load NLTK's stop word list to remove stop words from the text; Apply cleaning function: Apply the above cleaning function to the text data; S2.2 Use the Jieba word segmentation tool to perform word segmentation: Install the Jieba word segmentation tool; Import stuttering participle; Word segmentation: Use the stammering word segmentation to segment the cleaned text; S2.3 Use the part-of-speech tagging tool of the NLTK library to perform part-of-speech tagging: Download the part-of-speech tagging resource (if not downloaded yet); Part-of-speech tagging: Use NLTK's part-of-speech tagger to tag the text after word segmentation; S2.4 Use tools such as SnowballStemmer to extract stems: Install SnowballStemmer (usually installed with NLTK); Import SnowballStemmer; Stemming: Initialize SnowballStemmer and apply it to the tokenized text (or words after part-of-speech tagging).
4. According to claim 1, a method for classifying positive and negative network public opinion is characterized in that: Step S3: Feature extraction: Calculate the TF-IDF value using TfidfVectorizer of scikit-learn, initialize and set parameters, fit the converted text data to obtain a feature matrix, and load the sentiment dictionary (such as HowNet). The specific steps for extracting sentiment vocabulary features in the text are as follows: S3.1 Use scikit-learn's TfidfVectorizer to calculate TF-IDF values: Initialize TfidfVectorizer: Use the TfidfVectorizer class in Python's scikit-learn library for initialization; Set parameters: Set the parameters of TfidfVectorizer according to requirements; Fitting and converting text data: Input the preprocessed text data into TfidfVectorizer for fitting and conversion to obtain the TF-IDF feature matrix of the text; S3.2 Load the sentiment dictionary and extract sentiment vocabulary features: Load sentiment dictionary: select and load the appropriate sentiment dictionary; Extract sentiment vocabulary features: Extract sentiment vocabulary features in the text based on the occurrence and sentiment intensity of sentiment vocabulary in the text; this can be achieved by traversing the vocabulary in the text and matching it with the sentiment dictionary; the matched sentiment vocabulary can be quantified according to its sentiment intensity as part of the feature.
5. According to claim 1, a method for classifying positive and negative network public opinion is characterized in that: Model construction and training in step S4: Use TensorFlow / PyTorch to build a CNN model. Input the feature vector into the model, set the hyperparameters, and then train it with annotated public opinion data, use cross entropy loss and Adam optimization, and integrate sentiment classification labels. The specific steps are as follows: S4.1 Build a convolutional neural network (CNN) model using TensorFlow or PyTorch: Choose a deep learning framework: Based on your needs and familiarity, choose TensorFlow and PyTorch as deep learning frameworks; Build CNN model: Design and build convolutional neural network model using the building tools of the selected framework; S4.2 Input the feature vector into the model and set the hyperparameters: Input feature vector: The extracted feature vector (such as TF-IDF feature matrix and sentiment vocabulary features) is used as the input data of the model; Set hyperparameters: According to the requirements of the model and the characteristics of the training data, set appropriate hyperparameters, which have an important impact on the performance of the model and the training effect; S4.3 Use labeled online public opinion data for training: Prepare training data: Use labeled online public opinion data as training data for the model; the data should include text content and corresponding sentiment rating labels; Training model: Input the training data into the CNN model and use the cross entropy loss function to evaluate the difference between the model's prediction results and the true labels; Use the Adam optimizer to update the model's parameters to minimize the loss function; Incorporate sentiment rating labels: During the training process, sentiment rating labels are incorporated into the training data so that the model can learn the features corresponding to different sentiment intensities. This helps the model to more accurately identify the sentiment intensity of the text in subsequent sentiment classification tasks.
6. According to claim 1, a method for classifying positive and negative network public opinion is characterized in that: The classification prediction in step S5: preprocess and extract features of the network public opinion text to be classified according to the above steps, and then input the trained model to obtain positive and negative classification results including sentiment rating (-3, -2, -1, 0, 1, 2, 3). The specific steps are: S5.1 Text preprocessing: Clean the classified online public opinion texts; Use the Jieba word segmentation tool to segment the cleaned text; Perform part-of-speech tagging on the text after word segmentation to further understand the structure and meaning of the text; Perform stemming to restore words to their basic form for better feature matching; S5.2 Feature extraction: Use scikit-learn's TfidfVectorizer class to calculate the TF-IDF value of the text and obtain the TF-IDF feature matrix of the text; Load sentiment dictionaries, such as HowNet sentiment dictionary, and extract sentiment vocabulary features based on the occurrence and sentiment intensity of sentiment vocabulary in the text; Combine the TF-IDF feature matrix and sentiment vocabulary features into a complete feature vector; S5.3 Model input and prediction: Input the extracted feature vector into the trained convolutional neural network (CNN) model; The model processes the input feature vector and extracts high-level features of the text through structures such as convolutional layers, pooling layers, and fully connected layers; The model output contains positive and negative classification results of sentiment rating (-3, -2, -1, 0, 1, 2, 3); These sentiment ratings represent the sentiment intensity of the text, with negative values representing negative sentiment, positive values representing positive sentiment, and 0 representing neutral sentiment; S5.4 Interpretation and application of results: Based on the output of the model, explain the sentiment classification results of the text; Apply classification results to practical scenarios, such as public opinion monitoring and sentiment analysis, to assist decision-making or provide insights.