Text Sentiment Analysis Method Based on Transfer Learning and Improved Bag-of-Words Model
Through transfer learning and improving the bag of word model, a text sentiment analysis model was constructed, which solved the problem of poor model effect and high training cost under different types of product review data, and achieved low-cost and efficient sentiment analysis.
Patent Information
- Application Number
- CN202211490263.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-11-25
AI Technical Summary
In the prior art sentiment analysis model trained under review data of different types of products, there are problems such as poor model effect and high training cost.
Using a method based on transfer learning and improving bag-of-word model, the feature extractor is pre-trained through the bert base chinese model, and encoding it using the K-means clustering algorithm and fuzzy theory to construct a text sentiment analysis model.
Reduces the cost of sentiment analysis, can effectively handle reviews of new categories of products, reduces computational costs and the limitations of small data sets, and reduces catastrophic forgetting problems during model training.
Smart Images

Figure CN116089605B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular, to a text sentiment analysis method, system, computer device, and storage medium based on transfer learning and an improved bag-of-words model. Background Art
[0002] With the development of Internet technology, the number of online shopping groups is gradually increasing. As of June 2022, the scale of online shopping users reached 841 million, accounting for 80% of the overall Internet users. Facing such a large online shopping group, hundreds of millions of comments are generated on e-commerce platforms every day. These comments have high reference value for both merchants and consumers. Properly handling these valuable comments plays an important role in creating a good shopping environment. With the emergence of breakthrough technologies in the field of natural language understanding in recent years, the academic research on text sentiment analysis has become more extensive and in-depth. Traditional text sentiment analysis methods are to train a sentiment analysis model under the comment data of a certain type of commodity, so as to obtain the sentiment analysis results of the comments of this type of commodity.
[0003] However, when training a sentiment analysis model under the comment data of a certain type of commodity, if we want to apply this model to other types of commodities, due to the differences in the distribution of comment data, the effect of the model will become worse. At the same time, because deep models require a large number of parameters to be trained, if we want to retrain a model, it will cost a great deal. Therefore, when analyzing different types of commodities, it is necessary to train corresponding analysis models, which has the problem of high cost. Summary of the Invention
[0004] Based on this, in order to solve the above technical problems, a text sentiment analysis method based on transfer learning and an improved bag-of-words model is provided, which can perform sentiment analysis on the comments of different categories of commodities and reduce the cost of sentiment analysis.
[0005] A text sentiment analysis method based on transfer learning and an improved bag-of-words model, the method includes:
[0006] Collecting each comment data of different types of commodities and constructing each of the comment data into a data set;
[0007] Preprocessing the data set to obtain a processed comprehensive comment data set;
[0008] Pre-training a feature extractor according to the comprehensive comment data set, wherein the bertbase chinese model is used as the feature extractor and pre-trained on the comprehensive comment data set by using MLM;
[0009] Construct a dataset of specific product reviews, input the dataset of specific product reviews into the bert basechinese model, and extract feature vectors;
[0010] Input the feature vectors into the improved Bag ofvisual words. The improved Bag ofvisualwords clusters the feature vectors through the K-means clustering algorithm and encodes them according to fuzzy theory to obtain output vectors, and performs normalization processing on the output vectors to obtain a text sentiment analysis model;
[0011] Perform text sentiment analysis through the text sentiment analysis model.
[0012] In one embodiment, the construction of each of the review data into a dataset includes:
[0013] Save each of the review data in the form of csv, and each piece of data includes a category, a positive / negative label, and a review.
[0014] In one embodiment, the preprocessing of the dataset to obtain a processed comprehensive review dataset includes:
[0015] Extract the review parts of each of the review data from the dataset;
[0016] Use regular expressions to remove meaningless symbols and non-Chinese content from each of the review parts to obtain a comprehensive review dataset.
[0017] In one embodiment, the input of the dataset of specific product reviews into the bert basechinese model to extract feature vectors includes:
[0018] Tokenize the data in the input dataset of specific product reviews through the Tokenizer tool, and add Tokens to the tokenized samples;
[0019] Obtain the dictionary of the bert base chinese model during pre-training, and map each of the Tokens to the corresponding ID according to the dictionary;
[0020] Convert the equal-length samples mapped to IDs in the dataset of specific product reviews into a numerical matrix through the bertbase chinese model, and extract the semantic features of the sentences and the context information of the Tokens in the dataset of specific product reviews, and output through the output layer.
[0021] In one embodiment, the pre-training using MLM on the comprehensive review dataset includes:
[0022] Select target Tokens with a target proportion from the samples with the Tokens added;
[0023] Replace the first threshold number of the target Tokens with masks, replace the second threshold number of the target Tokens with random Tokens, and retain the third threshold number of the target Tokens.
[0024] In one embodiment, the construction of the specific product review dataset includes:
[0025] Obtain the review data of a specific product and construct a preliminary specific product review dataset;
[0026] Preprocess the review data in the preliminary specific product review dataset by means of regular expressions to obtain a processed specific product review dataset;
[0027] Divide the specific product review dataset to construct a training set, a validation set, and a test set.
[0028] In one embodiment, the improved Bag of visual words clusters the feature vectors through the K-means clustering algorithm and then encodes them according to fuzzy theory to obtain an output vector, and normalizes the output vector to obtain a text sentiment analysis model, including:
[0029] Extract training samples from the training set, perform semantic feature extraction through the bert base chinese model, and use the K-means clustering method for the extracted features to obtain a list of clustering centers;
[0030] Encode the extracted features using the improved Bag of visual words, and each sample is encoded as a numerical vector;
[0031] Convert the feature vector into a probability value.
[0032] A text sentiment analysis system based on transfer learning and an improved bag-of-words model, the system includes:
[0033] A data acquisition module for collecting each review data of different types of products and constructing each of the review data into a dataset;
[0034] A preprocessing module for preprocessing the dataset to obtain a processed comprehensive review dataset;
[0035] A pre-training module for pre-training a feature extractor based on the comprehensive review dataset. Herein, the bertbasechinese model is used as the feature extractor, and pre-training is performed on the comprehensive review dataset using MLM;
[0036] A feature extraction module for constructing a specific product review dataset, inputting the specific product review dataset into the bert base chinese model, and extracting feature vectors;
[0037] A model training module for inputting the feature vectors into the improved Bag of visual words. The improved Bag of visual words clusters the feature vectors through the K-means clustering algorithm and then encodes them according to the fuzzy theory to obtain output vectors, and normalizes the output vectors to obtain a text sentiment analysis model;
[0038] A sentiment analysis module for performing text sentiment analysis through the text sentiment analysis model.
[0039] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0040] Collect the review data of each product of different types, and construct each review data into a dataset;
[0041] Preprocess the dataset to obtain a processed comprehensive review dataset;
[0042] Pre-train a feature extractor based on the comprehensive review dataset. Herein, the bert base chinese model is used as the feature extractor, and pre-training is performed on the comprehensive review dataset using MLM;
[0043] Construct a specific product review dataset, input the specific product review dataset into the bert basechinese model, and extract feature vectors;
[0044] Input the feature vectors into the improved Bag of visual words. The improved Bag of visualwords clusters the feature vectors through the K-means clustering algorithm and then encodes them according to the fuzzy theory to obtain output vectors, and normalizes the output vectors to obtain a text sentiment analysis model;
[0045] Perform text sentiment analysis through the text sentiment analysis model.
[0046] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0047] Collect the respective review data of different types of commodities, and construct each of the review data into a data set;
[0048] Preprocess the data set to obtain a processed comprehensive review data set;
[0049] Pre-train a feature extractor according to the comprehensive review data set, where the bert base chinese model is used as the feature extractor, and pre-training is performed on the comprehensive review data set using MLM;
[0050] Construct a specific commodity review data set, input the specific commodity review data set into the bert basechinese model, and extract feature vectors;
[0051] Input the feature vectors into the improved Bag ofvisual words. The improved Bag ofvisualwords clusters the feature vectors through the K-means clustering algorithm and then encodes them according to the fuzzy theory to obtain output vectors, and normalizes the output vectors to obtain a text sentiment analysis model;
[0052] Perform text sentiment analysis through the text sentiment analysis model.
[0053] The above-mentioned text sentiment analysis method, system, computer device and storage medium based on transfer learning and improved bag-of-words model collect various review data of different types of commodities, and construct each of the review data into a data set; preprocess the data set to obtain a processed comprehensive review data set; pre-train a feature extractor according to the comprehensive review data set, wherein the bert base chinese model is used as the feature extractor, and MLM is used for pre-training on the comprehensive review data set; construct a specific commodity review data set, input the specific commodity review data set into the bert base chinese model, and extract feature vectors; input the feature vectors into the improved Bag of visual words, and the improved Bag of visual words clusters the feature vectors through the K-means clustering algorithm and encodes them according to the fuzzy theory to obtain an output vector, and normalize the output vector to obtain a text sentiment analysis model; perform text sentiment analysis through the text sentiment analysis model. Through transfer learning and the Bag of visual words method, it is possible to well process the reviews of newly emerging categories of commodities. At the same time, when retraining the model, since there are fewer parameters to be learned, not only can the computational cost be reduced, but also the limitation of small data sets can be successfully overcome; in addition, there is no need to fine-tune the feature extractor anymore, and only the clustering centers need to be updated according to the training data. This training strategy can well retain the knowledge learned by the model before while learning new knowledge, reduce the "catastrophic forgetting" problem, and reduce the cost of text sentiment analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 FIG. is an application environment diagram of a text sentiment analysis method based on transfer learning and improved bag-of-words model in an embodiment;
[0055] Figure 2 FIG. is a schematic flowchart of a text sentiment analysis method based on transfer learning and improved bag-of-words model in an embodiment;
[0056] Figure 3 FIG. is a schematic diagram of the process of training a text sentiment analyzer in an embodiment;
[0057] Figure 4 FIG. is a schematic diagram of the process of building a text sentiment analysis model in an embodiment;
[0058] Figure 5 FIG. is a structural block diagram of a text sentiment analysis system based on transfer learning and improved bag-of-words model in an embodiment;
[0059] Figure 6 FIG. is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0060] In order to make the objectives, technical solutions and advantages of the present application more clearly understood, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0061] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe the number of thresholds, but the number of these thresholds is not limited by these terms. These terms are only used to distinguish the first number of thresholds from another number of thresholds. For example, without departing from the scope of the present application, the first number of thresholds may be referred to as the second number of thresholds, and similarly, the second number of thresholds may be referred to as the first number of thresholds. Both the first number of thresholds and the second number of thresholds are the number of thresholds, but they are not the same number of thresholds.
[0062] The text sentiment analysis method based on transfer learning and improved bag-of-words model provided by the embodiments of the present application can be applied to, for example, Figure 1 the application environment shown. As Figure 1 shown, the application environment includes a computer device 110. The computer device 110 can collect various review data of different types of commodities and construct the review data into a data set; the computer device 110 can preprocess the data set to obtain a processed comprehensive review data set; the computer device 110 can pre-train a feature extractor according to the comprehensive review data set, where the bert base chinese model is used as the feature extractor and pre-trained on the comprehensive review data set using MLM; the computer device 110 can construct a specific commodity review data set, input the specific commodity review data set into the bert base chinese model, and extract feature vectors; the computer device 110 can input the feature vectors into the improved Bag of visual words, and the improved Bag of visual words clusters the feature vectors through the K-means clustering algorithm and encodes them according to the fuzzy theory to obtain an output vector, and normalizes the output vector to obtain a text sentiment analysis model; the computer device 110 can perform text sentiment analysis through the text sentiment analysis model. Among them, the computer device 110 can be, but is not limited to, various personal computers, laptop computers, smart phones, robots, unmanned aerial vehicles, tablet computers and other devices.
[0063] In one embodiment, as Figure 2 shown, a text sentiment analysis method based on transfer learning and improved bag-of-words model is provided, including the following steps:
[0064] Step 202: Collect the review data of different types of products, and construct the review data into a dataset.
[0065] The computer device can collect the review data of different types of products. In this embodiment, the collected review data can include the review data of 10 categories of products. There are more than 60,000 pieces of product review data, including about 30,000 positive reviews and 30,000 negative reviews. For example, the collected product review data can include books (3,851 pieces), tablets (10,000 pieces), mobile phones (2,323 pieces), fruits (10,000 pieces), shampoos (10,000 pieces), water heaters (575 pieces), Mengniu (2,033 pieces), clothes (10,000 pieces), computers (3,992 pieces), hotels (10,000 pieces). After the computer device collects the review data, it can construct a dataset.
[0066] Step 204: Preprocess the dataset to obtain the processed comprehensive review dataset.
[0067] Among them, the preprocessing can be an operation performed on the pre-trained feature extractor. The pre-trained feature extractor only needs to use the review part in the dataset. Therefore, the processed comprehensive review dataset only contains the review part.
[0068] Step 206: Pre-train the feature extractor according to the comprehensive review dataset. Among them, the bert base chinese model is used as the feature extractor, and MLM is used for pre-training on the comprehensive review dataset.
[0069] The computer device can use the bert base chinese model provided by Hugging Face as the feature extractor. Among them, the bert base chinese model has been pre-trained in a large Chinese corpus, and there are about 110 million parameters to be learned. Therefore, it can be imagined that if the transfer learning method is not used, the computational cost required to re-train a Chinese sentiment analysis model each time is huge. Specifically, the bertbase chinese feature extractor has been pre-trained in a large general corpus. Here, in order to improve the effect of the feature extractor, it is further pre-trained on the comprehensive review dataset, and the pre-training method used is MLM.
[0070] Step 208: Construct a specific product review dataset, input the specific product review dataset into the bert basechinese model, and extract the feature vectors.
[0071] After the feature extractor is further pre-trained successfully, a specific product review dataset can be reconstructed in the computer device. Among them, the specific product review dataset is different from the comprehensive review dataset, and the specific product review dataset can be used to extract feature vectors. In this embodiment, when training the sentiment analysis model of specific product reviews, the specific product review dataset is input into the further pre-trained bert base chinese model to extract features. During the training process, the parameters of the bertbase chinese model do not need to be learned again, that is, transfer learning.
[0072] Step 210, input the feature vectors into the improved Bag of visual words. The improved Bag of visual words clusters the feature vectors through the K-means clustering algorithm and then encodes them according to the fuzzy theory to obtain an output vector, and normalizes the output vector to obtain a text sentiment analysis model.
[0073] The traditional Bag of visual words model contains feature extraction, obtaining cluster centers by k-means clustering, encoding, and normalization. The improved Bag of visual words, that is, the improved Bag of visual words, encodes according to the fuzzy theory during encoding.
[0074] Step 212, perform text sentiment analysis through the text sentiment analysis model.
[0075] In this embodiment, through transfer learning and the Bag of visual words method, it is possible to well process the reviews of newly emerging categories of products. At the same time, when retraining the model, since there are fewer parameters to be learned, not only can the computational cost be reduced, but also the limitation of small datasets can be successfully overcome; in addition, there is no need to fine-tune the feature extractor again, and only the cluster centers need to be updated according to the training data. This training strategy can well retain the knowledge learned by the model before while learning new knowledge, reduce the "catastrophic forgetting" problem, and reduce the cost of text sentiment analysis.
[0076] In one embodiment, a text sentiment analysis method based on transfer learning and an improved Bag of visual words model may further include the process of constructing a dataset. The specific process includes: saving each review data in the form of csv, and each data includes a category, a positive and negative label, and a review.
[0077] Among them, each row in the review data saved in the form of csv is a product review data, and each data includes three parts: category, positive and negative label, and review.
[0078] In one embodiment, a text sentiment analysis method based on transfer learning and an improved bag-of-words model may further include a data preprocessing process. The specific process includes: extracting the comment part of each comment data from the dataset; using regular expressions to remove meaningless symbols and non-Chinese content in each comment part to obtain a comprehensive comment dataset.
[0079] In one embodiment, a text sentiment analysis method based on transfer learning and an improved bag-of-words model may further include a process of pre-training a feature extractor. The specific process includes: tokenizing the data in the input specific product comment dataset through the Tokenizer tool and adding Tokens to the tokenized samples; obtaining the dictionary of the bert basechinese model during pre-training and mapping each Token to the corresponding ID according to the dictionary; converting the equal-length samples mapped to IDs in the specific product comment dataset into a numerical matrix through the bertbasechinese model and extracting the semantic features of the sentences and the context information of the Tokens in the specific product comment dataset, and outputting through the output layer.
[0080] Hugging Face provides a tool called Tokenizer that can tokenize the input Chinese comment data by character unit, add special tokens to the tokenized samples, and then map each token to the corresponding id according to the dictionary obtained by the bert base chinese model during pre-training. Among them, a token refers to the smallest unit after the text is segmented, and in this embodiment, it can be a character.
[0081] Among them, [CLS] is placed at the beginning of the sentence, and the representation vector C obtained through the feature extractor can be used for subsequent classification tasks; [SEP] is placed at the end of the sentence; [UNK] refers to unknown characters; [MASK] is used to cover some words in the sentence. After covering the words with [MASK], the [MASK] vector output by the bert base chinese model is used to predict what the word is, which is also one of the tasks of the model pre-training; [PAD] is used to pad sentences shorter than the maximum length so that the data can be input into the model with equal length.
[0082] In this embodiment, the bert base chinese model consists of three parts: an embedding layer (EmbeddingLayer), an encoder of Transformer, and an output layer. Among them, the embedding layer is used to convert the equal-length samples mapped to ids into a numerical matrix representation of [512, 768] dimensions; the encoder of Transformer is used to extract the semantic features of the input sentence and the context information of tokens, and it is a dynamic encoder; the output layer processes the output of the encoder to complete different downstream tasks.
[0083] In one embodiment, a text sentiment analysis method based on transfer learning and an improved bag-of-words model may further include a process of pre-training using the MLM task. The specific process includes: selecting target tokens with a target proportion from the samples with Tokens added; selecting a first threshold number of target tokens to be replaced with mask, selecting a second threshold number of target tokens to be replaced with random tokens, and selecting a third threshold number of target tokens to be retained.
[0084] When further pre-training the bert base chinese model, since the chinese text sentiment analysis does not handle the case of sentence pairs, only the semantic understanding ability of the model needs to be further pre-trained using the MLM task. Among them, the target proportion can be 15%; the first threshold number can be 80%; the second threshold number can be 10%; the third threshold number can be 10%.
[0085] Specifically, after selecting 15% of the tokens from the samples with Tokens added, not all of them are replaced with the [mask] token. The actual operation is: from the selected 15% part, 80% of them are replaced with [mask]; 10% are replaced with a random token; and the remaining 10% retain the original token. Among them, the mask_token_list list is used to save the original tokens replaced by [mask], and the mask_position_list list is used to save the positions of the original tokens replaced by [mask] in the samples.
[0086] In one embodiment, a text sentiment analysis method based on transfer learning and an improved bag-of-words model may further include a process of processing a specific product review dataset. The specific process includes: obtaining the review data of a specific product and constructing a preliminary specific product review dataset; preprocessing the review data in the preliminary specific product review dataset in a regular expression manner to obtain a processed specific product review dataset; dividing the specific product review dataset to construct a training set, a validation set, and a test set.
[0087] In this embodiment, the specific commodities can be divided into two categories: digital products and snacks. The review dataset of digital products has 4,000 reviews, with approximately 2,000 positive and negative reviews each. The review dataset of snack products has 5,000 reviews, with approximately 2,500 positive and negative reviews each. For the review data in the two datasets, meaningless symbols and non-Chinese content are removed using regular expressions, and then the two datasets are divided according to the ratios of 80%, 10%, and 10% respectively to construct a training set, a validation set, and a test set.
[0088] In one embodiment, a text sentiment analysis method based on transfer learning and an improved bag-of-words model may further include improving the processing process of the Bag of visual words. The specific process includes: extracting training samples from the training set, performing semantic feature extraction through the bert base chinese model, and using the K-means clustering method for the extracted features to obtain a list of clustering centers; encoding the extracted features using the improved Bag of visual words, and each sample is encoded as a numerical vector; converting the feature vector into a probability value.
[0089] Among them, when training the sentiment analyzer for the reviews of a certain type of commodity, 50% of the samples are randomly selected from the training set of the review dataset of this type of commodity, and the further pre-trained bert base chinese model is used to perform semantic feature extraction on these samples.
[0090] After extracting the features, the K-means clustering algorithm is adopted to cluster these feature vectors. After clustering, K clustering centers are obtained, and these clustering center vectors are saved in the Centre_List list. Here, K is equal to 300, and the length of the Centre_List list is also 300.
[0091] During encoding, according to the fuzzy theory, after calculating the Euclidean distance between the local feature and each clustering center, it is no longer encoded only with 0 and 1, but is encoded with a number between (0, 1] according to the formula m(D i ,C j )=exp(-(D(i,j)-min(D)) 2 / σ). This is exactly the core of the improvement of the traditional Bag of visual words method and the key to the good effect of this method in natural language processing. In the formula m(D i ,C j )=exp(-(D(i,j)-min(D)) 2 / σ), m(D i ,C j) represents the i-th local feature of the sample and its encoding at the j-th position; D(i,j) represents the Euclidean distance between the i-th local feature of the sample and the j-th cluster center; σ is a hyperparameter that can be adjusted according to the model's performance; min(D) represents the shortest Euclidean distance between the i-th local feature of the sample and all cluster centers.
[0092] After obtaining the encodings of all positions of the i-th local feature of the sample, the local feature can be represented as a 300-dimensional numerical vector, denoted by A(D i ,C k ), that is, the i-th local feature of the sample and its numerical vector at the k-th position.
[0093] After obtaining the vector representations of all local features of a sample, the formula D p = can be used to add up these vectors to obtain the output D p of the sample, where n is the number of local features.
[0094] After encoding, a sample obtains a 300-dimensional vector representation. The softmax function is used to normalize the output values, converting all values in the vector to probability (between 0 and 1) values, and the sum of all probability values equals 1.
[0095] In one embodiment, the process of training a text sentiment analyzer is as Figure 3 shown. By constructing a comprehensive Chinese review dataset for an e-commerce platform, followed by data preprocessing, and further pre-training a feature extractor; then, a specific product review dataset can be constructed, preprocessed, and divided into a training set, a validation set, and a test set; among them, the training set can be used for feature extraction in the subsequent text sentiment analysis process and for improving the Bag of visual words encoding; after initializing the parameters to be trained in the model, the model can be trained based on the data in the training set and the validation set, and the optimal parameter model can be saved after model validation, and then the model can be tested with the data in the test set, and finally a text sentiment analyzer can be obtained.
[0096] In one embodiment, a text sentiment analysis method based on transfer learning and an improved bag-of-words model may further include the process of building a text sentiment analysis model, as Figure 4 shown. Randomly select 50% of the samples from the training set, use the bert base chinese model to extract semantic features, and use the K-means clustering method for the extracted features to obtain a list of cluster centers;
[0097] Next, semantic feature extraction is performed on all samples in the training set using the bert base chinese model, and the extracted features are encoded using the improved Bag of visual words method. Each sample is encoded as a 300-dimensional numerical vector;
[0098] Next, the fully connected layer converts the input numerical vector containing all feature information into the probabilities of final classification into each category. Here, it is a binary classification task, and the softmax function is used as the activation function. The output of the fully connected layer contains two neurons. When training a new model, only the parameters of the fully connected layer here need to be learned.
[0099] Among them, to control the parameters not to be learned, the backpropagation operation of the network in pytorch is based on the Variable object. There is a parameter requires_grad in Variable. Setting requires_grad = False, the network will not calculate the gradient for this layer. When validating and testing the model, directly use the cluster centers generated during training without performing K-means clustering again.
[0100] It should be understood that although the steps in the above flow chart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flow chart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential either, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0101] In one embodiment, as Figure 5 shown, a text sentiment analysis system based on transfer learning and improved bag-of-words model is provided, including: a data collection module 510, a preprocessing module 520, a pre-training module 530, a feature extraction module 540, a model training module 550, and a sentiment analysis module 560, where:
[0102] The data collection module 510 is used to collect various review data of different types of commodities and construct the review data into a data set;
[0103] The preprocessing module 520 is used to preprocess the data set to obtain a processed comprehensive review data set;
[0104] A pre-training module 530 for pre-training a feature extractor according to a comprehensive review dataset, where the bert basechinese model is used as the feature extractor, and MLM is used for pre-training on the comprehensive review dataset;
[0105] A feature extraction module 540 for constructing a specific product review dataset, inputting the specific product review dataset into the bert base chinese model, and extracting feature vectors;
[0106] A model training module 550 for inputting the feature vectors into an improved Bag ofvisual words. The improved Bagofvisual words clusters the feature vectors through the K-means clustering algorithm and encodes them according to the fuzzy theory to obtain output vectors, and normalizes the output vectors to obtain a text sentiment analysis model;
[0107] A sentiment analysis module 560 for performing text sentiment analysis through the text sentiment analysis model.
[0108] In one embodiment, the data collection module 510 is further configured to save each comment data in the form of csv, and each data includes a category, a positive / negative label, and a comment.
[0109] In one embodiment, the preprocessing module 520 is further configured to extract the comment part of each comment data from the dataset; use regular expressions to remove meaningless symbols and non-Chinese content in each comment part to obtain a comprehensive review dataset.
[0110] In one embodiment, the feature extraction module 540 is further configured to tokenize the data in the input target dataset through the Tokenizer tool, and add Tokens to the tokenized samples; obtain the dictionary of the bert base chinese model during pre-training, and map each Token to the corresponding ID according to the dictionary; convert the equal-length samples mapped to IDs in the specific product review dataset into a numerical matrix through the bertbase chinese model, and extract the semantic features of the sentences and the context information of the Tokens in the specific product review dataset, and output through the output layer.
[0111] In one embodiment, the model training module 550 is further configured to pre-train the bert base chinese model using the MLM task; select target Tokens with a target proportion from the samples with Tokens added; select a first threshold number of target Tokens and replace them with masks, select a second threshold number of target Tokens and replace them with random Tokens, and select a third threshold number of target Tokens to remain.
[0112] In one embodiment, the data acquisition module 510 is further configured to obtain comment data of a specific commodity and construct a preliminary specific commodity comment data set; preprocess the comment data in the preliminary specific commodity comment data set by using regular expressions to obtain a processed specific commodity comment data set; divide the specific commodity comment data set to construct a training set, a validation set, and a test set.
[0113] In one embodiment, the model training module 550 is further configured to extract training samples from the training set, perform semantic feature extraction through the bertbase chinese model, use the K-means clustering method for the extracted features to obtain a list of cluster centers; encode the extracted features by using an improved Bag of visual words, and each sample is encoded as a numerical vector; convert the feature vector into a probability value.
[0114] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 6 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a text sentiment analysis method based on transfer learning and an improved bag-of-words model. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device may be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0115] Those skilled in the art can understand that Figure 6 the structure shown in
[0116] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:
[0117] Collect the review data of various types of commodities, and construct each review data into a data set;
[0118] Preprocess the data set to obtain a processed comprehensive review data set;
[0119] Pre-train a feature extractor based on the comprehensive review data set. Among them, use the bert base chinese model as the feature extractor and perform pre-training on the comprehensive review data set using MLM;
[0120] Construct a specific commodity review data set, input the specific commodity review data set into the bert base chinese model, and extract feature vectors;
[0121] Input the feature vectors into the improved Bag of visual words. The improved Bag of visual words clusters the feature vectors through the K-means clustering algorithm and then encodes them according to the fuzzy theory to obtain an output vector, and performs normalization processing on the output vector to obtain a text sentiment analysis model;
[0122] Perform text sentiment analysis through the text sentiment analysis model.
[0123] In one embodiment, when the processor executes the computer program, the following steps are further implemented: save each review data in the form of csv, and each piece of data includes a category, a positive / negative label, and a review.
[0124] In one embodiment, when the processor executes the computer program, the following steps are further implemented: extract the review part of each review data from the data set; use regular expressions to remove meaningless symbols and non-Chinese content in each review part to obtain a comprehensive review data set.
[0125] In one embodiment, when the processor executes the computer program, the following steps are further implemented: tokenize the data in the input specific commodity review data set through the Tokenizer tool, and add Tokens to the tokenized samples; obtain the dictionary of the bert base chinese model during pre-training, and map each Token to the corresponding ID according to the dictionary; convert the equal-length samples mapped to IDs in the specific commodity review data set into a numerical matrix through the bert base chinese model, and extract the semantic features of the sentences and the context information of the Tokens in the specific commodity review data set, and output through the output layer.
[0126] In one embodiment, when the processor executes the computer program, the following steps are further implemented: select target Tokens with a target proportion from the samples with Tokens added; select a first threshold number of target Tokens and replace them with masks, select a second threshold number of target Tokens and replace them with random Tokens, and select a third threshold number of target Tokens to remain unchanged.
[0127] In one embodiment, when the processor executes the computer program, the following steps are further implemented: obtain the review data of a specific commodity and construct a preliminary specific commodity review data set; preprocess the review data in the preliminary specific commodity review data set in the way of regular expressions to obtain the processed specific commodity review data set; divide the specific commodity review data set to construct a training set, a validation set, and a test set.
[0128] In one embodiment, when the processor executes the computer program, the following steps are further implemented: extract training samples from the training set, perform semantic feature extraction through the bert base chinese model, and use the K-means clustering method for the extracted features to obtain a list of clustering centers; encode the extracted features using the improved Bag of visual words, and each sample is encoded as a numerical vector; convert the feature vector into a probability value.
[0129] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the following steps are implemented:
[0130] Collect the review data of each different type of commodity and construct each review data into a data set;
[0131] Preprocess the data set to obtain a processed comprehensive review data set;
[0132] Pre-train a feature extractor according to the comprehensive review data set. Among them, use the bert base chinese model as the feature extractor and perform pre-training on the comprehensive review data set using MLM;
[0133] Construct a specific commodity review data set, input the specific commodity review data set into the bert base chinese model, and extract feature vectors;
[0134] Input the feature vectors into the improved Bag of visual words. The improved Bag of visual words clusters the feature vectors through the K-means clustering algorithm and then encodes them according to the fuzzy theory to obtain an output vector, and performs normalization processing on the output vector to obtain a text sentiment analysis model;
[0135] Perform text sentiment analysis through a text sentiment analysis model.
[0136] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: saving each comment data in the form of csv, and each piece of data includes a category, a positive / negative label, and a comment.
[0137] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: extracting the comment part of each comment data from the dataset; removing meaningless symbols and non-Chinese content in each comment part by using regular expressions to obtain a comprehensive comment dataset.
[0138] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: tokenizing the data in the input specific product comment dataset through the Tokenizer tool, and adding Tokens to the tokenized samples; obtaining the dictionary of the bert base chinese model during pre-training, and mapping each Token to the corresponding ID according to the dictionary; converting the equal-length samples mapped to IDs in the specific product comment dataset into a numerical matrix through the bert base chinese model, and extracting the semantic features of the sentences and the context information of the Tokens in the specific product comment dataset, and outputting through the output layer.
[0139] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: selecting target Tokens with a target proportion from the samples with Tokens added; selecting a first threshold number of target Tokens to be replaced with mask, selecting a second threshold number of target Tokens to be replaced with random Tokens, and selecting a third threshold number of target Tokens to be retained.
[0140] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: obtaining the comment data of a specific product and constructing a preliminary specific product comment dataset; preprocessing the comment data in the preliminary specific product comment dataset by using regular expressions to obtain a processed specific product comment dataset; dividing the specific product comment dataset to construct a training set, a validation set, and a test set.
[0141] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: extracting training samples from the training set, performing semantic feature extraction through the bert base chinese model, using the K-means clustering method for the extracted features to obtain a list of clustering centers; encoding the extracted features by using the improved Bag of visual words, and each sample is encoded as a numerical vector; converting the feature vector into a probability value.
[0142] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0143] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0144] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A text sentiment analysis method based on transfer learning and an improved bag-of-words model, characterized in that, The method includes: Collecting the review data of various types of commodities, and constructing each of the review data into a data set; Preprocessing the data set to obtain a processed comprehensive review data set; Pre-training a feature extractor according to the comprehensive review data set, where the bertbase chinese model is used as the feature extractor, and pre-training is performed on the comprehensive review data set using MLM; Constructing a specific commodity review data set, inputting the specific commodity review data set into the bertbase chinese model, and extracting feature vectors; Input the feature vector into the improved Bag of visual words. The improved Bag of visual words clusters the feature vector through the K-means clustering algorithm and then encodes it according to the fuzzy theory to obtain an output vector, and performs normalization processing on the output vector to obtain a text sentiment analysis model; during encoding, according to the fuzzy theory, after calculating the Euclidean distance between the local feature and each cluster center, it no longer only takes 0 and 1 for encoding, but according to the formula m(D i ,C j ) = exp(-(D(i,j)-min(D)) 2 / σ), and encodes it with a number between (0, 1]; m(D i ,C j ) represents the encoding of the i-th local feature of the sample at the j-th position; D(i,j) represents the Euclidean distance between the i-th local feature of the sample and the j-th cluster center; σ is a hyperparameter; min(D) represents the shortest Euclidean distance between the i-th local feature of the sample and all cluster centers; obtain the encodings of all positions of the i-th local feature of the sample, and represent the local feature as a 300-dimensional numerical vector, denoted by A(D i ,C k ), that is, the numerical vector of the i-th local feature of the sample at the k-th position; after obtaining the vector representation of all local features of a sample, through the formula add up these vectors to obtain the output D p of the sample, where n is the number of local features; after encoding is completed, a sample obtains a 300-dimensional vector representation, and the softmax function is used to perform normalization operation on the output value, and all values in the vector are converted into probability values (between 0 and 1), and the sum of all probability values is equal to 1; Performing text sentiment analysis through the text sentiment analysis model.
2. The text sentiment analysis method based on transfer learning and improved bag-of-words model according to claim 1, characterized in that, The constructing each of the review data into a data set includes: Saving each of the review data in the form of csv, and each piece of data includes a category, a positive / negative label, and a review.
3. The text sentiment analysis method based on transfer learning and improved bag-of-words model according to claim 1, characterized in that The preprocessing the data set to obtain a processed comprehensive review data set includes: Taking out the review part of each of the review data from the data set; Removing meaningless symbols and non-Chinese content in each of the review parts by using regular expressions to obtain a comprehensive review data set.
4. The text sentiment analysis method based on transfer learning and improved bag-of-words model according to claim 1, characterized in that, The inputting the specific commodity review data set into the bert base chinese model and extracting feature vectors includes: Performing word segmentation on the data in the input specific commodity review data set through the Tokenizer tool, and adding Tokens to the segmented samples; Obtaining the dictionary of the bert base chinese model during pre-training, and mapping each of the Tokens to the corresponding ID according to the dictionary; Converting the equal-length samples mapped to IDs in the specific commodity review data set into a numerical matrix through the bert base chinese model, and extracting the semantic features of the sentences and the context information of the Tokens in the specific commodity review data set, and outputting through the output layer.
5. The text sentiment analysis method based on transfer learning and improved bag-of-words model according to claim 4, characterized in that, The pre-training on the comprehensive review data set using MLM includes: Selecting target Tokens with a target proportion from the samples with the Tokens added; Selecting a first threshold number of the target Tokens to be replaced with mask, selecting a second threshold number of the target Tokens to be replaced with random Tokens, and selecting a third threshold number of the target Tokens to be retained.
6. The text sentiment analysis method based on transfer learning and improved bag-of-words model according to claim 1, characterized in that, The constructing the specific commodity review data set includes: Obtaining the review data of a specific commodity and constructing a preliminary specific commodity review data set; Preprocessing the review data in the preliminary specific commodity review data set by using regular expressions to obtain a processed specific commodity review data set; Dividing the specific commodity review data set to construct a training set, a validation set, and a test set.
7. The text sentiment analysis method based on transfer learning and improved bag-of-words model according to claim 6, characterized in that, The improved Bag of visual words clusters the feature vectors through the K-means clustering algorithm and encodes according to the fuzzy theory to obtain an output vector, and normalizes the output vector to obtain a text sentiment analysis model, including: Extract training samples from the training set, perform semantic feature extraction through the bertbase chinese model, and use the K-means clustering method on the extracted features to obtain a list of cluster centers; Encode the extracted features using the improved Bag of visual words, and each sample is encoded as a numerical vector; Convert the feature vector into a probability value.
8. A text sentiment analysis system based on transfer learning and an improved bag-of-words model, characterized in that, The system includes: A data collection module for collecting various review data of different types of commodities and constructing each of the review data into a data set; A preprocessing module for preprocessing the data set to obtain a processed comprehensive review data set; A pre-training module for pre-training a feature extractor according to the comprehensive review data set. Among them, the bertbasechinese model is used as the feature extractor, and MLM is used for pre-training on the comprehensive review data set; A feature extraction module for constructing a specific commodity review data set, inputting the specific commodity review data set into the bertbase chinese model, and extracting a feature vector; A model training module for inputting the feature vectors into an improved Bag of Visual Words. The improved Bag of Visual Words clusters the feature vectors through the K-means clustering algorithm and then encodes them according to the fuzzy theory to obtain an output vector, and normalizes the output vector to obtain a text sentiment analysis model. During encoding, according to the fuzzy theory, after calculating the Euclidean distance between the local features and each cluster center, instead of only taking 0 and 1 for encoding, it is encoded according to the formula m(D i ,C j ) = exp(-(D(i,j)-min(D)) 2 / σ), using numbers between (0, 1] for encoding; m(D i ,C j ) represents the encoding of the i-th local feature of the sample at the j-th position; D(i,j) represents the Euclidean distance between the i-th local feature of the sample and the j-th cluster center; σ is a hyperparameter; min(D) represents the shortest Euclidean distance between the i-th local feature of the sample and all cluster centers; to obtain the encodings of all positions of the i-th local feature of the sample, the local feature is represented as a 300-dimensional numerical vector, denoted by A(D i ,C k ), that is, the numerical vector of the i-th local feature of the sample at the k-th position. After obtaining the vector representations of all local features of a sample, through the formula add up these vectors to obtain the output D p of the sample, where n is the number of local features. After encoding is completed, a sample obtains a 300-dimensional vector representation. The softmax function is used to normalize the output values, converting all values in the vector to probability values (between 0 and 1), and the sum of all probability values is equal to 1; A sentiment analysis module for performing text sentiment analysis through the text sentiment analysis model.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
A method for sentiment analysis of film reviews based on deep learning and natural language processing
AU2020100710A4
Visual word bag feature weighting method and system based on classification drive
CN103399870A