Online comment sentiment classification method based on BERT, BiGRU and multi-head attention mechanism

By combining BERT, BiGRU and multi-head attention mechanisms, the problems of high computational complexity and resource consumption when dealing with long text emotion classification are solved, and efficient extraction and accurate classification of complex semantics and key emotional information are achieved.

CN120030166APending Publication Date: 2025-05-23BEIJING TECH & BUSINESS UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510113985.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When existing deep learning models deal with emotion classification of multi-part texts, they have high computational complexity and resource consumption, making it difficult to effectively capture complex semantics and key emotional information in the text, resulting in limited accuracy of emotion classification.

Method used

Combining BERT's deep bidirectional context understanding ability, global dependence modeling of multi-head attention, and BiGRU's long-term dependence capture ability on sequence data, an online commentary emotion classification model is constructed.

Benefits of technology

It realizes efficient and accurate classification of online commentary emotions, can understand complex semantics more accurately, extract key emotional information, and improve the accuracy of emotional classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030166A_ABST
    Figure CN120030166A_ABST
Patent Text Reader

Abstract

The invention discloses an online comment sentiment classification method based on BERT, BiGRU and a multi-head attention mechanism. The method comprises the steps that data preprocessing is conducted on a text data set; constructing an emotion classification model based on BERT, BiGRU and a multi-head attention mechanism; loading a training set training model after data preprocessing, and training the data in batches through multi-round iteration; and classifying the emotion types of the online comments based on the model. According to the method, efficient and accurate classification of online comment emotions is realized by combining the deep bidirectional context understanding ability of BERT, the global dependency modeling of multi-head attention and the long-term dependency relationship capturing ability of BiGRU on sequence data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of natural language processing and sentiment analysis, and in particular to an online review sentiment classification method based on BERT, BiGRU and a multi-head attention mechanism. Background Art

[0002] With the popularization of the Internet and the rise of social media, more and more tourists choose to write online reviews on social platforms to share their travel experiences and feelings. These platforms generate a large amount of review data every day. Mining the emotional tendencies of these review data can help scenic spot managers understand the overall satisfaction of tourists and improve the service quality of scenic spots accordingly, thereby enhancing the attractiveness and competitiveness of scenic spots.

[0003] Although existing deep learning models, such as BERT, Long Short-Term Memory (LSTM), and Gated Recurrent Unit (GRU), have made significant progress in sentiment analysis tasks, they still face many challenges in processing sentiment classification of multi-sentence long texts. Although the BERT model can effectively capture the deep semantics of texts through its deep bidirectional attention mechanism, its computational complexity and resource consumption limit its application scope when faced with long texts. In addition, although traditional RNNs and their variants, such as LSTM and GRU, have certain advantages in processing sequence data, they often find it difficult to extract key sentiment information from texts when faced with comments with rich sentiment levels and complex structures, resulting in limited accuracy in sentiment classification. These limitations limit the effectiveness of existing models in practical applications and also provide new research directions and room for improvement for the development of sentiment analysis technology. Summary of the invention

[0004] The present invention provides an online review sentiment classification method based on BERT, Bidirectional Gated Recurrent Unit (BiGRU) and multi-head attention mechanism to overcome the shortcomings of the prior art. The method achieves efficient and accurate classification of online review sentiment by combining BERT's deep bidirectional context understanding ability, multi-head attention's global dependency modeling and BiGRU's ability to capture long-term dependencies of sequence data.

[0005] The online review sentiment classification method based on BERT, BiGRU and multi-head attention mechanism provided by the present invention specifically includes the following steps:

[0006] Step 1: Data preprocessing. Divide the text data set into a training set, a validation set, and a test set in a ratio of 8:1:1, and then preprocess the text data, including text segmentation, ID conversion, sentence truncation and padding, and mask creation, to obtain data that can be used by the model. The specific steps are steps 1.1 to 1.4:

[0007] Step 1.1 Use the BertTokenizer of the BERT pre-trained model to tokenize each piece of text data, add a special classification tag [CLS] before the segmentation result of each piece of text data, and record the length of the sentence after segmentation.

[0008] Step 1.2 converts each word in the word segmentation result list into its corresponding ID in the BERT vocabulary.

[0009] Step 1.3 truncates or pads the sentence according to the specified pad_size. If the sentence length is greater than or equal to pad_size, truncate the sentence and keep only the first pad_size words; if the sentence length is less than pad_size, pad 0 after the sentence until pad_size is reached.

[0010] Step 1.4 creates a mask based on the actual length of the sentence, with the positions of actual words represented by 1 and the padded placeholder positions represented by 0.

[0011] Step 2: Construct a sentiment classification model. Construct a sentiment classification model based on BERT, BiGRU and multi-head attention mechanism. The specific steps are steps 2.1 to 2.4:

[0012] Step 2.1 uses the BERT pre-trained model as a text encoder to obtain a deep feature vector representation of the text.

[0013] Step 2.2: Based on the features output by BERT, a multi-head attention mechanism is applied to further model the global dependencies between different words in a sentence, thereby enhancing the representation capability of semantic features.

[0014] The attention-weighted features in step 2.3 are input into the bidirectional GRU network to capture contextual information, especially long-range dependencies in sequence data.

[0015] Step 2.4 adds Dropout to the output layer of the bidirectional GRU to reduce overfitting, and uses the output of the last time step of the bidirectional GRU to map it to the classification space through a fully connected layer.

[0016] Step 3: Model training. Load the training set training model after data preprocessing, and train the data in batches through multiple rounds of iterations: in each batch, perform forward propagation to calculate the prediction results, use the cross entropy loss function to calculate the loss value between the predicted output and the actual label, and then perform backpropagation to calculate the gradient, and update the model parameters through the AdamW optimizer, while dynamically adjusting the learning rate using linear warm-up and decay strategies. During the training process, evaluate the performance of the model on the validation set at regular intervals, and save the model parameters with the smallest validation set loss as the best model. If the loss of the validation set does not improve within multiple batches, the early stopping mechanism is triggered to terminate the training. The entire process outputs training and validation logs for monitoring the performance of the model.

[0017] Step 4: Online review sentiment classification. Classify the sentiment types of online reviews based on the model and save the classification results. The specific steps are step 4.1 to step 4.4:

[0018] Step 4.1 performs data preprocessing on the online comments to be classified, including text segmentation, ID conversion, sentence truncation and padding, and mask creation to convert the text data into the format required by the model.

[0019] Step 4.2 converts the preprocessed text data into tensors.

[0020] Step 4.3 loads the parameters of the best model during the training process and inputs the converted tensor data into the model for sentiment classification prediction.

[0021] Step 4.4 The model outputs the sentiment category label for each comment and saves the result in a CSV file.

[0022] Compared with the prior art, the method of the present invention has the following beneficial effects:

[0023] (1) By combining BERT’s deep bidirectional contextual representation, it can more accurately understand the complex semantics and implicit emotions in the text.

[0024] (2) By introducing the multi-head attention mechanism, the model is allowed to learn information in parallel in different representation subspaces, thereby enhancing the expressiveness of the model. The attention mechanism dynamically focuses on the most important part of the text, improving the ability to extract key information.

[0025] (3) By introducing BiGRU, the model can capture long-distance dependencies in the text, thereby enhancing the modeling ability of sequence information and achieving accurate identification of sentiment tendencies.

[0026] (4) When processing online review sentiment classification, the method of the present invention can not only utilize the powerful semantic representation ability of BERT, but also use the sequence modeling ability of BiGRU to more deeply understand and classify the sentiment information in the text. This combination enables the model to perform well in processing complex sentiment expressions and long text dependencies. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a flowchart of an online review sentiment classification method based on BERT, BiGRU and multi-head attention mechanism provided by the present invention;

[0028] Figure 2 It is a structural diagram of the sentiment classification model constructed by the present invention based on BERT, BiGRU and multi-head attention mechanism; DETAILED DESCRIPTION

[0029] In order to clarify the purpose and technical solution of the present invention, the embodiments of the present invention are further described below in conjunction with the accompanying drawings.

[0030] This paper proposes an online review sentiment classification method based on BERT, BiGRU and multi-head attention mechanism. Figure 1 The specific implementation process mainly includes four parts: data preprocessing, sentiment classification model construction, model training, and online comment sentiment classification. The specific implementation steps are as follows:

[0031] Step 1: Data preprocessing. Divide the text data set into a training set, a validation set, and a test set in a ratio of 8:1:1, and then preprocess the text data, including text segmentation, ID conversion, sentence truncation and padding, and mask creation, to obtain data that can be used by the model. The specific steps are:

[0032] Step 1.1 Use the BertTokenizer of the BERT pre-trained model to tokenize each piece of text data, add a special classification tag [CLS] before the segmentation result of each piece of text data, and record the length of the sentence after segmentation.

[0033] Step 1.2 converts each word in the word segmentation result list into its corresponding ID in the BERT vocabulary.

[0034] Step 1.3 truncates or pads the sentence according to the specified pad_size. If the sentence length is greater than or equal to pad_size, truncate the sentence and keep only the first pad_size words; if the sentence length is less than pad_size, pad 0 after the sentence until pad_size is reached.

[0035] Step 1.4 creates a mask based on the actual length of the sentence, with the positions of actual words represented by 1 and the padded placeholder positions represented by 0.

[0036] Step 2: Construct a sentiment classification model. Construct a sentiment classification model based on BERT, BiGRU and multi-head attention mechanism. The model structure is as follows: Figure 2 The specific steps are:

[0037] Step 2.1 uses the BERT pre-trained model as a text encoder to obtain a deep feature vector representation of the text.

[0038] Step 2.2: Based on the features output by BERT, a multi-head attention mechanism is applied to further model the global dependencies between different words in a sentence, thereby enhancing the representation capability of semantic features.

[0039] The attention-weighted features in step 2.3 are input into the bidirectional GRU network to capture contextual information, especially long-range dependencies in sequence data.

[0040] Step 2.4 adds Dropout to the output layer of the bidirectional GRU to reduce overfitting, and uses the output of the last time step of the bidirectional GRU to map it to the classification space through a fully connected layer.

[0041] Specifically, BERT uses a bidirectional Transformer encoder, which, through its self-attention mechanism, can simultaneously consider the left and right context information of each word when processing the input text, thereby effectively capturing the deep semantic relationship in the text. The input of BERT consists of three embeddings: tag embedding, sentence embedding, and position embedding. Among them, the tag embedding represents the embedding of the current word, the sentence embedding represents the index embedding of the sentence where the current word is located, and the position embedding represents the position index of the current word in the sentence. These embeddings are spliced ​​together to form the final input representation. The output of BERT is a set of vector representations generated for each input word. These vectors integrate the semantic information of the entire text and provide powerful language understanding capabilities for downstream tasks.

[0042] In this embodiment, the BERT pre-trained model is used to extract these semantically rich vector representations.

[0043] Multi-Head Attention is an extended form of the self-attention mechanism. It divides the input data into multiple heads, each of which performs a self-attention calculation to generate an attention matrix and corresponding output. These outputs are then concatenated together and the final output is obtained through a linear transformation. In this way, the multi-head attention mechanism can make the model more effective in processing long sequence data because it can extract feature information from multiple dimensions and enhance the model's expressiveness. The specific calculation process is as follows:

[0044] (1) Input transformation: The input sequence first passes through three different linear transformation layers to obtain query, key, and value matrices respectively. These linear transformations are usually implemented through fully connected layers.

[0045] (2) Head division: The obtained query, key, and value matrices are divided into multiple heads (i.e., multiple subspaces), each head has different linear transformation parameters and focuses on different parts of the input sequence.

[0046] (3) Self-attention calculation: For each head, a scaled dot product attention operation is performed. Specifically, the dot product of the query and the key is calculated, scaled, biased, and the attention weights are obtained using the softmax function. These weights are used to weight the value matrix to generate a weighted sum as the output of each head.

[0047] (4) Concatenation and fusion: The outputs of all heads are concatenated together to form a long vector. Then, a final linear transformation is performed on the concatenated vector to integrate the information from different heads and obtain the final multi-head attention output.

[0048] In this embodiment, the number of heads in the multi-head attention layer is 3, the drop rate is 0.2, and batch priority is set.

[0049] BiGRU is an extension based on GRU, which is a variant of recurrent neural network. It mainly has two gates: update gate and reset gate. The update gate determines how much information of the previous time step needs to be retained in the hidden state of the current time step, and the reset gate determines to what extent the hidden state of the previous time step is ignored. By introducing the update gate and reset gate mechanism, GRU can better capture important features in sequence data. BiGRU consists of two independent GRUs, one processes the sequence in forward order to capture the previous information; the other processes the sequence in reverse order to capture the subsequent information. This bidirectional structure enables BiGRU to better capture long-distance dependencies in the sequence and improve the model's ability to understand sequence data.

[0050] The hidden layer state of BiGRU at the current moment is determined by the current input, the hidden layer state output of the forward GRU at the previous moment, and the hidden layer state output of the reverse GRU at the previous moment. The calculation formula is shown in equations (1) to (3). First, the input x at time t is combined t and the forward hidden layer state output at time t-1 Get the forward hidden layer state at time t Then, the input x at time t is integrated t And the reverse hidden layer state output at time t-1 Get the reverse hidden layer state at time t The hidden layer state h of BiGRU at time t t It is obtained by weighted summation of the hidden layer states in two directions, where w t and v t denote the weights of the forward hidden layer state and the reverse hidden layer state, b t Represents the offset of the hidden layer state at time t:

[0051]

[0052]

[0053]

[0054] In this embodiment, the input dimension of GRU is consistent with the output feature dimension of BERT. The hidden layer size of GRU is set to 150, the number of GRU stacking layers is set to 2, and the first dimension of the input tensor is the batch size.

[0055] Step 3: Model training. Load the training set training model after data preprocessing, and train the data in batches through multiple rounds of epoch iterations: in each batch, perform forward propagation to calculate the prediction results, use the cross entropy loss function to calculate the loss value between the predicted output and the actual label, then perform backpropagation to calculate the gradient, and update the model parameters through the AdamW optimizer, while dynamically adjusting the learning rate using linear warm-up and decay strategies. During the training process, evaluate the performance of the model on the validation set at regular intervals, and save the model parameters with the smallest validation set loss as the best model. If the loss of the validation set does not improve within multiple batches, the early stopping mechanism is triggered to terminate the training. The entire process outputs training and validation logs for monitoring the performance of the model.

[0056] In this example, the epoch is 10, the batch size is 16, the sentence padding length is 256, and the initial learning rate is 5e-5. In order to avoid overfitting and overparameterization of the model during training, the dropout is set to 0.1.

[0057] Step 4: Online review sentiment classification. Classify the sentiment types of online reviews based on the model and save the classification results. The specific steps are:

[0058] Step 4.1 performs data preprocessing on the online comments to be classified, including text segmentation, ID conversion, sentence truncation and padding, and mask creation to convert the text data into the format required by the model.

[0059] Step 4.2 converts the preprocessed text data into tensors.

[0060] Step 4.3 loads the parameters of the best model during the training process and inputs the converted tensor data into the model for sentiment classification prediction.

[0061] Step 4.4 The model outputs the sentiment category label for each comment and saves the result in a CSV file.

[0062] In order to verify the online comment sentiment classification method based on BERT, BiGRU and multi-head attention mechanism proposed in the present invention, the present invention conducted an experiment on online comment sentiment classification on the ChnSentiCorp_htl_all text dataset. The parameters of the best model in the training process were loaded to ensure that the model used the parameters with the best performance on the validation set. The model performance was evaluated on the test set, with accuracy, precision, recall and F1-score as evaluation indicators.

[0063] Accuracy refers to the ratio of the number of samples correctly predicted by the model to the total number of predicted samples. Precision refers to the ratio of samples predicted by the model to samples that are actually positive. Recall refers to the ratio of samples correctly predicted by the model to samples that are actually positive. The F1 score is the harmonic mean of precision and recall, which aims to take both into account. The calculation formula is as follows:

[0064]

[0065]

[0066]

[0067]

[0068] Among them, TP is the number of correctly predicted positive samples, TN is the number of correctly predicted negative samples, FP is the number of negative samples incorrectly predicted as positive, and FN is the number of positive samples incorrectly predicted as negative.

[0069] Experimental environment: This experiment is implemented using PyTorch1.11.0 and trained and tested on the NVIDIA RTX 3090 GPU.

[0070] Experimental results: In order to verify the effectiveness of this method, the method proposed in this invention is compared with other sentiment classification methods. The results are shown in Table 1, where the best result is marked in bold.

[0071] Table 1 Comparison results of the method proposed in this paper with other sentiment classification methods

[0072]

[0073] As can be seen from Table 1, the recall rate of the BERT model is relatively low, and the model may not capture all positive samples well, indicating that the model may have some omission problems. Compared with the pure BERT model, after adding GRU, the accuracy and recall rate of the model are improved, but the precision rate is reduced. This shows that GRU helps the model capture sequence information and improves the recognition ability of positive samples, but at the same time may introduce some misjudgments. The accuracy of the bidirectional GRU is not much different from that of the unidirectional GRU, but the precision rate is improved. The bidirectional GRU can capture the forward and backward sequence information at the same time, and has more advantages in sentiment analysis tasks. After adding the multi-head attention mechanism, the model has significant improvements in precision, recall and F1 score, especially in precision, reaching the highest of 0.9409 among all models. This shows that the multi-head attention mechanism can help the model identify positive samples more accurately and reduce misjudgments.

[0074] In summary, the method proposed in the present invention has the best comprehensive performance in various indicators. It can not only effectively identify positive samples, but also reduce misjudgment.

[0075] The specific implementation modes of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above implementation modes, and corresponding changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.

Claims

1. A method for online review sentiment classification based on BERT, BiGRU and multi-head attention mechanism, characterized in that: Perform data preprocessing on text datasets; build sentiment classification models based on BERT, BiGRU, and multi-head attention mechanisms; train the model using the training set after data preprocessing, and train the data in batches through multiple rounds of iterations; Classify the sentiment types of online reviews based on the model; specifically include the following steps: 1) Perform data preprocessing on the text dataset; including: dividing the text dataset into training set, validation set and test set in a ratio of 8:1:1; preprocessing the text data, including text segmentation, ID conversion, sentence truncation and padding, and mask creation, to obtain data that can be used by the model. 2) Build a sentiment classification model based on BERT, BiGRU and multi-head attention mechanism; including: Use the BERT pre-trained model as a text encoder to obtain a deep feature vector representation of the text; Based on the features output by BERT, a multi-head attention mechanism is applied to further model the global dependencies between different words in a sentence, thereby enhancing the representation capability of semantic features. The attention-weighted features are input into a bidirectional GRU network to capture contextual information, especially long-range dependencies in sequence data; Dropout is added to the output layer of the bidirectional GRU to reduce overfitting, and the output of the last time step of the bidirectional GRU is used to map it to the classification space through a fully connected layer; 3) Load the training set training model after data preprocessing, and train the data in batches through multiple rounds of iterations; 4) Classify the sentiment types of online comments based on the model and save the classification results; including: Perform data preprocessing on online comments to be classified, including text segmentation, ID conversion, sentence truncation and padding, and mask creation to convert text data into the format required by the model; Convert the preprocessed text data into tensors; Load the parameters of the best model during training and input the converted tensor data into the model for sentiment classification prediction; The model outputs the sentiment category label for each review and saves the results in a CSV file.

2. The online review sentiment classification method based on BERT, BiGRU and multi-head attention mechanism according to claim 1, characterized in that: The specific process of model training in step 3) is as follows: in each batch, forward propagation is performed to calculate the prediction results, the cross entropy loss function is used to calculate the loss value between the predicted output and the actual label, and then back propagation is performed to calculate the gradient, and the model parameters are updated through the AdamW optimizer, while the learning rate is dynamically adjusted using the linear warm-up and decay strategy. During the training process, the performance of the model on the validation set is evaluated at regular intervals, and the model parameters with the smallest validation set loss are saved as the best model. If the loss of the validation set does not improve within multiple batches, the early stopping mechanism is triggered to terminate the training. The entire process outputs training and validation logs for monitoring the performance of the model.

Citation Information

Cited By

  • Conversation processing method and device based on artificial intelligence

    CN120316248A

  • Artificial intelligence-based dialogue processing method and device

    CN120316248B