A Network Public Opinion Sentiment Analysis Method Based on Multi-Model Fusion Transfer Learning
By adopting multi-model fusion transfer learning method in network public sentiment analysis, and using models such as ERNIE, Text CNN and BiGRU for feature extraction and parameter migration, the problem of scarcity of labeled data is solved, and the accuracy and analysis performance of the model are improved.
Patent Information
- Application Number
- CN202310322355.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-29
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-03-29
AI Technical Summary
The existing methods of online public opinion sentiment analysis require a large amount of labeled data to effectively learn deep-level features, but labeled data is scarce in practical applications, resulting in low model accuracy.
Using a multi-model fusion transfer learning method, the text data is converted into dynamic word vectors through ERNIE pre-trained model, combined with a joint model of Text CNN and BiGRU for feature extraction, and the analysis performance of the model is improved by using parameter transfer and attention mechanisms.
It effectively solves the problem of scarcity of labeled data, improves the accuracy and analysis performance of the online public opinion sentiment analysis model, and can better understand people's emotional tendencies towards hot and sensitive issues.
Smart Images

Figure CN116467406B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing, and particularly to a network public opinion sentiment analysis method based on multi-model fusion transfer learning. Background Art
[0002] Text sentiment analysis is an important branch of natural language processing. It targets text data and is a process of analyzing, processing, summarizing, and reasoning subjective texts with emotional colors. Network public opinion sentiment analysis is one of the main research directions of text sentiment analysis. The sentiment analysis of relevant text information in public opinion events is generally divided into methods based on sentiment dictionaries, traditional machine learning methods, and deep learning methods. The most commonly used deep learning-based sentiment analysis methods still require a large amount of labeled data to learn deeper features. At present, there is only a small amount of labeled data in the analysis of network public opinion. In recent years, with the development of transfer learning, the problem of lack of a large amount of labeled data has been solved. Transfer learning is to transfer the knowledge of one domain (source domain) to another domain (target domain) so that the target domain can achieve better learning results. The core is to find the similarity between the source domain and the target domain and make reasonable use of it. The more factors shared by two different domains, the easier the transfer learning is, and vice versa. Summary of the Invention
[0003] Object of the Invention: There is a problem of scarce labeled data in the current network public opinion sentiment analysis work, but the existing sentiment analysis methods still require a large amount of labeled data to learn deeper features. Therefore, in view of the above problems, a multi-model fusion transfer learning is proposed, which can solve the problem of scarce labeled data by transfer learning with the help of domains with rich labeled data, and improve the accuracy of the model.
[0004] To achieve the above object, the technical solution adopted by the present invention is:
[0005] A network public opinion sentiment analysis method based on multi-model fusion transfer learning, comprising the following steps:
[0006] S1: Clean the data of the source domain and the target domain, and then use the ERNIE pre-trained model to convert the text data obtained by data cleaning into dynamic word vectors;
[0007] S2: Use a joint model of Text CNN and BiGRU for feature extraction, where the Text CNN model is responsible for extracting local features, and the BiGRU model is responsible for extracting global features, and merge the local features and global features in the pre-trained model;
[0008] S3: Train the source domain model, and use parameter transfer to transfer the parameters of the feature layer of the pre-trained model to the target domain model as the initial parameters of the feature layer of the target model;
[0009] S4: Use the attention mechanism to assign weights to the output matrix of the feature extraction layer to further highlight the weights of key words, and then fuse the feature vectors obtained by the attention mechanism;
[0010] S5: Use the target domain data to train the target model, and calculate the analysis result of the text through the Softmax classifier;
[0011] S6: Compare and evaluate the analysis result with the analysis result of the single network analysis method in terms of accuracy, precision, recall rate, and F1 value.
[0012] In the step S1, data cleaning of the source domain and target domain data means removing data containing a lot of @ symbols, blank lines, spaces, etc. from the data to improve the data quality, and then using the ERNIE pre-trained model to convert the text data obtained by data cleaning into dynamic word vectors. The core technical advantage of ERNIE is that it creatively combines big data pre-training with multi-source rich knowledge, and through continuous learning technology, continuously absorbs new knowledge in aspects such as vocabulary, structure, and semantics in a large amount of text data, realizing the continuous evolution of the model effect.
[0013] The formula for generating dynamic word vectors using the ERNIE model is as follows:
[0014]
[0015] Among them, Q is the query matrix, K is the key matrix, and V is the value matrix. The three are obtained by multiplying the text sequence X obtained in the knowledge integration stage by the corresponding weight matrix;
[0016] For the self-attention vector Z, it will enter the next round of self-attention operation again after being mapped by the feed-forward neural network, and finally obtain the output of ERNIE, which is also the input of the feature layer model. The formula is as follows:
[0017] R = W·Z + b
[0018] Among them, W is the weight matrix and b is the bias value.
[0019] In step S2, a combined model of Text CNN and BiGRU is used for feature extraction. The characteristic of the convolutional neural network is that it can extract local features of the matrix. The local feature is a sliding window of the word vector matrix composed of several words. The advantage is that it can automatically combine and screen text features to obtain semantic information at different abstraction levels. Therefore, the model is used to extract local features. BiGRU is composed of a forward GRU and a backward GRU. The output of the network is obtained by superimposing the forward output and the backward output. BiGRU can learn the temporal relationship between the previous moment, the next moment and the current state, which can improve the learning ability of the model. Therefore, the BiGRU model is used to extract global features.
[0020] The formula for using the Text CNN model to extract local features is as follows:
[0021] S i = MAX{C} = MAX(C1, C2,..., C n-h+1 )
[0022] S = [S1, S2,..., S m
[0023] where C is the feature vector obtained by the convolutional layer, and S is the feature vector obtained by the operation of the Text CNN model;
[0024] The formula for using the BiGRU model to extract global features is as follows:
[0025]
[0026]
[0027] where, is the forward hidden layer state, is the backward hidden layer state, R i is the dynamic word vector output by the ERNIE model.
[0028] In step S3, first, the feature vectors of the source domain are merged and classified for output. Secondly, the source domain model is trained, and the parameters of the feature layer of the pre-trained model are migrated to the target domain model by using parameter migration as the initial parameters of the feature layer of the target model.
[0029] In step S4, the attention mechanism is used to assign weights to the output matrix of the feature extraction layer to further highlight the weights of key words. The formula is as follows:
[0030]
[0031]
[0032] Among them, Q is the query vector, K is the key vector, and V is the value vector.
[0033] Fuse the feature vectors obtained by the attention mechanism. The formula is as follows:
[0034] M = [M S , M H
[0035] In the step S5, the target domain data is used to train the target model, and the analysis result of the text is calculated through the Softmax classifier. The formula is as follows:
[0036] y = softmax(w * M + b)
[0037] Among them, w is the weight of the fully connected layer, and b is the bias term.
[0038] In the step S6, the accuracy rate represents the percentage of the correctly predicted results in the total samples. The precision rate represents the proportion of the correctly predicted positive samples in the actually predicted positive samples. The recall rate represents the proportion of the correctly predicted positive samples in the positive samples. The corresponding formulas are as follows:
[0039]
[0040]
[0041]
[0042] Among them, TP represents correctly predicting a positive sample as positive, FN represents incorrectly predicting a positive sample as negative, FP represents incorrectly predicting a negative sample as positive, and TN represents correctly predicting a negative sample as negative. If only considering precision or only considering recall rate, neither can be used as an index to evaluate the quality of a model. The F1 value is the harmonic mean of Precision and Recall. Therefore, the F1 value is used to reconcile the two, compatible with precision and recall rate. The formula is as follows:
[0043]
[0044] Compared with the existing technologies, the beneficial effects of the present invention are:
[0045] The present invention uses the ERNIE model to convert text into word vectors. The ERNIE model innovatively combines big data pre-training with multi-source rich knowledge. Through the continuous learning technology, it continuously absorbs new knowledge in aspects such as vocabulary, structure, and semantics in a large amount of text data, and can solve the problem of static representation of word vectors.
[0046] The present invention uses a combined model of Text CNN and BiGRU as a feature extractor, which can better extract deep features, so as to better understand people's sentiment tendencies towards hot and sensitive issues.
[0047] The present invention uses transfer learning to transfer the parameters of the feature extraction layer trained in the source domain to the feature extraction layer in the target domain as its initial parameters, which can solve the problem of scarce labeled data in network public opinion sentiment analysis. After the transfer, the attention mechanism is used to improve the analysis performance of the model. Brief Description of the Drawings
[0048] Figure 1 is a schematic flowchart of the method of the present invention;
[0049] Figure 2 is a schematic structural diagram of the method of the present invention;
[0050] Figure 3 is a schematic diagram of the transfer learning strategy. Detailed Embodiments
[0051] The present invention discloses a network public opinion sentiment analysis method based on multi-model fusion transfer learning, as Figures 1 to 3 shown, which includes the following steps:
[0052] Step 1: Clean the data in the source domain and the target domain, and then use the ERNIE pre-trained model to convert the text data obtained by data cleaning into dynamic word vectors. The core technical advantage of ERNIE is that it creatively combines big data pre-training with multi-source rich knowledge. Through continuous learning technology, it continuously absorbs new knowledge in aspects such as vocabulary, structure, and semantics in a large amount of text data, realizing the continuous evolution of the model effect, just like humans continuously learning. The ERNIE model generates dynamic word vectors mainly including two parts: knowledge integration and Transformer encoder.
[0053] The specific steps of knowledge integration are as follows:
[0054] 1. Adopt basic-level masking. It takes the sentence as a sequence of basic language units. For English, the basic unit is the word, and for Chinese, the basic unit is the Chinese character. During the pre-training process, use basic-level mask, that is, randomly mask one character in the sentence.
[0055] 2. Adopt entity-level masking. Named entities include place names, personnel, organizations, and items, etc. At this stage, first analyze the named entities in the sentence, and then mask a part with entity-level mask.
[0056] 3. Adopt phrase-level masking, where a phrase is a group of characters acting as a conceptual unit. At this stage, first analyze the phrases in the sentence, and then mask a part with phrase-level masks.
[0057] After the masking training in the knowledge integration stage, a corresponding text sequence X = [X1, X2,..., X n is obtained, where n is the length of the text sequence.
[0058] The specific steps of the Transformer encoder are as follows:
[0059] 1. Multiply the text vectors obtained in the knowledge integration stage by three weight matrices W Q , W K , W V respectively to obtain the query matrix Q, the key matrix K, and the value matrix V. The formulas are as follows:
[0060] Q = X · W Q
[0061] K = X · W K
[0062] V = X · W V
[0063] 2. Use the self-attention mechanism to calculate the attention between the target word and all words and phrases in the text sequence. Among them, softmax is used to normalize the scores. The formula is as follows:
[0064]
[0065] 3. After the self-attention vector is mapped through the feed-forward neural network, it enters the next round of self-attention operation again. Finally, the obtained R i (1, 2,..., n) is the output of ERNIE and also the input of the feature layer model. The formula is as follows:
[0066] R = W · Z + b
[0067] where W is the weight matrix and b is the bias value.
[0068] Step 2: Use a combined model of Text CNN and BiGRU for feature extraction. The convolutional neural network can extract local features of a matrix. The local feature is a sliding window of the word vector matrix composed of several words. The advantage is that it can automatically combine and filter text features to obtain semantic information at different abstraction levels. Therefore, use the Text CNN model to extract local features. BiGRU consists of a forward GRU and a backward GRU. The output of the network is obtained by superimposing the forward output and the backward output. BiGRU can learn the temporal relationship between the previous moment, the next moment and the current state, which can improve the learning ability of the model. Therefore, use the BiGRU model to extract global features.
[0069] The specific steps to extract local features using the Text CNN model are as follows:
[0070] 1. In the present invention, the convolutional kernels of the convolutional layer have three different sizes. The width of the convolutional kernel is the same as the dimension of the word vector, which is d, and the heights are 2, 3, and 4 respectively. There are multiple convolutional kernels of each size. For a text with an input length of n, the convolutional layer performs a convolutional operation on the text input vector by adopting h different-sized sliding windows to learn text features. At position i, the convolutional feature value is obtained through the convolutional kernel, and the feature C i is calculated by the following formula:
[0071] C i = f(WR i:i+h-1 + b)
[0072] where W represents the convolutional kernel, which is an n×d-dimensional weight matrix, b represents the bias, R i:i+h-1 represents the sliding window composed of the i-th row to the i+h-1-th row of the input matrix, which is used to extract features at different positions of the text, h represents the height of the convolutional kernel, f represents the activation function, and C i represents the i-th feature extracted by a convolutional kernel. The entire convolutional kernel extracts features from the text along the sliding window respectively, and a feature vector C = [C1, C2,..., C n-h+1 is obtained.
[0073] 2. In the present invention, the maximum pooling method extracts the maximum value in the above feature vector C to replace the entire feature vector. The feature vectors obtained by convolution with other convolutional kernels of the same size also undergo maximum pooling processing, and then these maximum values are fused to form a new feature vector S. The formula is as follows:
[0074] S i = MAX{C} = MAX(C1, C2,..., C n-h+1 )
[0075] S = [S1, S2,..., S m
[0076] The formula for extracting global features using the BiGRU model is as follows:
[0077]
[0078]
[0079] Among them, is the forward hidden layer state, is the backward hidden layer state, and R i is the dynamic word vector output by the ERNIE model.
[0080] Step 3: Train the source domain model, and use parameter transfer to transfer the parameters of the feature layer of the pre-trained model to the target model as the initial parameters of the feature layer of the target model. Specifically, the present invention is mainly divided into two parts: the source model and the target model, and a dual-channel model of Text CNN and BiGRU is used for feature extraction. Among them, the Text CNN model extracts local features of the text, and the BiGRU model extracts global feature information. First, use the dual-channel model of Text CNN and BiGRU to obtain feature parameters for the news data with rich labels. The pre-trained model can obtain the general language of the text, and then transfer the parameters of the feature extraction layer of the pre-trained model to the target extraction layer of the target model. These parameters will be used as the initial values of the feature extraction layer of the target task, and then use the public opinion data to train the model.
[0081] Step 4: Use the attention mechanism to assign weights to the output matrix of the feature extraction layer to further highlight the weights of key words, and then fuse the feature vectors obtained by the attention mechanism. Specifically, the present invention uses the self-attention mechanism for the local feature matrix S (S = [S1, S2,..., S m ) and the global feature matrix H (H = [H1, H2,..., H n ) extracted by the BIGRU model respectively, and then performs fusion. The specific steps are as follows:
[0082] 1. For each input, first map it to three different spaces to obtain three vectors: the query vector Q, the key vector K, and the value vector V. The formula is as follows:
[0083] Q S = S * W i Q , Q H = H * W i Q
[0084] K S = S * W i K ,K H = H * W i K
[0085] V S = S * W i V ,V H = S * W i H
[0086] 2. Use the scaled dot - product as the attention scoring function to obtain the output vector sequence. The formula is as follows:
[0087]
[0088]
[0089] Among them, Q is the query vector, K is the key vector, and V is the value vector;
[0090] 3. Fuse the feature vectors obtained by the attention mechanism. The formula is as follows:
[0091] M = [M S , M H
[0092] Step 5: Use the target - domain data to train the target model, and calculate the analysis result of the text through the Softmax classifier. The formula is as follows:
[0093] y = softmax(w * M + b)
[0094] Among them, w is the weight of the fully - connected layer, and b is the bias term.
[0095] Step 6: The accuracy rate represents the percentage of the correctly predicted results in the total samples. The precision rate represents the proportion of the correctly predicted positive samples in the actually predicted positive samples. The recall rate represents the proportion of the correctly predicted positive samples in the positive samples. The corresponding formulas are as follows:
[0096]
[0097]
[0098]
[0099] Among them, TP means correctly predicting a positive sample as positive, FN means incorrectly predicting a positive sample as negative, FP means incorrectly predicting a negative sample as positive, and TN means correctly predicting a negative sample as negative. If only considering precision or only considering recall cannot be used as an indicator to evaluate the quality of a model, the F1 value is the harmonic mean of Precision and Recall. Therefore, the F1 value is used to reconcile the two, taking into account both precision and recall. The formula is as follows:
[0100]
[0101] The above is a detailed introduction to the embodiments of the present invention in conjunction with the accompanying drawings. The specific implementation manners herein are only used to help understand the method of the present invention. For those of ordinary skill in the art, according to the idea of the present invention, changes and modifications can be made within the scope of the specific implementation manners and applications. Therefore, this specification of the present invention should not be construed as a limitation to the present invention.
Claims
1. A method for network public opinion sentiment analysis based on multi-model fusion transfer learning, characterized in that: Including the following steps: S1: Clean the source domain and target domain data, and then use the ERNIE pre-trained model to convert the text data obtained after data cleaning into dynamic word vectors; S2: Use a joint model of Text CNN and BiGRU for feature extraction. Among them, the Text CNN model is responsible for extracting local features, and the BiGRU model is responsible for extracting global features. The local features and global features are merged in the pre-trained model; S3: Train the source domain model, and use parameter transfer to transfer the parameters of the feature layer of the pre-trained model to the target domain model as the initial parameters of the feature layer of the target model; S4: Use the attention mechanism to assign weights to the output matrix of the feature extraction layer to further highlight the weights of key words, and then fuse the feature vectors obtained by the attention mechanism; S5: Use the target domain data to train the target model, and calculate the analysis result of the text through the Softmax classifier; S6: Compare and evaluate the analysis results with those of the single network analysis method in terms of accuracy, precision, recall, and F1 value.
2. The network public opinion sentiment analysis method based on multi-model fusion transfer learning according to claim 1, characterized in that: In step S1, when cleaning the source domain and target domain data, it means removing a lot of @ symbols, blank lines, and blank space data in the data to improve the data quality. Then use the ERNIE pre-trained model to convert the text data obtained after data cleaning into dynamic word vectors. The core technical advantage of ERNIE is that it creatively combines big data pre-training with multi-source rich knowledge. Through continuous learning technology, it continuously absorbs new knowledge in vocabulary, structure, and semantics from massive text data, realizing the continuous evolution of the model effect. The formula for generating dynamic word vectors using the ERNIE model is as follows: Among them, Q is the query matrix, K is the key matrix, and V is the value matrix. The three are obtained by multiplying the text sequence X obtained in the knowledge integration stage by the corresponding weight matrix; For the self-attention vector Z, it will enter the next round of self-attention operation again after being mapped by the feed-forward neural network, and finally obtain the output of ERNIE, which is also the input of the feature layer model. The formula is as follows: R = W·Z + b Among them, W is the weight matrix and b is the bias value.
3. The network public opinion sentiment analysis method based on multi-model fusion transfer learning according to claim 1, wherein: In step S2, when using a joint model of Text CNN and BiGRU for feature extraction, the characteristic of the convolutional neural network is that it can extract local features of the matrix. The local features are the sliding windows of the word vector matrix composed of several words. The advantage is that it can automatically combine and screen text features to obtain semantic information at different abstraction levels. Therefore, use the model to extract local features; BiGRU is composed of a forward GRU and a backward GRU. The output of the network is obtained by superimposing the forward output and the backward output. BiGRU can learn the temporal relationship between the previous moment and the next moment and the current state, which can improve the learning ability of the model. Therefore, use the BiGRU model to extract global features; The formula for using the Text CNN model to extract local features is as follows: S i = MAX{C} = MAX(C1, C2,..., C n-h+1 ) S = [S1, S2,..., S m Among them, C is the feature vector obtained by the convolutional layer, and S is the feature vector obtained by the operation of the Text CNN model; The formula for using the BiGRU model to extract global features is as follows: Among them, is the forward hidden layer state, is the backward hidden layer state, and R i is the dynamic word vector output by the ERNIE model.
4. The network public opinion sentiment analysis method based on multi-model fusion transfer learning according to claim 1, wherein: In step S3, first, the feature vectors of the source domain are merged and classified for output. Secondly, the source domain model is trained, and the parameters of the feature layer of the pre-trained model are transferred to the target domain model by parameter transfer as the initial parameters of the feature layer of the target model.
5. The network public opinion sentiment analysis method based on multi-model fusion transfer learning according to claim 1, characterized in that: In step S4, the attention mechanism is used to assign weights to the output matrix of the feature extraction layer to further highlight the weights of key words. The formula is as follows: Among them, Q is the query vector, K is the key vector, and V is the value vector; The feature vectors obtained by the attention mechanism are fused. The formula is as follows: M = [M S , M H Among them, M S and M H are feature vectors obtained through the attention mechanism.
6. The network public opinion sentiment analysis method based on multi-model fusion transfer learning according to claim 1, characterized in that: In step S5, the target domain data is used to train the target model, and the analysis result of the text is calculated by the Softmax classifier. The formula is as follows: y = softmax(w * M + b) Among them, w is the weight of the fully connected layer, and b is the bias term.
7. The network public opinion sentiment analysis method based on multi-model fusion transfer learning according to claim 1, wherein: In step S6, the accuracy rate represents the percentage of the correctly predicted results in the total samples. The precision rate represents the proportion of the correctly predicted positive samples in the actually predicted positive samples. The recall rate represents the proportion of the correctly predicted positive samples in the positive samples. The corresponding formulas are as follows: Among them, TP represents correctly predicting a positive sample as positive, FN represents incorrectly predicting a positive sample as negative, FP represents incorrectly predicting a negative sample as positive, and TN represents correctly predicting a negative sample as negative. If only considering the precision rate or only considering the recall rate, neither can be used as an indicator to evaluate the quality of a model. The F1 value is the harmonic mean of Precision and Recall. Therefore, the F1 value is used to reconcile the two, compatible with the precision rate and the recall rate. The formula is as follows: Among them, P represents the precision rate, and R represents the recall rate.
Citation Information
Patent Citations
Cross-domain text sentiment classification method based on parameter migration and attention sharing mechanism
CN113326378A
Semantic sentiment analysis method fusing in-depth features and time sequence models
US11194972B1