A Mongolian aspect-level sentiment analysis method based on pre-trained capsule network
By pre-training the capsule network model on Mongolian texts and combining attention mechanism and residual links, the problem of scarcity of corpus in Mongolian-level sentiment analysis is solved, which significantly improves the accuracy of sentiment analysis.
Patent Information
- Application Number
- CN202210264289.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-17
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-03-17
AI Technical Summary
Due to the scarcity of marked emotional corpus in Mongolian, it is difficult for the existing technology to effectively conduct Mongolian-level sentiment analysis, resulting in unsatisfactory model results.
Using a method based on pre-training capsule network, a capsule network model that integrates attention mechanism and residual links is pre-trained on large-scale Mongolian texts to obtain the general semantic extraction ability of Mongolian texts, and aspect-level emotion occlusion pre-training is carried out to improve the model's ability to express Mongolian semantics and emotional characteristics.
Through pre-training and emotional enhancement pre-training, the accuracy of Mongolian sentiment analysis results was improved, the problem of scarcity of marked corpus was overcome, and better Mongolian-level sentiment analysis effects were achieved.
Smart Images

Figure CN114742064B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of artificial intelligence, relates to text sentiment analysis, and particularly to a Mongolian aspect-level sentiment analysis method based on a pre-trained capsule network. Background Art
[0002] Sentiment analysis, also known as tendency analysis or opinion mining, is an important information analysis and processing technology. Text sentiment analysis aims to mine and judge the polarity of human emotions contained in the text. Text sentiment analysis can be roughly divided into three types of tasks according to the granularity: paragraph-level text sentiment analysis, sentence-level text sentiment analysis, and aspect-level sentiment analysis. Aspect-level sentiment analysis is a fine-grained sentiment analysis task that aims to identify the sentiment polarity of a specified aspect in a sentence. A sentence may contain multiple different aspects, and the sentiment polarity of each aspect may be different.
[0003] As an important branch of natural language processing, text sentiment analysis is used in scenarios such as public opinion analysis and recommendation systems. With the development of the Internet and e-commerce, the industry's demand for aspect-level sentiment analysis is greater than the coarser-grained chapter-level and sentence-level text sentiment analysis. Both Chinese and English sentiment analysis tasks have made great progress. However, since most current text sentiment analysis uses supervised learning, this method requires a large amount of annotated sentiment corpus to be prepared in advance. However, there is a lack of annotated sentiment corpus in Mongolian. It is very time-consuming and costly to obtain sentiment corpus by manual annotation. The results of the models used for Mongolian sentiment analysis are not ideal.
[0004] Therefore, using unlabeled Mongolian text to mine and train a model to perform Mongolian aspect-level sentiment analysis has become an important issue that needs to be solved urgently. Summary of the invention
[0005] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a Mongolian aspect-level sentiment analysis method based on a pre-trained capsule network, which uses a capsule network that integrates an attention mechanism and residual links as the model body, and performs pre-training on large-scale Mongolian texts to obtain the general semantic extraction capability of Mongolian texts. After that, aspect-level sentiment masking pre-training is performed, so that the model has a strong expression capability for Mongolian semantics and sentiment features, thereby realizing aspect-level sentiment analysis of Mongolian texts.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is:
[0007] A Mongolian aspect-level sentiment analysis method based on a pre-trained capsule network comprises the following steps:
[0008] Step 1, collecting and arranging corpus, the corpus includes unannotated Mongolian text corpus, annotated Chinese emotion corpus and Chinese-Mongolian parallel corpus, the annotated Chinese emotion corpus includes aspect-level Chinese emotion corpus;
[0009] Step 2, constructing a Chinese-Mongolian neural machine translation model and training it using the Chinese-Mongolian parallel corpus;
[0010] Step 3, translating the annotated Chinese emotion corpus into Mongolian emotion corpus using the Chinese-Mongolian neural machine translation model trained in step 2;
[0011] Step 4, constructing a model for pre-training, wherein the model for pre-training adopts a deep attention capsule network;
[0012] Step 5, pre-training the model constructed in step 4 using the unlabeled Mongolian text corpus to obtain a pre-trained Mongolian language model;
[0013] Step 6, using the Mongolian emotional corpus to perform emotion enhancement pre-training on the model obtained in step 5;
[0014] Step 7, using the aspect-level Mongolian text sentiment corpus obtained by translating the aspect-level Chinese sentiment corpus, fine-tune the model after sentiment enhancement in step 6.
[0015] Compared with the prior art, the present invention has the following beneficial effects:
[0016] First, the present invention designs a new model structure mongoCapsNet, which fully integrates the advantages of the attention mechanism and the capsule network in the form of alternating modules composed of multi-head self-attention layers and capsule layers. At the same time, residual connections are added to the model to enhance the generalization ability and robustness of the model. The model learns the interdependence between words in the text through the attention mechanism, and better extracts the aspect words and their corresponding emotions through the capsule network. Secondly, the problem of scarcity of Mongolian annotated texts is solved by pre-training. Finally, emotion-enhanced pre-training is used to improve the model's ability to extract text emotional features. Through these improvements, the accuracy of Mongolian sentiment analysis results is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a flow chart of the present invention.
[0018] Figure 2 This is the structure diagram of the Mongolian-Chinese translation model.
[0019] Figure 3 This is the structure diagram of the deep attention capsule network model.
[0020] Figure 4 is the interaction graph between capsules. DETAILED DESCRIPTION
[0021] The embodiments of the present invention are described in detail below with reference to the accompanying drawings and examples.
[0022] The present invention is a method for Mongolian aspect-level sentiment analysis based on a pre-trained capsule network, and the method flow chart is as follows: Figure 1 shown.
[0023] Step 1: Collect and organize the corpus.
[0024] The corpus collected by the present invention includes large-scale unannotated Mongolian text corpus, annotated Chinese sentiment corpus and Chinese-Mongolian parallel corpus, wherein the annotated Chinese sentiment corpus includes aspect-level Chinese sentiment corpus.
[0025] Collecting and organizing corpus includes text data crawling and text preprocessing, including:
[0026] 1) Text data crawling
[0027] There are many websites written in Mongolian on the Internet, and there are a lot of unlabeled Mongolian texts on them. Write a python script to crawl these unlabeled texts as pre-training corpus. In addition, it is necessary to collect Chinese sentiment analysis corpus for subsequent Mongolian sentiment enhancement pre-training corpus and Mongolian aspect-level sentiment analysis corpus.
[0028] 2) Text preprocessing
[0029] Text preprocessing mainly includes cleaning, word segmentation and building a Mongolian dictionary. The crawled original Mongolian text is cleaned, and the cleaning content includes: removing useless tags, special symbols, and punctuation marks; the cleaned text is processed by word segmentation, and the Mongolian text is divided into a list composed of tokens; then a Mongolian dictionary is built, that is, the word frequency of the word segmented list is counted and sorted by word frequency, and each Mongolian word is mapped to a frequency-sorted sequence number to build a dictionary from Mongolian words to natural numbers.
[0030] Step 2: Build a Chinese-Mongolian neural machine translation model and train the machine translation model using Chinese-Mongolian parallel corpus.
[0031] The Chinese-Mongolian neural machine translation model uses a classic encoder-decoder architecture. In this invention, the ERNIE model with good performance in the Chinese NLP field is used as the encoder, and the decoder uses the transformer decoder. The Mongolian-Chinese machine translation model structure is as follows: Figure 2 shown.
[0032] ERNIE is a pre-trained language model that works well in the Chinese language field and can extract the semantics of Chinese more accurately. The extracted semantics are passed to the decoder in the form of an intermediate state, and the decoder generates the target language, that is, Mongolian. For example, if a Chinese sentence "You are so beautiful" is input from the input end, the sentence will be segmented into 'you', 'really', and 'beautiful'. These three words will be mapped into natural numbers according to the established dictionary, and then converted into real number vectors through the embedding layer, that is, S = {v1,v2,v3,…,v n ,},v i ∈R d The text sequence is subjected to feature extraction through ERNIE, and the extracted features are passed to the decoder in the form of an intermediate state. The decoder <start>and the intermediate state to start predicting word by word until the output terminal symbol <end>, indicating that the translation of this sentence is complete. When training the model, use parallel corpus pairs from Chinese to Mongolian, compare the probability values output by the decoder with the real Mongolian words, calculate the loss, and then train the model parameters through stochastic gradient descent until the model converges. After the model converges, BLEU can be used to evaluate the model.
[0033] Step 3: Use the Chinese-Mongolian neural machine translation model trained in step 2 to translate the annotated Chinese emotion corpus into Mongolian emotion corpus.
[0034] Step 4: Use the deep attention capsule network to build a model structure for pre-training. The model structure is as follows: Figure 3 As shown, the expression is as follows:
[0035] 1) The first layer of the model is the embedding layer, which converts each token in the input text sequence into a real vector, and the length of the vector is the hyperparameter d;
[0036] 2) Position encoding of the embedded sequence;
[0037] 3) Use a multi-head self-attention mechanism to extract dependencies within the text from the position-encoded sequence;
[0038] 4) The output of multi-head self-attention is fed into the capsule network module, which consists of two layers of capsules. Each layer contains 20 capsules, each capsule has 10 neurons, and the interaction between the two layers of capsules uses a dynamic routing algorithm.
[0039] 5) The structure of multi-head self-attention plus double-layer capsule composed of 3) and 4) is used as a module, overlapping n layers;
[0040] 6) Then, the final capsule layer is added for classification to classify the text semantic features and sentiment features extracted by the model, thereby determining the sentiment polarity of the text.
[0041] The model uses the capsule network's ability to highly generalize position and posture information to highly abstract the semantic information between words in the text. Furthermore, residual connections can be added to the deep attention capsule network to enhance the generalization and robustness of the model.
[0042] In this model, the interaction between capsules is carried out through a dynamic routing algorithm. Different from traditional neurons, each capsule outputs a vector, while neurons output scalars. The direction of the vector output by the capsule represents the position and posture information of a feature, and the modulus of the vector represents the probability of the feature existing, which is a real number between 0 and 1. After the output of the lower-level capsule is routed to the higher-level capsule, the vector modulus needs to be compressed to between 0 and 1 by squeezing. The interaction between the two layers of capsules is as follows: Figure 4 As shown. According to this figure, the formula for squeezing between capsules is:
[0043] u1=W1v1,u2=W2v2
[0044] s=c1u1+c2u2
[0045]
[0046] Among them, W1 and W2 are learnable parameter matrices, v1 and v2 are vectors of low-level capsule outputs, c1 and c2 are weights of low-level feature vectors, s is the weighted sum of u1 and u2, and u1 and u2 are learnable parameters W i and the input vector v i The vector obtained by multiplication encodes the relative position relationship between low-level features and high-level features. These relative relationships are contained in the learnable parameter matrix W i v3 is the output vector calculated by the high-level capsule based on v1 and v2. The modulus of v3 represents the probability of the existence of a higher-level feature, and the direction of the vector represents the pose information of the feature.
[0047] by Figure 4 For example, the purpose of the dynamic routing algorithm of the present invention is to obtain the values of c1 and c2, and its pseudo code is as follows:
[0048]
[0049] for r=1to T do
[0050]
[0051] s r =c1u1+c2u2
[0052] a r =squash(s r )
[0053]
[0054]
[0055] in, is an operation that assigns an initial value of 0, where The 0 in the upper right corner indicates that the round is the 0th round, i.e. the first round. The lower right corner indicates the serial number, which corresponds to the serial number of the input capsule. T is the number of loop iterations, and r is the current number of iterations. The result is a real number between 0 and 1, representing the weight. r =c1u1+c2u2 is the weighted sum, s r is an intermediate result. r ) is a squeeze operation, which transforms the vector s r The modulus length is compressed to between 0 and 1 and assigned to a r .a r ·u i is the inner product of the vectors, and the result is a real number. Will gradually change , thereby changing the size of c1 and c2.
[0056] Step 5: Pre-train the model constructed in step 4 using the unlabeled Mongolian text corpus collected in step 1 to obtain a pre-trained Mongolian language model.
[0057] The first way to pre-train is to randomly mask the words in the sentence. That is, mask some words and predict the next sentence as follows:
[0058] 1) Randomly select a certain proportion (e.g. 15%) of tokens to mask and use tokens from the context to predict the masked tokens in a self-supervised manner;
[0059] 2) Predict the next sentence to explicitly model the logical relationship between text pairs. Although the random masking token method can encode bidirectional context to represent words, it cannot explicitly model the logical relationship between text pairs. In order to help understand the relationship between two text sequences, the task of next sentence prediction is added to pre-training. When generating sentence pairs for the pre-training task, there is a 50% probability that they are consecutive sentences labeled "true"; with the other 50% probability, the second sentence is randomly extracted from the corpus and marked as "false". For example, when inputting a sentence "The weather is really nice today", the model will randomly mask some of the words, such as the randomly masked words are "weather" with the symbol <mask>So the original sentence becomes "Today <mask>Very good", input the model and let the model guess <mask>It represents that word.
[0060] Another way to pre-train the model is to input two sentences at the same time and let the model determine whether the two sentences are adjacent. Then, the model parameters are updated by performing gradient descent based on the loss with the true value.
[0061] These two methods can use unlabeled text data to allow the model to learn the association information between text words and words, and sentences and sentences.
[0062] Step 6, in step 5, by randomly masking some words and predicting the next sentence, the model learns the common Mongolian semantic expression. On this basis, in order to make the model have better abstract ability for emotional features, the present invention uses Mongolian emotional corpus to perform emotional enhancement pre-training on the model obtained in step 5.
[0063] In the general pre-trained language model, some words with sentiment bias are ignored during the pre-training process. Since the final task of the present invention is sentiment analysis, it is necessary to focus on learning the text features of some sentiment words. For example, the present invention adopts the sentiment MASK strategy to enhance the sentiment of the model, and the steps are as follows:
[0064] 1) Mask aspect word-emotion word pairs. It should be noted that in a sentence, at most two pairs are masked and they are random;
[0065] 2) Mask sentiment words. It should be noted that in a sentence, the number of masked tokens cannot exceed 10% of the total number of tokens in the current sentence;
[0066] 3) MASK common words. It should be noted that the prerequisite for its execution is that the token ratio of MASK sentiment words does not reach 10%. At the same time, the number of MASK common words is only to supplement the remaining number that did not reach 10% in the second step.
[0067] By masking sentiment words and aspect words, the model can guess the polarity of sentiment words. For example, "Your car has a beautiful appearance" becomes "Your car <mask>real <mask>”, the model needs to predict that the obscured <mask>It is a positive emotion.
[0068] Step 7: The pre-trained model after sentiment enhancement can express Mongolian semantics and sentiment features well. On this basis, a downstream aspect-level Mongolian sentiment classifier is added. The aspect-level Mongolian text sentiment corpus obtained by translating the aspect-level Chinese sentiment corpus is used to fine-tune the aspect-level sentiment classification of the model after sentiment enhancement in step 6, so that it converges to the local optimal solution of the task. Finally, a two-layer capsule network is added to the pre-trained model for aspect-level sentiment classification.
[0069] For example, the sentence "Your car has a beautiful appearance, but poor power." When the sentence is sent to the model, it will be processed into 'appearance': 'Your car has a beautiful appearance, but poor power. ', that is, in the form of (aspect word: sentence), and then the model is asked to predict whether the sentiment about the aspect word 'appearance' in the sentence is positive or negative. The pre-trained model obtained in step 6 is fine-tuned for aspect-level sentiment classification using a large number of pre-processed sentiment corpora in the form of (aspect word: sentence), and the model parameters are slightly adjusted through the gradient descent algorithm, so that the model is suitable for Mongolian aspect-level sentiment classification. It is worth noting that the sentiment corpus in the formal training is Mongolian, and Chinese is used as an example here for ease of understanding.< / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / end> < / start>
Claims
1. A Mongolian aspect-level sentiment analysis method based on a pre-trained capsule network, characterized in that: The steps include: Step 1, collecting and arranging corpus, the corpus includes unannotated Mongolian text corpus, annotated Chinese emotion corpus and Chinese-Mongolian parallel corpus, the annotated Chinese emotion corpus includes aspect-level Chinese emotion corpus; Step 2, constructing a Chinese-Mongolian neural machine translation model and training it using the Chinese-Mongolian parallel corpus; Step 3, translating the annotated Chinese emotion corpus into Mongolian emotion corpus using the Chinese-Mongolian neural machine translation model trained in step 2; Step 4: construct a model for pre-training. The model for pre-training adopts a deep attention capsule network with the following structure: 1) The first layer of the model is the embedding layer, which converts each token in the input text sequence into a real vector, and the length of the vector is the hyperparameter d; 2) Position encoding of the embedded sequence; 3) Use a multi-head self-attention mechanism to extract dependencies within the text from the position-encoded sequence; 4) The output of multi-head self-attention is fed into the capsule network module, which consists of two layers of capsules. Each layer contains 20 capsules, each capsule has 10 neurons, and the interaction between the two layers of capsules uses a dynamic routing algorithm. 5) The structure of multi-head self-attention plus double-layer capsule composed of 3) and 4) is used as a module, overlapping n layers; 6) Then, the final capsule layer is added for classification, and the semantic features and sentiment features of the text extracted by the model are classified to determine the sentiment polarity of the text; Step 5, pre-training the model constructed in step 4 using the unlabeled Mongolian text corpus to obtain a pre-trained Mongolian language model; Step 6, using the Mongolian emotional corpus to perform emotion enhancement pre-training on the model obtained in step 5; Step 7, using the aspect-level Mongolian text sentiment corpus obtained by translating the aspect-level Chinese sentiment corpus, fine-tune the model after sentiment enhancement in step 6.
2. The Mongolian aspect-level sentiment analysis method based on the pre-trained capsule network according to claim 1 is characterized in that: In the step 1, collecting and organizing corpus includes text data crawling and text preprocessing; the text preprocessing includes cleaning, word segmentation and building a Mongolian dictionary; wherein the cleaning content includes: removing useless tags, special symbols, and punctuation marks; word segmentation is to divide the Mongolian text into a list composed of tokens; building a Mongolian dictionary is to count the word frequencies of the list and sort them according to the word frequencies, map each Mongolian word to a sequence number sorted by the word frequencies, and build a dictionary from Mongolian words to natural numbers.
3. The Mongolian aspect-level sentiment analysis method based on pre-trained capsule network according to claim 1 is characterized in that: In step 2, the Chinese-Mongolian neural machine translation model adopts an encoder-decoder architecture; wherein the encoder adopts the ERNIE pre-trained model and the decoder adopts the Transformer decoder.
4. The Mongolian aspect-level sentiment analysis method based on pre-trained capsule network according to claim 1, characterized in that: In step 4, residual connections are added to the deep attention capsule network to enhance the generalization ability and robustness of the model.
5. The Mongolian aspect-level sentiment analysis method based on pre-trained capsule network according to claim 1, characterized in that: The output of each capsule is a vector. The direction of the vector represents the position and posture information of a feature. The modulus of the vector represents the probability of the feature existing. The modulus is a real number between 0 and 1. After the output of the low-level capsule is routed to the high-level capsule, the vector modulus is compressed to between 0 and 1 by squeezing. The squeezing formula is: u1=W1v1,u2=W2v2 s=c1u1+c2u2 Among them, W1 and W2 are learnable parameter matrices, v1 and v2 are vectors of low-level capsule outputs, c1 and c2 are weights of low-level feature vectors, s is the weighted sum of u1 and u2, and u1 and u2 are learnable parameters W i and the input vector v i The vector obtained by multiplication encodes the relative position relationship between the low-level features and the high-level features. v3 is the output vector calculated by the high-level capsule based on v1 and v2. The modulus of v3 represents the probability of the existence of higher-level features, and the direction of the vector represents the pose information of the feature.
6. The Mongolian aspect-level sentiment analysis method based on pre-trained capsule network according to claim 1, characterized in that: When performing pre-training in step 5, some words are masked and the next sentence is predicted, as follows: 1) Randomly select the first proportion of tokens to mask and use tokens from the context to predict the masked tokens in a self-supervised manner; 2) Predict the next sentence to explicitly model the logical relationship between text pairs. When generating sentence pairs for the pre-training task, there is a 50% probability that they are consecutive sentences labeled "true"; with the other 50% probability, the second sentence is randomly extracted from the corpus and labeled "false".
7. The Mongolian aspect-level sentiment analysis method based on pre-trained capsule network according to claim 1, characterized in that: When performing pre-training in step 5, two sentences are inputted at the same time, and the model determines whether the two sentences are adjacent sentences, and then performs gradient descent based on the loss with the true value, thereby updating the model parameters.
8. The Mongolian aspect-level sentiment analysis method based on pre-trained capsule network according to claim 1, characterized in that: In step 6, the emotion MASK strategy is used to enhance the emotion of the model. The steps are as follows: 1) Mask aspect word-emotion word pairs. In a sentence, at most two pairs can be masked randomly. 2) MASK sentiment words. In a sentence, the number of tokens masked cannot exceed 10% of the total number of tokens in the current sentence. 3) MASK common words. The prerequisite for its execution is that the token ratio of the MASK sentiment words does not reach 10%, and the number of MASK common words supplements the remaining number that does not reach 10%.
9. The Mongolian aspect-level sentiment analysis method based on pre-trained capsule network according to claim 1, characterized in that: In step 7, the model is fine-tuned for aspect-level sentiment classification so that it converges to a local optimal solution for the task.
Citation Information
Patent Citations
Emotion classification method and system, storage medium and equipment
CN110826336A
Mongolian multi-modal sentiment analysis method based on T-M BERT pre-training model
CN114153973A