A Mongolian pre-training sentiment analysis method integrating roots, affixes and phonetic symbols

By combining the Mongolian BERT pre-training model of root affixes and phonetic symbols, and adding contrast learning to the pre-training task, the problem of lack of Mongolian sentiment analysis data and inconsistent language characteristics is solved, and the accuracy of sentiment analysis is significantly improved.

CN114742046BActive Publication Date: 2025-05-20INNER MONGOLIA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210252395.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-15
Publication Date
2025-05-20
Estimated Expiration
2042-03-15

AI Technical Summary

Technical Problem

In the Mongolian sentiment analysis task, due to the lack of data and the inconsistent language characteristics with the English pre-trained model, the sentiment analysis effect was poor.

Method used

A Mongolian BERT pre-training model is adopted that combines root affixes and phonetic symbols, and contrast learning is added to the pre-training task for data augmentation.

Benefits of technology

It effectively alleviates the problem of lack of Mongolian data, improves the accuracy of Mongolian sentiment analysis tasks, and can be applied to other Mongolian natural language processing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114742046B_ABST
    Figure CN114742046B_ABST
Patent Text Reader

Abstract

A Mongolian pre-training sentiment analysis method for fusing roots, affixes and phonetic symbols, pre-processing Mongolian corpus; constructing a Mongolian BERT pre-training model, wherein word embedding, root embedding, affix embedding and phonetic symbol embedding are constructed in its embedding layer; the embedding is spliced ​​to obtain a fused embedding, and then the fused embedding is added to the position embedding to form a model input; in the Mongolian BERT pre-training model, the fusion task of contrastive learning and MLM is pre-trained; Mongolian sentiment corpus is pre-processed; and sentiment analysis is performed on Mongolian sentiment corpus using a trained Mongolian BERT pre-training model fusing roots, affixes and phonetic symbols. The present invention pre-trains the BERT model by fusing roots, affixes and phonetic symbols, and in addition, integrates the contrastive learning method into the MLM task, and realizes data enhancement through contrastive learning, thereby improving the accuracy of the model and the accuracy of sentiment analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a Mongolian pre-trained sentiment analysis method that combines roots, affixes, and phonetic symbols. Background Art

[0002] Text sentiment analysis, also known as opinion mining, refers to the analysis of subjective text with emotional color, mining the emotional tendency contained therein, and classifying the emotional attitude. As a research hotspot in natural language processing, text sentiment analysis has great research significance in public opinion analysis, user profiling, and recommendation systems.

[0003] Since Mongolian is a minority language, the research on natural language processing related to Mongolian started relatively late, and the research progress of Mongolian sentiment analysis is relatively slow. Currently, all sentiment analysis methods based on deep learning require a large amount of corpus data as a driving force, and the Mongolian sentiment corpus is currently in a scarce stage. How to alleviate the problem of insufficient resources has become an important research topic in Mongolian sentiment analysis.

[0004] Most of the latest sentiment analysis methods based on deep learning are achieved by fine-tuning pre-trained models to obtain better sentiment analysis results. Since pre-trained models such as BERT were initially designed for English, and Mongolian is an agglutinative language with a different word-formation method from English, the sentiment analysis effect of Mongolian by fine-tuning the BERT model is very poor.

[0005] Therefore, while enhancing the data of the Mongolian corpus, combined with the language characteristics of Mongolian, using the Mongolian monolingual corpus to pre-train the BERT model, so as to improve the accuracy of downstream sentiment analysis tasks is an urgent problem to be solved. Summary of the Invention

[0006] In order to overcome the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a Mongolian pre-trained sentiment analysis method that combines roots, affixes, and phonetic symbols. The BERT model is pre-trained by combining roots, affixes, and phonetic symbols. In addition, the method of contrastive learning is incorporated into the MLM task to achieve data augmentation through contrastive learning, thereby improving the accuracy of the model and further improving the accuracy of sentiment analysis.

[0007] In order to achieve the above purpose, the technical solution adopted by the present invention is:

[0008] A Mongolian pre-trained sentiment analysis method that combines roots, affixes, and phonetic symbols, comprising the following steps:

[0009] Step 1, preprocess the Mongolian corpus;

[0010] Step 2: Construct a Mongolian BERT pre-training model. In its embedding layer, construct word embeddings, root embeddings, affix embeddings, and phonetic symbol embeddings; splice the embeddings to obtain fused embeddings, and then add the fused embeddings to positional embeddings to form the model input.

[0011] Step 3: In the Mongolian BERT pre-training model, perform pre-training on the fused task of contrastive learning and MLM.

[0012] Step 4: Preprocess the Mongolian sentiment corpus.

[0013] Step 5: Use the trained Mongolian BERT pre-training model that fuses roots, affixes, and phonetic symbols to perform sentiment analysis on the Mongolian sentiment corpus.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0015] The present invention utilizes the language characteristics of Mongolian itself, designs a dedicated Mongolian BERT pre-training model by fusing the roots, affixes, and phonetic symbols of Mongolian, and adds contrastive learning to the pre-training task of the model for data augmentation. Such a design not only alleviates the problem of scarce Mongolian data, but also effectively improves the accuracy of the Mongolian sentiment analysis task; similarly, this model can also be applied to other Mongolian natural language processing tasks, which helps to improve the accuracy of tasks such as machine translation and sentence semantic similarity; furthermore, the present invention can also provide a model reference for other minority languages. Description of the Drawings

[0016] Figure 1 is a schematic diagram of the overall process of the present invention.

[0017] Figure 2 is a schematic diagram of the pre-training process of the model of the present invention.

[0018] Figure 3 is a schematic diagram of the sentiment analysis process of the present invention.

[0019] Figure 4 is a schematic diagram of the root and affix splitting of the example sentences of the present invention.

[0020] Figure 5 is the root embedding of the present invention.

[0021] Figure 6 is the affix embedding of the present invention.

[0022] Figure 7 is the phonetic symbol embedding of the present invention.

[0023] Figure 8 is the fused embedding of the present invention.

[0024] Figure 9This is the overall structure of the pre-trained model of the present invention. Specific Embodiments

[0025] The following combines the accompanying drawings and embodiments to detail the implementation manner of the present invention.

[0026] As Figure 1 shown, the present invention is a Mongolian pre-trained sentiment analysis method that combines root affixes and phonetic symbols. The present invention is divided into two parts. The first part is the Mongolian BERT pre-training process that combines root affixes and phonetic symbols. The pre-training task is composed of the fusion of contrastive learning and the MLM task, as Figure 2 shown. The second part is to use the pre-trained model above to perform sentiment analysis tasks, as Figure 3 shown.

[0027] In the embodiments of the present invention, the specific steps include:

[0028] Step 1, perform preprocessing operations on the Mongolian corpus.

[0029] Exemplarily, the preprocessing of the present invention includes data cleaning and word segmentation. First, perform data cleaning operations on the Mongolian corpus, and then use BPE to perform word segmentation operations on the Mongolian language.

[0030] For the word segmentation operation, each Mongolian word is separated by root affixes as units. For example: translated as "has come", the word segmentation should divide it into the prefix (come) and the suffix (indicating the past). For words that cannot be segmented into root affixes, they remain unchanged. Taking "People have gone to work" as an example, its corresponding Mongolian expression is The result of word segmentation according to root affixes is as Figure 4 shown.

[0031] Step 2, construct a Mongolian BERT pre-training model composed of an embedding layer and a BERT Encoder structure. In order to combine the language characteristics of the Mongolian language itself, word embeddings, root embeddings, affix embeddings, and phonetic symbol embeddings are constructed in the embedding layer of the pre-training. The model splices these four embeddings to obtain a fused embedding, and then adds the fused embedding to the position embedding to form the input of the model.

[0032] Among them, the embedding layer of the traditional BERT is composed of a fused embedding, a position embedding, and a text embedding. The improvement of the present invention is to incorporate the underlying fused embedding into the root embedding, affix embedding, and phonetic symbol embedding in addition to the word embedding.

[0033] In one embodiment of the present invention, since the text embedding is added together with the position embedding to obtain the model input only when performing the NSP training task, and the use cases of the present invention do not involve the NSP task, there is no need to add text embedding in the embedding layer. That is, on the basis of word embedding, position embedding and text embedding in the traditional BERT embedding layer, the present invention removes the text embedding and adds root embedding, affix embedding and phonetic symbol embedding, and the embedding layer of the present invention is composed of the fused embedding and the position embedding. The BERT Encoder structure is the same as that of the traditional BERT, that is, it is formed by stacking multiple layers of Transformer Encoder.

[0034] Since Mongolian is a verb-centered language, the inflection of the verb end determines the context of a sentence. For example, the suffix for expressing negation in Mongolian is For a sentence containing this suffix, its emotional color is obvious. Therefore, constructing root embedding and affix embedding helps to improve the accuracy of downstream sentiment analysis tasks.

[0035] Where:

[0036] Word embedding is performed at the word granularity, which is the same as the token embedding in the original BERT model.

[0037] Root embedding is to embed the root words after word segmentation. In the present invention, the CNN and max pooling are used for this sequence to obtain the final root word sequence. Taking the word (people) as an example, as Figure 5 shown.

[0038] Affix embedding is to embed the affixes after word segmentation. In the present invention, the CNN and max pooling are used for this sequence to obtain the final affix sequence. Taking the word (people) as an example, as Figure 6 shown. For words that cannot be segmented into roots and affixes, their prototypes are directly used as roots and affixes for embedding.

[0039] Phonetic symbol embedding is to embed the international phonetic symbols corresponding to Mongolian words. In the present invention, the CNN and max pooling are used for this sequence to obtain the final phonetic symbol sequence. Taking the word (people) as an example, as Figure 7 shown. Since Mongolian is similar to Chinese and there are cases of "polysyllabic words", such as when expressing the meaning of "star", the pronunciation is while when expressing the meaning of "now", its pronunciation is Therefore, phonetic symbol embedding helps to improve the accuracy of model training.

[0040] The fusion embedding is obtained by concatenating the matrices obtained from word embedding, root embedding, affix embedding, and phonetic embedding after passing through a CNN, and then passing through a fully connected layer to obtain the fusion embedding corresponding to the Mongolian word. Taking the word (people) as an example, as Figure 8 shown.

[0041] The fusion embedding corresponding to each Mongolian word is added to the positional embedding as the model input.

[0042] It is easy to understand that in an embodiment of the present invention, when an NSP training task needs to be performed, the embedding layer is composed of a fusion embedding, a positional embedding, and a text embedding. The fusion embedding is added to the positional embedding and the text embedding to form the model input.

[0043] Step 3: In the Mongolian BERT pre-training model, perform pre-training for contrastive learning and the MLM task. In the contrastive learning task, the method of randomly dropping masks is used as a way of data augmentation to construct positive samples in contrastive learning. The same sample, that is, the input vector obtained from the embedding layer, is input into the Mongolian BERT pre-training model twice, and two different vectors s i and s' i are obtained by randomly dropping masks. These two vectors are used as a pair of positive samples, and another input in a randomly sampled batch is used as the negative sample s j .

[0044] The loss function L i for contrastive learning is:

[0045]

[0046] where ω is a hyperparameter and n is the size of a batch.

[0047] cos(s i , s' i ) is the cosine similarity between vector s i and vector s' i , and its formula is:

[0048]

[0049] cos(s i , s j ) is the cosine similarity between vector s i and vector s j , and its formula is:

[0050]

[0051] The MLM pre-training task randomly masks a portion (e.g., 15%) of the tokens. During the random masking process, a first proportion (e.g., 10%) of the words are replaced with other words, a second proportion (e.g., 10%) of the words remain unchanged, and the remaining words (80%) are replaced with the mask [MASK]. The loss function of the MLM pre-training task is:

[0052]

[0053] where θ are the parameters of the Encoder part in the Mongolian BERT pre-training model, θ′ are the parameters in the output layer connected to the Encoder in the MLM pre-training task, M is the set of masked words, m k is the masked word, p is the prediction probability of sample k, and |V| is the vocabulary size;

[0054] Therefore, the loss function that fuses contrastive learning and the MLM pre-training task is:

[0055]

[0056] Take (People have gone to work) as an example. The overall structure of the pre-training model is as Figure 9 shown.

[0057] Step 4, preprocess the Mongolian sentiment corpus.

[0058] Same as the preprocessing operation in Step 1, perform data cleaning on the Mongolian sentiment corpus, and then use BPE to segment the Mongolian language based on word roots and affixes. For words that cannot be segmented into word roots and affixes, keep them as they are.

[0059] Step 5, use the trained Mongolian BERT pre-training model that fuses word roots, affixes, and phonetic symbols to perform sentiment analysis on the Mongolian sentiment corpus.

[0060] Put the segmented corpus into the model for sentiment analysis, and add a softmax layer to the model to obtain sentiment classification. Take (Not worth this price) as an example. After sending it into the model, a classification result of "-1" (negative) will be obtained.

Claims

1. A Mongolian pre-training sentiment analysis method integrating roots, affixes and phonetic symbols, characterized in that: The steps include: Step 1, preprocessing the Mongolian corpus; Step 2, constructing a Mongolian BERT pre-trained model, wherein word embedding, root embedding, affix embedding and phonetic symbol embedding are constructed in its embedding layer; the embeddings are concatenated to obtain a fused embedding, and then the fused embedding is added to the position embedding to form a model input; Step 3, pre-training the fusion task of contrastive learning and MLM in the Mongolian BERT pre-training model; Step 4, preprocessing the Mongolian sentiment corpus; Step 5: Use the trained Mongolian BERT pre-trained model that integrates roots, affixes and phonetic symbols to perform sentiment analysis on the Mongolian sentiment corpus; In step 3, the positive sample is constructed in contrastive learning by using the random discard mask method, and the same sample, that is, the input vector obtained by the embedding layer, is input into the Mongolian BERT pre-training model twice, and two different vectors s are obtained by randomly discarding the mask. i and i ′ , will s i and i ′ As a positive sample pair, randomly sample another input in a batch as a negative sample s j , then the loss function L of contrastive learning is i for: Where ω is a hyperparameter and n is the size of a batch; cos(s i ,s i ′ ) is the vector s i and vector s i ′ The cosine similarity of is: cos(s i ,s j ) is the vector s i and vector s j The cosine similarity of is: The MLM pre-training task randomly masks a portion of the tokens. In the process of random masking, the first proportion of words are replaced by other words, the second proportion of words remain unchanged, and the remaining words are replaced by masks [MASK]. The loss function of the MLM pre-training task is: Where θ is the parameter of the Encoder part in the Mongolian BERT pre-training model, θ ′ is the parameter in the output layer connected to the Encoder in the MLM pre-training task, M is the set of masked words, and m k is the masked word, p is the predicted probability of sample k, and |V| is the dictionary size; Then, the loss function of the MLM pre-training task integrating contrastive learning is:

2. The Mongolian pre-training sentiment analysis method integrating roots, affixes and phonetic symbols according to claim 1 is characterized in that: The step 1, pre-processing comprises: Data cleaning and word segmentation.

3. The Mongolian pre-training sentiment analysis method integrating roots, affixes and phonetic symbols according to claim 2 is characterized in that: The word segmentation is to separate each Mongolian word into roots and affixes; the words that cannot be segmented into roots and affixes are kept as they are.

4. The Mongolian pre-training sentiment analysis method integrating roots, affixes and phonetic symbols according to claim 1 is characterized in that: The embedding layer consists of fused embedding and position embedding.

5. The Mongolian pre-training sentiment analysis method integrating roots, affixes and phonetic symbols according to claim 1 is characterized in that: The embedding layer is composed of fusion embedding, position embedding and text embedding, and the fusion embedding is added to the position embedding and text embedding to form the model input.

6. The Mongolian pre-training sentiment analysis method integrating roots, affixes and phonetic symbols according to claim 1 is characterized in that: In step 2, word embedding is to embed at the word granularity; root embedding is to embed the root after word segmentation; affix embedding is to embed the affix after word segmentation, and for words that cannot be segmented into roots and affixes, their prototypes are directly embedded as roots and affixes; phonetic symbol embedding is to embed the international phonetic symbols corresponding to Mongolian words.

7. The Mongolian pre-training sentiment analysis method integrating roots, affixes and phonetic symbols according to claim 1 is characterized in that: The step 2, fusion embedding, is to splice together the matrices obtained after word embedding, root embedding, affix embedding and phonetic symbol embedding through CNN, and obtain the fusion embedding corresponding to the Mongolian vocabulary after passing through a fully connected layer.

8. The Mongolian pre-training sentiment analysis method integrating roots, affixes and phonetic symbols according to claim 1 is characterized in that: In step 4, similar to the preprocessing operation in step 1, a data cleaning operation is performed on the Mongolian sentiment corpus, and then Mongolian is segmented based on roots and affixes using BPE, and words that cannot be segmented into roots and affixes are kept as they are.

9. The Mongolian pre-training sentiment analysis method integrating roots, affixes and phonetic symbols according to claim 1 is characterized in that: In step 5, the segmented corpus is put into the model for sentiment analysis, and a softmax layer is added to the model to obtain sentiment classification.

Citation Information

Patent Citations

  • Online public opinion evolution simulation method and system based on deep learning

    CN112395417A

  • Specific target sentiment analysis method based on word shielding data enhancement and adversarial learning

    CN113723076A