A low-resource language news sentence sentiment analysis method and system based on grid structure
By converting the grid structure into a flat structure and combining relative position coding and multi-head attention mechanism, the problem of information overload in low-resource language news is solved, and efficient identification and classification of emotional information is achieved.
Patent Information
- Application Number
- CN202410275955.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-12
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-03-12
AI Technical Summary
Low resource language official news information is overloaded, and how to quickly extract hot events and emotional information from complicated texts has become an important need.
A low-resource language news sentence sentiment analysis method based on grid structure and multi-head attention is proposed. By destructively converting directed acyclic grid into a flat structure, word information is introduced, and the correlation relationship and semantic information of syllables and words is recognized by the multi-head self-attention mechanism, and emotional classification is finally performed through the full connection layer.
By enhancing sequence semantic understanding and effectively encoding the position and direction information of syllables and words, the accuracy and efficiency of sentiment analysis of low-resource language news texts is improved.
Smart Images

Figure CN118093874B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a sentiment analysis method in the field of low-resource language processing. Aiming at the actual needs of sentiment analysis of news sentences in low-resource languages, a sentiment analysis method for news sentences in low-resource languages based on a grid structure and multi-head attention is proposed. Background Art
[0002] As we enter the era of big data, official news in low-resource languages is growing explosively and in a variety of formats, and information overload is a serious problem. How to quickly obtain hot events from the complex official news in low-resource languages has become an important need. Summary of the invention
[0003] In order to meet the above application needs, the present invention takes low-resource language news text data as the research object, and proposes a low-resource language news sentence sentiment analysis method based on grid structure and multi-head attention. First, the directed acyclic grid is losslessly converted into a flat structure to introduce word information in the syllable sequence. Secondly, the relative position encoding mechanism is adopted to effectively encode the position and direction information of syllables and words. Then, the multi-head self-attention mechanism is used to identify the association and semantic information of syllables and words in the text. Finally, the sentiment category of low-resource language news text is obtained through full connection layer classification.
[0004] The present invention provides a low-resource language news sentence sentiment analysis method based on grid structure and multi-head attention, comprising the following steps:
[0005] 1. Preprocess the crawled low-resource language news text;
[0006] 2. Map the text segmented by syllables and words into vector representation;
[0007] 3. Convert the directed acyclic grid into a flat structure to enhance the word information in the syllable sequence;
[0008] 4. Use relative position encoding mechanism to effectively encode the position and direction information of syllables and words;
[0009] 5. Use multi-head self-attention mechanism to identify the association and semantic information of syllables and words in text;
[0010] 6. Realize sentiment classification of low-resource language news through classification network;
[0011] 7. Train the retrieval network model with training data and update the parameters, then test it on the test set.
[0012] The present invention discloses a low-resource language news sentence sentiment analysis method based on grid structure and multi-head attention. In step 1, the original low-resource language news hot event sentences are captured by a web crawler and pre-processed. Positive sentiment sentences and negative sentiment sentences are extracted from the original data manually and annotated as experimental data sets.
[0013] In the low-resource language news sentence sentiment analysis method based on grid structure and multi-head attention of the present invention, in step 2, in order to alleviate the impact of low-resource language word segmentation operation on sentiment analysis tasks, syllables and words are used as input units of the model, and the segmentation of low-resource language text is achieved through syllable separators. The Word2vec framework is used to obtain the Word2vec vectors of syllables and words, and the BERT-BOD pre-training model is used to obtain the BERT syllable vector of the syllable.
[0014] The low-resource language news sentence sentiment analysis method based on grid structure and multi-head attention of the present invention, in the step 3, in order to achieve the purpose of introducing word information to enhance the sequence semantic information, the grid structure is losslessly converted into a flat structure composed of several units, the start position and the end position of the syllable in the flat grid are the same, and the start position and the end position of the word are jumping. The flat grid can be restored to the original grid structure by a simple algorithm. Specifically, it is first assumed that the start position and the end position of the grid identifier (syllable or word) are the same, and then the jump path is established using the start position and the end position of the remaining words.
[0015] The low-resource language news sentence sentiment analysis method based on grid structure and multi-head attention of the present invention, in the step 4, for the flat grid structure, the position relationship vector is calculated by continuous transformation of the start position and the end position information. The position relationship vector can not only represent the relationship between the two identifiers, but also indicate other detailed information such as the distance between syllables and words, which is very important for improving the performance of classification tasks.
[0016] The low-resource language news sentence sentiment analysis method based on grid structure and multi-head attention of the present invention, in step 5, adopts a fully connected Softmax classification network to realize the low-resource language news sentiment classification.
[0017] In the low-resource language news sentence sentiment analysis method based on grid structure and multi-head attention of the present invention, in step 6, the cross entropy loss function is used to measure the gap between the true distribution and the predicted distribution, and the parameters in the model are updated using the back propagation method.
[0018] Compared with the prior art, the invention has the following beneficial effects: by losslessly converting the grid into a flat structure, the semantic understanding of the sequence is enhanced. The relative position encoding mechanism is used to encode the position and direction information of syllables and words, which further improves the semantic understanding of the syllable sequence. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0020] Figure 1 It is a model structure diagram of a low-resource language news sentence sentiment analysis method based on a grid structure and multi-head attention of the present invention;
[0021] Figure 2 It is a grid structure diagram;
[0022] Figure 3 It is a flat grid structure diagram;
[0023] Figure 4 It is a schematic diagram of the performance of different models on the low-resource language news sentence sentiment analysis task. DETAILED DESCRIPTION
[0024] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0025] Figure 1 This is a model structure diagram of a low-resource language news sentence sentiment analysis method based on a grid structure and multi-head attention in the present invention, comprising the following steps:
[0026] 1. Preprocessing. (1) Delete non-low-resource language characters. Since low-resource language text contains a certain number of non-low-resource language characters such as Arabic numerals, English letters, Chinese and English punctuation marks, but the model input must be pure low-resource language characters, all non-low-resource language characters in the text data are deleted. (2) Low-resource language word segmentation. Low-resource languages do not have natural separators in the form of spaces, but the invention research involves the semantic understanding of low-resource language words. Therefore, text word segmentation is required in the preprocessing stage. The present invention uses a low-resource language word segmenter to achieve word segmentation. (3) Low-resource language syllable segmentation. In order to alleviate the impact of low-resource language word segmentation operations on sentiment analysis tasks, the present invention uses syllables and words as the input units of the model. Therefore, the preprocessing operation needs to complete the segmentation of low-resource language text according to syllable separators.
[0027] 2. Lossless conversion of the grid into a flat structure. The grid structure can use vocabulary information to avoid the propagation of word segmentation errors. The grid structure is a directed acyclic graph, where each node is a syllable or a potential word. The schematic diagram is as follows: Figure 2 As shown in the figure. The grid structure is obtained by matching the sentence sequence through the dictionary. It is a non-ordered sequence. The starting syllable and the ending syllable of the word in the grid determine the word position. The dictionary for building the grid relies on the Trie tree to obtain it. From the root node to the leaf node, all the syllables on the traversal path are connected to form a word. The Trie tree has the advantages of sharing common prefixes, saving storage space, and efficient retrieval.
[0028] In order to achieve the purpose of introducing word information to enhance the sequence semantic information, the grid structure is losslessly transformed into a flat structure composed of several units. A flat grid is defined as a series of units, where a unit is related to a syllable or word, a start position and an end position. The flat grid structure is as follows Figure 3 As shown in the figure. An identifier is a syllable or a word. The start position and the end position indicate the position of the identifier in the original sequence. The above information determines the detailed position of the identifier in the grid. The start position and the end position of the syllable in the flat grid are the same, and the start position and the end position of the word are jumping. The flat grid can be restored to the original grid structure using a simple algorithm. In order to restore the syllable sequence, it is first assumed that the start position and the end position of the identifier are the same. Then, the jump path is established using the start position and the end position of the remaining words. Since the transformation is reversible, the flat grid can retain the original structure of the grid.
[0029] 3. Character embedding. When grid characters are input into the model, they can be regarded as a character sequence. ,in Represents a set of characters. Characters Word2vec vector As shown in formula (1), BERT vector of syllables It is expressed as shown in formula (2).
[0030] (1)
[0031] (2)
[0032] in, Represents the character vector lookup table trained using the Word2vec framework. BERT represents the syllable vector lookup table of the BERT pre-trained model.
[0033] 4. Relative position encoding. The position relationship vector is calculated by continuously transforming the start position and the end position information. The position relationship vector indicates the relationship between the two identifiers and also indicates other detailed information such as the distance between syllables and words.
[0034] Use start[i] and end[i] to indicate the start and end positions of the identifier. and The four relative position representations of are shown in equations (3) to (6).
[0035] (3)
[0036] (4)
[0037] (5)
[0038] (6)
[0039] in, express and The distance between the head positions, , , The meaning is similar. The relative position encoding vector between identifiers It is a nonlinear transformation of the four distances, as shown in formula (7).
[0040] (7)
[0041] in, is a learnable parameter, Represents a splicing operation, The calculation is shown in equations (8) to (9).
[0042] (8)
[0043] (9)
[0044] in, yes, , , , ; k represents the dimension index of position encoding.
[0045] In the self-attention mechanism, the identifier The query vector With identifier The key vector The calculation is shown in equations (10) to (11).
[0046] (10)
[0047] (11)
[0048] in, and is an identifier The word vector and position vector of and is an identifier The word vector and position vector of is a learnable parameter.
[0049] Then calculate the self-attention score , the calculation formula is shown in equations (12)~(14).
[0050] (12)
[0051] (13)
[0052] (14)
[0053] in, is a learnable parameter.
[0054] In order to use the relative position encoding vector , in formula (14), the identifier The absolute position vector of is replaced by the relative position vector. Since the query vectors of all query positions are the same when relative position encoding is used, Replace with the parameters that need to be learned , similarly Replace with the parameters that need to be learned . In addition, the weight matrix of the replacement key for and Through the above processing, the identifier and Self-attention score under relative position encoding As shown in formula (15).
[0055] (15)
[0056] 5. Semantic learning. The above embedding layer vector and relative position encoding layer vector are input into the semantic learning layer for feature extraction. The semantic learning layer is composed of a Transformer Encoder network, which includes a self-attention sublayer and a feed-forward neural network FFN sublayer. Each sublayer is followed by a residual connection and layer normalization. FFN uses a nonlinear full-position multilayer perceptron. Transformer Enconder uses independent multi-head self-attention on the sequence and concatenates the multi-head results as the final result.
[0057] For simplicity, the multi-head index is ignored in the following formula, and the calculation formula for each head is shown in formulas (16) to (18).
[0058] (16)
[0059] (17)
[0060] (18)
[0061] in, It is the word vector or the output of the previous encoder. is the learning parameter, , is the dimension of each head. Substituting (18) , we can get the attention vector.
[0062] 6. Classification output. The classification layer puts the features extracted by the semantic learning layer into the classifier for fitting and testing. During the fitting process, the features and categories are passed into the network for learning. When all the training data are fitted, the fitted model is tested using the test data. The present invention inputs the final representation of the low-resource language news sentence obtained in the previous layer into the Softmax classification module through a fully connected network, and obtains the final category after classification. The classification calculation method is shown in formula (19).
[0063] (19)
[0064] in, Represents the output of the classification layer, indicating the probability of the text belonging to positive sentiment or negative sentiment. is the network weight matrix, Bias for the network.
[0065] This paper uses the cross entropy loss function to measure the gap between the true distribution and the predicted distribution, and uses back propagation to update the parameters in the model. The calculation method of the cross entropy loss function is shown in formula (20).
[0066] (20)
[0067] in, is the training dataset size, Indicates the classification category, Indicates The samples correspond to The expected output probability of each category is Indicates The samples correspond to During the model training process, the dropout method is used to randomly control some hidden layer nodes in the network to stop working to prevent the model from overfitting.
[0068] Embodiment 1:
[0069] The experimental results in this embodiment are obtained by testing on a self-built data set. The technical effects of the present invention are as follows:
[0070] Figure 4 Figure 1 is a schematic diagram of the performance of different models on the low-resource language news sentence sentiment analysis task.
[0071] (1) Sentiment analysis model based on CNN. The model input is the same as the method in this paper. Then, operations such as convolution and maximum pooling are used to learn semantic features. Finally, Softmax excitation is used to achieve sentiment classification.
[0072] (2) Sentiment analysis model based on BiLSTM+ATT. The model input is the same as the CNN model. The BiLSTM network is used to learn the sentence sequence and semantic information, and the self-attention mechanism is used to weight the output vector. Finally, the Softmax classification network is used to achieve sentiment classification.
[0073] (3) Sentiment analysis model based on AT-DPCNN. The model input is the same as the BiLSTM+ATT model. The attention network focuses on the important semantic features that affect the sentiment trend, and then uses the CNN network to extract the semantic features again. The pooling operation is used to extract complex sentence features, and finally the Softmax activation network is used to achieve sentiment classification.
[0074] from Figure 4It is not difficult to find that the model of the present invention has achieved the best performance in F1 value, and is significantly higher than the second best classification model based on AT_DPCNN. Compared with the sentiment analysis model based on CNN, the F1 value is improved by 3.99%. The sentiment analysis model based on BiLSTM+ATT has the lowest F1 value, which is only 82.02%. This is because the important semantic features that affect the sentiment trend in the sentiment analysis task of low-resource language news sentences determine the sentiment analysis results. Since the convolution used by the CNN network is one-dimensional convolution, it is similar to the feature representation of n-gram in the sentence. It focuses on the important local features in the sequence that are conducive to the classification of sentiment tendency. The introduction of important local features enhances the understanding of sentence semantics, thereby improving the accuracy of sentiment classification of low-resource language news sentences. In this experiment, the performance of the BiLSTM+ATT network in focusing on important local features is weaker than that of the CNN network, which makes the F1 of the sentiment analysis model based on BiLSTM+ATT low.
[0075] Comparing the sentiment analysis model based on CNN and the sentiment analysis model based on AT_DPCNN, we can see that the F1 values of the two models are not much different, and the accuracy of the sentiment analysis model based on AT_DPCNN is higher than that of the sentiment analysis model based on CNN. This is because the attention layer and pooling processing layer in the AT_DPCNN network more effectively focus on the important semantic features that affect the sentiment trend, thereby improving the accuracy of sentiment classification of news sentences in low-resource languages.
[0076] The model proposed in the present invention converts the grid structure into a flat structure without loss, realizes the introduction of word information in the syllable sequence, and enhances the understanding of the sequence semantics. At the same time, the use of the relative position encoding mechanism effectively encodes the position and direction information of syllables and words. In addition, the multi-head attention mechanism focuses on identifying the association and semantic information of syllables and words in the text on the basis of relative position encoding, further enhancing the understanding of sequence semantics.
[0077] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A low-resource language news sentence sentiment analysis method based on grid structure and multi-head attention, the method comprising the following steps: Preprocess the crawled low-resource language news text; Map the text segmented by syllables and words into vector representation; The directed acyclic grid is losslessly converted into a flat structure to realize the introduction of word information in the syllable sequence. The directed acyclic grid is characterized by: the grid structure is a directed acyclic graph, each node is a syllable, or a potential word; the grid structure is obtained by matching the sentence sequence through the dictionary, which is a non-ordered sequence, and the starting syllable and the ending syllable of the word in the grid determine the word position; the dictionary for building the grid depends on the Trie tree to obtain, traversing the leaf nodes from the root node, and connecting all the syllables on the traversal path to form a word; the directed acyclic grid is losslessly converted into a flat structure, including: the flat grid is defined as a series of units, one unit is related to a syllable or word, a starting position and an ending position; the starting position and the ending position of the syllable or word indicate its position in the original sequence, the starting position and the ending position of the syllable in the flat grid are the same, and the starting position and the ending position of the word are jumping; in order to restore the syllable sequence, it is first assumed that the starting position and the ending position of the identifier are the same, and then the jumping path is established using the starting position and the ending position of the remaining words. Since the conversion is reversible, the flat grid can retain the original structure of the grid; Adopting relative position encoding mechanism, effectively encoding the position and direction information of syllables and words; Use multi-head self-attention mechanism to identify the association and semantic information between syllables and words in text; Sentiment classification of news in low-resource languages through classification networks; The retrieval network model is trained and the parameters are updated according to the training data, and then tested on the test set.
2. The low-resource language news sentence sentiment analysis method based on grid structure and multi-head attention as claimed in claim 1, characterized in that: 100,000 original low-resource language news hot event sentences were captured and preprocessed by web crawlers; hot events were defined as news events that were continuously reported on the same website or reported multiple times on different websites; 30,000 positive sentiment sentences and 30,000 negative sentiment sentences were manually extracted from the original data and annotated as the experimental data set.
3. The low-resource language news sentence sentiment analysis method based on grid structure and multi-head attention as described in claim 2 is characterized by: (1) Delete non-low-resource language characters. Since low-resource language texts contain a certain number of Arabic numerals, English letters, and Chinese and English punctuation marks, but the model input must be pure low-resource language characters, all non-low-resource language characters in the text data need to be deleted; (2) Low-resource language syllable segmentation. In order to alleviate the impact of low-resource language word segmentation operations on sentiment analysis tasks, syllables and words are used as input units of the model. Therefore, the preprocessing operation needs to complete the segmentation of low-resource language texts according to syllable separators.
4. The low-resource language news sentence sentiment analysis method based on grid structure and multi-head attention as described in claim 3, characterized in that: In order to solve the problem that the native Transformer performs poorly in classification tasks, for the flat grid structure, the continuous transformation of the start and end position information to calculate the position relationship vector is deleted. The position relationship vector can not only represent the relationship between two identifiers, but also indicate the distance between syllables and words; start[i] and end[i] are used to represent the start and end positions of the identifier.