Bullet screen sentiment analysis model construction method based on two-channel semantic enhancement

By introducing a dual-channel semantic enhancement mechanism and DPCNN model into the barrage sentiment analysis model, the problems of insufficient semantic information extraction of barrage text and excessive feature dimensions are solved, and more accurate and generalized sentiment analysis is achieved.

CN120020816APending Publication Date: 2025-05-20XINJIANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311561003.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-20
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

The prior art has shortcomings in the problems of insufficient extraction of semantic information in barrage text, excessive dimensions of emotional classification characteristics, and loss of initial semantic information.

Method used

A barrage sentiment analysis model based on dual-channel semantic enhancement is proposed. Global semantic information and word order information are extracted through the GCN model and the Bi-LSTM model, and combined with the sentence vector enhancement mechanism and the DPCNN model, the feature dimension is reduced and high-level features are extracted.

Benefits of technology

Effectively extract the word order and global semantic information of barrage text, reduce the dimension of emotional classification characteristics, improve the accuracy of sentiment analysis and the generalization performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020816A_ABST
    Figure CN120020816A_ABST
Patent Text Reader

Abstract

The invention discloses a bullet screen sentiment analysis model construction method based on double-channel semantic enhancement based on a deep learning method. The method comprises the following steps: firstly, acquiring a bullet screen text and completing data annotation to form a bullet screen data set; then, data preprocessing operations such as stop word removal and meaningless punctuation mark deletion are carried out; then, word embedding is carried out through BERT; then, using a GCN module and a Bi-LSTM module to form two channels to respectively extract global semantic features and word order features, and fusing the features extracted by the two channels; then, enhancing the fused features by using a sentence vector enhancement mechanism; thirdly, extracting high-level features in the enhanced features by using a DPCNN model, and reducing the dimensionality of the high-level features; and finally, performing sentiment classification on the low-dimensional and high-level features extracted by the DPCNN by using a multi-layer perceptron. According to the method, the problems of insufficient semantic information extraction, semantic loss and too high classification dimension in an existing bullet screen sentiment analysis model construction method can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field:

[0001] The present invention relates to the field of sentiment analysis, and particularly to a method for constructing a barrage sentiment analysis model based on dual-channel semantic enhancement, which is used to extract rich semantic information from barrage texts and perform sentiment classification. Background Art:

[0002] Sentiment analysis (SA) is one of the hot research directions in natural language processing and an important dimension of public opinion analysis. Its main goal is to classify a given text into multiple sentiment categories (such as positive, negative, or neutral), extract and identify people's sentiment tendencies towards a certain entity, and is of great significance to fields such as content recommendation, web public opinion mining, and social media analysis.

[0003] Barrage is a new type of short text that scrolls across videos and is widely favored by young users. During video viewing, viewers can send barrages to communicate in real time around the video content and convey their own ideas and emotions at the same time. The huge volume of barrage texts contains rich emotional information, and mining this information can provide important reference value for fields such as video content recommendation, computational advertising, and computational communication. However, barrages themselves have the characteristics of semantic sparsity and polysemy, and contain a large number of non-verbal emotional elements such as emoticons and kaomoji, which pose challenges to barrage sentiment analysis.

[0004] In recent years, due to the powerful learning and representation ability of deep learning models, the research perspective in the field of short text sentiment analysis has gradually shifted from machine learning to deep learning methods. As a new type of short text, the sentiment analysis of bullet comments also mainly focuses on deep learning methods. Bullet comments are spread in the online world in the form of text, but the text cannot be directly recognized by a computer. Therefore, it is necessary to map the text to data that can be recognized by a computer (such as ASCII code) in advance before performing operations on it. Similarly, in deep learning, it is necessary to map the text to word vectors in advance before it can be recognized by a deep learning model. This process is called "word embedding", and the model used to complete word embedding is called a "word embedding model". In sentiment analysis tasks, it is generally believed that the richer the semantic information extracted by the word embedding model, the more fully the subsequent calculations can mine the semantic information, which is more conducive to improving the accuracy of sentiment analysis. However, there are usually three problems in the existing technologies: (1) insufficient extraction of semantic information; (2) in the calculation steps of the model, some of the initial semantic information extracted by the word embedding model is easily lost; (3) the feature dimension for sentiment classification is too high, resulting in too much redundant information, thus affecting the generalization performance of the model. Therefore, the existing technologies need a method for constructing a bullet comment sentiment analysis model that can fully extract the semantic information of bullet comment text, alleviate the problem of loss of initial semantic information, and at the same time reduce the feature dimension of sentiment classification without losing semantic information. Summary of the Invention:

[0005] The purpose of the present invention is to address the deficiencies of the existing technologies and propose a bullet comment sentiment analysis model based on Dual-Channel Semantic Enhancement (DCSE), aiming to solve the problems of insufficient extraction of semantic information from bullet comment text and too high feature dimension for sentiment classification. At the same time, the present invention proposes a Sentence Vector Enhancement Mechanism (SVE), which is embedded in the model as a module of DCSE, aiming to address the problem of loss of initial semantic information.

[0006] The method of the present invention is implemented as follows:

[0007] Step S01: Obtain bullet comment text and construct a bullet comment corpus;

[0008] Step S02: Install and configure the Doccano data annotation platform, and perform sentiment annotation on the bullet comment text in the corpus to form a bullet comment dataset;

[0009] Step S03: Preprocess the bullet comment text in the bullet comment dataset;

[0010] Step S04: Perform word embedding using the BERT pre-trained model fine-tuned with bullet screen text;

[0011] Step S05: Take non-verbal emotional elements such as emojis and Japanese emoticons contained in the bullet screen dataset as nodes, and construct a semantic network graph based on the co-occurrence relationship of non-verbal emotional elements according to their co-occurrence relationship;

[0012] Step S06: Use the GCN model to perform graph convolution operation on the semantic network graph constructed in Step S05 to extract the global semantic information of the nodes;

[0013] Step S07: Use the Bi-LSTM model to process the word vectors generated in Step S04 to extract word order information;

[0014] Step S08: Integrate the global semantic information of the character nodes extracted in Step S06 with the word order information extracted in Step S07 to obtain comprehensive features;

[0015] Step S09: Use the sentence vector enhancement mechanism to enhance the features of the comprehensive features after fusion in Step S08;

[0016] Step S10: Use the DPCNN model to process the features obtained in Step S09, extract high-level features and reduce the dimension;

[0017] Step S11: Use a Multilayer Perceptron (MLP) to classify the low-dimensional high-level features obtained in Step S10.

[0018] Preferably, in the said Step S01, view the HTML source code of the page where the video is located, search for the cid code of the video, and obtain the bullet screen data file of the video through the website https: / / comment.bilibili.com / [cid].xml. Since the bullet screen data not only contains bullet screen text, but also contains various attribute tags of the bullet screen (color, font size, bullet screen mode, etc.), it is necessary to extract the bullet screen text in the bullet screen data to form a bullet screen corpus.

[0019] Preferably, in the said Step S02, Doccano is an open-source data annotation platform that can provide support for data annotation tasks such as text classification, named entity recognition, and sequence labeling. This platform provides the function of multi-person collaboration for data annotation, improving the efficiency of data annotation.

[0020] Preferably, in step S03, in order to avoid the influence of unnecessary noise on the model, it is first necessary to remove stop words and meaningless punctuation marks from the bullet screen text in the dataset, but keep non-verbal emotional elements such as emojis and kaomojis, because these elements contain rich emotional connotations. Among them, stop words refer to certain words or terms that need to be filtered out before processing natural language text to save storage space and improve information retrieval efficiency. Then, the average text length of the bullet screen is statistically calculated, and the part of the bullet screen that exceeds the average length is truncated.

[0021] Preferably, in step S04, the BERT pre-trained model is a word embedding model that can dynamically extract word vectors with rich semantic information from the bullet screen text. For the same word, BERT can generate different word vectors to represent the word according to the change of the context, which is a capability that traditional word embedding models represented by Word2vec, FastText, Glove, etc. do not have. Therefore, using BERT as the word embedding model for the bullet screen text can better fit its polysemous characteristics. In addition, BERT supports fine-tuning the pre-trained model using the collected bullet screen text, so that the generated word vectors are more targeted at the bullet screen text.

[0022] Preferably, in step S05, co-occurrence means whether two characters have ever appeared together in the same paragraph or the same text. If the two have appeared together, it means that the two together constitute the text context. The co-occurrence relationship between two characters indicates that they have semantic associations. Therefore, a semantic network graph can be established according to the co-occurrence relationship between characters to reflect the semantic relationship between each character in the dataset. In the present invention, by constructing a semantic network graph based on the co-occurrence relationship of non-verbal emotional elements, the purpose of enriching the semantics of emojis and kaomoji symbols can be achieved.

[0023] Preferably, in step S06, the GCN model can aggregate the node features by using the information of the edges in the semantic network graph, thereby generating a new node feature representation, and thus can realize the propagation of feature information along the edges in the graph to achieve the purpose of extracting semantic information from the co-occurrence relationship level of non-verbal emotional elements according to the semantic network graph;

[0024] Preferably, in step S07, the Bi-LSTM model is composed of two LSTM models in opposite directions, and can extract the semantic information of the bullet screen text from both positive and negative directions at the word order level, so as to effectively handle the two situations of the front and back placement of modifiers in the bullet screen text. Due to the superposition of the two LSTM models, the output features of the Bi-LSTM model have a relatively high dimension, which is twice the output dimension size of the LSTM model;

[0025] Preferably, in step S08, since the word vector generated in step S04 is extracted using the GCN model and the Bi-LSTM model in step S06 and step S07, respectively, the two types of features need to be fused to obtain a comprehensive feature, and ensure that the comprehensive feature can simultaneously contain the features of word order and global semantics extracted by the co-occurrence relationship of non-text emotional elements, so as to maximize the semantic information contained in the comprehensive feature;

[0026] Preferentially, in step S09, the sentence vector enhancement mechanism is proposed based on the group-by-group enhancement mechanism, which is mainly used to semantically enhance the comprehensive features integrated in step S08, strengthen the weight of key information in the features, and weaken the weight of non-key information. Among them, the group-by-group enhancement mechanism is a method of grouping word vectors, taking the average value of each group of feature vectors and calculating the similarity with each vector in the group, so as to further calculate the weight that should be assigned to each word vector. Because the group-by-group enhancement mechanism uses the average value of each group of word vectors as a reference to calculate semantic similarity, the mechanism lacks the ability to enhance word vectors based on the initial semantics of the barrage text. Therefore, the sentence vector enhancement mechanism uses the sentence vector of the barrage text generated by BERT to improve the group-by-group enhancement mechanism, and the sentence vector is added to the average vector to form a new evaluation vector to achieve the purpose of enhancing the word vector based on the initial semantics. Since the word vector generated by the BERT pre-training model used in step S04 already has rich semantic information, in the processing of step S06 and step S07, although the characteristics of the word vector are enhanced from the two aspects of word co-occurrence relationship and word order, the initial semantic information contained in the word vector generated by BERT will be lost because the values ​​of each element in the word vector have changed significantly. The lost semantic information plays an important role in alleviating the common polysemy phenomenon in barrage texts. Therefore, it is necessary to alleviate the semantic loss problem through the sentence vector enhancement mechanism;

[0027] Preferentially, in step S10, the comprehensive features obtained after fusion in step S08 have a higher dimension, while step S09 only enhances the semantic features without reducing the feature dimension. On the one hand, too high feature dimensions bring more parameters and increase the computational load during model training. On the other hand, high feature dimensions usually contain redundant feature information, which reduces the generalization ability of the model. Therefore, the DPCNN model is used to reduce the feature dimension while extracting high-level information. Compared with the classic CNN model, the computation time of each layer of the DPCNN model can be reduced exponentially in a "pyramid shape", which can increase the network depth without increasing the computational cost, and can extract information more efficiently when processing text data;

[0028] ​Preferably, in step S11, a multi-layer perceptron is used to process the high-level features extracted in step S10, and a Sigmoid activation function is embedded in the multi-layer perceptron to enhance its non-linear fitting ability.

[0029] The beneficial effects of the present invention are as follows: First, the present invention proposes a topological structure of a bullet screen sentiment analysis model based on dual-channel semantic enhancement, which can extract word order information and global semantic information from bullet screen texts, thereby providing comprehensive semantic support for sentiment classification; Second, based on the group-by-group enhancement mechanism, the present invention proposes a sentence vector enhancement mechanism to address the problem of initial semantic loss; Finally, the present invention uses the DPCNN model to process high-dimensional feature information, extracts high-level features without significantly increasing the computational amount, and reduces the feature dimension for classification. Description of the drawings:

[0030] Figure 1 It is a flowchart of a method for constructing a bullet screen sentiment analysis model based on dual-channel semantic enhancement;

[0031] Figure 2 It is an architecture diagram of a bullet screen sentiment analysis model based on dual-channel semantic enhancement;

[0032] Figure 3 It is a structure diagram of the sentence vector enhancement mechanism; Specific implementation manners:

[0033] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the following further describes the implementation solutions of the present invention in detail with reference to the drawings: In this method, first, the collected bullet screen data file is processed to extract the bullet screen text therein to form a corpus; Second, the Doccano data annotation platform is installed and configured, and the bullet screen text is uploaded to this platform for sentiment annotation to form a data set; Then, preprocessing operations such as removing stop words and meaningless punctuation marks are performed on the bullet screen text in the data set; Subsequently, the BERT pre-trained model is fine-tuned using the bullet screen text to generate word vectors, and the Bi-LSTM model and the GCN model are respectively used to extract word order information and global semantic information, and the features extracted by the two are fused; Next, the sentence vector enhancement mechanism is used to enhance the semantics of the fused feature vectors, and the DPCNN model is used to reduce the dimension of the enhanced features; Finally, a multi-layer perceptron is used for sentiment classification. As Figure 1 shown, the bullet screen sentiment analysis method based on dual-channel semantic enhancement specifically includes the following steps:

[0034] Step S01: Obtain the barrage text and construct a barrage corpus. Select a certain video on the Bilibili website, search for its cid code (an identifier used to obtain video barrage and other information on this website) in the web page source code parsed by the browser, and obtain the xml file containing the barrage data of the movie by accessing the website https: / / comment.bilibili.com / [cid].xml (the [cid] in the website needs to be replaced with the cid code of the video, and the xml file pointed to by this website stores the barrage information of a specific video), and download this file to the local for saving. Since the barrage text used in the present invention exists in the d tag of the xml file, a Python script is written using regular expressions to extract the barrage text content in the d tag to form a barrage corpus;

[0035] Step S02: Data annotation. In the terminal, use the pip command to install the Doccano data annotation platform. After installation, initialize it and create an administrator account, and then start the Doccano platform locally using the system's port number 8000. Create a text classification task in Doccano, upload the barrage corpus in Step S01 to the Doccano database, and click the "Start Annotation" button to start annotating the sentiment tendency of the barrage text. The final sentiment label of the barrage is determined by five annotators through voting. After annotation, export the barrage text to form a barrage dataset.

[0036] Step S03: Text preprocessing. Preprocess the barrage text in the barrage dataset obtained in Step S02. Use the Jieba module in Python to segment the barrage text, loop through each word segmented, and at the same time compare it with the Harbin Institute of Technology stop word list. If the word exists in the stop word list, delete it. Similarly, compare it with the constructed punctuation list (input by the keyboard) and delete meaningless punctuation. Finally, traverse all the barrage texts in the dataset and calculate the average length l mean , and truncate the part of some barrages that exceeds the average length l mean to achieve the purpose of unifying the text input length.

[0037] Step S04: Word embedding. Fine-tune the barrage text in the barrage dataset obtained in Step S03 on the BERT pre-trained model and obtain the word vector matrix A k ∈R n×e , which can be expressed by the formula:

[0038]

[0039] where n represents the length of the barrage text, X i$(i = 1, 2, \ldots, n)$ represents the word vectors generated by BERT, and $e$ represents the dimension of the word vectors. The operator represents concatenating multiple vectors.

[0040] Step S05: Construction of the semantic network graph. First, the non-text sentiment elements in the bullet screen dataset are de-duplicated and combined in pairs, and stored in a dictionary. Secondly, each bullet screen text in the dataset is traversed to extract the co-occurrence relationship of the non-text sentiment elements. If there is a co-occurrence relationship between two non-text sentiment elements, the co-occurrence times are recorded in the dictionary. Finally, using the Networkx module in the third-party library of Python, based on the co-occurrence information recorded in the dictionary, a semantic network graph based on the co-occurrence relationship of non-text sentiment elements is generated. In the graph, the nodes represent non-text sentiment elements, and the edges between the nodes represent co-occurrence relationships. If there is an edge between the nodes, it means that there is a co-occurrence relationship between the two non-text sentiment elements represented by the nodes at both ends in the corpus, and the co-occurrence times are used as the weight of the edge; if there is no edge between the nodes, it means that there is no co-occurrence relationship between the characters corresponding to the nodes at both ends.

[0041] Step S06: Extraction of global semantic features. The word vectors generated in step S04 and the semantic network graph based on the co-occurrence relationship of non-text sentiment elements constructed in step S05 are input into the GCN model. The GCN extracts the global semantic information in the semantic network graph, which can be expressed by the formula:

[0042]

[0043] Among them, $C$ k represents the node feature vector of the $k$-th layer, represents the adjacency matrix of the semantic network graph, $D$ k represents the weight matrix of the $k$-th layer, and $\sigma$ represents the activation function.

[0044] Step S07: Extraction of word order features. The word vectors generated in step S04 are input into a two-layer Bi-LSTM model. Utilizing the characteristic of the Bi-LSTM to extract features bidirectionally, the word order features of the bullet screen text are extracted from both the forward and reverse directions respectively, which can be expressed by the formula (taking the forward LSTM as an example):

[0045] $E$ t $=\sigma(A$ t $D$ ae $+U$ t-1 $D$ ue $+z$ e )

[0046] $G$ t $=\sigma(A$ t $D$ ag $+U$ t-1 $D$ ug+z g )

[0047] J t = σ(A t D aj + U t-1 D uj +z j )

[0048]

[0049]

[0050]

[0051] Among them, E t is the output of the input gate, G t is the output of the forget gate, J t is the output of the output gate, A t is the word vector, U t-1 is the hidden state output by the previous layer, D represents the weight matrix corresponding to each vector, and z is the bias value corresponding to each formula. is the candidate memory cell, K t is the memory cell. is the hidden state output by the forward LSTM, and ⊙ is the element-wise multiplication.

[0052] The calculation steps of the reverse LSTM are the same as those of the forward LSTM. The only difference is that the reverse LSTM extracts information by looping forward from the last word of the text. Therefore, according to the above calculation steps, it can be similarly obtained that The output of the Bi-LSTM can be expressed by the formula:

[0053]

[0054] Among them, U t is the output of the Bi-LSTM at the t-th character position.

[0055] Step S08: Semantic fusion. Since the output vector of the GCN model is the set of word vectors of all characters (non-repeating characters), it is necessary to extract the word vector corresponding to each character of the bullet screen text from C obtained in step S06 and connect them in the text order to form a matrix that can represent the text information, which can be expressed by the formula:

[0056]

[0057] Among them, C i is the word vector in the global semantic feature C extracted by the GCN, M is the bullet screen feature matrix formed by connecting the corresponding word vectors in the node global semantic feature C, and n is the text length.

[0058] Fuse M and U by addition, which can be expressed by the formula:

[0059] P = M + U

[0060] where P is the comprehensive feature after fusion.

[0061] Step S09: Sentence vector enhancement. Use the sentence vector enhancement mechanism to enhance the comprehensive feature obtained in step S08. This method can re - distribute the weights of each word vector with a relatively low computational cost. First, divide the feature matrix into several groups on average, which can be expressed by the formula:

[0062] P = {P 1 , P 2 , P 3 , …, P m}

[0063] p i = mean(P i )

[0064] v = p i + s

[0065] where P i = {p i1 , p i2 , p i3 , …, p ij} is the sub - matrix after grouping, p ij , j ∈ {1, 2, 3, …, n / m} is the j - th word vector in the sub - matrix P i , the shape of the sub - matrix P i is (n / m)×k, k is the dimension size of the feature vector, p i is the average vector of all word vectors in the sub - matrix P i , s is the sentence vector generated by the BERT pre - trained model in step S04, and v is the evaluation vector after fusing the sentence vector and the average vector.

[0066] Secondly, traverse each sub - matrix, and use the evaluation vector v to perform dot - product operations with each word vector in the sub - matrix respectively. Regard the result of the dot - product as the gap between the word vector and the evaluation vector, which can be expressed by the formula:

[0067] y ij = v · p ij

[0068] where y ij represents the gap between the word vector p ij and the evaluation vector v.

[0069] Then, process y using a linear layer and the sigmoid function ij to obtain the updated weight of p ij , which can be expressed by the formula:

[0070] r ij = sigmoid(Linear(y ij ))

[0071]

[0072] where r ij is the updated weight of p ij .

[0073] Finally, concatenate the enhanced sub - matrices, which can be expressed by the formula:

[0074]

[0075] Step S10: Convolutional dimensionality reduction. Since the features output by the Bi - LSTM model have a very high dimension (SVE only re - assigns the weights of word vectors and does not change the feature dimension), which is not conducive to direct sentiment classification. Therefore, the DPCNN model can reduce the feature dimension through pooling operations while extracting high - level feature information, which can be expressed by the formula:

[0076] Q = dpcnn(P)

[0077] where Q represents the feature vector after dimensionality reduction.

[0078] Step S11: Sentiment classification. Use a fully - connected layer composed of three linear layers to classify the feature Q, and embed the Sigmoid activation function in it to enhance its fitting ability for non - linear functions, which can be expressed by the formula:

[0079] T = mlp(Q)

[0080] where T is a tensor with a shape of [batch size, 2]. In its two dimensions, the side with the larger value is the model prediction result.

Claims

1. A method for constructing a barrage sentiment analysis model based on dual-channel semantic enhancement, characterized in that: include: (1) Collect barrage texts to form a barrage corpus. Use the Doccano data annotation platform to annotate the sentiment polarity of the barrage texts. In this process, the final sentiment polarity of the barrage is determined by five annotators through voting, thus forming a barrage dataset; (2) Each module in the barrage sentiment analysis model based on dual-channel semantic enhancement (DCSE) has a specific combination order; (3) The proposed sentence vector enhancement mechanism (Sentence Vector Enhancement Mechanism, SVE) has a specific order of steps for the calculation of sentence vectors.

2. According to claim 1, each module in the barrage sentiment analysis model based on dual-channel semantic enhancement has a specific combination order, characterized in that: First, the barrage text is preprocessed, and non-text emotional elements such as emoticons and emojis in the barrage are retained; after preprocessing, the barrage text containing non-text emotional elements is embedded using a pre-trained model that has been fine-tuned on the barrage text; After word embedding, a semantic network diagram based on the co-occurrence relationship between non-textual sentiment elements is constructed, and the GCN module is used to extract global semantic features in the network diagram; after word embedding (while constructing the semantic network diagram and extracting global semantic features), the Bi-LSTM module is used to extract the word order features of the barrage text; after extracting the global semantic features and word order features, the two features are semantically fused by vector addition; after semantic fusion, the sentence vector enhancement mechanism is used to semantically enhance the fused features; after semantic enhancement, the fused vector is convoluted and reduced in dimension using the DPCNN module; finally, after convolution and dimensionality reduction, the reduced-dimensional features are classified using a multi-layer perceptron consisting of three linear layers and a Sigmoid activation function.

3. The specific combination order of sentence vector operation steps in the sentence vector enhancement mechanism according to claim 1, characterized in that: First, the feature matrix is ​​divided into several sub-matrices. Second, the sentence vector is added to the average vector of each sub-matrix. Then, the dot product operation is performed on the added vector and each word vector in the sub-matrix. Finally, a linear transformation is performed on the dot product result to obtain a new weight distribution.