Sentiment Analysis Model Training Method, Device, Storage Medium and Electronic Device
By generating enhanced samples in sentiment analysis model training and using weighted cross-entropy loss function, the problem of category imbalance in sentiment analysis model training is solved, and the emotion classification accuracy is improved.
Patent Information
- Application Number
- CN202510258501.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-06
AI Technical Summary
There is a problem of category imbalance during the training of sentiment analysis model, which leads to insufficient learning of a few categories, affecting the accuracy of emotion classification.
By determining the target samples and their number proportions of a few emotional categories in the original text set, if the proportion is less than the preset value, the semantic correlation between the target samples and the context text is calculated, the enhanced samples are generated, and iterative training is used to use the weighted cross-entropy loss function to build a preset sentiment analysis model.
The diversity and coverage of a few categories of samples have been increased, the model's discriminant performance on a few categories of samples has been improved, effectively alleviating the impact of category imbalance and improving the accuracy of emotional classification.
Smart Images

Figure CN119761381B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and particularly relates to a method, device, storage medium and electronic device for training an emotion analysis model. Background Art
[0002] The basic task of emotion analysis is to extract and identify emotion information from text. With the rise of pre-trained language models such as BERT, significant breakthroughs have been made in emotion analysis technology.
[0003] Currently, in emotion analysis tasks, samples of different emotion categories are usually used to train an emotion analysis model. However, there will be a problem of class imbalance in the number of samples of different emotion categories in the sample training set. For example, the number of samples of negative and neutral emotion categories is usually less than that of positive emotion categories. This imbalance will cause the model to learn insufficiently for minority categories during training and tend to predict majority categories, thus affecting the overall classification performance of the emotion analysis model and unable to guarantee the emotion classification accuracy. Summary of the Invention
[0004] In view of this, the present application provides a method, device, storage medium and electronic device for training an emotion analysis model, and the main purpose is to be able to solve the problem of class imbalance during the training process of the emotion analysis model, so as to improve the emotion classification accuracy of the model.
[0005] According to the first aspect of the present application, a method for training an emotion analysis model is provided, and the method includes:
[0006] Determine the target samples belonging to the minority emotion category in the original text set, and the quantity ratio of the target samples in the original text set;
[0007] If the quantity ratio is less than a preset ratio, determine the semantic relevance between the context text corresponding to the target samples and the target samples;
[0008] Based on the semantic relevance, determine the enhanced samples corresponding to the target samples, and use the enhanced samples and the original text set as training samples together. Among them, if it is determined according to the semantic relevance that the context text is semantically relevant to the target samples, generate the enhanced samples based on the context text; if it is determined according to the semantic relevance that the context text is semantically irrelevant to the target samples, perform synonym replacement on the target samples to generate the enhanced samples;
[0009] Determine the initial emotion analysis model and the target category weights corresponding to different emotion categories, where the target category weights corresponding to the minority emotion category are greater than the target category weights corresponding to other emotion categories;
[0010] Input the training samples into the initial sentiment analysis model for aspect-based sentiment analysis to obtain the predicted sentiment categories of the training samples;
[0011] Based on the predicted sentiment categories and true sentiment categories of the training samples, and the target category weights corresponding to different sentiment categories, calculate the weighted cross-entropy loss function value;
[0012] According to the weighted cross-entropy loss function value, perform iterative training on the initial sentiment analysis model to construct a preset sentiment analysis model.
[0013] According to a second aspect of the present application, there is provided a sentiment analysis model training device, which includes:
[0014] A first determination unit, configured to determine target samples belonging to minority sentiment categories in the original text set, and the quantity ratio of the target samples in the original text set;
[0015] The first determination unit is further configured to determine the semantic relevance between the context text corresponding to the target sample and the target sample if the quantity ratio is less than a preset ratio;
[0016] The first determination unit is further configured to determine, based on the semantic relevance, enhanced samples corresponding to the target samples, and use the enhanced samples and the original text set as training samples together. Among them, if it is determined according to the semantic relevance that the context text is semantically relevant to the target sample, generate the enhanced samples based on the context text; if it is determined according to the semantic relevance that the context text is not semantically relevant to the target sample, perform synonym replacement on the target sample to generate the enhanced samples;
[0017] A second determination unit, configured to determine an initial sentiment analysis model and target category weights corresponding to different sentiment categories, where the target category weights corresponding to the minority sentiment categories are greater than the target category weights corresponding to other sentiment categories;
[0018] An analysis unit, configured to input the training samples into the initial sentiment analysis model for aspect-based sentiment analysis to obtain the predicted sentiment categories of the training samples;
[0019] A calculation unit, configured to calculate the weighted cross-entropy loss function value based on the predicted sentiment categories and true sentiment categories of the training samples, and the target category weights corresponding to different sentiment categories;
[0020] A training unit, configured to perform iterative training on the initial sentiment analysis model according to the weighted cross-entropy loss function value to construct a preset sentiment analysis model.
[0021] According to a third aspect of the present application, there is provided a storage medium on which a computer program is stored, and when the program is executed by a processor, the above-mentioned method for training an emotion analysis model is implemented.
[0022] According to a fourth aspect of the present application, there is provided an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, and when the processor executes the program, the above-mentioned method for training an emotion analysis model is implemented.
[0023] By means of the above technical solutions, compared with the prior art, an emotion analysis model training method, device, storage medium and electronic device provided by the present application can determine the semantic relevance between context text and target samples, and perform data augmentation according to the semantic relevance to determine augmented samples corresponding to the target samples, thereby increasing the diversity and coverage of minority class samples. At the same time, the present application introduces a weighted cross-entropy loss function, and the class weights in the weighted cross-entropy loss function can be dynamically adjusted according to the class distribution, that is, the class weights corresponding to minority emotion classes are made greater than the class weights corresponding to other emotion classes, so that the role of majority classes in the overall loss can be reduced, and the discrimination performance of the model for minority class samples can be improved, thereby effectively reducing the impact of class imbalance in the model training process and improving the emotion classification accuracy of the model.
[0024] The above description is only an overview of the technical solutions of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically exemplified below. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0026] Figure 1 A flowchart showing a method for training an emotion analysis model provided by an embodiment of the present application is shown;
[0027] Figure 2 A schematic diagram of the architecture of an emotion analysis model provided by an embodiment of the present application is shown;
[0028] Figure 3 A flowchart showing the feature extraction and analysis provided by an embodiment of the present application is shown;
[0029] Figure 4 A schematic diagram of the analysis result based on squid provided by an embodiment of the present application is shown;
[0030] Figure 5 Shows the schematic diagram of attention weights based on squid provided by the embodiment of the present application;
[0031] Figure 6 Shows the schematic diagram of the analysis result based on appetizer provided by the embodiment of the present application;
[0032] Figure 7 Shows the schematic diagram of attention weights based on appetizer provided by the embodiment of the present application;
[0033] Figure 8 Shows the schematic diagram of the analysis result based on meal provided by the embodiment of the present application;
[0034] Fig. 9 Shows the schematic diagram of attention weights based on meal provided by the embodiment of the present application;
[0035] Fig.10 Shows the structural schematic diagram of a sentiment analysis model training device provided by the embodiment of the present application. Detailed implementation manners
[0036] The present application will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.
[0037] There will be a problem of class imbalance in the number of samples of different sentiment categories in the sample training set. This imbalance will cause the model to have insufficient learning of minority categories during training, tend to predict majority categories, thereby affecting the overall classification performance of the sentiment analysis model and unable to guarantee the sentiment classification accuracy.
[0038] To solve the above problems, the embodiments of the present invention provide a sentiment analysis model training method, device, storage medium and electronic device, which are mainly applicable to aspect-based sentiment analysis, such as Figure 1 As shown, the method includes:
[0039] Step 10, determine the target samples belonging to the minority sentiment categories in the original text set, and the quantity ratio of the target samples in the original text set.
[0040] Among them, the original text set includes multiple texts used as samples. The description languages used for the multiple texts can be Chinese, English, or other languages. The types of the multiple texts can specifically be review texts, opinion texts, news texts, descriptive texts, etc. The sentiment categories specifically include negative, positive, and neutral. It should be noted that the sentiment categories in the embodiments of the present invention can be divided according to actual business requirements and are not limited to the above negative, positive, and neutral sentiment categories. In addition, the target samples are texts in the original text set that belong to the minority sentiment categories. For example, there are 100 texts in the original text set, and only 10 of them belong to the negative category. These 10 negative texts are the target samples.
[0041] For the embodiments of the present invention, aspect-based sentiment analysis aims to extract specific aspects from texts, such as the attributes of a certain product, and determine their sentiment polarities (sentiment classification). Obtain an original text set consisting of N reviews , where represents the i th text, represents the aspect concerned in this text, represents the corresponding sentiment polarity label (sentiment classification). The main purpose of the embodiments of the present invention is to learn a mapping function f so that:
[0042] f (( S i , A i ); θ ) = y i
[0043] where θ are model parameters.
[0044] To achieve this goal, the embodiments of the present invention construct a preset sentiment analysis model (XAI-BERT-MSO) to enable it to effectively represent the input text S i and the aspect A i , and capture the semantic relationship between them. After learning the joint representation of the text and the aspect in the high-dimensional space, the model can accurately classify the sentiment polarity.
[0045] Since the sentiment polarity (sentiment classification) is usually represented in text form, in order for the model to process and learn the sentiment labels in the aspect-based sentiment analysis task, the embodiments of the present invention convert these text labels into numerical forms, that is, perform label encoding. Based on this, the embodiments of the present invention define a mapping function to map the sentiment classification labels (such as positive, negative, and neutral) to numerical values respectively = 0, 1, 2, where represents the numerical label of the i th piece of text.
[0046] For the convenience of model training, the embodiments of the present invention further convert the numerical label into a one - hot encoded vector y i , where K = 3, K represents the number of emotion categories, and the specific form is:
[0047] y i = δ (1, y i + 1) , δ (2, y i + 1) , δ (3, y i + 1) ] ⊤
[0048] where δ ( k , y i + 1) represents the Kronecker delta function, which takes the value of 1 when k = y i + 1, and 0 otherwise. This encoding method in the embodiments of the present invention is concise and effective, and can facilitate the model to calculate the loss and perform backpropagation.
[0049] Furthermore, in order to solve the problem of class imbalance in the training process of the emotion analysis model, the embodiments of the present invention need to count the proportion of the number of target samples belonging to the minority emotion categories in the original text set. When this proportion is less than a certain value, sample augmentation is performed.
[0050] Before performing sample augmentation, the embodiments of the present invention also need to perform data cleaning on the original text set, that is, clean and standardize the original text data to eliminate noise and improve the performance and generalization ability of the model. For this process, the method includes: converting multiple pieces of text in the original text set into lowercase forms respectively to obtain multiple pieces of converted text; deleting the uniform resource locators, social media identifiers and redundant spaces in the multiple pieces of converted text to obtain multiple pieces of pre - processed text; determining the multiple pieces of pre - processed text as the cleaned text set. At the same time, determine the target samples belonging to the minority emotion categories in the cleaned text set, and the proportion of the number of the target samples in the cleaned text set. Among them, the uniform resource locator may refer to a URL link.
[0051] Specifically, for each piece of text S i Preprocessing is performed, including converting each piece of text to lowercase, removing URL links and social media identifiers, and finally eliminating extra spaces. After the above processing, the preprocessed text is obtained , and its specific expression is as follows:
[0052]
[0053] Among them, represents the combined operation of the above preprocessing steps. Through these processes in the embodiments of the present invention, the text data is standardized, laying a foundation for subsequent feature extraction and model training.
[0054] Step 20: If the quantity ratio is less than the preset ratio, determine the semantic relevance between the context text corresponding to the target sample and the target sample.
[0055] Among them, the preset ratio can be set according to actual business requirements, and the embodiments of the present invention do not make specific limitations thereon. For example, the preset ratio is set to 20%.
[0056] In the sentiment analysis task of the embodiments of the present invention, class imbalance is a common problem. Especially, the number of samples in the negative and neutral sentiment categories is usually less than that in the positive category. This imbalance will cause the model to have insufficient learning of the minority categories during training, tend to predict the majority category, and thus affect the overall classification performance.
[0057] To overcome the above problems, when the quantity ratio of the target samples in the minority sentiment categories is less than the preset ratio in the embodiments of the present invention, context sentences with semantic relevance are introduced to increase the number and semantic information of the minority category samples and improve the model's recognition ability for the minority categories. To determine the context sentences semantically relevant to the target sample, it is necessary to calculate the semantic relevance between the above context text and the target sample. For the calculation process of the semantic relevance, the method includes: determining the embedding vector corresponding to the context text and the embedding vector corresponding to the target sample; calculating the similarity between the context text and the target sample based on the embedding vector corresponding to the context text and the embedding vector corresponding to the target sample; and determining the semantic relevance between the context text and the target sample according to the similarity.
[0058] Specifically, extract the context sentences before and after each target sample in the minority sentiment categories from the original text set and , that is, the context text. When extracting the context text, boundary cases need to be considered. When i = 1, only extract the text ; wheni = N When, only extract the text .
[0059] Furthermore, in order to ensure the semantic relevance between the context text and the target sample, the embodiments of the present invention use the CLS vector of the BERT model to embed the sentences, calculate the similarity between the context text and the target sample, and then determine the semantic relevance between the context text and the target sample according to the similarity.
[0060] Step 30, based on the semantic relevance, determine the augmented sample corresponding to the target sample, and use the augmented sample and the original text set as the training samples together.
[0061] Wherein, if it is determined that the context text and the target sample are semantically relevant according to the semantic relevance, the augmented sample is generated based on the context text; if it is determined that the context text and the target sample are not semantically relevant according to the semantic relevance, the target sample is replaced with a synonym to generate the augmented sample.
[0062] For the embodiments of the present invention, after calculating the similarity between the context text and the target sample, if the similarity is greater than the preset similarity, it is determined that the context text and the target sample are semantically relevant, and the context text and the target sample are concatenated to generate the augmented sample corresponding to the target sample; if the similarity is less than or equal to the preset similarity, it is determined that the context text and the target sample are not semantically relevant, and the target sample is replaced with a synonym, and the target sample after the synonym replacement is concatenated with the target sample to generate the augmented sample corresponding to the target sample. Among them, the preset similarity can be set according to actual business needs, and the embodiments of the present invention do not make specific limitations on this.
[0063] Specifically, when the similarity is greater than the preset similarity, the target sample is concatenated with the context text to form the augmented sample , and the expression form is as follows:
[0064] =
[0065] Augmented sample Retains the sentiment label of the target sample , to ensure label consistency.
[0066] The embodiments of the present invention can ensure the semantic coherence of the generated text by selecting the context sentences semantically relevant to the target sample for concatenation.
[0067] For example, assume the original target sample is " The steak was overcooked and dry.”, its sentiment label is negative, and the context text is “ The ambiance of the restaurant was delightful. ”, and the context text is “ However, the dessert was a highlight of the meal. ”. The augmented sample generated by data augmentation is “ The ambiance of the restaurant was delightful. The steak was overcooked and dry. However, the dessert was a highlight of the meal. ”
[0068] This augmented sample contains both the negative sentiment evaluation of the target sample and combines the positive and neutral sentiment evaluations in the context, thereby being able to provide more comprehensive sentiment clues for the model.
[0069] Furthermore, in order to avoid excessive increase in the number of samples of the above categories, the embodiments of the present invention set that each sample of a minority category generates at most one augmented sample. At the same time, the embodiments of the present invention only perform data augmentation on categories with a quantity ratio lower than a preset ratio (such as 20%).
[0070] Traditional data augmentation methods (such as randomly inserting or deleting words, etc.) may damage the semantic integrity of the original text, and even introduce noise, affecting the learning effect of the model. Compared with traditional augmentation methods, the embodiments of the present invention can introduce context text related to semantics, enhance the naturalness and coherence of the samples, avoid semantic distortion caused by random replacement or insertion, and can ensure semantic integrity. At the same time, the context text in the embodiments of the present invention provides additional clues, which can help the model more comprehensively understand the sentiment expression, improve the discriminative ability of the model for minority categories, and provide rich semantic information. In addition, by controlling the generation ratio and diversity of the augmented samples, the embodiments of the present invention can help reduce overfitting, improve the performance of the model on the test set, and thus can improve the generalization ability of the model. Further, the augmentation method in the embodiments of the present invention is simple to operate, without additional data collection or complex preprocessing, and only needs to perform a simple splicing operation on the existing data to achieve the purpose of data augmentation.
[0071] Step 40: Determine the initial sentiment analysis model and the target category weights corresponding to different sentiment categories.
[0072] Among them, the initial sentiment analysis model includes a classification layer and an input layer and an encoding layer in a bidirectional encoder pre-trained model (BERT). In addition, the target category weights corresponding to the minority sentiment categories are greater than the target category weights corresponding to other sentiment categories.
[0073] For the embodiments of the present invention, when training the model, in addition to determining the training samples, it is also necessary to build the initial sentiment analysis model to be trained and initialize the model parameters. The overall architecture of this model is as Figure 2As shown, it includes a classification layer, an input layer of a pre-trained BERT model, and an encoding layer. The classification layer includes an attention fusion mechanism and a fully connected layer. The encoding layer includes multiple layers of encoders, and each layer of encoder adopts a multi-head attention mechanism. In the embodiment of the present invention, only the classification layer and the encoding layer are trained during the iterative training of the model, that is, the pre-trained BERT model is fine-tuned.
[0074] Furthermore, in actual application scenarios, there is often a class imbalance phenomenon in the training data, that is, the number of samples in some sentiment categories is significantly less than that in other categories. To solve this problem, the present invention introduces the class weights of different sentiment categories into the cross-entropy loss function, that is, adopts a weighted cross-entropy loss function, assigns higher weights to the minority classes, so as to reduce the dominant role of the majority classes in the overall loss, thereby improving the discrimination performance of the model for minority class samples. In order to fully reflect this mechanism during the optimization process, the embodiment of the present invention takes the class weights as hyperparameters to be tuned, and together with other parameters (such as learning rate, batch size), constitutes a parameter search space.
[0075] During the process of model construction and training, the selection of hyperparameters such as learning rate, batch size, and the above-mentioned class weights is crucial for model performance and robustness. In order to efficiently optimize in the parameter search space, the embodiment of the present invention adopts a grid search method to perform an exhaustive search in a predefined parameter set, and each group of candidate parameter combinations (including value combinations of different learning rates, batch sizes, and class weights) will go through the same training and validation process. By evaluating the performance of all candidate parameter combinations, the parameter configuration with the best average performance on the validation set can be selected from them.
[0076] Based on this, the method includes: constructing multiple groups of candidate parameter combinations based on the learning efficiency and batch size of model training, and the class weights corresponding to the different sentiment categories; adopting a grid search method to perform an exhaustive search in the multiple groups of candidate parameter combinations, and during the search process, using cross-validation to evaluate the performance of each group of candidate parameter combinations to obtain the performance evaluation results of each group of candidates; screening out the target candidate parameter combination according to the performance evaluation results of each group of candidate parameters; and determining the target class weight, target batch size, and target learning efficiency based on the target candidate parameter combination.
[0077] When evaluating the performance of each group of candidate parameter combinations, divide the training samples into multiple mutually exclusive subsets with the same number of samples, where the class ratio of each subset is the same as that of the training samples; for each group of candidate parameter combinations among the multiple groups of candidate parameter combinations, use the first subset among the multiple subsets as the first validation set, and use the other subsets except the first subset as the first training set, and based on the first training set, the first validation set, and each group of candidate parameter combinations, train and validate the initial sentiment analysis model to obtain the model performance index on the first validation set; use the second subset among the multiple subsets as the second validation set, and use the other subsets except the second subset as the second training set, and based on the second training set, the second validation set, and each group of candidate parameter combinations, train and validate the initial sentiment analysis model to obtain the model performance index on the second validation set; repeat the process of replacing the training set and the validation set and the process of training and validating the model until the model performance index with the last subset as the validation set is obtained; calculate the average value of the performance indexes corresponding to each group of candidate parameter combinations based on the multiple model performance indexes corresponding to each group of candidate parameter combinations; determine the performance evaluation result corresponding to each group of candidate parameter combinations according to the average value of the performance indexes.
[0078] Specifically, in order to make full use of limited data resources and ensure the reliability of model evaluation, the embodiments of the present invention adopt stratified K k-fold cross-validation to comprehensively and objectively evaluate the performance of each candidate parameter combination. That is, divide the training samples D according to the class distribution into K mutually exclusive and equal-sized subsets , and ensure that the class ratio of each subset is the same as that of the overall training samples. For each fold k , the training and validation are realized through the following process: First, use the th subset as the validation set , and combine the remaining -1 subsets as the training set . Then, train the model on the training set for the current candidate parameter combination to obtain the model parameters , and then use the trained model parameters to predict the validation set to obtain the predicted probability of each sample , where represents the iThe predicted probabilities that a sample belongs to the corresponding category. Then, record the model performance metrics (such as accuracy, precision, recall, etc.) on the validation set, and accumulate the metrics for this fold to evaluate the stability and overall performance of the candidate parameter combination under different data partitions.
[0079] By repeating the above process for all K folds, the model performance metrics of the current candidate parameter combination on all K folds can be obtained. After averaging these model performance metrics, the comprehensive performance of the current candidate parameter combination can be obtained. By introducing K k-fold cross-validation in the embodiments of the present invention, not only is the representative distribution of the model in terms of class ratios ensured, but also the dependence of the model on a specific data partition is effectively reduced, so that a more reliable and objective performance evaluation can be presented in the results.
[0080] The above process is executed for each group of candidate parameter combinations. After the average values of the performance metrics of all candidate parameter combinations are calculated, the embodiments of the present invention compare and select the candidate parameter combinations based on these average values of the performance metrics, so as to find the parameter configurations that perform excellently on the validation set, including the most suitable class weights (target class weights), learning rates, batch sizes, and other hyperparameters. Once the optimal parameter configuration is determined, the model randomly enters the final training stage.
[0081] The entire training process of the embodiments of the present invention can be divided into two key stages, namely the model optimization stage and the model training stage. Among them, the model optimization stage is the content introduced above, that is, using the K k-fold cross-validation framework to systematically search for and evaluate multiple groups of hyperparameter configurations including class weights. Try each candidate parameter combination one by one through grid search, and record the performance metrics of each group of parameters in each k-fold cross-validation. This stage ensures the rationality of hyperparameter selection and the generalization performance of the model under the conditions of limited model data and class imbalance, and finally obtains the optimal parameter configuration suitable for training on the complete dataset; the model training stage means that once the best hyperparameters are determined, the model will be retrained using the complete dataset to maximize the use of available samples and enhance the prediction ability. At the same time, a weighted cross-entropy loss function is used to maintain robustness to class imbalance. Finally, an early stopping strategy is implemented to stop training when the validation performance no longer improves to prevent performance degradation due to overfitting. The specific content of the model training stage is shown in Steps 50, 60, and 70.
[0082] Step 50: Input the training samples into the initial sentiment analysis model for aspect-based sentiment analysis to obtain the predicted sentiment categories of the training samples.
[0083] Among them, the initial sentiment analysis model includes a classification layer, an input layer, and an encoding layer in a bidirectional encoder pre-trained model.
[0084] For the embodiments of the present invention, in order to effectively represent the semantic relationship between the text and the aspect term, the embodiments of the present invention adopt a pre-trained BERT model for feature extraction. The input of the model includes the global information of the text and the semantic information of the aspect term. For the above process, as Figure 3 shown, it includes:
[0085] Step 51: Display and mark the aspect term in the training sample to obtain the marked training sample.
[0086] Among them, when displaying and marking, special markers are used to wrap the aspect term. It should be noted that the special markers in the embodiments of the present invention can be set according to actual business requirements. For example, the special marker is set as ' <aspect>’ is taken as an example below. Only take ‘ <aspect>’ for illustration purposes, but not limited to ‘ <aspect>’ is the limit.
[0087] In order to enhance the model's attention to aspect words , the embodiments of the present invention adopt a method of displaying and marking aspect words in the text. For example, the aspect words are marked with a special mark ' <aspect>'Wrap to enable the model to clearly identify the position and scope of aspect words and enhance the attention to them. To ensure that ' <aspect>' is not split by the tokenizer of the BERT model. In the embodiments of the present invention, the special token ' <aspect>' is added to the preset vocabulary of the BERT model as an indivisible special token.
[0088] For aspect words consisting of multiple words (such as " battery life "), the embodiments of the present invention adopt the method of holistic marking, that is, special tokens are added before and after the entire aspect word. For example, " battery life " is marked as <aspect> battery life < / aspect> . For the i th training sample , after marking, it can be expressed as .
[0089] For example, the training sample is " the steak was overcooked and dry. ". After the aspect word is visibly marked, is obtained as " the <aspect> steak< / aspect> was overcooked and dry. "
[0090] This way of visible marking in the embodiments of the present invention enables the model to pay more attention to the aspect word during the encoding process, thus helping to more accurately capture aspect-related sentiment information.
[0091] Step 52: Input the marked training sample into the input layer for processing to obtain the embedding vector corresponding to the marked training sample.
[0092] For the embodiments of the present invention, in the input layer, the marked training sample is tokenized to obtain a training sub-word sequence, and a start token and an end token are respectively added to the start position and the end position of the training sub-word sequence to obtain an extended training sub-word sequence; based on the index in the preset vocabulary, the extended training sub-word sequence is converted to obtain a training input index sequence; the word embedding vector, position embedding vector, and segment embedding vector corresponding to the training input index sequence are determined; the word embedding vector, the position embedding vector, and the segment embedding vector are added together to obtain the embedding vector corresponding to the marked training sample;
[0093] Specifically, first, in the input layer, the tokenizer of the BERT model is used to tokenize and encode the marked training sample. For a given training sample , it is decomposed into a training sub-word sequence , where L is the sequence length. In the embodiments of the present invention, the BERT model uses the WordPiece tokenization algorithm to split rare words into known sub-words, thereby improving the word coverage rate. Then, a start token [CLS] and an end token [SEP] are respectively added to the start position and the end position of the sequence to obtain an extended training sub-word sequence , and the expression form is as follows:
[0094]
[0095] Next, the extended training sub-word sequence is converted into indices in a preset vocabulary to generate a training input index sequence. , and the expression form is as follows:
[0096]
[0097] The input of the BERT model consists of word embeddings, position embeddings, and segment embeddings:
[0098]
[0099] Among them, represents the embedding vector corresponding to the tokenized training sample, represents the word embedding vector, represents the position embedding vector, represents the segment embedding vector. In the single-sentence input task, is equal to 0.
[0100] Step 53: Input the embedding vector corresponding to the tokenized training sample into the encoding layer for processing to obtain the hidden state corresponding to the tokenized training sample, and determine the global sample feature vector and the aspect-word sample feature vector based on the hidden state corresponding to the tokenized training sample.
[0101] Among them, the encoding layer includes multiple layers of encoders, and each layer of encoder adopts a multi-head attention mechanism, and a position bias term is introduced in each layer of encoder. It should be noted that the number of layers of the encoder can be set according to actual business requirements, and the embodiments of the present invention do not make specific limitations on this.
[0102] For the embodiments of the present invention, after obtaining the initial representation of the input layer, the encoding layer uses the pre-trained BERT model to perform in-depth feature extraction on the sequence data. Since the embodiments of the present invention display-mark the aspect words in the input text, this has a positive impact on the attention mechanism of the encoding layer.
[0103] When determining the hidden state corresponding to the tagged training sample, the embedding vector corresponding to the tagged training sample is input into the first encoder of the multi-layer encoder for encoding to obtain the output vector of the first encoder; the output vector of the first encoder is input into the second encoder of the multi-layer encoder for encoding to obtain a query matrix and a key matrix, and based on the query matrix, the key matrix, and the position bias term, the attention weights corresponding to the second encoder are calculated, where if the position corresponding to the position bias term belongs to the set of aspect word position indices of the display tag, it is determined that the position bias term is the initial bias parameter; if the position corresponding to the position bias term does not belong to the set of aspect word position indices of the display tag, it is determined that the position bias term is 0; based on the attention weights corresponding to the second encoder, the output vector of the second encoder is determined; continue to input the output vector of the second encoder into the third encoder of the multi-layer encoder for encoding until the last encoder completes encoding, and the output vector of the last encoder is determined as the hidden state corresponding to the tagged training sample.
[0104] Specifically, the encoding layer of BERT consists of multiple layers of transformer encoders, and each layer captures the relationships between the tokens in the sequence through the multi-head attention mechanism. Assume the output of the l -1 layer encoder is , and the calculation process of the l layer encoder is:
[0105]
[0106] In the self-attention mechanism, the calculation formula for the attention weights is:
[0107]
[0108] where Q is the query matrix, Q = , K is the key matrix, K = , and are projection matrices.
[0109] Since the aspect words are display-tagged, the model can more effectively focus on these positions in the attention calculation. To further enhance this effect, the embodiment of the present invention adds a position bias term to the attention score, making the model more inclined to focus on the aspect words of the display tag. Specifically, the modified calculation formula for the attention weights is:
[0110]
[0111] where is the position bias term, and when the position j When belonging to the set of aspect word position indices of the display token, , otherwise . is a learnable bias parameter used to control the preference degree for the display token position.
[0112] In the embodiment of the present invention, by introducing this position bias term into the attention mechanism, the encoding layer can effectively capture the relationship between the aspect word of the display token and the context, and the generated hidden state can better reflect the aspect-related features.
[0113] Extract the feature representation of the aspect word from the hidden state output by the encoding layer. Specifically, locate the display token ' <aspect>The position index set of the " ’ and the wrapped aspect words is obtained, and then the hidden states at these positions are averaged to obtain the feature representation of the aspect words, that is, the aspect word sample feature vector, which is specifically represented as follows:
[0114]
[0115] Among them, represents the hidden state of the BERT encoder at position j ; represents the position index set of the aspect words indicating the display markers, represents the aspect word sample feature vector.
[0116] In addition, the hidden state of the global [CLS] marker is extracted as the global text representation, that is, the global sample feature vector, which is specifically represented as follows:
[0117]
[0118] Among them, represents the hidden state at the position of the [CLS] marker in the i th training sample. j ;
[0119] Step 54: Input the global sample feature vector and the aspect word sample feature vector into the classification layer for sentiment classification to obtain the predicted sentiment category corresponding to the training sample.
[0120] Among them, the classification layer includes an attention fusion mechanism and a fully connected layer.
[0121] For the embodiments of the present invention, after rich feature representations are extracted in the encoding layer, the classification layer is responsible for mapping these features to the sentiment category space. For this process, step 54 specifically includes: calculating the correlation between the global sample feature vector and the aspect word sample feature vector based on the initial weight matrix corresponding to the attention fusion mechanism, as well as the global sample feature vector and the aspect word sample feature vector; based on the correlation, performing feature fusion on the global sample feature vector and the aspect word sample feature vector to obtain a fused feature vector; inputting the fused feature vector into the fully connected layer for sentiment classification to obtain the predicted sentiment category corresponding to the training sample.
[0122] Specifically, the global sample feature vector output by the encoding layer is fused with the aspect word sample feature vector . In order to make full use of the above two parts of information, the embodiments of the present invention adopt an attention-based fusion mechanism, that is, calculating the correlation between the global sample feature vector and the aspect word sample feature vector:
[0123]
[0124] Among them, represents the relevance, is a learnable weight matrix, is the Sigmoid activation function. Then, the generated fused feature vector can be specifically expressed as:
[0125]
[0126] Among them, is the fused feature vector.
[0127] This adaptive fusion method in the embodiments of the present invention enables the model to dynamically adjust the attention to the global text and aspect words according to the characteristics of the input data.
[0128] Next, the classification layer maps the fused feature vector to the sentiment category space through a fully connected layer:
[0129]
[0130] Among them, is the weight matrix, is the bias vector, K is the number of sentiment categories.
[0131] Finally, the Softmax function is used to convert the linear output into a category probability distribution:
[0132]
[0133] Among them, represents the predicted sentiment category corresponding to the training sample.
[0134] Through this design, the classification layer of the embodiments of the present invention not only utilizes the rich features provided by the encoding layer, but also further strengthens the attention to the display markers through the attention mechanism, improving the accuracy of sentiment classification.
[0135] Step 60: Calculate the weighted cross-entropy loss function value based on the predicted sentiment category and the true sentiment category of the training sample, and the target category weights corresponding to different sentiment categories.
[0136] Among them, the target category weights are the optimal parameter configurations selected based on the grid search strategy and hierarchical K-fold cross-validation.
[0137] In order to address the problem of class imbalance during the training process, the embodiments of the present invention introduce a weighted cross-entropy loss row number, which is defined as follows:
[0138]
[0139] Among them, N is the total number of training samples, K is the number of sentiment categories, represents the category k weight, which is a hyperparameter selected through grid search, is the i true sentiment category of the th training sample (represented by one-hot encoding), k is the predicted probability value of the model for the
[0140] In the embodiment of the present invention, by searching and tuning the weight categories the model can better balance the influence of various samples in loss calculation during the training process, so as to show higher robustness and generalization ability when facing imbalanced data.
[0141] Step 70: According to the weighted cross-entropy loss function value, perform iterative training on the initial sentiment analysis model to construct a preset sentiment analysis model.
[0142] For the embodiment of the present invention, after calculating the weighted cross-entropy loss function value, based on this weighted cross-entropy loss function value, perform iterative training on the encoding layer and the classification layer in the initial sentiment analysis model. When the loss function value is the smallest or reaches a preset number of iterations, output the finally trained preset sentiment analysis model.
[0143] After the model is trained, the preset sentiment analysis model can be used to perform sentiment analysis prediction on the text to be predicted, and based on the interpretability component, provide transparent insights into how aspect words and context clues affect the sentiment prediction. Based on this, the method includes: obtaining the text to be predicted; inputting the text to be predicted into the preset sentiment analysis model for sentiment classification to obtain the sentiment category corresponding to the text to be predicted, and visualizing the attention distribution during the model analysis process and the contribution degree of the aspect words marked in the text to be predicted to the sentiment category. Among them, the interpretability component can specifically be an XAI component.
[0144] Specifically, when performing sentiment analysis, first preprocess the text to be predicted to obtain the preprocessed text to be predicted, and display the marked aspect words in the preprocessed text to be predicted to generate the marked text to be predicted. Then use the same BERT tokenizer as in the training stage to tokenize the marked text to be predicted to obtain a subword sequence, and add special tokens [CLS] HE [SEP] at the start and end positions of the subword sequence to obtain an extended subword sequence. Next, convert the extended subword sequence into indices in the vocabulary to generate an input index sequence, and convert the input index sequence into embedding vectors. Further, input the embedding vectors corresponding to the marked text to be predicted into the encoding layer for encoding to obtain hidden states, then extract the global feature vectors, and at the same time average the hidden states at the marked positions to obtain the aspect word feature vectors. Further, use the attention fusion mechanism of the classification layer to fuse the global feature vectors and the aspect word feature vectors to generate fused feature vectors. Finally, map the fused feature vectors to the sentiment category space through a fully connected layer to generate a category probability distribution, and select the category with the highest probability as the final prediction result.
[0145] In the above prediction process, the text to be predicted undergoes the same preprocessing and feature extraction steps as in the training stage, ensuring that the input format is exactly aligned with the model's expectations. By explicitly marking the aspect words, the model can more accurately capture the sentiment information related to specific aspects, thereby effectively improving the accuracy of prediction. Using the model parameters optimized in the training stage, forward propagation is performed on the samples to be predicted, the probability distribution of the sentiment categories is output, and the final prediction result is selected based on the maximum probability principle, enabling accurate aspect sentiment analysis of the text to be predicted.
[0146] The following uses examples to illustrate the sentiment category prediction process and visualization process of the embodiments of the present invention.
[0147] Example 1: The text to be predicted is " For an appetizer, their calamari is a winner. ", and the aspect words are " appetizer " and " calamari ". In the prediction result, for " appetizer " it is neutral, and for " calamari " it is positive.
[0148] In this example, the sentence contains two aspect words " appetizer " (appetizer) and " calamari " (squid). The model can perform sentiment analysis on each aspect word separately through explicit marking <aspect> and< / aspect> , thereby avoiding the confusion of sentiment polarities.
[0149] For the aspect word " calamari The pre - set sentiment analysis model uses display tags to focus on "squid" and related context words. The phrase in the sentence " is a winner " expresses strong positive sentiment, and the model successfully captures this sentiment and associates it with " calamari ", correctly predicting positive sentiment. Figure 4 The image in shows the analysis result, indicating that " winner " has the highest contribution to positive sentiment, which is 0.52, while the contribution of " calamari " is 0.21. Figure 5 The attention weight heatmap of shows that the model's attention is highly concentrated on " calamari " and " winner ", and the display tags further enhance this attention.
[0150] For the aspect word " appetizer ", the sentence only mentions appetizers without directly expressing any sentiment, so the sentiment polarity is classified as neutral. The model accurately identifies "appetizers" through display tags and focuses on its context. Due to the lack of obvious sentiment words, the model finally predicts the aspect word "appetizers" as having neutral sentiment. Figure 6 The analysis result of shows that "appetizers" has the highest neutral contribution. It can be seen from Figure 7 that the model's attention is mainly concentrated on "appetizers" and its display tags, effectively reducing interference from other sentiment words (such as " winner ").
[0151] Example 1 shows that when processing sentences containing multiple aspect words, the model uses display tags to conduct separate sentiment analysis on each word, which can avoid confusion of sentiment polarity. This method enhances the model's attention to specific aspects and improves the accuracy of sentiment classification.
[0152] Example 2: The text to be predicted is " Way too much money for such a terrible meal. ", the aspect word is " meal ", and the prediction result is negative.
[0153] This text to be predicted contains a single aspect word "meal", and the model clearly identifies its position through display tags. Therefore, the word " terrible " obtains a relatively high attention weight, indicating strong negative sentiment. Finally, the model associates this sentiment with "meal" and predicts the sentiment of "meal" as negative. Figure 8 The analysis result of shows that " terrible " has the highest contribution to negative sentiment, which is 0.23. It can be seen from Fig. 9 that the model's attention is also highly concentrated on " terrible " and " meal ", and the display tags enhance the attention to the aspect word.
[0154] Example 2 represents a relatively simple sentiment analysis task involving a single aspect term, and the model performs excellently, verifying its effectiveness in handling a single aspect term.
[0155] Display tagging plays a crucial role in enhancing the model's performance. It can help accurately locate aspect terms, thereby reducing confusion in sentences containing multiple aspects. By tagging these aspects, the model can focus more intensively on relevant terms, thus facilitating better sentiment analysis. In addition, display tagging enables the model to associate the correct polarity with each aspect, minimizing the risk of misclassification.
[0156] Figure 4-Figure 9 Shows how word contributions and attention heatmaps can clarify the model's decision-making process. For positive or negative sentiment, strongly emotional words (such as " winner ” or " terrible ”) and their corresponding aspect tags have higher correlations. In contrast, for neutral sentiment, the model mainly focuses on the aspect term and the surrounding context, avoiding interference from other emotional words. In addition, display tags always obtain higher attention weights, highlighting their importance in guiding aspect-based classification.
[0157] In summary, the preset sentiment analysis model provided by the embodiments of the present invention can not only capture aspect-level sentiment clues with high accuracy, but also improve the transparency of the classification process insights, making it more interpretable and practical in actual applications.
[0158] A method for training a sentiment analysis model provided by the embodiments of the present invention introduces a weighted cross-entropy loss function. The class weights in the weighted cross-entropy loss function can be dynamically adjusted according to the class distribution, that is, the class weights corresponding to the minority sentiment classes are made greater than the class weights corresponding to other sentiment classes. Thereby, it can reduce the role of the majority classes in the overall loss, improve the discriminant performance of the model for minority class samples, and thus effectively reduce the impact of class imbalance in the model training process and improve the sentiment classification accuracy of the model. In addition, by displaying and tagging aspect terms in the input sequence in the embodiments of the present invention, the sentiment analysis model can better capture the complex semantic relationships between aspect words and the context. This design alleviates the problems of long-distance dependence and semantic sparsity, thereby improving the accuracy of sentiment classification. Further, the embodiments of the present invention combine explainable artificial intelligence with the sentiment analysis model, can visualize the attention distribution and sentiment weights, and provide clear insights into the decision-making process. This design method not only increases the transparency of the model, but also provides strong support for users to understand the model behavior and enhances users' trust.
[0159] Further, as Figure 1 and Figure 3 For the specific implementation of the method shown, this embodiment provides an apparatus for training an emotion analysis model, as Fig.10 shown. The apparatus includes: a first determination unit 101, a second determination unit 102, an analysis unit 103, a calculation unit 104, and a training unit 105.
[0160] The first determination unit 101 may be configured to determine target samples belonging to a minority emotion category in the original text set, and the quantity ratio of the target samples in the original text set.
[0161] The first determination unit 101 may also be configured to, if the quantity ratio is less than a preset ratio, determine the semantic relevance between the context text corresponding to the target samples and the target samples.
[0162] The first determination unit 101 may also be configured to, based on the semantic relevance, determine enhanced samples corresponding to the target samples, and use the enhanced samples and the original text set as training samples together. Among them, if it is determined according to the semantic relevance that the context text is semantically relevant to the target samples, enhanced samples are generated based on the context text; if it is determined according to the semantic relevance that the context text is not semantically relevant to the target samples, synonym replacement is performed on the target samples to generate the enhanced samples.
[0163] The second determination unit 102 may be configured to determine an initial emotion analysis model and target category weights corresponding to different emotion categories, where the target category weights corresponding to the minority emotion categories are greater than the target category weights corresponding to other emotion categories.
[0164] The analysis unit 103 may be configured to input the training samples into the initial emotion analysis model to perform aspect-based emotion analysis, and obtain the predicted emotion categories of the training samples.
[0165] The calculation unit 104 may be configured to calculate the value of the weighted cross-entropy loss function based on the predicted emotion categories and the true emotion categories of the training samples, and the target category weights corresponding to different emotion categories.
[0166] The training unit 105 may be configured to perform iterative training on the initial emotion analysis model according to the value of the weighted cross-entropy loss function, and construct a preset emotion analysis model.
[0167] In some embodiments, the first determination unit 101 may be specifically configured to determine the embedding vector corresponding to the context text and the embedding vector corresponding to the target sample; calculate the similarity between the context text and the target sample based on the embedding vector corresponding to the context text and the embedding vector corresponding to the target sample; and determine the semantic relevance between the context text and the target sample according to the similarity.
[0168] The first determination unit 101 may also be specifically configured to, if the similarity is greater than a preset similarity, splice the context text and the target sample to generate an enhanced sample corresponding to the target sample; if the similarity is less than or equal to the preset similarity, perform a synonym replacement on the target sample, and splice the target sample after the synonym replacement and the target sample to generate an enhanced sample corresponding to the target sample.
[0169] In some embodiments, the apparatus further includes: a cleaning unit.
[0170] The cleaning unit may be configured to convert each of the multiple texts in the original text set into a lowercase form to obtain the multiple converted texts; delete the uniform resource locators, social media identifiers, and redundant spaces in the multiple converted texts to obtain the multiple preprocessed texts; and determine the multiple preprocessed texts as the cleaned text set.
[0171] The first determination unit 101 may also be specifically configured to determine the target samples belonging to the minority sentiment categories in the cleaned text set, and the quantity ratio of the target samples in the cleaned text set.
[0172] In some embodiments, the initial sentiment analysis model includes a classification layer and an input layer and an encoding layer in a bidirectional encoder pre-trained model, and the analysis unit 103 includes: a tagging module, an input module, an encoding module, and a classification module.
[0173] The tagging module may be configured to display and tag aspect words in the training samples to obtain the tagged training samples, where special tags are used to wrap the aspect words when displaying and tagging.
[0174] The input module may be configured to input the tagged training samples into the input layer for processing to obtain the embedding vectors corresponding to the tagged training samples.
[0175] The encoding module can be used to input the embedding vector corresponding to the marked training sample into the encoding layer for processing, obtain the hidden state corresponding to the marked training sample, and determine the global sample feature vector and the aspect word sample feature vector based on the hidden state corresponding to the marked training sample.
[0176] The classification module can be used to input the global sample feature vector and the aspect word sample feature vector into the classification layer for sentiment classification, and obtain the predicted sentiment category corresponding to the training sample.
[0177] In some embodiments, the input module can specifically be used to perform word segmentation on the marked training sample in the input layer to obtain a training sub-word sequence, and add a start marker and an end marker to the start position and the end position of the training sub-word sequence respectively to obtain an extended training sub-word sequence; convert the extended training sub-word sequence based on the indices in the preset vocabulary to obtain a training input index sequence; determine the word embedding vector, the position embedding vector, and the segment embedding vector corresponding to the training input index sequence; and add the word embedding vector, the position embedding vector, and the segment embedding vector to obtain the embedding vector corresponding to the marked training sample.
[0178] In some embodiments, the encoding layer includes multiple layers of encoders, and each layer of encoder adopts a multi-head attention mechanism. A position bias term is introduced in each layer of encoder. The encoding module can specifically be used to input the embedding vector corresponding to the marked training sample into the first layer of encoder in the multiple layers of encoders for encoding to obtain the output vector of the first layer of encoder; input the output vector of the first layer of encoder into the second layer of encoder in the multiple layers of encoders for encoding to obtain a query matrix and a key matrix, and calculate the attention weights corresponding to the second layer of encoder based on the query matrix, the key matrix, and the position bias term, where if the position corresponding to the position bias term belongs to the set of aspect word position indices of the display marker, determine that the position bias term is the initial bias parameter; if the position corresponding to the position bias term does not belong to the set of aspect word position indices of the display marker, determine that the position bias term is 0; determine the output vector of the second layer of encoder based on the attention weights corresponding to the second layer of encoder; continue to input the output vector of the second layer of encoder into the third layer of encoder in the multiple layers of encoders for encoding until the last layer of encoder completes encoding, and determine the output vector of the last layer of encoder as the hidden state corresponding to the marked training sample.
[0179] In some embodiments, the classification layer includes an attention fusion mechanism and a fully connected layer. The classification module can be specifically configured to calculate the correlation between the global sample feature vector and the aspect word sample feature vector based on the initial weight matrix corresponding to the attention fusion mechanism, the global sample feature vector, and the aspect word sample feature vector; perform feature fusion on the global sample feature vector and the aspect word sample feature vector based on the correlation to obtain a fused feature vector; input the fused feature vector into the fully connected layer for sentiment classification to obtain the predicted sentiment category corresponding to the training sample.
[0180] In some embodiments, the second determination unit 102 includes a construction module, a search module, a screening module, and a determination module.
[0181] The construction module can be configured to construct multiple groups of candidate parameter combinations based on the learning efficiency and batch size of model training, and the class weights corresponding to the different sentiment categories.
[0182] The search module can be configured to perform an exhaustive search among the multiple groups of candidate parameter combinations in a grid search manner, and during the search process, perform performance evaluation on each group of candidate parameter combinations in a cross-validation manner to obtain the performance evaluation results of each group of candidates.
[0183] The screening module can be configured to screen out the target candidate parameter combination according to the performance evaluation results of each group of candidate parameters.
[0184] The determination module can be configured to determine the target class weight, the target batch size, and the target learning efficiency based on the target candidate parameter combination.
[0185] In some embodiments, the search module may be specifically configured to divide the training samples into multiple mutually exclusive subsets with the same number of samples, where the class ratio of each subset is consistent with the class ratio of the training samples; for each group of candidate parameter combinations in the multiple groups of candidate parameter combinations, use the first subset among the multiple subsets as the first validation set, and the other subsets except the first subset as the first training set, and based on the first training set, the first validation set, and each group of candidate parameter combinations, train and validate the initial sentiment analysis model to obtain the model performance metrics on the first validation set; use the second subset among the multiple subsets as the second validation set, and the other subsets except the second subset as the second training set, and based on the second training set, the second validation set, and each group of candidate parameter combinations, train and validate the initial sentiment analysis model to obtain the model performance metrics on the second validation set; repeat the process of replacing the training set and the validation set and the process of training and validating the model until the model performance metrics of the last subset used as the validation set are obtained; calculate the average value of the performance metrics corresponding to each group of candidate parameter combinations based on the multiple model performance metrics corresponding to each group of candidate parameter combinations; determine the performance evaluation results corresponding to each group of candidate parameter combinations according to the average value of the performance metrics.
[0186] In some embodiments, the preset sentiment analysis model is embedded with an interpretability component, and the analysis unit 103 may be specifically configured to obtain the text to be predicted; input the text to be predicted into the preset sentiment analysis model for sentiment classification to obtain the sentiment category corresponding to the text to be predicted, and visualize the attention distribution during the model analysis process and the contribution degree of the aspect words marked in the text to be predicted to the sentiment category.
[0187] It should be noted that for other corresponding descriptions of each functional unit involved in the sentiment analysis model training device provided in the embodiments of the present invention, reference may be made to Figure 1 and Figure 3 the corresponding descriptions therein, which will not be elaborated here.
[0188] Based on the above methods as Figure 1 and Figure 3 shown, correspondingly, this embodiment also provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the sentiment analysis model training method as Figure 1 and Figure 3 shown.
[0189] Based on such understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.), and includes several instructions for causing an electronic device (such as a personal computer, a server, or a network device, etc.) to execute the methods of various implementation scenarios of the present application.
[0190] Based on the above-mentioned methods as Figure 1 and Figure 3 shown, and Fig.10 the virtual device embodiments shown, in order to achieve the above object, the embodiments of the present application also provide an electronic device, which can specifically be a personal computer, a tablet computer, a server, or other network devices, etc. The device includes a storage medium and a processor; the storage medium is used for storing a computer program; the processor is used for executing the computer program to implement the sentiment analysis model training methods as Figure 1 and Figure 3 shown.
[0191] Optionally, the above-mentioned physical device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, and so on. The user interface may include a display screen (Display), an input unit such as a keyboard (Keyboard), etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), etc.
[0192] Those skilled in the art can understand that the above-mentioned physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine some components, or have different component arrangements.
[0193] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing the hardware and software resources of the above-mentioned physical device, and supports the operation of an information processing program and other software and / or programs. The network communication module is used to implement the communication between the components inside the storage medium, as well as the communication between the storage medium and other hardware and software in the information processing physical device.
[0194] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform, or can also be implemented by hardware.
[0195] The embodiments of the present invention introduce a weighted cross-entropy loss function. The class weights in the weighted cross-entropy loss function can be dynamically adjusted according to the class distribution, that is, the class weights corresponding to the minority sentiment classes are made greater than the class weights corresponding to other sentiment classes. Thereby, the role of the majority classes in the overall loss can be reduced, and the discrimination performance of the model for minority class samples can be improved, so as to effectively mitigate the impact of class imbalance in the model training process and improve the sentiment classification accuracy of the model. In addition, the embodiments of the present invention can enable the sentiment analysis model to better capture the complex semantic relationships between aspect words and the context by explicitly marking aspect terms in the input sequence. This design alleviates the problems of long-distance dependence and semantic sparsity, thereby improving the accuracy of sentiment classification. Further, the embodiments of the present invention combine interpretable artificial intelligence with the sentiment analysis model, can visualize the attention distribution and sentiment weights, and provide clear insights for the decision-making process. This design method not only increases the transparency of the model, but also provides strong support for users to understand the model behavior and enhances users' trust.
[0196] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the drawings are not necessarily essential for implementing the present application. Those skilled in the art can understand that the modules in the device in the implementation scenario can be distributed in the device in the implementation scenario according to the description of the implementation scenario, or can be correspondingly changed and located in one or more devices different from the present implementation scenario. The modules in the above implementation scenario can be combined into one module, or further split into multiple sub-modules.
[0197] The above serial numbers of the present application are only for description and do not represent the advantages or disadvantages of the implementation scenarios. The above disclosure is only several specific implementation scenarios of the present application. However, the present application is not limited thereto, and any changes that can be thought of by those skilled in the art should fall within the protection scope of the present application.< / aspect> < / aspect> < / aspect> < / aspect> < / aspect> < / aspect> < / aspect>
Claims
1. A sentiment analysis model training method, characterized in that: include: Determine target samples belonging to minority sentiment categories in the original text collection, and the quantity ratio of the target samples in the original text collection; If the quantity ratio is less than a preset ratio, determining the semantic relevance between the context text corresponding to the target sample and the target sample; Based on the semantic relevance, an enhanced sample corresponding to the target sample is determined, and the enhanced sample and the original text set are used as training samples together, wherein if it is determined according to the semantic relevance that the context text is semantically related to the target sample, the enhanced sample is generated based on the context text; if it is determined according to the semantic relevance that the context text is semantically irrelevant to the target sample, the target sample is replaced with a synonym to generate the enhanced sample; Determine the initial sentiment analysis model and target category weights corresponding to different sentiment categories, wherein the target category weights corresponding to the minority sentiment categories are greater than the target category weights corresponding to other sentiment categories; Inputting the training sample into the initial sentiment analysis model to perform aspect-based sentiment analysis to obtain a predicted sentiment category of the training sample; Calculating a weighted cross entropy loss function value based on the predicted emotion category and the real emotion category of the training sample and the target category weights corresponding to the different emotion categories; According to the weighted cross entropy loss function value, the initial sentiment analysis model is iteratively trained to construct a preset sentiment analysis model; Wherein, determining the target category weights corresponding to different emotion categories includes: Based on the learning efficiency and batch size of the model training and the category weights corresponding to the different emotion categories, construct multiple sets of candidate parameter combinations; An exhaustive search is performed among the plurality of candidate parameter combinations by using a grid search method, and during the search process, a performance evaluation is performed on each candidate parameter combination by using a cross-validation method to obtain a performance evaluation result of each candidate parameter combination; Screening out a target candidate parameter combination according to the performance evaluation results of each group of candidate parameters; Based on the target candidate parameter combination, the target category weight, target batch size and target learning efficiency are determined.
2. The method according to claim 1, characterized in that The determining of the semantic relevance between the context text and the target sample includes: Determine an embedding vector corresponding to the context text and an embedding vector corresponding to the target sample; Based on the embedding vector corresponding to the context text and the embedding vector corresponding to the target sample, calculating the similarity between the context text and the target sample; Determining the semantic relevance between the context text and the target sample according to the similarity; The determining, based on the semantic relevance, an enhanced sample corresponding to the target sample includes: If the similarity is greater than a preset similarity, the context text is concatenated with the target sample to generate an enhanced sample corresponding to the target sample; If the similarity is less than or equal to the preset similarity, synonym replacement is performed on the target sample, and the target sample after synonym replacement is concatenated with the target sample to generate an enhanced sample corresponding to the target sample.
3. The method according to claim 1, characterized in that Before determining the target samples belonging to the minority sentiment category in the original text set and the quantity ratio of the target samples in the original text set, the method further includes: Converting multiple texts in the original text set into lowercase form to obtain multiple converted texts; Deleting uniform resource locators, social media identifiers, and redundant spaces from the converted multiple texts to obtain multiple preprocessed texts; Determine the preprocessed multiple texts as a cleaned text set; The step of determining target samples belonging to a minority sentiment category in the original text set and the quantity ratio of the target samples in the original text set includes: Determine the target samples belonging to the minority sentiment category in the cleaned text collection and the quantity ratio of the target samples in the cleaned text collection.
4. The method according to claim 1, characterized in that: The initial sentiment analysis model includes a classification layer and an input layer and a coding layer in a bidirectional encoder pre-training model, and the training sample is input into the initial sentiment analysis model to perform aspect-based sentiment analysis to obtain a predicted sentiment category of the training sample, including: Displaying and marking aspect words in the training samples to obtain marked training samples, wherein when displaying the marks, the aspect words are wrapped with special marks; Inputting the labeled training samples into the input layer for processing to obtain an embedding vector corresponding to the labeled training samples; Inputting the embedding vector corresponding to the labeled training sample into the encoding layer for processing, obtaining the hidden state corresponding to the labeled training sample, and determining the global sample feature vector and the aspect word sample feature vector based on the hidden state corresponding to the labeled training sample; The global sample feature vector and the aspect word sample feature vector are input into the classification layer for sentiment classification to obtain the predicted sentiment category corresponding to the training sample.
5. The method according to claim 4, characterized in that Inputting the labeled training samples into the input layer for processing to obtain an embedding vector corresponding to the labeled training samples includes: Performing word segmentation processing on the marked training samples in the input layer to obtain a training subword sequence, and adding a start tag and an end tag to the start position and the end position of the training subword sequence respectively to obtain an expanded training subword sequence; Based on the indexes in the preset vocabulary, converting the expanded training subword sequence to obtain a training input index sequence; Determine a word embedding vector, a position embedding vector, and a segment embedding vector corresponding to the training input index sequence; Adding the word embedding vector, the position embedding vector and the segment embedding vector to obtain the embedding vector corresponding to the labeled training sample; and / or The encoding layer includes a multi-layer encoder, each layer of the encoder adopts a multi-head attention mechanism, and a position bias term is introduced in each layer of the encoder. The embedding vector corresponding to the marked training sample is input into the encoding layer for processing to obtain the hidden state corresponding to the marked training sample, including: Inputting the embedding vector corresponding to the marked training sample into the first layer encoder of the multi-layer encoder for encoding, to obtain the output vector of the first layer encoder; Inputting the output vector of the first layer encoder into the second layer encoder of the multi-layer encoder for encoding, obtaining a query matrix and a key matrix, and calculating the attention weight corresponding to the second layer encoder based on the query matrix, the key matrix and the position bias item, wherein if the position corresponding to the position bias item belongs to the aspect word position index set of the display mark, then determining the position bias item as an initial bias parameter; if the position corresponding to the position bias item does not belong to the aspect word position index set of the display mark, then determining the position bias item to be 0; Determining an output vector of the second layer encoder based on the attention weight corresponding to the second layer encoder; Continue to input the output vector of the second layer encoder into the third layer encoder of the multi-layer encoder for encoding until the last layer encoder completes encoding, and determine the output vector of the last layer encoder as the hidden state corresponding to the marked training sample; and / or The classification layer includes an attention fusion mechanism and a fully connected layer. The global sample feature vector and the aspect word sample feature vector are input into the classification layer for sentiment classification to obtain the predicted sentiment category corresponding to the training sample, including: Based on the initial weight matrix corresponding to the attention fusion mechanism, as well as the global sample feature vector and the aspect word sample feature vector, calculating the correlation between the global sample feature vector and the aspect word sample feature vector; Based on the correlation, the global sample feature vector and the aspect word sample feature vector are subjected to feature fusion to obtain a fused feature vector; The fused feature vector is input into the fully connected layer for sentiment classification to obtain the predicted sentiment category corresponding to the training sample.
6. The method according to claim 1, characterized in that The preset sentiment analysis model is embedded with an explainability component, and the method further includes: Get the text to be predicted; The text to be predicted is input into the preset sentiment analysis model for sentiment classification, the sentiment category corresponding to the text to be predicted is obtained, and the attention distribution in the model analysis process and the contribution of the aspect words marked in the text to be predicted to the sentiment category are visualized.
7. The method according to claim 1, characterized in that The grid search method is used to perform an exhaustive search among the multiple groups of candidate parameter combinations, and during the search process, a cross-validation method is used to perform a performance evaluation on each group of candidate parameter combinations to obtain a performance evaluation result of each group of candidate parameters, including: Dividing the training samples into multiple mutually exclusive subsets with the same number of samples, wherein the category ratio of each subset is consistent with the category ratio of the training samples; For each candidate parameter combination in the multiple candidate parameter combinations, use the first subset in the multiple subsets as a first validation set, and use the other subsets except the first subset as a first training set, and train and validate the initial sentiment analysis model based on the first training set and the first validation set, as well as each candidate parameter combination, to obtain a model performance indicator on the first validation set; Using a second subset of the multiple subsets as a second validation set, and using the other subsets except the second subset as second training sets, and training and validating the initial sentiment analysis model based on the second training set and the second validation set, and each set of candidate parameter combinations, to obtain a model performance indicator on the second validation set; Repeat the process of replacing the training set and the validation set and the model training and validation process until the last subset is obtained as the model performance indicator of the validation set; Based on multiple model performance indicators corresponding to each set of candidate parameter combinations, an average performance indicator corresponding to each set of candidate parameter combinations is calculated; The performance evaluation result corresponding to each group of candidate parameter combinations is determined according to the performance indicator average value.
8. A sentiment analysis model training device, characterized in that: include: A first determination unit is used to determine target samples belonging to a minority sentiment category in an original text set, and a quantity ratio of the target samples in the original text set; The first determining unit is further configured to determine the semantic relevance between the context text corresponding to the target sample and the target sample if the quantity ratio is less than a preset ratio; The first determination unit is further used to determine an enhanced sample corresponding to the target sample based on the semantic relevance, and use the enhanced sample and the original text set as training samples, wherein if the context text is determined to be semantically relevant to the target sample according to the semantic relevance, the enhanced sample is generated based on the context text; if the context text is determined to be semantically irrelevant to the target sample according to the semantic relevance, the target sample is replaced with a synonym to generate the enhanced sample; A second determination unit is used to determine the initial sentiment analysis model and target category weights corresponding to different sentiment categories, wherein the target category weights corresponding to the minority sentiment categories are greater than the target category weights corresponding to other sentiment categories; An analysis unit, configured to input the training sample into the initial sentiment analysis model to perform aspect-based sentiment analysis to obtain a predicted sentiment category of the training sample; A calculation unit, configured to calculate a weighted cross entropy loss function value based on the predicted emotion category and the real emotion category of the training sample, and the target category weights corresponding to the different emotion categories; A training unit, used for iteratively training the initial sentiment analysis model according to the weighted cross entropy loss function value to construct a preset sentiment analysis model; The second determination unit is specifically used to construct multiple groups of candidate parameter combinations based on the learning efficiency and batch size of model training, and the category weights corresponding to the different emotion categories; use a grid search method to perform an exhaustive search in the multiple groups of candidate parameter combinations, and during the search process, use a cross-validation method to perform performance evaluation on each group of candidate parameter combinations to obtain a performance evaluation result of each group of candidate parameters; according to the performance evaluation result of each group of candidate parameters, screen out a target candidate parameter combination; based on the target candidate parameter combination, determine the target category weight, target batch size and target learning efficiency.
9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. An electronic device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Emotion classification method based on unbalanced data
CN114881175A
Multi-label sentiment classification method and system based on non-autoregression model
CN118113871A