Interactive text fine-grained sentiment analysis method assisted by large language model
By integrating large language model, CNN and Transformer technology in natural language processing technology, combining dynamic weight adjustment and emotional labeling system construction, fine-grained emotion recognition and data quality problems are solved, and efficient sentiment analysis and adaptive processing of network interactive texts are realized.
Patent Information
- Application Number
- CN202411867747.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art is difficult to learn robust emotion representation from massive heterogeneous data, realize fine-grained emotion recognition, while ensuring data quality and model adaptability, especially when facing cross-domain adaptability, emotional granularity division and data quality problems.
The fine-grained emotional analysis method of interactive text assisted by large language models is adopted, and large language models, convolutional neural networks (CNNs) and Transformer technologies are integrated to realize adaptive processing of texts of different complexity through dynamic weight adjustment, and an emotional label system is built, emotional space coding, and design and judgment models to improve data quality and model flexibility.
It significantly improves the fine-grained emotional recognition capabilities of online interactive texts, ensures data quality and model adaptability, can more accurately reflect the direction of public opinion, and improves government governance capabilities.
Smart Images

Figure CN120087369A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a method for fine-grained sentiment analysis of interactive text assisted by a large language model. Background Art
[0002] With the rapid development of Internet technology and the extensive construction of government online interactive platforms, massive amounts of public appeals and government response text data are accumulating. Developing fine-grained sentiment analysis methods suitable for online interactive texts is of great significance for comprehensively grasping the trend of public opinion and improving government governance capabilities. However, fine-grained sentiment analysis faces three key challenges: cross-domain adaptability, sentiment granularity division, and data quality issues. First, texts in different fields have significant differences in language style, topic focus, and emotional expression, leading to domain adaptation problems. Online interactive texts involve multiple fields and topics, making it difficult to directly apply traditional single-field sentiment analysis methods. Second, public appeals may contain multiple emotions at the same time, and it is necessary to consider the diversity and coexistence of emotions. Traditional single-label sentiment classification methods, such as those based on sentiment dictionaries and machine learning, often perform significantly worse when dealing with cross-domain data and complex emotions. How to learn robust sentiment representation from massive heterogeneous data and achieve more fine-grained sentiment recognition is a difficult problem that needs to be solved urgently. Finally, data quality issues are also an important challenge. With the widespread application of large language models in the field of natural language processing, their potential in data annotation tasks has been widely recognized. However, large models may produce hallucinations or inaccuracies when generating annotation results, which has a serious impact on subsequent sentiment analysis tasks. The inability to effectively utilize the advantages of large models while overcoming their limitations to obtain high-quality annotated datasets is a key issue in current research. In addition, existing sentiment analysis models often have difficulty adapting to text inputs of different lengths and complexities. The length and complexity of online interactive texts vary greatly, ranging from a few short sentences to long and detailed appeals;
[0003] In summary, how to learn robust emotion representation from massive heterogeneous data and achieve more fine-grained emotion recognition while ensuring data quality and model adaptability is a difficult problem that needs to be solved urgently. This paper aims to solve these problems and proposes a fine-grained emotion recognition method for network interactive text that integrates a large language model and a dynamic adaptive architecture. Summary of the invention
[0004] In view of the deficiencies in the prior art, the present invention provides a fine-grained sentiment analysis method for interactive text assisted by a large language model, which solves the problems raised in the background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions: A fine-grained sentiment analysis method for interactive text assisted by a large language model, which integrates a large language model, a convolutional neural network (CNN) and Transformer technology, and realizes adaptive processing of texts with different complexities through dynamic weight adjustment, including the following steps:
[0006] Step 1: High cross-domain adaptability and optimized sentiment granularity division;
[0007] Step 2: Improve data quality and enhance model flexibility;
[0008] Step 3: Solve data imbalance;
[0009] Step 4: Construction of an emotion label system;
[0010] Step 5: Emotion space encoding;
[0011] Step 6: Emotion annotation experiment;
[0012] Step 7: Research and judgment model design and implementation.
[0013] The construction of the emotion label system: According to the arousal-valence theory of emotion, construct an emotion label system based on the arousal-valence theory of emotion. This step involves in-depth research on the theory of emotion psychology and combines the characteristics of the network interaction field to design an emotion classification system that has both theoretical basis and practical feasibility. The system not only considers the positivity and negativity of emotions, but also includes the dimension of emotion intensity, making emotion classification more refined and comprehensive.
[0014] Preferably, for the emotion space encoding, an 8-dimensional multi-label vector space {0,1}^8 is used to encode the emotion space. The encoding method allows for a more precise representation of complex emotion states. Each dimension represents a basic emotion category, and through the combination of 0 and 1, rich emotion states can be represented. The multi-label method is superior to the traditional single-label classification and can capture multiple emotions that may coexist in the text.
[0015] Preferably, for the emotion annotation experiment, two large model experiments are designed for emotion annotation. The first large model conducts preliminary emotion annotation, and another large model with better performance is used to review and correct the results annotated by the first large model, mapping the emotions reflected in the original text into a 56-dimensional target emotion space.
[0016] Preferably, for the design and implementation of the research and judgment model, 40% of the data is randomly selected from the labeled dataset. The expert team conducts manual review on the data that has been reviewed and corrected by the Claude 3.5 large model, and performs binary classification labeling for each piece of data to determine whether each piece of data generates hallucinations. 1 indicates that the data still generates hallucinations after correction, which is regarded as mislabeling, and the data after correction is regarded as the correct labeling result;
[0017] Use the binary classification labeled dataset as the training dataset to train the CNN+transformer model. The CNN+transformer model is an adaptive architecture with a weight ratio dynamically generated according to the text complexity;
[0018] The evaluation indicators for the effect of the research and judgment model include: Precision, Accuracy, and F1-score. The trained research and judgment model is applied to the remaining 60% of the labeled dataset, and finally a high-quality dataset labeled by the large model without the influence of hallucinations is obtained;
[0019] The research and judgment model adopts an adaptive architecture that dynamically fuses CNN and Transformer, and dynamically adjusts the proportion of CNN and Transformer according to the complexity of the input text. The specific implementation method is as follows:
[0020] a) Design a length perception module to evaluate the complexity of the input text;
[0021] b) Use the output of this module to dynamically adjust the number or weight of the CNN and Transformer layers;
[0022] c) Design a fusion module, such as a gating mechanism or an attention mechanism, to combine the outputs of the two architectures;
[0023] d) Introduce a meta-controller to learn the optimal architecture combination strategy during the training process;
[0024] Text complexity evaluation process:
[0025] Path 1: Evaluate according to the number of characters:
[0026] Count the number of characters in all network interaction request records to obtain the minimum value a1, the quartile b1, the average c1, the three-quarter quantile d1, and the maximum value e1. According to the interval where the request character count falls, different complexity values complex_value1 (0.25, 0.50, 0.75, 1.00) are assigned;
[0027] Path 2: Evaluate according to emotional words:
[0028] Count the number of emotion words in all records of network interaction demands, and obtain the minimum value a2, quartiles b2, mean c2, three - quarter quantile d2, and maximum value e2;
[0029] According to the interval in which the number of demand emotion words falls, assign different complexity values complex_value2(0.25, 0.50, 0.75, 1.00).
[0030] Preferably, for the data balancing project, use a large model with better performance to generate data, achieve balanced annotation results for various categories, solve the data imbalance problem, increase the number of minority - class samples through intelligent generation, and make the number of samples in each emotion category tend to be the same. This method can not only improve the overall performance of the model but also enhance the model's recognition ability for minority classes.
[0031] The training and evaluation of the emotion recognition model:
[0032] a) Select multiple model architectures for comparative experiments, including RoBERTa, BiLSTM, CNN, and their combined models;
[0033] b) Use indicators such as precision, accuracy, and F1 - value to comprehensively evaluate the performance of each model. By comparing the performance of different models on various indicators, select the most suitable model architecture for network interaction text sentiment analysis.
[0034] Preferably, for the dataset source and pre - processing, in the data pre - processing stage, the following processing was performed on the original dataset:
[0035] In the data pre - processing stage, the following processing was performed on the original dataset: (1) Data cleaning: Perform basic cleaning operations such as deduplication, removal of special characters, and conversion of letters to lowercase on the text data to improve data quality;
[0036] (2) Text encoding: Use the pre - trained word embedding model RoBERTa to directly encode each text sample into a fixed - length sentence vector as the input feature of the model;
[0037] (3) Data partitioning: For each dataset, randomly partition 80% of the samples as the training set and 20% as the test set for model training and evaluation. At the same time, to ensure the reliability and reproducibility of the experimental results, this paper fixed the random seed during data partitioning, ensuring a fair comparison of different methods on the same dataset. Through the above pre - processing steps, the original dataset was converted into a format suitable for model training and optimization, laying a data foundation for subsequent label annotation and model construction.
[0038] Preferably, through data balancing processing: 1) Class distribution balance: The Gini coefficient is used to measure the balance degree of class distribution. The Gini coefficient before balancing is 0.866, and it is reduced to 0.181 after balancing, indicating that the class distribution significantly tends to be balanced;
[0039] 2) Generated sample quality: The newly generated samples are evaluated using a judgment model. The results show that 95.4% of the generated samples are considered to conform to the target sentiment class and express naturally. This proves that the large language model has excellent capabilities in generating high-quality text that conforms to specific sentiments;
[0040] 3) Semantic diversity: The cosine similarity is used to evaluate the semantic differences between the generated samples and the original samples. The results show that the average cosine similarity between the generated samples and the original samples is 0.871, indicating that the generated content not only maintains the relevance with the original data but also introduces new expression methods, increasing the diversity of the dataset;
[0041] 4) Impact on model performance: In this study, a basic RoBERTa classification model was trained using the datasets before and after balancing respectively, and evaluated on the same test set.
[0042] The utility model provides a fireproof optical fiber cable:
[0043] The present invention will significantly improve the fine-grained sentiment recognition ability of network interaction texts, which can not only provide more accurate and comprehensive public opinion information support for government decision-making, but also promote the improvement of government governance capabilities and the effective resolution of public demands. The implementation of the present invention will have an important impact on improving the quality of government public services, enhancing the effect of government-citizen interaction, and promoting social harmony and stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 Schematic diagrams of five emotion recognition models of the present invention;
[0045] Figure 2 Schematic diagram of the research idea framework of the present invention;
[0046] Figure 3 Flowchart of the method for fine-grained sentiment recognition of network interaction texts integrating large language models of the present invention;
[0047] Figure 4 Text fine-grained sentiment annotation framework based on multi-dimensional sentiment space and large language model of the present invention;
[0048] Figure 5 Schematic diagram of the process of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0050] The terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / " generally means an "or" relationship between the associated objects before and after.
[0051] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.
[0052] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "installed", "connected", "set" should be understood in a broad sense. For example, it can be fixedly connected and set, or detachably connected and set, or integrally connected and set. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0053] As Figures 1-5 shown, the network interaction text fine-grained sentiment recognition method integrating large language models proposed by the present invention achieves remarkable technical effects through a series of innovative technologies:
[0054] The large language model-assisted interactive text fine-grained sentiment analysis method integrates large language models, convolutional neural networks (CNNs), and Transformer technologies, and realizes adaptive processing of texts with different complexities through dynamic weight adjustment, including the following steps:
[0055] Step 1: High cross-domain adaptability and optimized sentiment granularity division;
[0056] Step 2: Improve data quality and enhance model flexibility;
[0057] Step 3: Solve data imbalance;
[0058] Step 4: Construction of the emotion label system;
[0059] Step 5: Emotion space encoding;
[0060] Step 6: Emotion annotation experiment;
[0061] Step 7: Design and implementation of the judgment model.
[0062] Construction of the emotion label system: According to the arousal - valence theory, construct an emotion label system based on the arousal - valence theory. This step involves in - depth research on emotion psychology theory and combines the characteristics of the online interaction field to design an emotion classification system that has both theoretical basis and practical feasibility. The system not only considers the positivity and negativity of emotions but also includes the dimension of emotion intensity, making emotion classification more refined and comprehensive.
[0063] Emotion space encoding: Use an 8 - dimensional multi - label vector space {0, 1}^8 to encode the emotion space. The encoding method allows for a more precise representation of complex emotion states. Each dimension represents a basic emotion category, and through the combination of 0 and 1, rich emotion states can be represented. The multi - label method is superior to the traditional single - label classification and can capture multiple emotions that may co - exist in the text.
[0064] Emotion annotation experiment: Design two large - model experiments for emotion annotation. The first large - model conducts preliminary emotion annotation, and another large - model with better performance reviews and corrects the results annotated by the first large - model, mapping the emotions reflected in the original text into a 56 - dimensional target emotion space.
[0065] Design and implementation of the judgment model: Randomly extract 40% of the data from the annotated dataset. The expert team conducts manual review on the data that has been reviewed and corrected by the claude3.5 large - model, and conducts binary - classification annotation for each piece of data to determine whether each piece of data generates hallucinations. 1 indicates that the data still generates hallucinations after correction and is regarded as mislabeled, and 0 indicates that the data after correction is the correct annotation result;
[0066] Use the binary - classification annotated dataset as the training dataset to train the CNN + transformer model. The CNN + transformer model is an adaptive architecture with weight ratios dynamically generated according to text complexity;
[0067] The evaluation metrics for the judgment model include: Precision, Accuracy, and F1-score. The trained judgment model is applied to the remaining 60% of the labeled dataset, and finally a high-quality dataset labeled by the large model without the influence of hallucinations is obtained.
[0068] The judgment model adopts an adaptive architecture that dynamically fuses CNN and Transformer, and dynamically adjusts the proportion of CNN and Transformer according to the complexity of the input text. The specific implementation method is as follows:
[0069] a) Design a length perception module to evaluate the complexity of the input text;
[0070] b) Use the output of this module to dynamically adjust the number or weights of the CNN and Transformer layers;
[0071] c) Design a fusion module, such as a gating mechanism or an attention mechanism, to combine the outputs of the two architectures;
[0072] d) Introduce a meta-controller to learn the optimal architecture combination strategy during the training process;
[0073] The process of text complexity evaluation:
[0074] Path 1: Evaluate according to the number of characters:
[0075] Count the number of characters in all network interaction demand records to obtain the minimum value a1, the quartile b1, the average value c1, the three-quarter digit d1, and the maximum value e1. According to the interval where the number of demand characters falls, assign different complexity values complex_value1 (0.25, 0.50, 0.75, 1.00);
[0076] Path 2: Evaluate according to emotional words:
[0077] Count the number of emotional words in all network interaction demand records to obtain the minimum value a2, the quartile b2, the average value c2, the three-quarter digit d2, and the maximum value e2;
[0078] According to the interval where the number of demand emotional words falls, assign different complexity values complex_value2 (0.25, 0.50, 0.75, 1.00);
[0079] Data balancing project. Use a large model with better performance to generate data to achieve balanced annotation results for each category. In solving the data imbalance problem, increase the number of minority class samples through intelligent generation, so that the number of samples for each emotional category tends to be the same. This method can not only improve the overall performance of the model, but also enhance the model's recognition ability for minority classes.
[0080] Emotion Recognition Model Training and Evaluation:
[0081] a) Select multiple model architectures for comparative experiments, including RoBERTa, BiLSTM, CNN, and their combined models;
[0082] b) Use metrics such as precision, accuracy, and F1-score to comprehensively evaluate the performance of each model. By comparing the performance of different models on various metrics, select the most suitable model architecture for network interaction text sentiment analysis.
[0083] Dataset Source and Preprocessing. In the data preprocessing stage, the research performed the following processing on the original dataset:
[0084] In the data preprocessing stage, this paper performed the following processing on the original dataset: (1) Data cleaning: Perform basic cleaning operations on the text data, such as removing duplicates, special characters, and converting letters to lowercase, to improve data quality;
[0085] (3) Text encoding: Use the pre-trained word embedding model RoBERTa to directly encode each text sample into a fixed-length sentence vector as the input feature of the model;
[0086] (3) Data partitioning: For each dataset, randomly partition 80% of the samples as the training set and 20% as the test set for model training and evaluation. At the same time, to ensure the reliability and reproducibility of the experimental results, this paper fixed the random seed during data partitioning, ensuring a fair comparison of different methods on the same dataset. Through the above preprocessing steps, the original dataset was transformed into a format suitable for model training and optimization, laying a data foundation for subsequent label annotation and model construction.
[0087] Through data balancing processing, 1) Class distribution balance: Use the Gini coefficient to measure the balance degree of class distribution. The Gini coefficient before balancing was 0.866, and it decreased to 0.181 after balancing, indicating that the class distribution significantly tended to be balanced;
[0088] 2) Generated sample quality: Evaluate the newly generated samples using a judgment model. The results show that 95.4% of the generated samples were considered to conform to the target sentiment category and express naturally. This proves that the large language model has excellent capabilities in generating high-quality text that conforms to specific emotions;
[0089] 3) Semantic diversity: Use cosine similarity to evaluate the semantic difference between the generated samples and the original samples. The results show that the average cosine similarity between the generated samples and the original samples is 0.871, indicating that the generated content not only maintains the relevance to the original data but also introduces new expressions, increasing the diversity of the dataset;
[0090] 4) Impact on model performance: In this study, a basic RoBERTa classification model was trained using the datasets before and after balancing, and evaluated on the same test set.
[0091] Dataset source and preprocessing:
[0092] Two representative online interaction platforms, the website of the Shenzhen Municipal People's Government and the Shenzhen section of the People's Daily Online Message Board, were selected as the data sources. First, these two platforms represent the voices of the government official and the mainstream media respectively. The data sources are authoritative and reliable, and can truly reflect the actual situation of government-people interaction. Second, both platforms have a large amount of government-people interaction data, covering a wide range of public demands, providing rich text materials for carrying out fine-grained sentiment analysis. Moreover, as a representative of China's special economic zones and innovative cities, Shenzhen has typical significance in urban governance and has sufficient data on online interaction platforms, facilitating large-scale computational analysis. Table 3 presents the basic information statistics of these two data sources.
[0093] Table 3 Basic Information of the Sentiment Domain Dataset
[0094]
[0095]
[0096] As can be seen from Table 1, the dataset of the Leadership Letter Box module of the Shenzhen Municipal People's Government is sourced from the official government website, representing a relatively formal and authoritative government-people interaction channel, with a data volume of 16,196 items; while the dataset of the Shenzhen section of the People's Daily Online Message Board comes from the mainstream media platform, covering a wider range of public demands, with a data volume of 13,217 items. The time spans of these two datasets are the same, both from December 31, 2022, to January 1, 2024. In terms of the granularity of sentiment label division, both datasets are divided into 8 intervals, corresponding to 56 emotion labels. Under the same sentiment granularity, datasets from different sources can more fairly compare the performance of the proposed method, excluding the influence caused by differences in sentiment labels.
[0097] In the data preprocessing stage, the research processed the original dataset as follows: In the data preprocessing stage, this paper processed the original dataset as follows: (1) Data cleaning: Basic cleaning operations such as deduplication, removal of special characters, and conversion of letters to lowercase were performed on the text data to improve data quality. (2) Text encoding: The pre-trained word embedding model RoBERTa was used to directly encode each text sample into a fixed-length sentence vector as the input feature of the model. (3) Data partitioning: For each dataset, 80% of the samples were randomly selected as the training set and 20% as the test set for model training and evaluation. At the same time, to ensure the reliability and reproducibility of the experimental results, this paper fixed the random seed during data partitioning, ensuring a fair comparison of different methods on the same dataset. Through the above preprocessing steps, the original dataset was converted into a format suitable for model training and optimization, laying a data foundation for subsequent label annotation and model construction.
[0098] Experiment on Fine-Grained Multi-Task Sentiment Annotation of Network Interaction Texts:
[0099] For single-label and multi-label classification tasks, different experimental schemes were designed, and the optimal annotation optimization path was selected through comparative analysis to obtain high-quality fine-grained sentiment annotation results for the network interaction dataset.
[0100] In the single-label classification task, we compared and tested two annotation optimization paths on two datasets, namely the Shenzhen Municipal People's Government - Leadership Letter Box module and the People's Daily Online Message Board - Shenzhen module: using the large model alone and the combination of large model + large model.
[0101] Table 4 Comparison of the Effects of Sentiment Annotation Experimental Paths in Single-Label Classification Tasks
[0102]
[0103]
[0104] As shown in Table 4, the combination of large model + large model performs outstandingly in terms of model performance, and the Precision, Accuracy, and F1-score on the leadership letter box and message board modules are all better than those of the large model combination. Taking the leadership letter box module as an example, the Precision, Accuracy, and F1-score after the dual large models all reach 0.99, which is 0.14, 0.31, and 0.24 higher than those of the large model. This indicates that using rule correction can more accurately perform fine-grained division of sentiment labels. In addition, the combined model also has better generalization ability than using the large model alone. Considering the above factors, the combination path of large model + large model is a better choice.
[0105] Table 5 Comparison of the Effects of Sentiment Annotation Paths in Multi-Label Classification Tasks
[0106]
[0107] In the multi-label classification task, as shown in Table 5. We also compared two annotation optimization paths. Similar to the single-label task, the large model + large model experiment has great advantages in model performance.
[0108] In the single-label classification task, in the leadership mailbox module, in Experiment 2 compared with Experiment 1, the emotional range decreased by 1, and the emotion labels decreased by 19. For example, "anxious", "fidgety", etc. were aligned with "anxiety"; in the People's Daily Online message board module, in Experiment 2 compared with Experiment 1, the emotional range remained unchanged, and the emotion labels decreased by 19. In the multi-label classification task, in the leadership mailbox module, in Experiment 2 compared with Experiment 1, the emotional range remained unchanged, and the emotion labels decreased by 20; in the People's Daily Online message board module, in Experiment 2 compared with Experiment 1, the emotional range decreased by 1, and the emotion labels decreased by 29. From the experimental results, whether in the single-label classification task or the multi-label classification task, the method of integrating large language models and rule correction can effectively reduce the emotional range and the number of emotion labels, more accurately perform fine-grained emotion annotation on network interaction texts, reduce unnecessary label redundancy, improve the consistency and accuracy of annotation, and such results are also beneficial to the execution of subsequent emotion analysis tasks. At the same time, the experimental results also show that although large models have advantages in common sense reasoning, knowledge supplementation, etc., there are still some limitations. For example, emotion labels that are disconnected from the emotion label system will be generated. This indicates that in actual emotion analysis tasks, it is still necessary to combine professional emotion theories and artificial rules for correction and optimization. Through detailed experiments and analyses in this section, the effectiveness and superiority of the combined path of large model + rule correction in the fine-grained emotion annotation task of the network interaction dataset are verified. While ensuring a relatively high annotation quality, this path demonstrates better generalization ability. This lays a solid foundation for the subsequent construction of a high-quality network interaction emotion analysis dataset.
[0109] Experiment on the emotion recognition model of the high-quality annotation dataset:
[0110] The performance of five mainstream models in multi-classification and multi-label classification tasks was compared and tested.
[0111] Performance comparison of emotion recognition models in multi-classification tasks:
[0112] On the dataset after emotional range division and emotion annotation, the performance of five emotion recognition models in multi-classification tasks was evaluated. Table 7 summarizes the performance of each model on the dataset and gives the most matching model for different datasets.
[0113] Table 7 Performance comparison of emotion recognition models in multi-classification tasks
[0114]
[0115] In the single-label classification task, the RoBERTa model performs best on the leadership mailbox module, with the Accuracy, Precision, and F1 values reaching 57.76%, 54.03%, and 54.65% respectively; the RoBERTa+CNN model performs best on the People's Daily Online message board module, with the Accuracy, Precision, and F1 values reaching 51.68%, 42.19%, and 43.08% respectively, and also has a good performance on the leadership mailbox module. From the results in Table 8, it can be seen that the RoBERTa+CNN model uses RoBERTa to extract high-quality semantic representations, and then performs local feature extraction and classification through CNN, which can maximize the mining of the emotional semantics of the text, and achieves a good emotion recognition effect by combining the semantic representation ability of the pre-trained language model and the feature extraction ability of the convolutional neural network.
[0116] The above single-label multi-classification tasks are all emotion classifications with 32 categories. The classifier needs to handle a lot of details, resulting in an increase in the complexity of the model. The text may contain multiple emotion labels at the same time, and the label with the highest possibility needs to be selected, which increases the difficulty and complexity of classification. In this case, an accuracy of about 0.6 already indicates that the model has a strong recognition ability, and the Precision and F1-score metrics also illustrate the comprehensive performance of the model in balancing precision and recall. Compared with the multi-class emotion recognition model (28 emotion categories, the model download volume is 2,103,593, and the accuracy is 0.45) with high reference and high download on Huggingface currently, the emotion recognition model in this study has achieved better results. Performance comparison of emotion recognition models for multi-label classification tasks
[0117] The performance of five emotion recognition models was evaluated for multi-classification tasks. Table 8 summarizes the performance of each model on the dataset and gives the most matching model for different datasets.
[0118] Table 8 Performance comparison of emotion recognition models for multi-label classification tasks
[0119]
[0120]
[0121] In the multi-label classification task, the RoBERTa+BiLSTM model performs excellently. Especially in the People's Daily Message Board module, the Accuracy of the model reaches 0.4233, significantly higher than other models. Its Precision and F1-score are 0.6301 and 0.6321 respectively, which also shows strong performance. The advantage of the RoBERTa+BiLSTM model lies in combining the powerful semantic representation ability of the RoBERTa model and the ability of BiLSTM to capture sequence information. This combination enables the model to better understand the context and emotional features of the text, thus performing well in complex multi-label sentiment recognition tasks. The number of binary label classifications in the two sections is 45 and 40 respectively, and the number of categories is large. Each text contains multiple labels simultaneously, further increasing the complexity of the task. Considering the diversity of emotional labels and the imbalance of the dataset, although the model accuracy is not high, it is already a good result in such a difficult task.
[0122] Generally speaking, the RoBERTa+BiLSTM model demonstrates strong overall performance, especially performing well in multi-label classification tasks. Although the RoBERTa+BiLSTM model is not the best in some tasks, its performance often approaches that of the optimal model, demonstrating its advantages in complex tasks. RoBERTa+CNN performs well in single-label classification tasks and can effectively capture local features, while RoBERTa performs excellently on specific datasets (Leadership Letter Box module). When conducting sentiment analysis, the selection of an appropriate model needs to consider the specific characteristics of the task and the characteristics of the dataset. RoBERTa+BiLSTM performs excellently in dealing with complex emotional relationships and multi-label tasks, RoBERTa+CNN has more advantages in single-label tasks, and RoBERTa can also demonstrate powerful performance in specific tasks. This comprehensive analysis helps to better select and apply models to improve the accuracy and effectiveness of sentiment analysis.
[0123] Step 1: Construct an emotional label system based on the arousal-valence theory:
[0124] To achieve cross-domain emotional representation, a unified emotional label system is constructed using the arousal-valence theory.
[0125] Specifically, the classic two-dimensional emotional space model proposed by Russell is adopted, and emotions are divided into the following 8 intervals:
[0126] High-arousal positive emotion area: Corresponding to emotional states with high arousal and high positive valence, the emotion labels are excitement, agitation, joy, ecstasy, fascination, surprise, and enthusiasm.
[0127] High and medium arousal positive emotion area: corresponding to emotional states of high and medium arousal, and high and medium positive valence, with emotional labels of pleasure, happiness, joy, satisfaction, pride, optimism, and cheerfulness.
[0128] Medium arousal positive emotion area: corresponding to emotional states of medium arousal and low positive valence, with emotional labels of geniality, relaxation, contentment, comfort, confidence, expectation, and coziness.
[0129] Low arousal positive emotion area: corresponding to emotional states of low arousal and slightly positive valence, with emotional labels of peace of mind, leisurely, calm, serene, contented, warm, and tranquil.
[0130] High arousal negative emotion area: corresponding to emotional states of high arousal and high negative valence, with emotional labels of anxiety, panic, anger, despair, breakdown, shock, and contempt.
[0131] High and medium arousal negative emotion area: corresponding to emotional states of high and medium arousal and high and medium negative valence, with emotional labels of disgust, frustration, sadness, annoyance, irritability, pessimism, and fear.
[0132] Medium arousal negative emotion area: corresponding to emotional states of medium arousal and low negative valence, with emotional labels of depression, disappointment, gloom, worry, tension, confusion, and helplessness.
[0133] Low arousal negative emotion area: corresponding to emotional states of low arousal and slightly negative valence, with emotional labels of tiredness, loneliness, burnout, apathy, numbness, depression, and dullness.
[0134] Step 2: Use a large model to achieve emotion annotation.
[0135] Select GPT-4 as the large model for emotion classification. Utilize its advantages in common sense reasoning, knowledge completion, etc., and conduct emotion partitioning on the dataset based on its BROKE framework. Randomly select 40% of the already annotated dataset, and use Claude3.5sonnet to review and correct the annotation results of GPT-4 to obtain the final emotion annotation results of the large model.
[0136] Step 3: Introduce a judgment model to obtain high-quality emotion annotation results.
[0137] The expert team re-annotates the authenticity of 40% of the dataset in step 2, giving a binary classification label: 1 indicates that the result annotated by the large model has hallucinations and is incorrect, and 0 indicates that the result annotated by the large model is correct. Use the binary classification label data to train the judgment model. The judgment model is a binary classification model that adopts an adaptive architecture that dynamically fuses CNN and Transformer. This model dynamically adjusts the proportion of CNN and Transformer according to the input length to adapt to text inputs of different lengths. Use the trained judgment model to perform binary classification and authenticity judgment on the remaining 60% dataset. The judgment model receives the original text input of the network interaction, and the model outputs a binary classification label indicating whether the result is true. Use the datasets in the 40% dataset and the 60% dataset that are judged to be true by the judgment model as the high-quality sentiment annotation datasets.
[0138] Specific implementation of the judgment model:
[0139] a) Length perception module: Evaluate the complexity of the input text.
[0140] b) Dynamic adjustment module: Adjust the number or weight of the CNN and Transformer layers according to the text complexity.
[0141] c) Fusion module: Use a gating mechanism or an attention mechanism to combine the outputs of the two architectures.
[0142] d) Meta controller: Learn the optimal architecture combination strategy during training.
[0143] Text complexity evaluation process:
[0144] Path 1: Evaluate according to the number of characters
[0145] Count the number of characters in all network interaction request records to obtain the minimum value a1, the quartile b1, the average value c1, the three-quarter digit d1, and the maximum value e1.
[0146] According to the interval in which the number of request characters falls, assign different complexity values complex_value1 (0.25, 0.50, 0.75, 1.00).
[0147] Path 2: Evaluate according to emotion words
[0148] Count the number of emotion words in all network interaction request records to obtain the minimum value a2, the quartile b2, the average value c2, the three-quarter digit d2, and the maximum value e2.
[0149] According to the interval in which the number of request emotion words falls, assign different complexity values complex_value2 (0.25, 0.50, 0.75, 1.00).
[0150] Comprehensive evaluation:
[0151] Calculate the sum of complex_value1 + complex_value2.
[0152] According to the sum value, classify the text into four categories and dynamically adjust the weights of CNN and Transformer:
[0153] * Uncomplicated text (≤0.5): CNN weight 0.8, Transformer weight 0.2
[0154] * Medium complexity text (0.5 - 1): CNN weight 0.6, Transformer weight 0.4
[0155] * Medium - upper complexity text (1 - 1.5): CNN weight 0.4, Transformer weight 0.6
[0156] * Complicated text (>1.5): CNN weight 0.2, Transformer weight 0.8
[0157] Step Four: Implementation of data balancing.
[0158] The present invention uses an intelligent data generation method based on Claude 3.5 Sonnet to solve the data imbalance problem. The specific implementation steps are as follows:
[0159] a) Data analysis: Conduct statistical analysis on the obtained high - quality labeled dataset, determine the number of samples in each sentiment category, identify the category with the largest number of samples, and set its number of samples as the target number N.
[0160] b) Imbalance identification: For each sentiment category i, calculate the difference between it and the target number Di = N - Ni, where Ni is the current number of samples in this category, and determine the categories that need data augmentation (categories with Di>0).
[0161] c) Prompt engineering: Design a specific prompt template for Claude 3.5 Sonnet, including the following elements:
[0162] 1. Task description: Generate network interaction texts that conform to specific sentiment categories;
[0163] 2. Sentiment category description: Elaborate on the characteristics and manifestation methods of the target sentiment;
[0164] 3. Text length and complexity requirements: Ensure that the generated texts are consistent with the style of the original dataset;
[0165] 4. Domain features: Provide specific vocabulary, expression methods, and common themes in the field of network interaction.
[0166] d) Data generation:
[0167] For each sentiment category i that needs to be enhanced:
[0168] 1. Use Claude 3.5 Sonnet to generate Di new text samples;
[0169] 2. After each generation, add the generated text as an example to the prompt to enhance the diversity of subsequent generations.
[0170] e) Quality control:
[0171] Design a rule-based filtering system for preliminary screening of the generated text:
[0172] 1. Check whether the text length is within the predetermined range;
[0173] 2. Verify whether it contains specific domain keywords;
[0174] 3. Use a simple sentiment analysis tool to predict the sentiment tendency of the generated text.
[0175] Through this data balancing method based on Claude 3.5 Sonnet, the present invention can effectively solve the data imbalance problem while ensuring the quality and diversity of the generated data. This method not only improves the balance of the dataset, but also enhances the model's ability to recognize various emotions, providing a high-quality and balanced data basis for the subsequent training of the emotion recognition model.
[0176] Step Five: Construct an emotion recognition model based on the network interaction dataset. In order to construct an efficient and accurate emotion recognition model on the publicly demanded classification dataset with high-quality emotion annotation that has been data-balanced, this paper designs the following five model architectures for comparative experiments. Figure 3 Shows the overall architectures of these five emotion recognition models.
[0177] Judgment of model performance evaluation:
[0178] The performance evaluation of the judgment model is a key step to ensure the acquisition of high-quality emotion annotation data in the end. This study adopts a multi-dimensional analysis method to comprehensively evaluate the model performance. Table 7 shows the detailed performance indicators of the judgment model on the test set. It can be seen from Table 7 that the model shows overall excellent performance, with an accuracy rate of 0.8354, a precision rate of 0.8638, and an F1 score of 0.7915.
[0179] Table 7 Performance Evaluation of the Judgment Model
[0180]
[0181] Based on the performance of the judgment model, the research conducted a quality assessment on the remaining 60% of the data (17,648 pieces) that were individually labeled by the large model in Subsection 3.1 and did not participate in the large model review and correction. By analyzing the confidence distribution of the model, a stratification strategy was adopted as follows:
[0182]
[0183] For the 11,719 pieces of data with a confidence level higher than 0.9, they were directly included in the high-quality data set; for the 3,530 pieces of data with a confidence level between 0.7 and 0.9, manual sampling review was conducted, and approximately 90% of them passed the quality inspection; while the 1,764 pieces of data with a confidence level lower than 0.7 were directly excluded. Through the overall statistics of the 40% of the data that were individually labeled by the large model and participated in the large model correction and manual review in Subsection 3.1, the composition of the high-quality sentiment annotation data set was finally obtained as shown in the table:
[0184]
[0185] The multi-stage data quality control method combines the advantages of manual expert judgment and machine learning models, which not only significantly improves the quality of sentiment annotation but also greatly enhances the efficiency of data processing. Finally, 97% of the original data was retained in the data set, ensuring both data quality and maintaining the representativeness and diversity of the data well. The 3% of the data that was excluded may contain some special or difficult-to-process cases, which may lead to poor performance of the final model when dealing with very complex or rare sentiment expressions.
[0186] Analysis of the experimental results of data balancing engineering:
[0187] This study obtained a data set containing 28,507 pieces of high-quality labeled data. However, preliminary analysis found that there were significant differences in the distribution of these data among different sentiment categories, and large differences in sentiment categories were bound to affect the performance of subsequent sentiment recognition models. To solve this problem, the research adopted a data balancing strategy based on the large language model. First, a statistical analysis of the distribution of each sentiment category in the original data set was conducted, and the results are shown in Table 9:
[0188] Table 9 Distribution of sentiment categories in the high-quality sentiment annotation data set
[0189]
[0190]
[0191]
[0192] As can be seen from Table 9, there is an obvious class imbalance problem in the high-quality sentiment annotation dataset obtained by the multi-stage combination method. The number of emotion samples of confusion, anger, and anxiety is relatively large, accounting for 31.3%, 25.1%, and 20.5% respectively, while the number of emotion samples such as calm, excited, and relaxed is extremely small. This is also related to the characteristics of online political participation texts because most online political participation texts are public complaints and dissatisfaction. This imbalance will cause the emotion recognition model to be biased towards the classes with a larger number of samples during the training process, thus affecting the recognition ability of the minority classes.
[0193] To alleviate this problem, a data generation strategy based on the Claude 3.5 Sonnet large language model was studied. The specific steps are as follows: For the classes with the number of samples below the average level (839), the large language model is used to generate additional samples. During the generation process, a specific prompt was designed based on the BROKE framework called by the large model, and the final prompt for data generation was determined through multiple debuggings with example data to ensure that the generated text conforms to the target emotion class and context. Secondly, for the newly generated text data by the large model, a judgment model is used to screen the generated data, and only the samples judged as "correct" are retained. Repeat steps 2 and 3 until the number of samples in each class is close to the average level. After data balancing, the distribution of the online political participation dataset after data balancing is shown in Table 10:
[0194] Table 10 Distribution of Emotion Classes in the Balanced Dataset
[0195]
[0196]
[0197] Through data balancing, the number of samples in each emotion category was successfully adjusted to a similar level. It is worth noting that in order to maintain data quality, the study moderately reduced the number of samples in categories with originally large numbers (such as confusion, anger, anxiety), rather than simply increasing the number of samples in minority categories. This strategy not only ensures the balance of the dataset but also avoids the quality risks that may be brought about by excessive generation. To evaluate the effect of data balancing, the study conducted an overall evaluation from the following four aspects: 1) Category distribution balance: The Gini coefficient was used to measure the balance degree of category distribution. The Gini coefficient before balancing was 0.866, and it decreased to 0.181 after balancing, indicating that the category distribution significantly tended to be balanced. 2) Quality of generated samples: The newly generated samples were evaluated using a judgment model. The results showed that 95.4% of the generated samples were considered to conform to the target emotion category and had natural expressions. This proves that the large language model has excellent capabilities in generating high-quality texts that conform to specific emotions. 3) Semantic diversity: The cosine similarity was used to evaluate the semantic differences between the generated samples and the original samples. The results showed that the average cosine similarity between the generated samples and the original samples was 0.871, indicating that the generated content not only maintained the relevance to the original data but also introduced new expressions, increasing the diversity of the dataset. 4) Impact on model performance: In this study, a basic RoBERTa classification model was trained using the datasets before and after balancing respectively, and evaluated on the same test set. The results are shown in Table 11:
[0198] Table 11 Comparison of model performance before and after data balancing
[0199]
[0200] Performance evaluation of the fine-grained emotion recognition model:
[0201] In this subsection, multiple models were used to train the balanced data to select the best fine-grained emotion recognition model. The study selected five classic deep learning models to participate in the experiment: RoBERTa, CNN, BiLSTM, RoBERTa+CNN, and RoBERTa+BiLSTM. To ensure a fair comparison, all models used the same preprocessing steps and training parameters. The dataset was randomly divided into a training set (80%) and a test set (20%). The study used the Adam optimizer, with a learning rate of 2e-5, a batch size of 32, and 50 training epochs. Table 12 shows the overall performance of the five models on the test set:
[0202] Table 12 Comparison of the overall performance of different models on the test set
[0203]
[0204] In the experimental results of the sentiment classification task, the RoBERTa model demonstrated excellent performance, with an accuracy rate of 0.7985, significantly higher than that of CNN (0.7011) and BiLSTM (0.7120). This indicates that RoBERTa is more superior in capturing the complex semantic information of online governance texts. This advantage is mainly due to the deep context understanding ability of the RoBERTa model, which can better handle the subtle differences between complex emotions. In contrast, although CNN and BiLSTM have certain advantages in capturing local features and sequence information of texts, they are not as good as RoBERTa with the Transformer architecture in dealing with long-distance dependencies and multi-layer semantic interactions. In terms of precision, the score of RoBERTa was 0.8046, far higher than that of CNN (0.7496) and BiLSTM (0.7666), highlighting its powerful feature extraction ability. Further analyzing the F1 score, RoBERTa still performed the best. In contrast, the F1 scores of CNN (0.7118) and BiLSTM (0.7278) were lower, indicating that although these two models may perform better in some sentiment categories, they failed to achieve the effect of RoBERTa in terms of global balance.
[0205] The performance of the combined models RoBERTa+CNN and RoBERTa+BiLSTM was lower than expected. Neither their accuracy rate, precision, nor F1 score exceeded that of the individual RoBERTa, and they even showed disadvantages in all evaluation metrics, especially with accuracy rates of 0.6412 and 0.6492 respectively. This result may reflect the challenges of multi-model fusion in the sentiment classification task. The theoretical basis of the combined model is to integrate the advantages of different models. However, in practical applications, the powerful semantic understanding ability of RoBERTa is already sufficient to capture the complex relationships between sentiment categories. Adding other models may introduce noise or increase unnecessary complexity, which instead affects the overall performance.
[0206] The study selected the best-performing RoBERTa to explore its performance in each emotion category, as shown in Table 13. Observing the 34 emotion categories, it can be seen that there are certain differences in the performance of Accuracy, Precision and F1-score among the categories. For example, emotion categories such as "despair" and "shock" perform well, while categories such as "anxiety" and "helplessness" perform relatively weakly. At the same time, this study conducted group analysis based on 8 emotion intervals and positive and negative emotions, as shown in Tables 14-15. By dividing the 34 emotions into 8 emotion intervals, it can be seen that the performance of negative and positive emotions is different in different arousal levels, and that the performance of negative emotions gradually weakens as the arousal level increases. In the neutral arousal emotional zone, negative emotions perform better, while in other intervals, positive emotions may perform more prominently. However, when emotions are divided into two categories, positive and negative, negative emotions perform better overall, which may be because negative emotions are easier to be clearly defined and identified. This grouping can help us better understand the relationship between emotion categories.
[0207] Table 13 Classification effect of the model on each category
[0208]
[0209]
[0210]
[0211] Table 14 Classification results of the model in 8 emotional intervals
[0212]
[0213]
[0214] Table 15 Classification results of the model in the positive and negative polarity emotion range
[0215]
[0216] In summary, RoBERTa's excellent performance in the online political text sentiment classification task demonstrates its significant advantages in processing complex, multi-level semantic information. Although the combined model should theoretically enhance the classification ability, in this task, the use of RoBERTa alone can provide the best results, and further model combinations may reduce performance due to insufficient feature fusion or increased complexity.
[0217] Through the integration of large language models and traditional machine learning methods, this study proposed a novel method for fine-grained sentiment recognition of online political participation texts, successfully improving the accuracy, efficiency, and applicability of sentiment analysis, and providing a powerful tool for government departments to accurately grasp public opinion and enhance governance capabilities. Based on the experiments and in-depth analysis of this study, the following four main conclusions were drawn:
[0218] (1) Multi-dimensional fine-grained sentiment annotation system: improving the accuracy of online political participation text analysis. The effectiveness and applicability of the fine-grained sentiment annotation system. The 8 emotion intervals and 56 emotion label systems constructed based on the arousal-valence theory of emotion demonstrated significant effectiveness and wide applicability. This multi-dimensional sentiment label system can not only effectively capture the fine-grained sentiment in online political participation texts, but also significantly improve the accuracy and granularity of sentiment analysis. The experimental results show that compared with previous fine-grained sentiment classification methods, the high-quality data annotation method proposed in this study improved the overall performance by 21.65% in identifying complex emotion expressions.
[0219] (2) GPT-4 assisted sentiment annotation of online political participation texts: efficiency improvement and human-machine collaboration strategy. GPT-4.0 demonstrated excellent performance in the fine-grained sentiment annotation task of online political participation texts, but also exposed some limitations. By designing a scientific prompt strategy template, high-quality sentiment annotation data was successfully obtained using large language models. Compared with manual annotation, GPT-4.0 not only improved the annotation efficiency (the speed was increased by about 10 times), but also significantly improved the annotation consistency, with the Cohen's Kappa coefficient increasing from 0.72 to 0.85. However, this study found that large language models still have difficulties in dealing with professional terms and implicit expressions in specific fields, with an error rate of 21.83%. To solve this problem, this study proposed a hybrid annotation strategy, combining the efficiency of large language models and the expertise of human experts, and increasing the accuracy of such special cases to 99%. This strategy not only improves the annotation quality, but also provides new ideas for the optimization of large language models in professional field applications.
[0220] (3) Collaboration between RoBERTa Judgment Model and Large Language Model: Enhancing Sentiment Annotation Quality and Data Balance. The proposed RoBERTa-based judgment model effectively solves the "hallucination" problem that may occur in large language models. At the same time, using the large language model for data balance optimization significantly alleviates the class imbalance problem. The judgment model can accurately identify and screen high-quality sentiment annotation results, improving the quality of the annotation data by 27.96%. In terms of data balance, this study not only successfully reduces the sample distribution differences among various categories but also improves the overall model performance by 38.08%. More notably, this data quality control and balance optimization method demonstrates strong transferability, providing a new solution to the common problems of data quality and class imbalance in the field of natural language processing.
[0221] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising such element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0222] The embodiments of the present application have been described above with reference to the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of the present application and without departing from the spirit and scope protected by the claims of the present application, can still make many forms, all of which fall within the protection scope of the present application.
Claims
1. A fine-grained sentiment analysis method for interactive text assisted by a large language model, characterized in that: It combines large language models, convolutional neural networks (CNNs), and Transformer technologies, and achieves adaptive processing of texts of different complexities through dynamic weight adjustment, including the following steps: Step 1: High cross-domain adaptability and optimized sentiment granularity division; Step 2: Improve data quality and enhance model flexibility; Step 3: Solve data imbalance; Step 4: Constructing the emotional label system; Step 5: Emotional space coding; Step 6: Sentiment labeling experiment; Step 7: Analyze model design and implementation.
2. The method for fine-grained sentiment analysis of interactive text assisted by a large language model according to claim 1, characterized in that: The emotional labeling system is constructed as follows: According to the emotional arousal-valence theory, an emotional labeling system is constructed. Based on the emotional arousal-valence theory, an emotional labeling system is constructed. This step involves in-depth research on emotional psychology theory, combined with the characteristics of the field of network interaction, to design an emotional classification system that has both a theoretical basis and is practical. This system not only takes into account the positivity and negativity of emotions, but also includes the dimension of emotional intensity, making emotional classification more refined and comprehensive.
3. The method for fine-grained sentiment analysis of interactive text assisted by a large language model according to claim 1, characterized in that: The emotion space encoding adopts 8-dimensional multi-label vector space {0,1}^8 to encode the emotion space. The encoding method allows more accurate representation of complex emotion states. Each dimension represents a basic emotion category. The combination of 0 and 1 can represent rich emotion states. The multi-label method is superior to the traditional single-label classification and can capture multiple emotions that may exist simultaneously in the text.
4. The method for fine-grained sentiment analysis of interactive text assisted by a large language model according to claim 1, characterized in that: The sentiment labeling experiment is designed with two large model experiments for sentiment labeling experiments. The first large model performs preliminary sentiment labeling, and another large model with better performance is used to review and correct the labeling results of the first large model, mapping the emotions reflected in the original text into a 56-dimensional target sentiment space.
5. The method for fine-grained sentiment analysis of interactive text assisted by a large language model according to claim 1, characterized in that: The design and implementation of the judgment model randomly extracts 40% of the data from the labeled data set. The expert team manually reviews the data after the Claude3.5 large model review and correction, and performs binary classification labeling for each data to determine whether each data produces hallucinations. 1 means that the data after correction still produces hallucinations and is regarded as an incorrect labeling, and 0 means that the data after correction is correctly labeled. The binary classification annotated dataset is used as the training dataset to train the CNN+transformer model. The CNN+transformer model is an adaptive architecture with dynamically generated weights based on the complexity of the text. The evaluation indicators of the judgment model include: Precision, Accuracy, and F1-score. The trained judgment model is applied to the remaining 60% of the labeled data set, and finally a high-quality data set with large model annotations without hallucinations is obtained. The judgment model adopts an adaptive architecture that dynamically integrates CNN and Transformer, and dynamically adjusts the proportion of CNN and Transformer according to the complexity of the input text. The specific implementation method is as follows: a) Design a length-aware module to evaluate the complexity of the input text; b) Use the output of this module to dynamically adjust the number or weights of CNN and Transformer layers; c) Design fusion modules, such as gating mechanisms or attention mechanisms, to combine the outputs of the two architectures; d) Introducing a meta-controller to learn the optimal architecture combination strategy during training; Text complexity evaluation process: Path 1: Evaluate based on the number of characters: Count the number of characters in all online interactive appeal records, and get the minimum value a1, quartile b1, average c1, three-quarter digit d1, and maximum value e1. According to the interval in which the number of appeal characters falls, assign different complexity values complex_value1 (0.25, 0.50, 0.75, 1.00); Path 2: Evaluate based on sentiment words: Count the number of emotional words in all online interactive appeal records, and obtain the minimum value a2, quartile b2, average c2, three-quarter digit d2, and maximum value e2; According to the interval in which the number of appeal emotional words falls, different complexity values complex_value2 (0.25, 0.50, 0.75, 1.00) are assigned.
6. The method for fine-grained sentiment analysis of interactive text assisted by a large language model according to claim 1, characterized in that: The data balancing project uses a large model with better performance to generate data, achieves balanced labeling results for each category, and solves the data imbalance problem by increasing minority class samples through intelligent generation, so that the number of samples in each emotion category tends to be consistent. This method can not only improve the overall performance of the model, but also enhance the model's recognition ability for minority categories.
7. The method for fine-grained sentiment analysis of interactive text assisted by a large language model according to claim 1, characterized in that: The emotion recognition model training and evaluation: a) Select a variety of model architectures for comparative experiments, including RoBERTa, BiLSTM, CNN and their combined models; b) Comprehensively evaluate the performance of each model using indicators such as precision, accuracy, and F1 value. By comparing the performance of different models on various indicators, select the model architecture that is most suitable for sentiment analysis of online interactive text.
8. The method for fine-grained sentiment analysis of interactive text assisted by a large language model according to claim 1, characterized in that: Source and preprocessing of the data set. In the data preprocessing stage, the original data set was processed as follows: In the data preprocessing stage, this paper processes the original data set as follows: (1) Data cleaning: basic cleaning operations such as deduplication, removal of special characters, and lowercase conversion of letters are performed on the text data to improve data quality; (2) Text encoding: Use the pre-trained word embedding model RoBERTa to directly encode each text sample into a fixed-length sentence vector as the input feature of the model; (3) Data partitioning: For each data set, 80% of the samples are randomly divided as training sets and 20% as test sets for model training and evaluation. At the same time, in order to ensure the reliability and reproducibility of the experimental results, this paper fixes the random seed when partitioning the data to ensure fair comparison of different methods on the same data set. Through the above preprocessing steps, the original data set is converted into a format suitable for model training and optimization, laying a data foundation for subsequent label annotation and model construction.
9. The method for fine-grained sentiment analysis of interactive text assisted by a large language model according to claim 1, characterized in that: The data is balanced: 1) Class distribution balance: The Gini coefficient is used to measure the balance of class distribution. The Gini coefficient before balance is 0.866, and after balance is reduced to 0.181, indicating that the class distribution is significantly balanced; 2) Generated sample quality: The newly generated samples were evaluated using the judgment model, and the results showed that 95.4% of the generated samples were considered to be in line with the target sentiment category and the expression was natural. This proves that the large language model has excellent ability to generate high-quality text that conforms to specific sentiments; 3) Semantic diversity: Cosine similarity is used to evaluate the semantic difference between the generated samples and the original samples. The results show that the average cosine similarity between the generated samples and the original samples is 0.871, indicating that the generated content not only maintains the relevance to the original data, but also introduces new expressions, increasing the diversity of the dataset; 4) Impact on model performance: This study used the datasets before and after equalization to train a basic RoBERTa classification model and evaluated it on the same test set.
Citation Information
Cited By
Emotion analysis method based on emergency scene
CN120804331A
Feature processing method of text sequence and related device
CN122509164A