Emotion evaluation method in client evaluation text content

Through deep learning model and BERT algorithm model, the emotional classification and analysis of customer evaluations is solved, and the problem of emotional recognition of text evaluation in the existing technology is realized, and the emotion automation evaluation and quantitative feedback of customer evaluations is improved, and the evaluation efficiency and accuracy are improved.

CN119988620APending Publication Date: 2025-05-13CHINA LIFE INSURANCE CO LTD HEBEI BRANCH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411933314.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art has problems such as large quantity, difficulty in evaluating and low efficiency in the emotional recognition of text evaluation. Especially in the insurance sales industry, the lack of quantitative means makes it difficult for text evaluation to become a reference basis for evaluating sales strategies.

Method used

By establishing a basic emotion database, using deep learning models and pre-trained BERT algorithm models, emotion classification and emotion intensity analysis are performed on customer evaluations, combined with natural language processing technology and multi-dimensional emotion data annotation, personalized customer portraits are constructed, and emotional labels and confidence are predicted through model regression algorithms.

Benefits of technology

It realizes emotional automation evaluation of customer evaluation, improves the accuracy and efficiency of emotional recognition, can quantify customer feedback, enhance the company's connection with customers, and reduce the impact of personal subjective judgments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988620A_ABST
    Figure CN119988620A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of text language analysis, in particular to a method for evaluating emotion in client evaluation text content, which comprises the following steps of: S1, establishing a basic emotion database through the client evaluation text content, performing similarity screening, invalid deletion and coding on data, and removing texts without evaluation value to obtain a basic emotion database; carrying out natural language preprocessing on the residual evaluation texts; s2, performing sentiment classification on the evaluation of the customer based on a deep learning model, tracking the change of the evaluation of the customer through time sequence analysis, judging the change of the customer on products and services at different time points by using clustering analysis, and constructing a personalized portrait of the customer, in the method, the change of the evaluation of the customer is analyzed and tracked through A I, and a theme label is distributed for an evaluation text; and meanwhile, when the BERT algorithm model is trained, the training effect of the model is enhanced through MLM mask, NSP prediction and customer personalized portrait, and the purpose of truly extracting semantics and emotions in the text is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text language analysis, and more specifically, to a method for evaluating emotions in text content of customer evaluations. Background Art

[0002] Text reviews are customers' opinions on a service or product expressed in written form. They not only convey customers' emotions, but also reflect specific product features, service quality, and customer needs. Therefore, emotional evaluation of text reviews can provide support for subsequent service improvements and secondary sales.

[0003] However, due to the relatively large number of text evaluations, manual screening and annotation is often time-consuming and laborious, and cannot achieve direct economic benefits. Conversely, when direct labeling feedback is used, the customer's wishes may be difficult to understand due to the complexity of the text content. Therefore, in the insurance sales industry, the collection and analysis of text evaluations often only exist as a service project for a small number of important customers. For ordinary sales personnel, the lack of quantitative means makes it difficult for text evaluations to become a reference for evaluating sales strategies. In addition, the recognition of text emotions mostly relies on fixed text patterns, such as fixed popular word combinations or emoticons, which lack connection with the context and have poor generalization ability. In summary, the current evaluation of text still has problems such as large quantity, difficult evaluation and low efficiency. Summary of the invention

[0004] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a method for evaluating emotions in customer evaluation text content to solve the problems raised in the above-mentioned background technology.

[0005] To achieve the above object, the present invention provides the following technical solution: a method for evaluating emotions in customer evaluation text content, comprising the following steps:

[0006] S1: Establish a basic sentiment database through the text content of customer reviews, perform similarity screening, invalid deletion and encoding on the data, add sentiment dictionary expansion, evaluate the validity of the text through the predefined vocabulary library, remove text with no evaluation value, and perform natural language preprocessing on the remaining evaluation texts to process the corpus into digital tokens;

[0007] S2: Based on the deep learning model, sentiment classification of customer reviews is performed, and purchase interests and complaint tendencies are identified through specific keywords and sentence structures. Changes in customer reviews are tracked through time series analysis, and cluster analysis is used to determine changes in customers' attitudes toward products and services at different time points. Personalized customer portraits are constructed, and multi-dimensional sentiment data is annotated. Layered recommendations are made based on sentiment annotations, and annotation results are verified.

[0008] S3: Based on the pre-trained BERT algorithm model, supplement the insurance knowledge related corpus, combine the manually annotated sentiment semantic library, conduct targeted reinforcement training of the model, use MLM masked language modeling and NSP prediction tasks to strengthen the training of the BERT model, capture the contextual relationship and the association between sentences, match the personalized portrait of customers in the corresponding time period, and perform Fine-Tuning for sentiment classification tasks based on the BERT model;

[0009] S4: The encoder part of BERT encodes the input Token corpus into tokens containing context information, and uses dynamic maximum pooling technology to convert different numbers of Tokens into representations of uniform size.

[0010] S5: The representative tags obtained by the encoder are further classified by the classifier, and the input content is nonlinearly transformed and normalized through the activation function. Some neurons are randomly discarded in each training iteration. For the class imbalance problem in the sentiment classification task, the weighted cross entropy loss function is used to weight the minority class samples.

[0011] S6: Based on the model regression algorithm, predict the sentiment label of each review, calculate the corresponding confidence, and revise the customer's star rating. After comparing the sentiment label with the actual evaluation, those with large differences are transferred to manual review, and the model is iterated according to the review annotations.

[0012] As a further solution of the present invention, in S1, the data is screened for similarity, invalidated and encoded, specifically:

[0013] Similarity screening: Use cosine similarity to calculate the similarity between the text content and the preset sentence, and directly assign the preset emotional label to similar comments;

[0014] Invalid deletion: remove text content that is too short or does not contain enough information;

[0015] Encoding: Convert text content into one-hot discrete data encoding form.

[0016] As a further solution of the present invention, in S1, natural language preprocessing is performed on the remaining evaluation text, including information density evaluation, sentence segmentation and Unicode conversion, wherein:

[0017] Information density evaluation: Count the ratio of the number of different words in the text to the total number of words, and measure the text diversity through Shannon entropy;

[0018] Sentence segmentation: split the review text into meaningful short sentences;

[0019] Unicode conversion: Convert text to unified Unicode encoding to ensure the consistency of text.

[0020] As a further solution of the present invention, in S1, the expansion of the sentiment dictionary is based on the statistical number of words, the number of times used is compared, high-frequency words are screened, the word attributes and necessity are determined jointly by AI and humans, and necessary words are added to the predefined vocabulary library.

[0021] As a further solution of the present invention, in S2, the identification of purchase interest and complaint tendency is specifically as follows:

[0022] Remove stop words, extract stop words that contribute less to semantic analysis, restore the text to its basic form, and identify and mark the part of speech of each word;

[0023] Keyword extraction, identifying common words that represent purchase interest and complaint tendency;

[0024] Text structure analysis, identifying the text expression intention through the structural forms of affirmative sentences, negative sentences, rhetorical questions, interrogative sentences and double negative sentences.

[0025] As a further solution of the present invention, in S2, the customer's personalized portrait contains multi-dimensional information, including basic information, interests and hobbies, preferences, consumption habits, lifestyle and pain points, and the changes in the customer's personalized portrait are matched and marked in chronological order.

[0026] As a further solution of the present invention, in S2, the recommendation strategy is adjusted according to the annotations corresponding to the customer big data portrait, and compensatory measures are recommended when the customer expresses dissatisfaction, and more value-added services and related products are recommended when the customer expresses satisfaction, and then the correctness of the annotations and recommendations is verified according to the feedback trends of the customers after the recommendation.

[0027] As a further solution of the present invention, in S3, the BERT model is trained using MLM masked language modeling and NSP prediction tasks, specifically:

[0028] Input a pair of sentences connected by a special separator, input a randomly selected mask;

[0029] For the MLM task, some words in the sentence are randomly masked, and the masked words are predicted based on the context information. The difference between the model predicted words and the actual words is measured by the loss function.

[0030] For the NSP task, the order of sentences is determined by judging the natural connection relationship between a pair of input sentences.

[0031] As a further solution of the present invention, in S3, the customer big data portrait of the corresponding period is matched, specifically:

[0032] Predict customer interest categories based on contextual information and customer profiles;

[0033] Determine the relevance of sentences and customer behaviors based on the match between the context;

[0034] The customer portrait is used as an additional feature input and combined with the text content into the model, and the model performance is further enhanced through multimodal learning methods.

[0035] As a further solution of the present invention, in S6, the confidence is the predicted probability of the model, and the corresponding star ratings are mapped according to positive, neutral, negative and complex emotion labels.

[0036] Technical effects and advantages of a method for evaluating emotions in text content of customer evaluations of the present invention:

[0037] The text content of insurance customer reviews is analyzed through AI technology, and the changing trends of the reviews are tracked, and the corresponding product and service changes are matched. Natural language processing technology and LDA algorithm are used to assign topic tags to customer review texts, evaluate customer personalized information, and then build personalized portraits that meet the needs of insurance customers. At the same time, when training the BERT algorithm model, MLM mask and NSP prediction are used as the basis, and the training effect of the BERT algorithm model is enhanced from multiple dimensions in conjunction with the customer personalized portrait. The flexibility and efficiency of the pooling operation are improved through dynamic maximum pooling technology, information loss is reduced, and the generalization ability of the model is enhanced. In addition, the WCE weighted cross entropy is used as the loss function to weight minority samples to balance the data of various types of sentiment words, so as to achieve the purpose of improving the BERT algorithm model and truly extracting semantics and emotions from the text through the model, avoiding time-consuming and labor-intensive situations in sentiment evaluation, and then enabling text evaluation to be directly used as the evaluation basis for sales strategies, quantifying customer feedback indicators, strengthening the connection between companies and customers, and reducing the impact of personal subjective judgment. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a schematic diagram of a method for evaluating emotions in text content of customer evaluations according to the present invention. DETAILED DESCRIPTION

[0039] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0040] Example

[0041] Figure 1 The present invention provides a method for evaluating emotions in text content of customer evaluations, which comprises the following steps:

[0042] S1: Establish a basic sentiment database through the text content of customer reviews, perform similarity screening, invalid deletion and encoding on the data, add sentiment dictionary expansion, evaluate the validity of the text through the predefined vocabulary library, remove text with no evaluation value, and perform natural language preprocessing on the remaining evaluation texts. Process the corpus into digital tokens, obtain the text content of customer reviews through API interfaces, data capture and crawler tools, and then build a non-relational database through inverted index algorithms, Snappy compression algorithms, etc., and use cosine similarity to screen the text for similarity to remove duplicate and redundant text content and reduce the interference caused by data redundancy. At the same time, use TF-IDF, TextRank sorting algorithms, etc. to sort out some texts that are too short and meaningless texts. Invalid words are deleted and the text content is converted into a data format that can be used for machine learning through one-hot encoding. Then, the difference detection algorithm is used to match each text content according to the words defined in the vocabulary library, and then valuable texts containing positive and negative words are screened out according to the number of matched words, which provides a basis for the subsequent quantification of text sentiment tendency. At the same time, relevant preprocessing is required for this part of valuable text to facilitate subsequent rapid use. The tokenization process can convert the text into digital form for the understanding and processing of computer models, further ensuring the realization of tasks such as sentiment classification and sentiment intensity analysis. In addition, in order to improve the accuracy of sentiment analysis, the coverage of sentiment vocabulary can be appropriately increased by expanding the sentiment dictionary to ensure that more sentiment expressions can be recognized;

[0043] S2: Based on the deep learning model, sentiment classification of customer reviews is performed. Purchase interest and complaint tendency are identified through specific keywords and sentence structures. Changes in customer reviews are tracked through time series analysis. Cluster analysis is used to determine changes in customer attitudes towards products and services at different time points. Customer personalized portraits are constructed. Multi-dimensional sentiment data annotation is performed. Layered recommendations are made based on sentiment annotations. The annotation results are verified. When analyzing individual words and understanding the context of sentences with deep learning models, the potential sentiment information in specific words and sentences can be captured through word vectors and attention mechanism models, so as to more accurately determine whether customers show purchase interest or complaint tendency for the product. Further, by analyzing the positive words, purchase intention phrases, specific content of the evaluation, and negative sentences in customer reviews, customers' purchase interest, complaint sentiment tendency, and whether there is the potential to be converted into actual purchases can be identified. At the same time, time series analysis can be used to track customer evaluation trends, discover sentiment fluctuation patterns, and obtain valuable information about customer loyalty and changes in brand awareness. In addition, customer sentiment scores and evaluation content can be used to identify the potential of customer complaints. Cluster analysis based on multiple dimensions such as content, purchase behavior, etc. can divide customers into different categories of groups, and discover the emotional changes of certain groups, thereby identifying customers' different attitudes and needs for products or services in different time periods. In summary, the long short-term memory network, convolutional neural network and K-means clustering can construct dynamic big data emotional and behavioral portraits of customers to obtain additional multi-dimensional personalized information of customers. In addition, based on sentiment analysis, the LDA algorithm can be used to perform multi-dimensional real-time annotation of different emotional tendencies, purchase intentions and complaint tendencies to form more refined emotional labels, and further cooperate with matrix decomposition, nearest neighbor search, Learning-to-Rank algorithm and LSTM algorithm for hierarchical recommendation. For example, customers with positive emotions can be recommended new products and promotional activities, while customers with negative emotions can be recommended after-sales services and product improvements, thereby greatly improving customer conversion rate and satisfaction. At the same time, combined with deep learning models and rule engines, it can be checked whether the emotional labels are consistent with the actual feedback of customers after hierarchical recommendations, and then it can be judged whether there is a labeling error;

[0044] S3: Based on the pre-trained BERT algorithm model, supplement the insurance knowledge related corpus, combine with the manually annotated sentiment semantic library, conduct targeted reinforcement training of the model, use MLM masked language modeling and NSP prediction tasks to strengthen the training of the BERT model, capture the contextual relationship and the association between sentences, match the personalized portrait of customers in the corresponding time period, and perform Fine-Tuning on the sentiment classification task based on the BERT model. By supplementing the professional knowledge corpus related to the insurance field, the performance of the BERT algorithm model in this field can be further improved. At the same time, combined with the manually annotated slang sentiment semantic library, it can provide the model with more accurate sentiment analysis, so that the model can not only understand the insurance-related text, but also the sentiment classification task. Semantic interpretation can also effectively identify the emotional color of slang in texts. On this basis, the MLM task can be used to mask certain words in the text, and the semantic understanding ability of the model can be enhanced by allowing the model to infer the correct words in the context. At the same time, the NSP task can help the model understand the logical relationship between sentences by predicting the correlation between two sentences, capturing the semantic correlation of the context, and further by additionally matching the current personalized portrait of the customer, the model's perception of customer needs can be combined with multi-dimensional information. In addition, Fine-Tuning can be used to optimize the model's performance in sentiment discrimination, so that it can accurately judge the customer's emotional tendencies in insurance texts and provide insurance companies with more accurate customer insights.

[0045] S4: The input Token corpus is encoded by the Encoder part of BERT and converted into tags containing context information. The dynamic maximum pooling technology is used to convert different numbers of Tokens into representations of uniform size. The multi-layer Transformer network of BERT can convert each Token into a high-dimensional vector representation containing its context information. Then, the self-attention mechanism can capture the relationship between the tokens in the input sequence, so that the token contains the information of the vocabulary itself while combining the semantic and syntactic relationship of the surrounding words to form a richer context-aware representation. In addition, the dynamic maximum pooling can select the most significant part of the sequence by pooling the token representation of each layer, and retain the most informative Token features by maximizing the selection, so that the original Token representations of different lengths can be converted into vector representations of uniform size. While enhancing the model's adaptability to inputs of different lengths, it also helps the model effectively extract key information in the context, ensuring the ability to integrate information, and the pooled vector can be used for subsequent classification, generation or other natural language processing tasks;

[0046] S5: The representative tags obtained by the Encoder are further classified by the Classifier, and the input content is nonlinearly transformed and normalized through the activation function. Some neurons are randomly discarded in each training iteration. For the class imbalance problem in the sentiment classification task, the weighted cross entropy loss function is used to weight the minority class samples. The Encoder encoder can extract the representative features of the input text and pass them as input to the Classifier, which further classifies the text sentiment according to the features. At the same time, activation functions such as ReLU or Sigmoid can be used to The input content is transformed nonlinearly to increase the expressive power of the model. In addition, the activation function can be used to normalize the data to ensure that the scales of different eigenvalues ​​are consistent, thereby improving the stability and convergence speed of training. The Dropout technology can randomly discard some neurons to reduce the dependence of the neural network on training data, prevent overfitting, and enable the model to perform better on unknown data. The weighted cross entropy loss function can be used to give higher weights to minority class samples to avoid the situation where traditional classification methods tend to predict a larger number of categories, thereby prompting the model to pay more attention to minority class samples and improve the prediction effect of the classifier under class imbalance.

[0047] S6: Based on the model regression algorithm, predict the sentiment label of each review, calculate the corresponding confidence, and revise the customer's star rating. After comparing the sentiment label with the actual evaluation, those with large differences are transferred to manual review, and the model is iterated according to the review annotations. The input text can be processed by the regression model algorithm to generate a continuous sentiment score to represent the sentiment intensity or tendency and match the corresponding sentiment label. At the same time, the credibility of the prediction result, that is, the confidence, is calculated through standard error, confidence interval, residual analysis or Bayesian regression. At this time, the star rating can be adjusted according to the sentiment label and confidence predicted by the model. If the model determines that a certain review has a strong positive emotion, but the actual star rating is medium, its star rating can be increased for update. At the same time, manual review and confirmation can be used for some cases where there are large differences between the predicted sentiment labels and the actual customer evaluation content to determine whether the model has errors or deviations. The results and annotations of the manual review can be fed back to the model as new training data to help the model adjust its parameters and improve its prediction ability, thereby improving the recognition accuracy of customer evaluation emotions, so as to achieve the purpose of continuous optimization and effective response to complex and diversified text data.

[0048] Among them, in S1, the data is screened for similarity, invalidly deleted and encoded, specifically:

[0049] Similarity screening: Use cosine similarity to calculate the similarity between text content and preset sentences. For similar comments, directly assign preset sentiment tags. Use TF-IDF, Word2Vec, BERT, etc. to convert text content and preset sentences into vector representations to calculate the cosine similarity between the text vector and the preset sentence vector. For comments with high similarity, preset sentiment tags can be directly assigned to improve the processing efficiency of the model and avoid the need to perform sentiment analysis on each comment.

[0050] Invalid deletion: Remove text content that is too short or does not contain enough information. Use sentence length features and machine learning to identify reviews with only one word or sentence and sentiment-ambiguous sentences to prevent such text from reducing the accuracy of sentiment analysis.

[0051] Encoding: Convert text content into one-hot discrete data encoding form.

[0052] In S1, natural language preprocessing is performed on the remaining evaluation text, including information density assessment, sentence segmentation, and Unicode conversion, where:

[0053] Information density assessment: Count the ratio of the number of different words in the text to the total number of words, and measure the text diversity through Shannon entropy. Information density assessment can improve information transmission efficiency, optimize text content, and reduce information overload;

[0054] Sentence segmentation: Segment the review text into meaningful short sentences. Sentence segmentation can improve the model's processing efficiency for text and help extract more accurate semantic information.

[0055] Unicode conversion: Convert text to unified Unicode encoding to ensure the consistency of the text. Through Unicode conversion, each character can be assigned a unique code point to ensure compatibility between different systems, platforms and languages.

[0056] Among them, in S1, the expansion of the sentiment dictionary compares the number of uses based on the number of words counted, screens high-frequency words, and AI and humans jointly determine the word attributes and necessity. The necessary words are added to the predefined vocabulary library, and the number of occurrences of each word is stored in a hash table. TF-IDF is used for comparison to identify high-frequency words and low-frequency words. Since some high-frequency words in the text have emotional tendencies, machine learning, natural language processing and other technologies can be used to analyze whether high-frequency words really have emotional attributes and their emotional intensity in specific contexts. At the same time, humans are used to review and adjust the emotional attributes of these words according to the actual situation and context, so as to avoid AI from misjudging some words due to its inability to fully understand certain complex contexts or cultural backgrounds. The high-frequency words reviewed by AI and humans can be added to the predefined vocabulary library of the sentiment dictionary for better sentiment analysis and emotion recognition.

[0057] Among them, in S2, the identification of purchase interest and complaint tendency is specifically as follows:

[0058] Remove stop words, extract stop words that contribute less to semantic analysis, restore the text to its basic form, identify and mark the part of speech of each word, and simplify the text content by removing stop words;

[0059] Keyword extraction, identifying common words that represent purchase interest and complaint tendency;

[0060] Text structure analysis can identify the text expression intention through the structural forms of affirmative sentences, negative sentences, rhetorical questions, interrogative sentences and double negative sentences. Through text simplification and keyword extraction, it can analyze the surface meaning, implicit intention, emotional color and attitude towards certain issues of the text in combination with the actual sentence patterns of the text, thereby obtaining a more accurate text emotion.

[0061] Among them, in S2, the customer's personalized portrait contains multi-dimensional information, including basic information, interests and hobbies, preferences, consumption habits, lifestyles, and pain points, and the changes in the customer's personalized portrait are matched and marked in chronological order. By collecting and analyzing multi-dimensional information such as customer behavior, needs, and interests, a dynamic customer portrait can be generated, and the portrait can help companies understand customers from multiple dimensions. As time goes by, the timeliness of the customer portrait can be guaranteed by regularly updating it. Further, when analyzing the text content of its evaluation in the subsequent period, the emotional bias in the text can be judged more comprehensively and accurately based on the current customer's personalized portrait.

[0062] Among them, in S2, the recommendation strategy is adjusted according to the annotations corresponding to the customer's big data portrait. When the customer expresses dissatisfaction, compensatory measures are recommended, and when the customer expresses satisfaction, more value-added services and related products are recommended. Then, according to the feedback trend of the customer after the recommendation, the correctness of the annotation and recommendation is verified. According to the customer's emotional feedback and behavioral data, the recommendation strategy can be adjusted in a targeted manner through AI combined with the preset plan. Then, the accuracy of the customer's emotional annotation can be verified through the feedback from the customer after the strategy adjustment, so as to provide more reliable data for the subsequent BERT algorithm model training.

[0063] In S3, the BERT model is trained using MLM masked language modeling and NSP prediction tasks, specifically:

[0064] Input a pair of sentences connected by a special separator, input a randomly selected mask;

[0065] For the MLM task, some words in the sentence are randomly masked, and the masked words are predicted based on the context information. The difference between the words predicted by the model and the actual words is measured by the loss function. The MLM task enables the model to understand the context in both directions and grasp the grammatical structure of the sentence and the relationship between words.

[0066] For the NSP task, by judging the natural connection relationship of a pair of input sentences and determining the order of the sentences, and by judging whether there is a natural connection between the sentences, the model can better understand long texts and improve its performance in processing complex language tasks.

[0067] Among them, in S3, the big data portrait of customers in the corresponding period is matched, specifically:

[0068] Predict the customer's interest category based on context information and customer profile. Use the BERT model to comprehensively predict the customer's possible interest direction by combining context information and the label corresponding to the current customer profile.

[0069] Based on the matching degree between the contextual sentences and customer behaviors, the correlation between them is judged, and a context-aware model is constructed through Euclidean distance or neural network model, and the matching degree of the customer to the current text content is predicted based on the customer's behavior sequence;

[0070] The customer portrait is used as an additional feature input and combined with the text content input into the model. The model performance is further enhanced through multimodal learning methods. Through feature-level fusion, model-level fusion, and decision-level fusion in multimodal learning, the structured data provided by the customer portrait and the unstructured data provided by the text content can be integrated together. Through multi-level and multi-dimensional information fusion, the performance of the model in sentiment analysis can be effectively improved, thereby more accurately judging customer emotional feedback.

[0071] Among them, in S6, the confidence is the prediction probability of the model, and the corresponding star rating is mapped according to the positive, neutral, negative and complex emotional labels. The corresponding star rating can be mapped by combining the prediction probability of the model, that is, the confidence, with the emotional label to more intuitively display the emotional analysis results, and the higher the confidence, the higher the star rating.

[0072] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0073] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for evaluating emotions in customer evaluation text content, characterized in that: The steps include: S1: Establish a basic sentiment database through the text content of customer reviews, perform similarity screening, invalid deletion and encoding on the data, add sentiment dictionary expansion, evaluate the validity of the text through the predefined vocabulary library, remove text with no evaluation value, and perform natural language preprocessing on the remaining evaluation texts to process the corpus into digital tokens; S2: Based on the deep learning model, sentiment classification of customer reviews is performed, and purchase interests and complaint tendencies are identified through specific keywords and sentence structures. Changes in customer reviews are tracked through time series analysis, and cluster analysis is used to determine changes in customers' attitudes toward products and services at different time points. Personalized customer portraits are constructed, and multi-dimensional sentiment data is annotated. Layered recommendations are made based on sentiment annotations, and annotation results are verified. S3: Based on the pre-trained BERT algorithm model, supplement the insurance knowledge related corpus, combine the manually annotated sentiment semantic library, conduct targeted reinforcement training of the model, use MLM masked language modeling and NSP prediction tasks to strengthen the training of the BERT model, capture the contextual relationship and the association between sentences, match the personalized portrait of customers in the corresponding time period, and perform Fine-Tuning for sentiment classification tasks based on the BERT model; S4: The encoder part of BERT encodes the input Token corpus into tokens containing context information, and uses dynamic maximum pooling technology to convert different numbers of Tokens into representations of uniform size. S5: The representative tags obtained by the encoder are further classified by the classifier, and the input content is nonlinearly transformed and normalized through the activation function. Some neurons are randomly discarded in each training iteration. For the class imbalance problem in the sentiment classification task, the weighted cross entropy loss function is used to weight the minority class samples. S6: Based on the model regression algorithm, predict the sentiment label of each review, calculate the corresponding confidence, and revise the customer's star rating. After comparing the sentiment label with the actual evaluation, those with large differences are transferred to manual review, and the model is iterated according to the review annotations.

2. A method for evaluating emotions in customer evaluation text content according to claim 1, characterized in that: In S1, the data is screened for similarity, invalid data is deleted, and encoded. Specifically: Similarity screening: Use cosine similarity to calculate the similarity between the text content and the preset sentence, and directly assign the preset emotional label to similar comments; Invalid deletion: remove text content that is too short or does not contain enough information; Encoding: Convert text content into one-hot discrete data encoding form.

3. The method for evaluating emotions in customer evaluation text content according to claim 1, characterized in that: In S1, natural language preprocessing is performed on the remaining evaluation text, including information density assessment, sentence segmentation, and Unicode conversion, where: Information density evaluation: Count the ratio of the number of different words in the text to the total number of words, and measure the text diversity through Shannon entropy; Sentence segmentation: split the review text into meaningful short sentences; Unicode conversion: Convert text to unified Unicode encoding to ensure the consistency of text.

4. The method for evaluating emotions in customer evaluation text content according to claim 1, characterized in that: In S1, the expansion of the sentiment dictionary is based on the number of words counted, the number of times used is compared, and high-frequency words are screened. The AI ​​and human jointly determine the word attributes and necessity, and the necessary words are added to the predefined vocabulary library.

5. The method for evaluating emotions in customer evaluation text content according to claim 1, characterized in that: In S2, the identification of purchase interest and complaint tendency is as follows: Remove stop words, extract stop words that contribute less to semantic analysis, restore the text to its basic form, and identify and mark the part of speech of each word; Keyword extraction, identifying common words that represent purchase interest and complaint tendency; Text structure analysis, identifying the text expression intention through the structural forms of affirmative sentences, negative sentences, rhetorical questions, interrogative sentences and double negative sentences.

6. The method for evaluating emotions in customer evaluation text content according to claim 1, characterized in that: In S2, the customer's personalized portrait contains multi-dimensional information, including basic information, interests and hobbies, preferences, consumption habits, lifestyle, and pain points, and the changes in the customer's personalized portrait are matched and marked in chronological order.

7. The method for evaluating emotions in customer evaluation text content according to claim 1, characterized in that: In S2, the recommendation strategy is adjusted according to the annotations corresponding to the customer's big data portrait. When the customer expresses dissatisfaction, compensatory measures are recommended. When the customer expresses satisfaction, more value-added services and related products are recommended. The correctness of the annotations and recommendations is then verified based on the customer's feedback trends after the recommendation.

8. The method for evaluating emotions in customer evaluation text content according to claim 1, characterized in that: In S3, the BERT model is trained using the MLM masked language modeling and NSP prediction tasks, specifically: Input a pair of sentences connected by a special separator, input a randomly selected mask; For the MLM task, some words in the sentence are randomly masked, and the masked words are predicted based on the context information. The difference between the model predicted words and the actual words is measured by the loss function. For the NSP task, the order of sentences is determined by judging the natural connection relationship between a pair of input sentences.

9. The method for evaluating emotions in customer evaluation text content according to claim 1, characterized in that: In S3, match the customer big data portrait of the corresponding period, specifically: Predict customer interest categories based on contextual information and customer profiles; Determine the relevance of sentences and customer behaviors based on the match between the context; The customer portrait is used as an additional feature input and combined with the text content into the model, and the model performance is further enhanced through multimodal learning methods.

10. The method for evaluating emotions in customer evaluation text content according to claim 1, characterized in that: In S6, the confidence is the predicted probability of the model, and the corresponding star ratings are mapped according to positive, neutral, negative and complex emotion labels.

Citation Information

Cited By

  • Story question and answer evaluation method and system based on multi-dimensional score and emotion association

    CN121168638A

  • Data labeling and processing method and system based on natural language model

    CN121481642A