A satisfaction evaluation method considering user attribute preference

CN115828914BActive Publication Date: 2026-09-29HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310026834.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-09
Publication Date
2026-09-29
Estimated Expiration
2043-01-09

AI Technical Summary

Technical Problem

但已有利用评论文本分析满意度的方法中至少存在如下问题:首先,大部分方法的实施需要足够的标注数据,无形地增加了分析消费者满意度的成本;其次,部分方法分析的结果主要是消费者满意度的分类情况,即不满意或者满意,而无法了解消费者对产品各方面具体的满意程度;再次,部分方法更多的是同等看待消费者对各产品属性的满意度情况,而忽视了消费者对各产品属性有着不同的偏好;最后,已有方法也较少地关注商家的真正需求,缺少对特定产品属性满意度情况的重点分析

Benefits of technology

[0045]1、本发明提出的一种考虑用户属性偏好的满意度评估方法,通过预先设置兴趣属性的种子词集,并考虑消费者属性偏好的加权,同时使用种子主题模型与文本情感分析相结合的方法,实现了商家可以根据自身需求了解消费者多方面满意度的功能效果,不仅降低了数据标注所带来的较高成本,还有效提升了满意度分析结果的精准度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115828914B_ABST
    Figure CN115828914B_ABST
Patent Text Reader

Abstract

The application discloses a satisfaction evaluation method considering user attribute preference, and the steps include: 1, obtaining the review text of the product to be analyzed and pre-processing; 2, constructing the seed word set of the interest attribute; 3, constructing the feature word set of the interest attribute and the attribute preference vector of the consumer based on the review text and the seed topic model; 4, constructing the consumer sentiment value vector of the interest attribute based on the attribute feature word set and the sentiment dictionary; 5, generating the consumer satisfaction result based on the attribute preference vector and the sentiment value vector. The application considers the influence of the user attribute preference, can effectively measure the consumer satisfaction of the specific product attribute, and has high application value for understanding the consumer feedback and promoting the product improvement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of big data analysis and natural language processing, specifically to a machine learning-based method for analyzing customer satisfaction with e-commerce products. Background Technology

[0002] Consumer satisfaction refers to the difference between a consumer's expected quality of a product before a purchase and their perceived quality after purchase. According to relevant research, consumer satisfaction has a significant positive correlation with consumer loyalty, repurchase rate, and market share. Confirming consumer satisfaction levels plays a crucial role in optimizing product design, improving service content, gaining market insights, and increasing customer stickiness.

[0003] Traditional methods for measuring consumer satisfaction primarily involve questionnaire analysis. Related patents mainly divide consumer satisfaction into different indicators, then develop corresponding scales for statistical analysis to obtain the final satisfaction level. One specific example involves an online method for statistically analyzing group satisfaction. However, questionnaire surveys often incur high data collection costs, and the resulting data volume is relatively low and somewhat unstable. In recent years, with the development of e-commerce and big data technologies, many patents have begun to quantify consumer satisfaction based on diversified data: for example, using historical business data and fitting user satisfaction models to objectively assess consumer satisfaction, specifically involving big data-based user satisfaction assessment methods and devices. However, such methods only obtain an overall satisfaction score and cannot fully understand consumers' satisfaction with various aspects. Another approach uses voice and facial data to infer consumer satisfaction by analyzing features such as voice and facial expressions, specifically involving a voice recognition-based emotion analysis system and an expression-based satisfaction analysis method. However, voice and facial data are often difficult to obtain and pose certain privacy risks.

[0004] Furthermore, online reviews are increasingly becoming key data for understanding satisfaction levels. In particular, review texts describe various aspects of consumers' product experience, providing businesses with more authentic and convenient market information. This is a crucial source of information for identifying potential customers, attracting new customers, and managing existing customers. Existing patents primarily utilize machine learning and other technologies to obtain consumer satisfaction data through sentiment analysis, specifically involving a machine learning-based method for analyzing customer satisfaction in e-commerce products. However, existing methods for analyzing satisfaction using review texts suffer from at least the following problems: First, most methods require sufficient labeled data, which increases the cost of analyzing consumer satisfaction. Second, some methods primarily categorize consumer satisfaction as either satisfied or dissatisfied, failing to reveal the specific level of satisfaction consumers have with various aspects of the product. Third, some methods treat consumer satisfaction with each product attribute equally, neglecting the fact that consumers have different preferences for each attribute. Finally, existing methods pay less attention to the true needs of businesses and lack focused analysis of satisfaction with specific product attributes. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention proposes a satisfaction assessment method that considers user attribute preferences. This method aims to effectively utilize review texts to understand consumers' satisfaction with various aspects of a product and takes into account differences in consumer preferences for different product attributes. This allows for a more accurate representation of consumer satisfaction with each product attribute, helping businesses to understand consumers' post-purchase reactions at a low cost, thereby promoting product improvement and service upgrades.

[0006] To achieve the above-mentioned objectives, the present invention adopts the following technical solution:

[0007] The present invention provides a satisfaction assessment method that considers user attribute preferences, characterized by comprising:

[0008] Step A: Obtain the original review text set D' of the product to be analyzed, and preprocess each review text in the original review text set D' to obtain the standard text set. Where, N D Represents the number of preprocessed comment texts, u∈[1,N] D ];d u Let represent the preprocessed comment text of the u-th comment, and w un d u The nth word in N u d u The total number of words in the text, n∈[1,N] u ];

[0009] Step B: Construct a seed word set for interest attributes Where, N S s represents the number of interest attributes. j Let represent the seed word set for the j-th interest attribute, and Let m be the word used to characterize the j-th interest attribute, where m∈[1,M] and j∈[1,N]. S ];

[0010] Step C: Construct a feature word set for interest attributes and a vector of consumer attribute preferences;

[0011] Step C1: Set N r A regular topic and N S A collection of themes consisting of seed themes r k Let s represent the k-th regular topic. j Let T represent the j-th seed topic, and let N be the number of topics in T. T =N r +N s k∈[1,N r ],j∈[1,N S ];

[0012] make Represents the k-th regular topic r k The word distribution, and Follows the parameter β r Dirichlet distribution Dir(β) r ), and have in, This indicates that the v-th word is assigned to the k-th regular topic r. k The probability of , and v∈[1,V], where V represents the total number of unique words in the text set D;

[0013] make Represents the j-th seed topic s j The word distribution, and Obtained by parameter β s Dirichlet distribution Dir(β) s ), and have in, This indicates that the v-th word is assigned to the j-th seed topic s. j The probability of , and v∈[1,V], where V represents the total number of unique words in the text set D;

[0014] Step C2: Let π t Indicates the word distribution from seed topics Word distribution in relation to conventional topics The probability of drawing a word and assigning it to the t-th topic, and π t It follows a beta distribution Beta(1,1) with all parameters 1, and t∈[1,N] T Let θ u This represents the text of the uth comment. u The topic distribution, and θ u It follows a Dirichlet distribution Dir(α) with parameter α;

[0015] From parameter θ u The multinomial distribution Mult(θ) u Sample and generate d from ) u The nth word w un Theme z un , i.e. z un ~Mult(θ) u Define indicator variable x; un Obtain the parameter as Bernoulli distribution If x un If the value is 0, then the parameter is multinomial distribution Generate the nth word w un ,Right now If x un If the value is 1, then the parameter is multinomial distribution Generate the nth word w un ,Right now

[0016] Step C3: When the number of topics in the set T is N T When taking G different values, calculate N using equation (1). T perturbation level for the g-th value g , g∈[1,G]:

[0017]

[0018] In equation (1), N u This indicates the uth text d u The total number of words in w uv d u The vth unique word in the text, p g (w uv ) represents N T d at the g-th value u The vth unique word w uv The probability of occurrence;

[0019] Calculate N using equation (2) TThematic consistency coh under the g-th value g , g∈[1,G]:

[0020]

[0021] In equation (2), p g (w v1 ) represents N T The v1-th word w under the g-th value v1 The probability of occurrence, p g (w v1 ,w v2 ) represents N T The v1-th word w under the g-th value v1 Other words w v2 The probability of them occurring together;

[0022] For all degrees of perplexity {coh g |g=1,2,…,G} and topic consistency {coh g The indicator trend chart is obtained by visualizing the data in the sequence |g=1,2,…,G}. Based on the indicator trend chart, H topics are selected, and the document-topic probability matrix under the H topics is obtained. and topic-keyword probability matrix The probability matrix The text of the uth comment in the article d u The probabilities under the H topics are denoted as... d u The probability under the h-th topic, h∈[1,H], u∈[1,N] D ]; probability matrix The probability of the word in the h-th topic is denoted as Let v represent the probability that the v-th word belongs to the h-th topic, where h∈[1,H] and v∈[1,V].

[0023] Step C4: Match the topics corresponding to the seed word set S of interest attributes with... By matching the latent topics in the data, a probability matrix of interest attribute topics and topic words is obtained. And the probability matrix of comment text - interest attribute topic Wherein, probability matrix The probability of the word with the j-th interest attribute is denoted as . Let v represent the probability that the v-th word belongs to the j-th interest attribute, where j∈[1,N]. S ], v∈[1,V]; probability matrix The text of comment #u in the middle. u The probability of belonging to each interest attribute is denoted as . d uThe probability of the j-th interest attribute, j∈[1,N] S ];u∈[1,N D ];

[0024] Step C5: Based on the probability matrix A probability threshold y is set, and words with a probability greater than or equal to the probability threshold y under each interest attribute are selected, thus forming a set of attribute feature words for the product to be analyzed. s' j Let s' represent the word sequence of the j-th interest attribute, and let s' j ={w' j1 ,w' j2 ,…,w' jm ,…,w' jM}, w' jm This represents the m-th word used to characterize the j-th interest attribute topic;

[0025] Based on probability matrix Get the text of the uth comment d u Chinese consumers' preferences for the j-th interest attribute Therefore, equation (3) is used to calculate the average preference a of all consumers for the j-th interest attribute. j This leads to the consumer's attribute preference vector.

[0026]

[0027] Step D: Construct a sentiment vector of consumers' interest attributes;

[0028] Step D1: Select any word w from a dictionary and obtain F meanings of word w, where the fth meaning of word w includes a positive score Pos. f Negative scoring f Thus, the sentiment value μ(w) of word w is calculated using equation (4);

[0029]

[0030] Step D2: Define a sliding window of size c. For the j-th interest attribute, iterate through the u-th comment text d. u Each word in the set and its features are compared with the word set s' j Match each word in the text, when the u-th comment text d u A certain word in the word w' jm If a match is successful, then follow the sliding window of size c to d. u Search for words with the part of speech as adjectives. And as the word w' jmThe sentiment words are used to calculate the initial sentiment value of the j-th interest attribute using equation (4).

[0031] According to the sliding window of size c, at d u Search for words in Chinese degree adverb And use equation (4) to calculate the degree adverb Emotional weight Therefore, the weighted sentiment value of the j-th interest attribute is calculated.

[0032] According to the sliding window of size c, at d u Search The negative words are counted, and the number of negative words is q, thus calculating the number of negative words in the comment text d. u The final sentiment value of the j-th interest attribute

[0033] Step D3: Calculate the average sentiment value e of the j-th interest attribute among all consumers using equation (5). j This yields the consumer sentiment vector of all interest attributes of the product to be analyzed.

[0034]

[0035] Step E: Generate consumer satisfaction results;

[0036] Step E1: Calculate the percentage of consumer attribute preference for the j-th interest attribute using equation (6). j This allows us to obtain the attribute influence weight vector for consumer satisfaction.

[0037]

[0038] Step E2: Calculate the consumer satisfaction score se for the j-th interest attribute using equation (7). j This allows us to obtain the satisfaction vector of consumers regarding each interest attribute of the analyzed product.

[0039] se j =l j ×e j (7)

[0040] Calculate the overall customer satisfaction (TSE) for the analyzed product using equation (8):

[0041]

[0042] The present invention provides an electronic device, including a memory and a processor, wherein the memory is used to store a program that supports the processor in executing the satisfaction evaluation method, and the processor is configured to execute the program stored in the memory.

[0043] The present invention discloses a computer-readable storage medium on which a computer program is stored, wherein the computer program is executed by a processor to perform the steps of the satisfaction evaluation method.

[0044] Compared with existing technologies, the beneficial effects of this invention are reflected in:

[0045] 1. The present invention proposes a satisfaction assessment method that considers user attribute preferences. By pre-setting a seed word set of interest attributes and considering the weighting of consumer attribute preferences, and by using a combination of seed topic model and text sentiment analysis, the method enables businesses to understand the multi-faceted satisfaction of consumers according to their own needs. This not only reduces the high cost of data annotation, but also effectively improves the accuracy of satisfaction analysis results.

[0046] 2. This invention takes the product attributes that merchants care about as its starting point, and constructs a seed word set of interest attributes in the early stage of satisfaction assessment. The seed word set is used to supervise the topic modeling process and guide the topic modeling results, thereby forming a more focused and complete set of attribute feature words for subsequent sentiment value calculation and satisfaction assessment. This significantly solves the problem that the assessment results are too general or lack focus.

[0047] 3. This invention fully considers the impact of differences in consumer attribute preferences on satisfaction results. It constructs consumer attribute preference vectors through comment text mining and uses these attribute preference vectors to weight the sentiment value calculation results, thereby obtaining a corrected satisfaction result. This result more accurately reflects the actual satisfaction of consumers, and there is also a certain degree of comparability between different interest attributes, thus improving the reference value of the evaluation results. Attached Figure Description

[0048] Figure 1 This is a flowchart illustrating a satisfaction assessment method that considers user attribute preferences according to the present invention.

[0049] Figure 2 This is a schematic diagram of the preprocessing flow for comment text in this invention;

[0050] Figure 3 This is a schematic diagram of the feature extraction process for interest attributes in this invention;

[0051] Figure 4 This is a schematic diagram illustrating the process of calculating the sentiment value of interest attributes in this invention. Detailed Implementation

[0052] In this embodiment, a satisfaction assessment method considering user attribute preferences includes:

[0053] 1. Obtain consumer review texts of the product to be analyzed and construct a seed keyword set of interest attributes. Further, this step includes: identifying the product to be analyzed, selecting product attributes of interest as the focus of satisfaction analysis, and constructing a seed keyword set; designing a web crawler script based on online reviews of the product, capturing review texts within a suitable time period and storing them in a database, and then preprocessing the captured review text data.

[0054] 2. Construct a feature word set and attribute preference vector for interest attributes based on comment text and seed word set. Further, this step includes: using the seed word set and comment text for a seed topic model; selecting an appropriate number of topics based on the evaluation metrics of the topic model; using the resulting topic-word matrix to construct the feature word set for the interest attributes; and simultaneously, calculating the mean of the resulting document-topic matrix to obtain the attribute preference vector, which represents the consumer's preference for each interest attribute.

[0055] 3. Construct consumer sentiment value vectors for interest attributes based on sentiment analysis using feature word sets and sentiment dictionaries. Further, this step includes: designing sentiment calculation rules for comment texts; for each comment text, matching attribute words and their corresponding sentiment words, degree adverbs, negation words, etc., from the feature word set; calculating the sentiment value for each interest attribute in each consumer comment according to the sentiment calculation rules; and obtaining the consumer sentiment value vector by calculating the average value, which is used to represent the consumer's emotional attitude towards the interest attributes.

[0056] 4. Generate consumer satisfaction results based on attribute preference vectors and sentiment value vectors. Further, this step includes: calculating the influence weight of each interest attribute in satisfaction based on the attribute preference vector, and generating consumer satisfaction results based on the influence weight vector and sentiment value vector. Specifically, this satisfaction assessment method considers user attribute preferences, such as... Figure 1 As shown, it includes:

[0057] Step S101: Obtain the review text data of the product to be analyzed and perform preprocessing.

[0058] Specifically, a web crawler script is designed based on online product reviews to capture review texts within a suitable timeframe and store them in a database. For example, the Python Requests library can be used to parse webpage data to obtain the original review text set D' of the product to be analyzed. Each review text in set D' is then preprocessed to obtain a standard text set. Where, N D Represents the number of preprocessed texts, u∈[1,N] D];d u This represents the preprocessed text of the u-th line, and w un d u The nth word, N u d u The total number of words in the text, n∈[1,N] u ];

[0059] Preprocessing steps are as follows Figure 2 As shown, this can be performed using Python's Natural Language Toolkit (NLTK). Specifically, it includes comment text regularization, which removes non-English characters from the text and converts all words to lowercase; text segmentation, which converts all text sentences into a word list to facilitate subsequent recognition and statistics of related words; setting stop word lists and high-frequency word lists with less meaning, and removing these words as well as deleting words with a length of less than 1; and part-of-speech tagging, which performs word form restoration based on the part of speech of words in the original sentence, and filters words that appear less than five times in the comment text through word frequency statistics.

[0060] Step S102: Construct a seed word set for each interest attribute in the product to be analyzed.

[0061] Specifically, a seed keyword set S for interest attributes is constructed. Merchants can select product attribute themes of interest as analysis objects based on the situation of the products to be analyzed, and set specific words included in them based on experience. The seed keyword set can be denoted as... N S s represents the number of interest attributes. j Let represent the seed word set for the j-th interest attribute, and Let m be the word used to characterize the j-th interest attribute, where m∈[1,M] and j∈[1,N]. S For example, if you need to understand consumers' reactions to product prices, you can set the attribute category "price" and set the corresponding seed keyword set, which can include "price", "value", "quality", "money", etc.

[0062] Step S103: Construct the feature word set and attribute preference vector of interest attributes.

[0063] Specifically, a seed topic model can be used to facilitate the construction of attribute feature word sets. This model mainly supervises its learning process by loading pre-defined seed words, thereby effectively improving topic-word distribution and document-topic distribution. The specific generation process is as follows:

[0064] The setting is N r A regular topic and NS A collection of themes consisting of seed themes r k Let s represent the k-th regular topic. j Let T represent the j-th seed topic, and let N be the number of topics in T. T =N r +N s k∈[1,N r ],j∈[1,N S ];

[0065] make Represents the k-th regular topic r k The word distribution, and Obtained by parameter β r The Dirichlet distribution, i.e. And have in, This indicates that the v-th word is assigned to the k-th regular topic r. k Let v ∈ [1, V], where V represents the total number of unique words in the text set D; let v = 1, V ... Represents the j-th seed topic s j The word distribution, and Obtained by parameter β s The Dirichlet distribution, i.e. And have in, This indicates that the v-th word is assigned to the j-th seed topic s. j The probability of , and v∈[1,V], where V represents the total number of unique words in the text set D;

[0066] Define π t Indicates the word distribution from seed topics Word distribution in relation to conventional topics The probability of drawing a word and assigning it to the t-th topic is given by π. t It follows a beta distribution with all parameters equal to 1, i.e., π. t ~Beta(1,1), and t∈[1,N] T Let θ u This represents the text of the uth comment. u The topic distribution, and θ u It follows a Dirichlet distribution with parameter α, i.e., θ u ~Dir(α); from parameter θ u d is generated by sampling from a multinomial distribution. u The nth word w un Theme z un , i.e. z un ~Mult(θ) uDefine indicator variable x; un Obtain the parameter as The Bernoulli distribution, i.e. If x un If the value is 0, then the parameter is The nth word w is generated from the multinomial distribution. un ,Right now If x un If the value is 1, then the parameter is The nth word w is generated from the multinomial distribution. un ,Right now

[0067] Furthermore, let N be the number of topic sets T. T Taking G different values, the Gibbs sampling method is applied to the inference process of the seed topic model. Various parameters are set, and supervised learning is performed using the attribute seed word set. Specifically, the topic model needs to select an appropriate number of topics, which can be balanced by calculating the perplexity score and the topic consistency score. Perplexity is a standard method for measuring the predictive ability of a language model; a lower score indicates a higher predictive ability. Perplexity is equivalent to the reciprocal of the geometric mean likelihood of each word. N is calculated using equation (1). T perturbation level for the g-th value g , g∈[1,G]:

[0068]

[0069] In equation (1), N u This indicates the uth text d u The total number of words in w uv d u The vth unique word in the text, p g (w uv ) represents N T d at the g-th value u The vth unique word w uv The probability of occurrence; topic consistency is a commonly used indicator to measure the interpretability of a topic. If a topic is more interpretable, the words that rank higher in that topic should co-occur more frequently in the document. N is calculated using equation (2). T Thematic consistency coh under the g-th value g , g∈[1,G]:

[0070]

[0071] In equation (2), p g (w v1 ) represents N T The v1-th word w under the g-th value v1The probability of occurrence, p g (w v1 ,w v2 ) represents N T The v1-th word w under the g-th value v1 Other words w v2 The probability of co-occurrence. Specifically, considering that models with a large number of topics generally fit the data better and can support finer-grained segmentation of the text, but the number of difficult-to-interpret topics also increases, to avoid over-clustering and improve topic interpretability, for all perplexity {coh}... g |g=1,2,…,G} and topic consistency {coh g The trend chart of the indicator is obtained by visualizing |g=1,2,…,G}. Based on the trend chart, the most suitable number of topics is selected as H.

[0072] Furthermore, the document-topic probability matrix is ​​obtained when the number of topics is H. The probability matrix The text of the uth comment in the article d u The probabilities under the H topics are denoted as... d u The probability under the h-th topic, h∈[1,H], u∈[1,N] D ]; probability matrix The probability of the word in the h-th topic is denoted as Let v represent the probability that the v-th word belongs to the h-th topic, where h∈[1,H] and v∈[1,V].

[0073] Furthermore, the topics corresponding to the attribute seed word set S are associated with... By matching the latent topics in the data, a probability matrix of interest attribute topics and topic words is obtained. And the probability matrix of comment text - interest attribute topic The probability matrix The probability of the word with the j-th interest attribute is denoted as . Let v represent the probability that the v-th word belongs to the j-th interest attribute, where j∈[1,N]. S ], v∈[1,V]; probability matrix The text of comment #u in the middle. u The probability of belonging to each interest attribute is denoted as . d u The probability of the j-th interest attribute, j∈[1,N] S ],u∈[1,N D ].

[0074] Furthermore, based on the probability matrix Set a probability threshold y and filter out words with a probability greater than or equal to the threshold under each interest attribute to construct an attribute feature word set S' = {s'1, s'2, ..., s'...} for the product to be analyzed. j ,…,s' NS}, s' j The word sequence representing the j-th interest attribute can be represented as s' j ={w' j1 ,w' j2 ,…,w' jm ,…,w' jM}, w' jm This represents the m-th word used to characterize the j-th interest attribute topic; based on a probability matrix. The text of the uth comment can be obtained. u Chinese consumers' preferences for the j-th interest attribute Therefore, equation (3) is used to calculate the average preference a of all consumers for the j-th interest attribute. j This leads to the consumer's attribute preference vector. The steps in S103 are as follows: Figure 3 As shown.

[0075]

[0076] Step S104: Construct the emotional value vector of consumers for each interest attribute.

[0077] Specifically, the SentiWordNet dictionary can be used to calculate the values ​​of specific sentiment words. This dictionary is mainly developed based on the English lexical semantic web WordNet. For each word and each definition, it provides a positive sentiment score Pos, a negative sentiment score Neg, and a neutral sentiment score Obj. The values ​​of these three scores are all in the range of 0 to 1 and add up to 1. For example, the word "estimable" is considered an adjective "calculable or estimable" in the set {estimable(J,3)}. Its sentiment value in SentiWordNet is Pos = 0.0, Neg = 0.0, and Obj = 1.0. However, if "estimable" belongs to the set {estimable(J,1)}, and its meaning corresponds to "worthy of respect", then its sentiment value is Pos = 0.75, Neg = 0.0, and Obj = 0.25. Furthermore, since the earlier a meaning appears in the order of a word within a certain part of speech, the more representative that meaning is, the first meaning can be assigned a weight of 1, the second a weight of 1 / 2, and so on. Therefore, by matching any word w in the SentiWordNet dictionary, we obtain F meanings of that word, and a positive score Pos is obtained for the f-th meaning of word w. f Negative scoring fj∈[1,N S The sentiment value μ(w) of word w can be calculated using equation (4);

[0078]

[0079] Furthermore, the role of degree adverbs is considered in the sentiment analysis process. A weighted approach is adopted, which involves searching for corresponding degree adverbs for different sentiment words and then multiplying the sentiment word score by different weights according to the degree of the degree adverb to obtain a more accurate sentiment value. The role of negation words is also considered in the sentiment analysis process. When the number of negation words in the comment text is odd, the inversion method is used to invert the overall sentiment value of the text sentence. When the number of negation words is even, the overall sentiment value of the text sentence remains unchanged.

[0080] Specifically, the emotion calculation process in step S104 is as follows: Figure 4 For the j-th interest attribute, iterate through the u-th comment text d. u The words in the set and the feature word set s' j Match each word in the text; define a sliding window of size c, and for the j-th interest attribute, iterate through the u-th comment text d. u Each word in the set and its features are compared with the word set s' j Match each word in the text, when the u-th comment text d u A certain word in the word w' jm If a match is successful, then follow the sliding window of size c to d. u Search for words with the part of speech as adjectives. And as the word w' jm The sentiment words are used to calculate the initial sentiment value of the j-th interest attribute using equation (4). Secondly, according to the sliding window of size c at d u Search for words in Chinese degree adverb And use equation (4) to calculate the degree adverb Emotional weight Therefore, the weighted sentiment value of the j-th interest attribute is calculated. If no degree adverb is present, its influence weight is 1 by default; furthermore, according to a sliding window of size c at d u Zhongsou The negative words are counted, and the number of negative words is q, thus calculating the number of negative words in the comment text d. u The final sentiment value of the j-th interest attribute Finally, the average sentiment value e of the j-th interest attribute among all consumers is calculated using equation (5). jThis yields the consumer sentiment vector of all interest attributes of the product to be analyzed.

[0081]

[0082] Step S105: Generate consumer satisfaction results based on the obtained attribute preference and sentiment value vectors.

[0083] Specifically, based on the attribute preference vector obtained in step S103 The percentage of consumer attribute preference for the j-th interest attribute is calculated using equation (6). j This allows us to obtain the attribute influence weight vector for consumer satisfaction.

[0084]

[0085] Combined with the sentiment value vector obtained in step S104 and the resulting influence weight vector Calculate the consumer satisfaction score se for the j-th interest attribute using equation (7). j This allows us to obtain the satisfaction vector of consumers regarding each interest attribute of the analyzed product.

[0086] se j =l j ×e j (7)

[0087] Calculate the overall customer satisfaction (TSE) for the analyzed product using equation (8):

[0088]

[0089] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in performing the above-described satisfaction evaluation method. The processor is configured to execute the program stored in the memory.

[0090] In this embodiment, a computer-readable storage medium stores a computer program, which is executed by a processor to perform the steps of the above-described satisfaction evaluation method.

[0091] In summary, this invention integrates an unsupervised learning method based on seed words, which reduces reliance on manual data labeling while enhancing the efficiency of consumer satisfaction analysis. Furthermore, it considers an attribute preference-weighted method for calculating consumer satisfaction based on sentiment lexicon analysis, which helps to more accurately represent consumer satisfaction with products, enabling businesses to better understand consumer reactions and promote improvements to related products or services.

Claims

1. A satisfaction assessment method that considers user attribute preferences, characterized in that, include: Step A: Obtain the original review text set of the product to be analyzed. For the original set of comment texts Each comment text in the document is preprocessed to obtain a standard text set. ,in, This indicates the number of preprocessed comment texts. ; Indicates the first The preprocessed comment text, and , express The first in One word, express The total number of words in the text ; Step B: Construct a seed word set for interest attributes ,in, Indicates the number of interest attributes. Indicates the first j A seed word set for each interest attribute, and , Indicates the characterization of the first j The first interest attribute m One word, , ; Step C: Construct a feature word set for interest attributes and a vector of consumer attribute preferences; Step C1: Set by A regular theme and A collection of themes consisting of seed themes , Indicates the first A typical topic, Indicates the first A seed theme, making Number of topics , , ; make Indicates the first A regular theme The word distribution, and Obtain the parameter as Dirichlet distribution , and have ,in, Indicates the first The word was assigned to the first A regular theme The probability, and , Represents a collection of texts The total number of unique words in the text; make Indicates the first Seed Themes The word distribution, and Obtain the parameter as Dirichlet distribution , and have ,in, Indicates the first The word was assigned to the first Seed Themes The probability, and , Represents a collection of texts The total number of unique words in the text; Step C2: Let Indicates the word distribution from seed topics Word distribution in relation to conventional topics Extract words from the middle and assign them to the first The probability of each topic, and It follows a beta distribution with all parameters equal to 1. ,and ;make Indicates the first Comment text Thematic distribution, and Obtain the parameter as Dirichlet distribution ; From the parameter is multinomial distribution Sampling and generating The Middle one word Theme ,Right now Define indicator variables Obtain the parameter as Bernoulli distribution ,like If the value is 0, then the parameter is multinomial distribution The generation of the first one word ,Right now ;like If the value is 1, then the parameter is multinomial distribution The generation of the first one word ,Right now ; Step C3: When the topic set quantity Pick When there are different values, use equation (1) to calculate In the perplexity for each value , : (1) In equation (1), Indicates the first Article text The total number of words in the text express The Middle One unique word. express In the Under each value The Middle A unique word The probability of occurrence; Calculate using equation (2) In the Subject consistency under various values , : (2) In equation (2), express In the The first value under the given condition one word The probability of occurrence express In the The first value under the given condition one word Other words The probability of them occurring together; For all levels of confusion Consistency with the theme Visualize the indicator trend chart to obtain the indicator trend chart, and select based on the indicator trend chart. One topic, and received Document-topic probability matrix under each topic and topic-keyword probability matrix The probability matrix The first in Comment text exist The probability under each topic is denoted as . , express In the The probability under each topic Probability matrix The first in The word probability of each topic is denoted as: , Indicates the first The word belongs to the first The probability of each topic ; Step C4: Set the seed word set for interest attributes The corresponding theme and By matching the latent topics in the data, a probability matrix of interest attribute topics and topic words is obtained. And the probability matrix of comment text - interest attribute topic , where the probability matrix The Middle The probability of a word with an interest attribute is denoted as . , Indicates the first The word belongs to the first The probability of each interest attribute. Probability matrix The Middle Comment text The probability of belonging to each interest attribute is denoted as . , express In the The probability under each interest attribute ; ; Step C5: Based on the probability matrix Set probability threshold And filter out those with a probability greater than or equal to the probability threshold under each interest attribute. The words are used to form a set of attribute feature words for the product to be analyzed. , Indicates the first A sequence of words with interest attributes, and , Indicates the first The one used to characterize the first j Words related to interest-related themes; Based on probability matrix Get the first Comment text Chinese consumers' views on the first j Preferences for individual interest attributes Therefore, equation (3) is used to calculate the value of all consumers for the first... j Average preference for each interest attribute This leads to the consumer's attribute preference vector. : = (3) Step D: Construct a sentiment vector of consumers' interest attributes; Step D1: Select any word from a dictionary. Get words of a kind of meaning, and words The This meaning includes positive scores With negative scores Thus, the word can be calculated using equation (4). Emotional value ; (4) Step D2: Define the size as The sliding window, for the first The interest attribute is traversed. Comment text Each word in the set of feature words, and Match each word in the string, when the first word is matched... Comment text A certain word and word If a match is successful, then classify by size. The sliding window in Search for words with the part of speech as adjectives. And as a word The emotional words are used to calculate the first one using equation (4). Initial sentiment value of each interest attribute ; According to size The sliding window in Search for words in Chinese degree adverb And use equation (4) to calculate the degree adverb Emotional weight , thus calculating the first Weighted sentiment score of each interest attribute ; According to size The sliding window in Search The number of negative words is obtained by counting them. Thus, the calculation is performed in the comment text. The Middle The final sentiment value of each interest attribute ; Step D3: Calculate the first step using equation (5). The average sentiment score of each interest attribute among all consumers This yields the consumer sentiment vector of all interest attributes of the product to be analyzed. ; = (5) Step E: Generate consumer satisfaction results; Step E1: Calculate the first step using equation (6). Percentage of consumer attribute preferences for each interest attribute This allows us to obtain the attribute influence weight vector for consumer satisfaction. ; (6) Step E2: Calculate the first step using equation (7). j Consumer satisfaction with individual interest attributes This allows us to obtain the satisfaction vector of consumers regarding each interest attribute of the analyzed product. ; (7) Equation (8) is used to calculate the overall consumer satisfaction with the analyzed product. : (8)。 2. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports the processor in executing the satisfaction evaluation method of claim 1, and the processor is configured to execute the program stored in the memory.

3. A computer-readable storage medium storing a computer program, characterized in that, The computer program is executed by the processor to perform the steps of the satisfaction evaluation method of claim 1.

Citation Information

Patent Citations

  • Recommendation method based on user score analysis

    CN111061962A

  • User comment text emotion mining model for automatically extracting fine-grained attributes

    CN114564956A