Multi-dimensional service evaluation word cloud chart generation method and device based on customer perception
By performing text compensation and subject word extraction for user comments, combined with semantic recognition processing of preset customer topic perception models, a multi-dimensional word cloud chart is generated, which solves the problem of word cloud chart data loss caused by simple expression and missing sentences in user comments, reduces the waste of computing resources and improves the accuracy of the chart.
Patent Information
- Application Number
- CN202510235196.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-28
AI Technical Summary
When generating word cloud charts, the user comments are simple and the sentences are missing, which leads to incomplete vocabulary extracted and difficult to distinguish the semantics of the vocabulary, resulting in the missing word cloud chart data, which requires regeneration and waste of computing resources.
By obtaining customer review text from the service online comment platform, text compensation is performed to improve data quality, extracting multi-dimensional topic phrases of comment text, determining preset customer topic perception models, performing semantic recognition processing, and constructing multi-dimensional word cloud charts.
It reduces the waste of computing resources, improves the accuracy of the generated multi-dimensional word cloud chart, and avoids the phenomenon of regeneration due to data loss.
Smart Images

Figure CN120163136A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technologies, and more particularly, to a method and apparatus for generating a multi-dimensional service evaluation word cloud chart based on customer perception. Background Art
[0002] Improving the service quality of merchants to better adapt to the current preferences of customers and identifying the short board of service quality that affects customer satisfaction can promote the circulation of offline goods. Currently, when generating a word cloud chart, the commonly adopted method is to use all user comments to collect all the words in the comments, and then draw a word cloud chart using all the words.
[0003] However, it has been found in practice that when generating a word cloud chart in the above manner, the following technical problems often exist:
[0004] There are often cases where the expressions in user comments are simple and sentences are missing, resulting in incomplete extraction of all words and difficulty in distinguishing the semantics of words. As a result, the generated word cloud chart is missing data and needs to be regenerated, which further leads to waste of computing resources.
[0005] The above information disclosed in this background art section is only used to enhance the understanding of the background of the inventive concept, and thus, it may include information that does not form the prior art known to those of ordinary skill in the art in this country. Summary of the Invention
[0006] This summary of the disclosure is intended to introduce concepts in a brief form, which will be described in detail in the following detailed implementation section. This summary of the disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0007] Some embodiments of the present disclosure propose a method and apparatus for generating a multi-dimensional service evaluation word cloud chart based on customer perception to solve one or more of the technical problems mentioned in the above background art section.
[0008] In a first aspect, some embodiments of the present disclosure provide a method for generating a multi-dimensional service evaluation word cloud chart based on customer perception. The method includes: obtaining customer review texts from a service online review platform to obtain a customer review text set; performing text compensation on each customer review text in the customer review text set to generate a compensated review text set; extracting text topic words from each compensated review text in the compensated review text set to generate a multi-dimensional topic word group set of review texts; determining a preset customer topic perception model corresponding to each multi-dimensional topic word group of review texts in the multi-dimensional topic word group set of review texts as the current topic perception model; performing semantic recognition processing on each multi-dimensional topic word of review texts in the multi-dimensional topic word group set of review texts based on the current topic perception model and the customer review text set to generate a recognized topic word information sequence; and constructing a multi-dimensional word cloud chart according to the recognized topic word information sequence.
[0009] In a second aspect, some embodiments of the present disclosure provide a device for generating a multi-dimensional service evaluation word cloud chart based on customer perception. The device includes: an obtaining unit configured to obtain customer review texts from a service online review platform to obtain a customer review text set; a text compensation unit configured to perform text compensation on each customer review text in the customer review text set to generate a compensated review text set; a topic word extraction unit configured to extract text topic words from each compensated review text in the compensated review text set to generate a multi-dimensional topic word group set of review texts; a model determination unit configured to determine a preset customer topic perception model corresponding to each multi-dimensional topic word group of review texts in the multi-dimensional topic word group set of review texts as the current topic perception model; a semantic recognition processing unit configured to perform semantic recognition processing on each multi-dimensional topic word of review texts in the multi-dimensional topic word group set of review texts based on the current topic perception model and the customer review text set to generate a recognized topic word information sequence; and a word cloud chart construction unit configured to construct a multi-dimensional word cloud chart according to the recognized topic word information sequence.
[0010] In a third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device storing one or more programs thereon, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect.
[0011] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium storing a computer program thereon, where the program, when executed by a processor, implements the method described in any implementation manner of the first aspect.
[0012] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the method for generating a multi-dimensional service evaluation word cloud chart based on customer perception according to some embodiments of the present disclosure, waste of computing resources can be reduced. Specifically, the reason for wasting computing resources is as follows: In user comments, there are often cases where the expressions are simple and sentences are missing, resulting in incomplete extraction of all words. At the same time, it is difficult to distinguish the semantics of words. As a result, the data of the generated word cloud chart is missing and needs to be regenerated. Based on this, in the method for generating a multi-dimensional service evaluation word cloud chart based on customer perception according to some embodiments of the present disclosure, first, customer comment texts are obtained from a service online review platform to obtain a set of customer comment texts. Then, considering that the original comment texts usually contain a large amount of non-standard content, such as spelling mistakes, special symbols, emojis, etc., which may interfere with text analysis and reduce the model effect, noise data (such as HTML tags, advertisements, redundant characters, etc.) will affect operations such as word segmentation and word frequency statistics. At the same time, a large number of meaningless characters or words will increase the time cost of data processing. Therefore, the original comments need to be cleaned. Thus, text compensation is performed on each customer comment text in the above-mentioned set of customer comment texts to generate a set of compensated comment texts. Thereby, the data quality can be improved, noise can be reduced, and the data accuracy can be improved through text compensation. Next, text topic words are extracted from each of the compensated comment texts in the above-mentioned set of compensated comment texts to generate a set of multi-dimensional topic word groups for comment texts. Here, through text topic word extraction, it can be used to determine the topic words and other words in the comments, so as to distinguish semantics and extract text features. In addition, a preset customer topic perception model corresponding to each multi-dimensional topic word group for comment texts in the above-mentioned set of multi-dimensional topic word groups for comment texts is determined as the current topic perception model. Then, based on the above current topic perception model and the above set of customer comment texts, semantic recognition processing is performed on each multi-dimensional topic word for comment texts in the above-mentioned set of multi-dimensional topic word groups for comment texts to generate a sequence of recognized topic word information. Here, by introducing a preset customer topic perception model, it can be used to judge the emotional information of topic words in comment texts according to the set of multi-dimensional topic word groups for comment texts. Thus, from the perspective of customers' comments, the topic words of customer comment texts are perceived, thereby improving the accuracy of the extracted topic words. So as to construct a multi-dimensional word cloud chart according to the above sequence of recognized topic word information. Thereby, the accuracy of the generated multi-dimensional word cloud chart can be improved. Furthermore, the computing resources wasted in regeneration can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the accompanying drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the elements and elements are not necessarily drawn to scale.
[0014] Figure 1 is a flowchart of some embodiments of a method for generating a multi-dimensional service evaluation word cloud chart based on customer perception according to the present disclosure;
[0015] Figure 2 is a result display diagram of model perplexity and model consistency data of some embodiments of a method for generating a multi-dimensional service evaluation word cloud chart based on customer perception according to the present disclosure;
[0016] Figure 3 is a schematic diagram of the model structure of a comment quality recognition model of some embodiments of a method for generating a multi-dimensional service evaluation word cloud chart based on customer perception according to the present disclosure;
[0017] Figure 4 is a schematic diagram of a comment performance analysis chart of some embodiments of a method for generating a multi-dimensional service evaluation word cloud chart based on customer perception according to the present disclosure;
[0018] Figure 5 is a schematic diagram of a positive sentiment word cloud chart and a negative sentiment word cloud chart of some embodiments of a method for generating a multi-dimensional service evaluation word cloud chart based on customer perception according to the present disclosure;
[0019] Figure 6 is a framework diagram of a customer satisfaction regression model of some embodiments of a method for generating a multi-dimensional service evaluation word cloud chart based on customer perception according to the present disclosure;
[0020] Figure 7 is a schematic diagram of the structure of some embodiments of a device for generating a multi-dimensional service evaluation word cloud chart based on customer perception according to the present disclosure;
[0021] Figure 8 is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed implementation manners
[0022] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0023] In addition, it should be noted that for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0024] It should be noted that concepts such as "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0025] It should be noted that the modifications of "one" and "multiple" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0026] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0027] The following will detail this disclosure with reference to the accompanying drawings and in conjunction with embodiments.
[0028] Figure 1 Flow 100 of some embodiments of a method for generating a multi-dimensional service evaluation word cloud chart based on customer perception according to this disclosure is shown. The method for generating a multi-dimensional service evaluation word cloud chart based on customer perception includes the following steps:
[0029] Step 101, obtain customer review texts from a service online review platform to obtain a customer review text set.
[0030] In some embodiments, the execution subject of the method for generating a multi-dimensional service evaluation word cloud chart based on customer perception can obtain customer review texts from a service online review platform in a wired or wireless manner to obtain a customer review text set. Among them, the service online review platform can be a public platform for commenting on merchants.
[0031] It should be noted that the above wireless connection methods can include but are not limited to 3G / 4G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra wideband) connections, and other currently known or future-developed wireless connection methods.
[0032] Step 102, perform text compensation on each customer review text in the customer review text set to generate a compensated review text set.
[0033] In some embodiments, the above execution subject can perform text compensation on each customer review text in the above customer review text set to generate a compensated review text set.
[0034] In some optional implementation manners of some embodiments, the above execution subject performs text compensation on each customer review text in the above customer review text set to generate a compensated review text set, which may include the following steps:
[0035] In the first step, character recognition is performed on each customer review text in the above customer review text set to remove meaningless characters, obtaining a post-removal review text set. Among them, the character identifiers in the customer review text can be indexed for removal to obtain the post-removal review text.
[0036] In the second step, using the pre-established stop word dictionary, stop word analysis is performed on each post-removal review text in the above post-removal review text set to generate a post-compensation review text set. Among them, stop word analysis is used to determine the stop words and the part-of-speech meaning information of the stop words in each post-removal review text.
[0037] In practice, the stop word dictionary contains a series of words and phrases related to the specific needs of users. During the data cleaning process, by referring to the user dictionary, key information in the text can be more accurately identified and processed. In word segmentation, the user dictionary can provide the lexical information required for word segmentation, thereby optimizing the word segmentation effect. For example, for some compound words or technical terms, the user dictionary can split them into smaller units to more accurately understand the meaning of the text. The user dictionary can also provide information such as the part of speech and meaning of the words, which helps to further improve the accuracy and efficiency of word segmentation.
[0038] Step 103, perform text topic word extraction on each post-compensation review text in the post-compensation review text set to generate a multi-dimensional topic word group set of review texts.
[0039] In some embodiments, the above execution subject can perform text topic word extraction on each post-compensation review text in the post-compensation review text set to generate a multi-dimensional topic word group set of review texts.
[0040] In some optional implementation manners of some embodiments, the above execution subject performs text topic word extraction on each post-compensation review text in the post-compensation review text set to generate a multi-dimensional topic word group set of review texts, which may include the following steps:
[0041] In the first step, perform topic word extraction on each post-compensation review text in the above post-compensation review text set to generate a current review text topic word group, obtaining a current review text topic word group set. Among them, the term frequency-inverse document frequency algorithm can be used to perform topic word extraction on each post-compensation review text in the above post-compensation review text set to generate a current review text topic word group, obtaining a current review text topic word group set.
[0042] In practice, considering that the bag-of-words model only records the number of occurrences of words and does not consider the weight of words or the positional relationship of words in the text, the term frequency (TF) and inverse document frequency (IDF) algorithms are used for word segmentation processing.
[0043] In the second step, duplicate removal is performed on the above-mentioned set of current review text topic phrases to generate a set of deduplicated text topic words.
[0044] In the third step, each deduplicated text topic word in the above-mentioned set of deduplicated text topic words is subject to topic division to generate a multi-dimensional topic phrase set of review text. Among them, the topic division can be to divide each deduplicated text topic word into any number (for example, 2 - 24 topic numbers) of different topics, and the number of deduplicated text topic words under each topic is not limited. Each multi-dimensional topic phrase of review text can correspond to a division method. Here, the multi-dimensional topic words of review text under each topic can be repeatedly divided.
[0045] As an example, an optional division method is to divide the topics into 14 topic numbers. The division results can be as shown in the following table:
[0046] Theme 1 Direct access, nice environment, catering, shopping, convenient, a lot, subway, transportation, convenient Theme 2 Center, provide, demand, nice, goods, experience, fashionable, meet, shopping, facilities Theme 3 Delicious food, discounts, clothes, brands, activities, great intensity, nice, many discounts Theme 4 Have a meal, capybara, Joy City, check in, cute, like, a lot, queue, flash mob, activity Theme 5 Experience, washroom, parking lot, queue, especially, clothes, flash mob, activity, each, cute Theme 6 Dishes, especially like, clean, taste, service, a lot, delicious, environment, nice Theme 7 Parking, shops, subway, things, snacks, check in, convenient, have a meal, activities, a lot Theme 8 Environment, air conditioner, a lot, check in, shops, often, flash mob, a lot, like, activity Theme 9 Check in, suitable, things, good place, nice, have a meal, like, big, weekend, a lot Theme 10 Queue, clear, nice, find, environment, brand, elevator, especially, a lot, things Theme 11 Often, like, brands, square, Tianshan, a lot, city, Youke, Parkson, activities Theme 12 Supermarket, like, nice, activities, department store, Taikoo, Jiuguang, a lot, brands Theme 13 Attitude, have a meal, security guard, service desk, garbage, parking fee, parking lot, park, points Theme 14 Also need, shops, nearby, points, Joy City, parking fee, clothes, activities, a lot, have a meal
[0047] In practice, in the process of adopting a technical solution to solve the problems mentioned in the background technology, there is often another technical problem 2: Even if all the words in the text can be found, it is difficult to conform to the inherent emotional value and customer satisfaction of the customer review text only by counting the importance of words according to word frequency. As a result, there is a deviation in topic words. Thus, there is also a large deviation in the finally generated multi-dimensional word cloud chart. Furthermore, it also needs to be regenerated and occupies more computing resources. In response to the above technical problem 2, the inventor decides to adopt the following solution.
[0048] Step 104, determine a preset customer topic perception model corresponding to each multi-dimensional topic phrase of review text in the multi-dimensional topic phrase set of review text as the current topic perception model.
[0049] In some embodiments, the above-mentioned execution entity can determine a preset customer topic perception model corresponding to each multi-dimensional topic phrase of review text in the multi-dimensional topic phrase set of review text as the current topic perception model.
[0050] In some optional implementation manners of some embodiments, the above-mentioned execution entity determines a preset customer topic perception model corresponding to each multi-dimensional topic phrase of review text in the multi-dimensional topic phrase set of review text as the current topic perception model, which may include the following steps:
[0051] In the first step, obtain an initial topic perception model from the database.
[0052] As an example, the initial topic perception model can be the LDA (Latent Dirichlet Allocation) unsupervised topic modeling algorithm.
[0053] In the second step, each multi-dimensional topic phrase of the above-mentioned review text multi-dimensional topic phrase set is used to perform model iteration on the above-mentioned initial topic perception model to obtain a set of topic quantity perception models. Among them, each topic quantity perception model can correspond to a multi-dimensional topic phrase of the review text in the above-mentioned review text multi-dimensional topic phrase set. Each multi-dimensional topic phrase of the review text corresponds to a different topic quantity. Each multi-dimensional topic phrase of the review text can be divided into multiple sub-groups of multi-dimensional topic words of the review text according to the topic quantity. Here, each multi-dimensional topic phrase in the above-mentioned review text multi-dimensional topic phrase set can be input into the above-mentioned initial topic perception model respectively for model iteration to train a topic quantity perception model.
[0054] In the third step, the model perplexity and model consistency data corresponding to each topic quantity perception model in the above-mentioned set of topic quantity perception models are determined. Among them, a preset test set can be used and input into each topic quantity perception model respectively to test the model perplexity and model consistency data of the topic quantity perception model.
[0055] As an example, the model perplexity and model consistency data can be as Figure 2 shown. Figure 2 The left graph in it is the graph corresponding to the model perplexity. Its horizontal axis is the topic quantity, and the vertical axis is the perplexity. The right graph is the graph corresponding to the model consistency data. Its horizontal axis is the topic quantity, and the vertical axis is the consistency value. Here, each point in the graph can correspond to the model perplexity and model consistency values of a subject quantity perception model.
[0056] In the fourth step, the topic quantity perception model corresponding to the model perplexity and model consistency data that meet the preset topic division conditions is determined as the current topic perception model. Among them, the multi-dimensional topic phrase of the review text corresponding to the above-mentioned current topic perception model is the current multi-dimensional topic phrase of the review text. Among them, the above-mentioned preset topic division conditions can be the highest comprehensive evaluation of the model perplexity and model consistency data. Specifically, it can also be seen from Figure 2 that the consistency scores of 14 topics reach local peaks, indicating an improvement in the internal semantic consistency of the topics. The model may capture a richer topic structure at this point. Although the perplexity curve has been showing an upward trend, the 14 topics are still within an acceptable range. Therefore, the multi-dimensional topic phrase of the review text with the subject quantity divided into 14 topics can be determined as the current multi-dimensional topic phrase of the review text. The corresponding subject quantity perception model is determined as the current topic perception model.
[0057] Step 105: Based on the current topic perception model and the customer review text set, perform semantic recognition processing on each multi-dimensional topic word in the multi-dimensional topic phrase set of the review text to generate an information sequence of recognized topic words.
[0058] In some embodiments, the above-mentioned execution entity may perform semantic recognition processing on each multi-dimensional topic word of the multi-dimensional topic word set of the review text based on the above-mentioned current topic perception model and the above-mentioned customer review text set, so as to generate a sequence of post-recognition topic word information.
[0059] In some optional implementation manners of some embodiments, the above-mentioned execution entity performs semantic recognition processing on each multi-dimensional topic word of the multi-dimensional topic word set of the review text based on the above-mentioned current topic perception model and the above-mentioned customer review text set, so as to generate a sequence of post-recognition topic word information, which may include the following steps:
[0060] First step, use the above-mentioned current topic perception model to perform word frequency recognition on each multi-dimensional topic word of the multi-dimensional topic word set of the review text, so as to generate a topic word frequency sequence. Among them, the above-mentioned current topic perception model can be used to extract the topic words in each compensated review text in the compensated review text set again to obtain a secondary extraction topic word set. Then, the topic words that are repeated in the secondary extraction topic word set and the above-mentioned multi-dimensional topic word set of the review text can be determined as potential topic words to obtain a potential topic word set. Finally, the word frequencies of each potential topic word in the potential topic word set can be summarized through python to obtain a topic word frequency sequence.
[0061] Second step, according to the above-mentioned topic word frequency sequence, perform sorting processing on each multi-dimensional topic word of the multi-dimensional topic word set of the review text, so as to generate a high-frequency vocabulary statistical table. Among them, the sorting processing can be performed in the order of the size of the word frequency.
[0062] Third step, determine each multi-dimensional topic word and the corresponding topic word frequency in the above-mentioned high-frequency vocabulary statistical table as post-recognition topic word information to obtain a sequence of post-recognition topic word information.
[0063] Step 106, construct a multi-dimensional word cloud chart according to the sequence of post-recognition topic word information.
[0064] In some embodiments, the above-mentioned execution entity may construct a multi-dimensional word cloud chart according to the above-mentioned sequence of post-recognition topic word information. Among them, the sequence of post-recognition topic word information can be input into a pre-set word cloud generator to generate a multi-dimensional word cloud chart.
[0065] As an example, the word cloud generator may include: WordClouds.com online word cloud generator, WordItOut word cloud generator, Easy Word Cloud generator, Micro Word Cloud generator, etc.
[0066] The optional implementation manners and related content in the above steps 104 - 106 are an inventive point of the embodiments of the present disclosure, which solve the above technical problem 2, that is, "even if all the words in the text can be found, it is difficult to conform to the inherent emotional value and customer satisfaction of the customer review text only by counting the word frequencies. As a result, there is a deviation in the topic words. Consequently, there is also a large deviation in the finally generated multi-dimensional word cloud chart. Further, it also needs to be regenerated, thus consuming more computing resources." The factors that lead to more computing resources consumption are usually as follows: even if all the words in the text can be found, it is difficult to conform to the inherent emotional value and customer satisfaction of the customer review text only by counting the word frequencies. As a result, there is a deviation in the topic words. Consequently, there is also a large deviation in the finally generated multi-dimensional word cloud chart. Further, it also needs to be regenerated. If the above factors are solved, the consumption of computing resources can be reduced. To achieve this effect, first, in order to better extract topic words according to customer emotions, an initial topic perception model is introduced. Here, in order to better distinguish the topic words in the text, by classifying the topic words, the topic words can be randomly assigned to various topic combination manners. Then, the optimal topic combination manner is determined through the initial topic perception model. Thus, accurate potential topic words and topic classification manners can be selected. Finally, it can be used to generate a more accurate multi-dimensional word cloud chart. Further, the consumption of computing resources can be reduced.
[0067] In practice, in the process of adopting the technical solution to solve the problems mentioned in the background art, there is often the following technical problem 3: since it is difficult to finely classify the customer's evaluation of the merchant service, it is difficult to construct an accurate positive sentiment word cloud chart and negative sentiment word cloud chart, and it is also difficult to determine the generation accuracy of the above model. As a result, errors occur in the stored multi-dimensional word cloud charts and customer satisfaction data tables, thus occupying storage space. In response to the above technical problem 3, the inventor decides to adopt the following solution.
[0068] Optionally, the above execution entity can also execute the following steps:
[0069] First step, use the pre-trained review quality recognition model to perform text recognition on each compensated review text in the above compensated review text set to generate a negative review text label group and a positive review text label group. Among them, the compensated review text can be input into the review quality recognition model to obtain a negative review text label and a positive review text label. Here, the negative review text label and the positive review text label can be classification labels of topic words.
[0070] As an example, the review quality recognition model can be constructed in the following way:
[0071] Specifically, considering the problem that the severe imbalance of the dataset affects model recognition, to solve this problem, the BCEWithLogitsLoss loss function is used, which is mainly used for multi-label binary classification tasks. And the label weights are introduced through the pos_weight parameter, thus solving the problem of label imbalance. The specific calculation method of the label weights is as follows: by calculating the number of positive class samples for each label, the weight of each label is calculated, and then it is applied in the loss function.
[0072] Based on BERT-Base-Chinese as the basic architecture, a multi-label classification head is customized to enable the model output layer to adapt to multi-label binary classification tasks. When the model is initialized, num_labels = 28 is set, and each label uses the sigmoid activation function for independent prediction.
[0073] Regarding tokenization, BertTokenizer is used for tokenization to ensure that the input text is converted into an input format suitable for the pre-trained model based on the Transformer encoder of BERT (Bi Direction Encoder Representations from Transformers). By setting padding = True and truncation = True, it is ensured that the input sequence has a unified length (the maximum length is 512) and does not exceed the maximum limit.
[0074] In addition, the model introduces early stopping callbacks to avoid overfitting during training. If the performance on the validation set no longer improves, the training will be automatically terminated. In practice, the weighted F1 score is used as the main metric for model evaluation, which is suitable for multi-label problems because it considers the performance of each label and can make adjustments for imbalanced data. The secondary metric is Hamming Accuracy, which measures the accuracy by calculating whether each label is correctly predicted.
[0075] Finally, in terms of model configuration and scalability: through the id2label and label2id mappings, the ID and name of the label are associated to facilitate the model to process and output labels, and the problem_type of the model is set to multi_label_classification to make it suitable for multi-label classification problems. Finally, the model will output the probability value of each label.
[0076] Here, the AdamW optimizer is adopted for model optimization, combined with the Weight Decay strategy, which can effectively prevent overfitting and accelerate convergence. Through the backpropagation mechanism, the model parameters are updated and the classification performance is optimized. Thus, the model is not only applicable to multi-label classification tasks, but also can be migrated to other text classification tasks, such as single-label classification or multi-task learning, by adjusting the classification head and loss function.
[0077] As an example, the model structure of the comment quality recognition model can be referred to Figure 3 as shown.
[0078] Optionally, the comment quality recognition model can be generated through the following steps:
[0079] First, after constructing 14 fine-grained dimensions of service quality, 7706 pieces of data are extracted for manual annotation. Each comment is marked with the relevant dimensions that appear, and scored according to the positive and negative tendencies of the sentiment. The sentiment tendency is divided into 2 categories: positive sentiment is labeled as 1, negative sentiment is labeled as 2, and those not mentioning this variable dimension are marked as 0. There are 3 classification values in total: {0, 1, 2}. During the annotation process, the established variable definitions and keyword libraries are strictly followed to ensure the accuracy of sentiment classification.
[0080] Then, feature variable decomposition is carried out: after annotation, in order to improve the accuracy of the model's learning of samples, all variables are decomposed into binary classification tasks. For example, the facility {0, 1, 2} is decomposed into facility positive {0, 1} and facility negative {0, 1}. This process enables the model to focus on the discrimination of a single sentiment direction during prediction, improving the classification accuracy. At this time, the 14 underlying feature variables are split into 28 underlying feature variables according to positive and negative.
[0081] After that, each input embedding is a combination of three embeddings. When BERT preprocesses the input text, it represents the relative position of each word in the sentence through position embeddings, distinguishes different sentences by taking the sentence as the task input through segment embeddings, and helps the model understand the structure of the input through token embeddings. The pre-training tasks of BERT are divided into two main tasks: Masked Language Model (MLM) and Next Sentence Prediction (NSP). In the MLM task, some words are randomly replaced with [MASK], and the model needs to predict the replaced words, enabling BERT to capture the context and grammatical relationships of the vocabulary. In the NSP task, two sentences are given, and the model needs to judge whether the second sentence logically follows the first sentence. This task helps BERT understand the logical and semantic relationships between sentences.
[0082] As an example, the data marked in this article is divided into a training set, a validation set, and a test set according to a ratio of 8:1:1, and the difference values TP, TN, FP, and FN between the data results returned by the statistical algorithm and the manually marked results are calculated. Among them: TP (True positive) True positive example: The number of samples where the model predicts positive and the label is actually positive; TN (True negative) True negative example: The number of samples where the model predicts negative and the label is actually negative; FP (False positive) False positive example: The number of samples where the model predicts positive and the label is actually negative; FN (False negative) False negative example: The number of samples where the model predicts negative and the label is actually positive; Based on TP, TN, FP, and FN, accuracy, precision, recall, and F1 value are calculated as evaluation indicators. Among them: Accuracy is the ratio of the number of samples predicted correctly by the model to the total number of samples.
[0083] In the second step, for each comment variable dimension corresponding to each current comment text topic word subgroup in the above current comment text topic phrase, the following steps are performed:
[0084] Step 1: Determine each compensated comment text in the above compensated comment text set that corresponds to the above current comment text topic word subgroup as the target comment text group.
[0085] Step 2: Determine the number of each target comment text in the above target comment text group as the total number of sentiment tendency comments.
[0086] Step 3: Determine the number of target comment texts corresponding to the above positive comment text label group in the above target comment text group as the number of positive sentiment tendency comments.
[0087] In the third step, according to the total number of sentiment tendency comments and the number of positive sentiment tendency comments corresponding to the comment variable dimension, a comment performance analysis chart for each comment variable dimension is established. Among them, the horizontal axis of the above comment performance analysis chart represents the importance of the comment variable dimension, and the vertical axis represents the satisfaction of the comment variable dimension.
[0088] As an example, refer to Figure 4 . Figure 4The horizontal axis in it represents performance, measured by the frequency of occurrence, that is, the proportion of the total sum of the sentiment tendencies of the variables in this dimension appearing in the overall population, denoted as I (importance). The vertical axis represents importance, measured by the positive review rate, that is, the proportion of the positive sentiment tendency of the variable in this dimension in the total sum of the sentiment tendencies of the variable in this dimension, denoted as P (positive). The formulas are as follows: I = a / c. P = g / a. Where c is the total number of all reviews in the sample, a is the total number of reviews with sentiment tendencies for a certain dimension variable (total number of sentiment tendency reviews), and g is the total number of reviews with positive sentiment tendencies for a certain dimension variable (number of positive sentiment tendency reviews). Import the data in the table into SPSS, and both the horizontal axis and the vertical axis use the average value of each score as the reference line, and are divided into an opportunity area, a maintenance area, a strength area, and a repair area.
[0089] Step 4: According to the above-mentioned comment performance analysis chart, establish positive sentiment word clouds and negative sentiment word clouds for each dimension of the comment variables, and store the above-mentioned comment performance analysis chart, each positive sentiment word cloud, and each negative sentiment word cloud. Among them, for each dimension of the comment variable, corresponding positive sentiment word clouds and negative sentiment word clouds can be generated.
[0090] As an example, taking the dimension of the comment variable "facilities" as an example, the generated positive sentiment word cloud and negative sentiment word cloud can be as Figure 5 shown.
[0091] Step 5: Determine the performance coordinates corresponding to each compensated comment text in the above-mentioned compensated comment text set in the above-mentioned comment performance analysis chart, and use the performance coordinates to label each compensated comment text to obtain a labeled comment text set. Here, the performance coordinates can be the horizontal and vertical coordinate values in the comment performance analysis chart. Labeling means marking the performance coordinates at the position of the corresponding word in the text.
[0092] Step 6: Input the above-mentioned labeled comment text set into a preset customer satisfaction regression model to generate a customer satisfaction data table corresponding to each dimension of the comment variable.
[0093] As an example, the framework diagram of the customer satisfaction regression model can be as Figure 6As shown, various data can be input into the customer satisfaction regression model to generate a customer satisfaction data table. Additionally, a person correlation test can be conducted for all comment variable dimensions. From the test results, it can be seen that there is no significant correlation between promotional activities and tangibility, personnel interaction, reliability, hedonicity, and policies. Among the variable coefficients with significant correlations, the values are all relatively low, and there is no high correlation, usually not causing a high degree of multicollinearity. Therefore, regression analysis can continue. To reduce the impact of collinearity between independent variables, the independent variables are standardized, the moderating variables are centered, and the interaction terms are constructed using the product of the standardized independent variables and the de-centered moderating variables. Subsequently, multiple linear regression analysis is performed. Using the hierarchical regression method, the control variables are placed in the first layer, all standardized independent variables are placed in the second layer, the moderating variables are placed in the third layer, and the interaction terms are placed in the fourth layer. That is, it corresponds to four models.
[0094] As an example, the equation of the customer satisfaction regression model can be: Consumer satisfaction = 0.069 for dining + 0.045 for promotional activities + 0.123 for tangibility + 0.083 for hedonicity + 0.057 for reliability + 0.055 for policies + 0.061 for personnel interaction + 0.053 for shopping mall type + 0.051 (hedonicity * shopping mall type) - 0.062 (dining * shopping mall type) + 0.425. Here, the significance levels of personnel interaction and tangibility are 0.107 and 0.124 respectively, not reaching the traditional significance level, showing a trend of marginal significance, and they are not temporarily included in the regression equation.
[0095] In practice, from the statistical results, there is no collinearity among the variables in the four models. In Model 1, the impacts of summer and autumn on the score are significant. The regression coefficient of autumn is lower than that of summer. The impact of winter on the score does not reach the commonly used significance level. The R-squared value of the model is 0.006, meaning that summer, autumn, and winter can explain 0.6% of the reasons for the change in the score. When conducting an F-test on the model, it is found that the model passes the F-test (F = 3.030, p < 0.05), that is to say, at least one of summer, autumn, and winter will have an impact on the score. The regression coefficient value of summer is 0.121 and shows significance (t = 2.804, p = 0.005 < 0.01). The regression coefficient value of autumn is 0.101 and shows significance (t = 2.431, p = 0.015 < 0.05), indicating that summer and autumn will have a significant positive impact on the score. The regression coefficient value of winter is 0.076 and does not show significance (t = 1.937, p = 0.053 > 0.05), indicating that winter will not have an impact on the score. In Model 2, after adding independent variables such as tangibility, policy, personnel interaction, reliability, hedonic, catering, and promotional activities, all control variables and independent variables are significant, indicating that the addition of independent variables improves the explanatory power of the model for the dependent variable, and the selection of independent variables is reliable. The R-squared value rises from 0.006 to 0.320, and the change in the F-value shows significance (p < 0.05). The independent variables contribute 31.4% of the explanatory power to the score and all have a significant positive impact on the score. The degrees of influence of the respective independent variables on the dependent variable from high to low are: tangibility, hedonic, personnel interaction, policy, catering, promotional activities, reliability. In Model 3, on the basis of Model 2, the type of shopping mall is added. The change in the F-value shows significance (p < 0.05). After adding the type of shopping mall, it has explanatory significance for the model. The R-squared value rises from 0.320 to 0.323. The type of shopping mall contributes 0.3% of the explanatory power to the score. The selection of the moderating variable is meaningful and will have a positive impact on the score. In Model 4, the interaction terms of the respective independent variables and the moderating variable are added. The change in the F-value shows significance (p < 0.05) and has explanatory significance for the model. The R-squared value rises from 0.323 to 0.331 and contributes 0.8% of the explanatory power to the score. After adding the moderating effect, the degrees of influence of the respective independent variables on the dependent variable also change. The degrees of influence from high to low are: tangibility, hedonic, catering, personnel interaction, policy, reliability, promotional activities. Among them, only the interaction term of catering and the type of shopping mall shows significance and has a significant negative impact on the score. The interaction terms of personnel interaction (p = 0.107) and hedonic (p = 0.085) and the type of shopping mall show marginal significance.
[0096] Step 7: Store each customer satisfaction data table and send each customer satisfaction data table to the target terminal for display.
[0097] The above optional steps and their related content are an inventive point of the embodiments of the present disclosure, which solves the above technical problem 3: "Since it is difficult to finely divide customers' evaluations of merchant services, it is difficult to construct accurate positive and negative sentiment word clouds, and it is also difficult to determine the generation accuracy of the above model. As a result, errors occur in the stored word cloud charts and customer satisfaction data tables, thus occupying storage space." The factors that lead to more storage space occupation are often as follows: Since it is difficult to finely divide customers' evaluations of merchant services, it is difficult to construct accurate positive and negative sentiment word clouds, and it is also difficult to determine the generation accuracy of the above model. As a result, errors occur in the stored word cloud charts and customer satisfaction data tables. If the above factors are solved, the occupation of storage space can be reduced. To achieve this effect, first, through text recognition, it can be used to distinguish the sentiment tendency of keywords in the text, that is, negative review text labels and positive review text labels. Then, the review performance analysis chart of each review variable dimension can be determined according to these labels. This characterizes the potential customer intentions reflected by customer reviews. So that the generated positive and negative sentiment word clouds can be more accurate. Here, performance coordinates are also used for marking, which can be used to provide deeper text association features for the introduced customer satisfaction regression model. Thus, the generated customer satisfaction data table conforms to user reviews. Furthermore, the accuracy of the generated customer satisfaction data table is improved. Therefore, there is no need for repeated generation and storage, reducing the occupation of storage resources.
[0098] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the method for generating a multi-dimensional service evaluation word cloud chart based on customer perception in some embodiments of the present disclosure, waste of computing resources can be reduced. Specifically, the reason for the waste of computing resources is as follows: In user comments, there are often cases where the expressions are simple and sentences are missing, resulting in incomplete extraction of all words. At the same time, it is difficult to distinguish the semantics of words. As a result, the data of the generated word cloud chart is missing and needs to be regenerated. Based on this, in some embodiments of the present disclosure, for the method for generating a multi-dimensional service evaluation word cloud chart based on customer perception, first, customer comment texts are obtained from a service online review platform to obtain a customer comment text set. Then, considering that the original comment texts usually contain a large amount of non-standard content, such as spelling mistakes, special symbols, emoticons, etc., which may interfere with text analysis and reduce the model effect, and noise data (such as HTML tags, advertisements, redundant characters, etc.) will affect operations such as word segmentation and word frequency statistics. At the same time, a large number of meaningless characters or words will increase the time cost of data processing. Therefore, it is necessary to clean the original comments. Therefore, text compensation is performed on each customer comment text in the above-mentioned customer comment text set to generate a compensated comment text set. Thus, the data quality can be improved, noise can be reduced, and the data accuracy can be improved through text compensation. Next, text topic words are extracted from each compensated comment text in the above-mentioned compensated comment text set to generate a multi-dimensional topic word group set of comment texts. Here, through the extraction of text topic words, it can be used to determine the topic words and other words in the comments, so as to distinguish semantics and extract text features. In addition, a preset customer topic perception model corresponding to each multi-dimensional topic word group of comment texts in the above-mentioned multi-dimensional topic word group set of comment texts is determined as the current topic perception model. Then, based on the above-mentioned current topic perception model and the above-mentioned customer comment text set, semantic recognition processing is performed on each multi-dimensional topic word of the multi-dimensional topic word group set of comment texts to generate a recognized topic word information sequence. Here, by introducing a preset customer topic perception model, it can be used to judge the emotional information of topic words in the comment texts according to the multi-dimensional topic word group set of comment texts. Thus, from the perspective of customers' comments, the topic words of the customer comment texts can be perceived, thereby improving the accuracy of the extracted topic words. So as to construct a multi-dimensional word cloud chart according to the above-mentioned recognized topic word information sequence. Thus, the accuracy of the generated multi-dimensional word cloud chart can be improved. Furthermore, the computing resources wasted by regeneration can be reduced.
[0099] Further referring to Figure 7 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a multi-dimensional service evaluation word cloud chart generation device based on customer perception. These device embodiments correspond to Figure 1 the method embodiments shown, and the device can be specifically applied to various electronic devices.
[0100] As Figure 7 shown, the multi-dimensional service evaluation word cloud chart generation device 700 based on customer perception in some embodiments includes: an acquisition unit 701, a text compensation unit 702, a topic word extraction unit 703, a model determination unit 704, a semantic recognition processing unit 705, and a word cloud chart construction unit 706. Among them, the acquisition unit 701 is configured to obtain customer review texts from a service online review platform to obtain a customer review text set; the text compensation unit 702 is configured to perform text compensation on each customer review text in the above customer review text set to generate a compensated review text set; the topic word extraction unit 703 is configured to perform text topic word extraction on each compensated review text in the above compensated review text set to generate a multi-dimensional topic word group set of review texts; the model determination unit 704 is configured to determine a preset customer topic perception model corresponding to each multi-dimensional topic word group of review texts in the above multi-dimensional topic word group set of review texts as the current topic perception model; the semantic recognition processing unit 705 is configured to perform semantic recognition processing on each multi-dimensional topic word of review texts in the above multi-dimensional topic word group set of review texts based on the above current topic perception model and the above customer review text set to generate a recognized topic word information sequence; the word cloud chart construction unit 706 is configured to construct a multi-dimensional word cloud chart according to the above recognized topic word information sequence.
[0101] It can be understood that the units described in the device 700 correspond to the respective steps in the method described with reference Figure 1 to. Thus, the operations, features, and beneficial effects described above for the method also apply to the device 700 and the units included therein, and will not be repeated here.
[0102] Next, with reference Figure 8 , which shows a schematic structural diagram of an electronic device (e.g., a computing device) 800 suitable for implementing some embodiments of the present disclosure. Figure 8 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present disclosure.
[0103] As Figure 8 shown, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which may perform various appropriate actions and processes according to a program stored in the read-only memory 802 or a program loaded from the storage device 808 into the random access memory 803. In the random access memory 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, the read-only memory 802, and the random access memory 803 are connected to each other through a bus 804. The input / output interface 805 is also connected to the bus 804.
[0104] Generally, the following devices can be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 can allow the electronic device 800 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 8 the electronic device 800 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices can be alternatively implemented or had. Figure 8 Each block shown in
[0105] Specifically, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such some embodiments, the computer program can be downloaded and installed from a network through the communication device 809, or installed from the storage device 808, or installed from the read-only memory 802. When the computer program is executed by the processing device 801, the above functions defined in the methods of some embodiments of the present disclosure are performed.
[0106] It should be noted that the computer-readable media described in some embodiments of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0107] In some embodiments, the client and the server may communicate using any currently known or future-developed network protocol such as HTTP (Hyper Text Transfer Protocol), and may be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0108] The above computer-readable medium may be included in the above electronic device; or it may exist independently without being assembled into the electronic device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to: obtain customer review text from an online service review platform to obtain a customer review text set; perform text compensation on each customer review text in the above customer review text set to generate a compensated review text set; extract text topic words from each compensated review text in the above compensated review text set to generate a multi-dimensional topic phrase set of review texts; determine a preset customer topic perception model corresponding to each multi-dimensional topic phrase of the review texts in the above multi-dimensional topic phrase set of review texts as the current topic perception model; perform semantic recognition processing on each multi-dimensional topic word of the review texts in the above multi-dimensional topic phrase set of review texts based on the above current topic perception model and the above customer review text set to generate a recognized topic word information sequence; construct a multi-dimensional word cloud chart according to the above recognized topic word information sequence.
[0109] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++; and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0110] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0111] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes an acquisition unit, a text compensation unit, a subject word extraction unit, a model determination unit, a semantic recognition processing unit, and a word cloud map construction unit. Among them, the names of these units do not constitute a limitation to the unit itself in some cases. For example, the acquisition unit can also be described as "the unit for acquiring customer review texts".
[0112] The functions described above can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0113] The above description is only some preferred embodiments of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, technical solutions formed by mutually replacing the above features with technical features having similar functions (but not limited to) disclosed in the embodiments of the present disclosure.
Claims
1. A method for generating a multi-dimensional service evaluation word cloud chart based on customer perception, comprising: Obtain customer review texts from the service online review platform to obtain a customer review text set; Performing text compensation on each customer review text in the customer review text set to generate a compensated review text set; Extracting text keywords from each post-compensation comment text in the post-compensation comment text set to generate a multi-dimensional keyword group set of comment texts; Determine a preset customer topic perception model corresponding to each review text multidimensional topic phrase in the review text multidimensional topic phrase set as the current topic perception model; Based on the current topic perception model and the customer comment text set, semantic recognition processing is performed on each comment text multidimensional topic word in the comment text multidimensional topic word group set to generate a recognized topic word information sequence; A multi-dimensional word cloud diagram is constructed based on the identified subject word information sequence.
2. The method according to claim 1, wherein: The method further comprises: The multidimensional word cloud diagram is stored, and the multidimensional word cloud diagram is sent to a display terminal for display.
3. The method according to claim 1, wherein: The text compensation is performed on each customer review text in the customer review text set to generate a compensated review text set, including: Performing character recognition on each customer review text in the customer review text set to remove meaningless characters, thereby obtaining a review text set after removal; Using a pre-established stop word dictionary, stop word analysis is performed on each removed comment text in the removed comment text set to generate a compensated comment text set, wherein the stop word analysis is used to determine the stop words in each removed comment text and the part-of-speech meaning information of the stop words.
4. The method according to claim 3, wherein: The extracting of text keywords from each post-compensation comment text in the post-compensation comment text set to generate a multi-dimensional keyword group set of comment texts includes: Performing subject word extraction on each post-compensation comment text in the post-compensation comment text set to generate a current comment text subject word group, thereby obtaining a current comment text subject word group set; De-duplication processing is performed on the current comment text subject word set to generate a de-duplication text subject word set; The individual deduplicated text keywords in the deduplicated text keyword set are divided into topics to generate a set of multi-dimensional keyword groups of the review text.
5. The method according to claim 1, wherein: The step of determining a preset customer topic perception model corresponding to each review text multidimensional topic phrase group in the review text multidimensional topic phrase group set as the current topic perception model includes: Obtain an initial topic-aware model from the database; Using each review text multidimensional subject word in the review text multidimensional subject word group set, the initial subject perception model is iterated to obtain a subject quantity perception model set, wherein each subject quantity perception model corresponds to a review text multidimensional subject word group in the review text multidimensional subject word group set, each review text multidimensional subject word group corresponds to a different number of subjects, and each review text multidimensional subject word group is divided into a plurality of review text multidimensional subject word subgroups according to the number of subjects; Determine the model perplexity and model consistency data corresponding to each topic quantity perception model in the topic quantity perception model set; The topic quantity perception model corresponding to the model perplexity and model consistency data that meet the preset topic division conditions is determined as the current topic perception model, wherein the multidimensional topic phrase group of the comment text corresponding to the current topic perception model is the current comment text topic phrase group.
6. The method according to claim 1, wherein: The method of performing semantic recognition processing on each multidimensional subject word of the review text in the multidimensional subject word group set of the review text based on the current subject perception model and the customer review text set to generate a sequence of recognized subject word information includes: Using the current topic perception model, frequency recognition is performed on each multidimensional topic word of the comment text in the multidimensional topic word group set of the comment text to generate a topic word frequency sequence; According to the frequency sequence of the subject words, sorting the multi-dimensional subject words of the review text in the main multi-dimensional subject word group set of the review text to generate a high-frequency word statistics table; The multi-dimensional subject words of each review text and the corresponding subject word frequencies in the high-frequency word statistics table are determined as the identified subject word information to obtain the identified subject word information sequence.
7. The method according to claim 6, wherein: The step of constructing a multidimensional word cloud chart according to the identified subject word information sequence includes: The high-frequency vocabulary statistics table is used to draw a word cloud for the multi-dimensional subject words of the review text included in each recognized subject word information in the recognized subject word information sequence to obtain a multi-dimensional word cloud chart.
8. A device for generating a multi-dimensional service evaluation word cloud chart based on customer perception, comprising: An acquisition unit is configured to acquire customer review texts from a service online review platform to obtain a customer review text set; A text compensation unit, configured to perform text compensation on each customer review text in the customer review text set to generate a compensated review text set; A subject word extraction unit is configured to extract text subject words from each post-compensation comment text in the post-compensation comment text set to generate a multi-dimensional subject word group set of the comment text; A model determination unit is configured to determine a preset customer topic perception model corresponding to each review text multidimensional topic phrase group in the review text multidimensional topic phrase group set as a current topic perception model; A semantic recognition processing unit is configured to perform semantic recognition processing on each review text multidimensional subject word in the review text multidimensional subject word group set based on the current subject perception model and the customer review text set, so as to generate a recognized subject word information sequence; The word cloud diagram construction unit is configured to construct a multi-dimensional word cloud diagram according to the identified subject word information sequence.
9. An electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A computer readable medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Integrated evaluation method for E-commerce service quality
CN108446813A
Artificial Intelligence Based Method and Apparatus for Constructing Comment Graph
US20180349355A1