Word cloud generation method and device based on commodity questionnaire comment text

Through the review effectiveness scoring and text similarity clustering algorithm based on the BERT model, synonymous evaluation phrases are merged to generate word clouds of advantages and suggestions, which solves the problems of word segmentation and insufficient semantic aggregation in the existing technology and realizes accurate analysis of multi-dimensional user reviews.

CN120705325APending Publication Date: 2025-09-26SHANGHAI QUZHI NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510810709.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

When generating word clouds, existing technologies use word segmentation to separate the relationship between topics and evaluation words, resulting in insufficient semantic aggregation and an inability to fully display different perspectives of reviews and users' multi-dimensional evaluations of products.

Method used

By building a review effectiveness scoring model based on the BERT model and manual rules, merging synonymous evaluation phrases, and using a clustering algorithm based on text similarity and cosine similarity of embedding word vectors to perform word segmentation and grouping, and through manual inspection and correction, generate a word cloud of advantages and suggestions.

Benefits of technology

It significantly improves the information aggregation and expression integrity of word clouds, accurately distinguishes the positive and negative evaluation dimensions of users on products, solves the parsing problems of typos, missing punctuation and ambiguous emotional orientation in complex review texts, and provides e-commerce platforms with efficient and intuitive multi-dimensional user demand analysis tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705325A_ABST
    Figure CN120705325A_ABST
Patent Text Reader

Abstract

The invention discloses a word cloud generation method and device based on a commodity questionnaire comment text. The method comprises the steps that a comment effectiveness scoring model is constructed through a BERT model and an artificial rule, an original comment text is cleaned, invalid characters are deleted, supplementary punctuations are predicted, and a standardized text is generated; segmenting the text into short sentences, removing stop words, extracting to-be-collected segmented words which are not matched with a preset word bank, screening the segmented words containing subject words and evaluation words, and storing the screened segmented words into an advantage / suggestion comment set; extracting an evaluation word segmentation list based on a simplified text rule, and converting the evaluation word segmentation list into representative key word segmentation through a similar word segmentation grouping dictionary; clustering and grouping in combination with text similarity and word vector cosine similarity, and generating a similar word segmentation dictionary after manual correction; and accumulating word frequencies of the same group and carrying out classified statistics to generate advantage word cloud and suggested word cloud. According to the method, the problems of semantic segmentation and synonym dispersion of traditional word cloud are solved, and high-precision visual support is provided for commodity optimization and user demand analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data mining technology, and in particular to a method and device for generating a word cloud based on product questionnaire review text. Background Art

[0002] Existing comment text analysis technologies primarily focus on identifying keywords in comments and analyzing the sentiment of these keywords. Methods used include dependency syntax, TD-IDF, Word2Vec, LDA algorithms, and BERT models to identify comment topics and sentiment and obtain evaluation segmentation for comment topics.

[0003] At present, existing technologies usually use a word segmenter to segment the review text, and then use a set of blocked words to exclude the words that do not meet the requirements. The words segmented by the word segmenter are all individual words, which will sever the relationship between the subject words and the evaluation words. Secondly, the focus of existing technologies is to judge the emotional tendencies of subject words and evaluations, and to score the emotions in the reviews, but rarely pay attention to the meaning of the evaluation words themselves. In addition, some technical studies take all the emotional words in the reviews as a whole and judge the similarity between the subject words and other subject words. The above methods do not meet the technical focus of generating word clouds. The review analysis word cloud needs to analyze the relationship between all the review words of a target product in a targeted manner, as well as how to comprehensively present different angles of the reviews.

[0004] Therefore, how to invent and develop a word cloud generation method to comprehensively display different angles of comments and show the various qualities of products that users care about has become an urgent problem to be solved. Summary of the Invention

[0005] To this end, the present invention provides a word cloud generation method and device based on product questionnaire review text. By merging synonymous evaluation phrases and accumulating word frequencies, the present invention solves the problems of word segmentation separating the relationship between the subject and the evaluation words and insufficient semantic aggregation in existing word cloud technology, and significantly improves the complete expressiveness of the word cloud for the multi-dimensional evaluation of the product.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for generating a word cloud based on product questionnaire review text, comprising:

[0007] Based on the BERT model and manual rules, a review effectiveness scoring model is constructed and trained. The review effectiveness scoring model is used to clean and screen the target product review text for effectiveness, generating standardized review text.

[0008] Segmenting the standardized review text by punctuation to generate short review text sentences; removing stop words from the short review text sentences, and performing matching screening through a set vocabulary to generate a review word set to be collected; screening the review word set to be collected according to set screening conditions, and adding the screened review word sets to the merit review set and the suggestion review set;

[0009] Simplifying the standardized review text according to text simplification rules to obtain a simplified review text; identifying an evaluation segmentation list from the simplified review text using a text similarity-based matching algorithm based on the advantage review set and the suggestion review set; converting the evaluation segmentation in the evaluation segmentation list into representative review segmentation according to a set similar segmentation grouping dictionary;

[0010] The representative review segmentation words are grouped by a clustering algorithm based on text similarity and cosine similarity of embedding word vectors, and the grouping results are manually checked and corrected to generate a similar segmentation grouping dictionary;

[0011] The word frequencies of the comment segmentations in the same group in the similar segmentation grouping dictionary are accumulated to obtain a word frequency list of the evaluation segmentations; based on the word frequency list, the word frequencies of the advantage key segmentations and the suggestion key segmentations are counted respectively to generate the advantage word cloud and the suggestion word cloud.

[0012] As a preferred solution for a word cloud generation method based on product questionnaire review text, in the process of performing text cleaning and effectiveness screening on the target product review text using the review effectiveness scoring model:

[0013] The expressions, meaningless characters and repeated texts in the target product review text are deleted, and the complete Chinese review is retained to obtain the validity review text; for the text without punctuation in the validity review text, the punctuation prediction model is used to predict and supplement the punctuation to generate the standardized review text.

[0014] As a preferred solution of a word cloud generation method based on product questionnaire review text, in the process of screening the review word set to be collected according to the set screening conditions:

[0015] The set screening conditions are: the length of the comment segmentation is within the set length range, and the comment segmentation contains the comment object and the corresponding comment word.

[0016] As a preferred solution of a word cloud generation method based on product questionnaire review text, in the process of simplifying the normalized review text according to the simplified text rules:

[0017] Punctuation, repeated text and meaningless text are removed from the standardized review text, and Chinese characters, key English characters and numeric characters are retained; the meaningless text is filtered through a stop word library.

[0018] As a preferred solution of a word cloud generation method based on product questionnaire review text, in the process of identifying the evaluation word list from the simplified review text by the text similarity-based matching algorithm:

[0019] The text similarity restriction condition in the text similarity-based matching algorithm is that the Levenshtein distance is not less than a set threshold.

[0020] The present invention also provides a word cloud generation device based on product questionnaire review text, based on the above word cloud generation method based on product questionnaire review text, comprising:

[0021] The review text cleaning module is used to build and train a review effectiveness scoring model based on the BERT model and manual rules; the review effectiveness scoring model is used to clean and screen the target product review text for effectiveness, generating standardized review text;

[0022] The new word discovery module is used to segment the standardized review text according to punctuation to generate short review text sentences; remove stop words from the short review text sentences, and perform matching and screening through a set vocabulary to generate a set of review segmentation words to be collected; screen the set of review segmentation words to be collected according to set screening conditions, and add the screened review segmentation words to the advantage review set and the suggestion review set;

[0023] A word segmentation matching module is configured to simplify the normalized review text according to simplified text rules to obtain a simplified review text; based on the merit review set and the suggestion review set, identify an evaluation word list from the simplified review text using a text similarity-based matching algorithm; and convert the evaluation word list into representative review word list according to a set similar word grouping dictionary;

[0024] A word segmentation grouping module is used to group the representative review words using a clustering algorithm based on text similarity and cosine similarity of embedding word vectors, and to manually check and correct the grouping results to generate a similar word segmentation grouping dictionary;

[0025] The word cloud generation module is used to accumulate the word frequencies of the comment segmentations in the same group in the similar segmentation grouping dictionary to obtain a word frequency list of the evaluation segmentations; based on the word frequency list, the word frequencies of the advantage key segmentations and the suggestion key segmentations are respectively counted to generate the advantage word cloud and the suggestion word cloud.

[0026] As a preferred solution for a word cloud generation device based on product questionnaire review text, in the review text cleaning module, in the process of text cleaning and effectiveness screening of the target product review text through the review effectiveness scoring model, the expressions, meaningless characters, and repeated texts in the target product review text are deleted, and the complete Chinese reviews are retained to obtain the effectiveness review text; for the text without punctuation in the effectiveness review text, the comment punctuation prediction model is used to predict and supplement punctuation to generate the standardized review text.

[0027] As a preferred solution for a word cloud generation device based on product questionnaire review text, in the new word discovery module, in the process of screening the review word set to be collected according to the set screening conditions, the set screening conditions are: the length of the review word is within the set length range, and the review word contains the review object and the corresponding review word.

[0028] As a preferred solution for a word cloud generation device based on product questionnaire review text, in the word segmentation matching module, in the process of simplifying the standardized review text according to the simplified text rules, the standardized review text is removed from punctuation, repeated text and meaningless text, and Chinese characters, key English characters and numeric characters are retained; the meaningless text is filtered through a stop word library.

[0029] As a preferred solution for a word cloud generation device based on product questionnaire review text, in the word segmentation matching module, in the process of identifying the evaluation word segmentation list from the simplified review text through the text similarity-based matching algorithm, the text similarity restriction condition in the text similarity-based matching algorithm is: the Levenstein distance is not less than a set threshold.

[0030] The present invention has the following advantages: by taking the segmentation of the combination of subject words and evaluation words as the basic unit, combined with the intelligent cleaning and dynamic new word extraction mechanism of the BERT model, the present invention effectively solves the problems of semantic fragmentation and redundant dispersion of synonyms in traditional word cloud technology, and can automatically merge synonymous expressions such as "very full of bubbles" and "sufficient amount of bubbles" into unified key segmentations and accumulate word frequencies when processing online review texts, significantly improving the information aggregation and expression integrity of the word cloud; at the same time, through the classification of advantage / suggestion review sets and manually corrected segmentation clustering strategies, accurately distinguishing the positive and negative evaluation dimensions of users on products, taking into account semantic accuracy and natural language habits, solving the parsing problems of typos, lack of punctuation and ambiguous emotional orientation in complex review texts, and providing e-commerce platforms and market research scenarios with efficient and intuitive multi-dimensional user demand analysis tools, facilitating product optimization decisions and precision marketing. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are merely exemplary, and those skilled in the art can, without inventive effort, derive other implementation drawings based on the provided drawings.

[0032] The structures, proportions, sizes, etc. illustrated in this specification are intended solely to complement the contents disclosed herein and to facilitate understanding and reading by persons skilled in the art. They are not intended to limit the conditions under which the present invention may be implemented and therefore have no substantive technical significance. Any structural modifications, changes in proportions, or adjustments in sizes, without affecting the efficacy and objectives of the present invention, shall remain within the scope of the technical contents disclosed herein.

[0033] Figure 1 This is a flow chart of a method for generating a word cloud based on product questionnaire review text provided in Example 1 of the present invention;

[0034] Figure 2 A word cloud schematic diagram is generated in a possible embodiment provided in Example 1 of the present invention;

[0035] Figure 3 This is a schematic diagram of the architecture of a word cloud generation device based on product questionnaire review text provided in Example 2 of the present invention. DETAILED DESCRIPTION

[0036] The following describes the implementation of the present invention using specific embodiments. Those skilled in the art will readily understand the other advantages and benefits of the present invention from the disclosure herein. Obviously, the embodiments described are only a portion of the present invention, not all of it. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.

[0037] Example 1

[0038] See also Figure 1 Embodiment 1 of the present invention provides a word cloud generation method based on product questionnaire review text, comprising the following steps:

[0039] S1. Build and train a review effectiveness scoring model based on the BERT model and manual rules; use the review effectiveness scoring model to clean and screen the target product review text for effectiveness, generating standardized review text;

[0040] S2. Segment the normalized review text by punctuation to generate review text short sentences; remove stop words from the review text short sentences, and perform matching screening through a set vocabulary to generate a review word set to be collected; screen the review word set to be collected according to set screening conditions, and add the screened review word sets to the advantage review set and the suggestion review set;

[0041] S3. Simplify the normalized review text according to the simplified text rules to obtain a simplified review text; based on the merit review set and the suggestion review set, identify an evaluation segmentation list from the simplified review text using a text similarity-based matching algorithm; and convert the evaluation segmentations in the evaluation segmentation list into representative review segmentations according to a set similar segmentation grouping dictionary;

[0042] S4. Grouping the representative review segmentation words by a clustering algorithm based on text similarity and cosine similarity of embedding word vectors, and manually checking and correcting the grouping results to generate a similar segmentation grouping dictionary;

[0043] S5. Accumulate the word frequencies of the comment segmentations in the same group in the similar segmentation grouping dictionary to obtain a word frequency list of the evaluation segmentations; based on the word frequency list, count the word frequencies of the advantage key segmentations and the suggestion key segmentations respectively to generate an advantage word cloud and a suggestion word cloud.

[0044] In this embodiment, in step S1, a review effectiveness scoring model is constructed and trained based on the BERT model and manual rules; the review text of the target product is cleaned and effectiveness screened using the review effectiveness scoring model to generate standardized review text;

[0045] Specifically, a review effectiveness scoring model is constructed through the BERT model (such as RoBERTa-base); the review effectiveness scoring model is trained through manual rules to obtain a trained review effectiveness scoring model; the target product review text is subjected to binary classification screening through the trained review effectiveness scoring model, invalid reviews (such as text with a score lower than 0.8) are deleted, and manual rules are used to remove emoticons, meaningless characters (such as "#¥%") and repeated text; at the same time, a punctuation prediction model based on the BERT+CRF architecture is used to automatically insert punctuation into long texts without punctuation (for example, "moisturizes well and absorbs quickly" is corrected to "moisturizes well and absorbs quickly") to generate grammatically standardized review text.

[0046] In this embodiment, in step S2, the normalized review text is segmented by punctuation to generate short review text sentences; the short review text sentences are de-stopped and matched and screened through a set vocabulary to generate a set of review word segments to be collected; the set of review word segments to be collected is screened according to set screening conditions, and the screened review word segments are added to the advantage review set and the suggestion review set;

[0047] Specifically, the normalized text is segmented into short sentences by punctuation marks (such as full stops, commas), de-stopped (such as "de", "le"), and then matched and screened through a preset vocabulary (including known topic words and evaluation words) to identify the word segments to be collected that are not matched (such as "humidity retention for 8h↑") and generate a set of review word segments to be collected; qualified word segments are screened according to preset conditions (the length of the review word segment is within the set length range and contains both a topic word and an evaluation word), and the qualified word segments are classified and stored in the advantage review set (such as "long-lasting moisturizing effect") or the suggestion review set (such as "high price").

[0048] In a possible embodiment, the new word discovery results are shown in Table 1:

[0049]

[0050] Table 1 New word discovery results

[0051] In this embodiment, in step S3, the normalized review text is simplified according to the simplified text rules to obtain a simplified review text; based on the advantage review set and the suggestion review set, an evaluation word segment list is identified from the simplified review text through a matching algorithm based on text similarity; the evaluation word segments in the evaluation word segment list are converted into representative review word segments according to a set similar word segment grouping dictionary;

[0052] Specifically, the normalized review text is simplified according to the simplified text rules. The main operations include removing all punctuation, only retaining Chinese characters, key English characters, and numerical characters, identifying and reducing repeated text, and using a stop word library to reduce meaningless text in the text. After simplification, a simplified review text is obtained; then, an evaluation word segment is identified from the simplified text using a matching algorithm based on text similarity and the pre-prepared advantage review set and suggestion review set. For example, "long moisturizing effect" is matched to "long-lasting moisturizing effect"; where the text similarity limit condition in the matching algorithm is: Levenshtein distance >= specified threshold (tentatively set to 0.8 in this invention).

[0053] During the matching process, the sets of merit and suggestion reviews are pre-compressed into a dictionary of the form {'keyword': review word list}, for example: {"long-lasting moisturizing effect":["moisturizing for a long time","moisturizing for a long time"]}. This narrows the set of review word segments used in each review. The identified review word lists are then filtered, with duplicate removal and the longest segmentation selected for similar text. After the review word lists are processed, a pre-prepared dictionary of similar segmentation groups is read to convert the identified segmentations into representative review word segments for their group. Product-specific associated segmentation lists can also be set to additionally match featured product reviews, and a product-specific blacklist of segmentations can be used to remove irrelevant segmentations.

[0054] In this embodiment, in step S4, the representative review segmentation words are grouped by a clustering algorithm based on text similarity and cosine similarity of embedding word vectors, and the grouping results are manually checked and corrected to generate a similar segmentation grouping dictionary;

[0055] Specifically, the present invention adopts a two-dimensional clustering strategy: first, the text similarity (edit distance) of the segmented words and the cosine similarity of the word vectors generated by BERT are calculated (threshold ≥ 0.85), and a hierarchical clustering algorithm is used to merge similar segmented words (such as "fast absorption" and "rapid absorption" are grouped together), and then the semantic deviation groups are split through manual inspection (such as distinguishing "cute packaging" from "fashionable packaging"), and finally a grouping dictionary in the format of {key segmented words: synonym list} is generated.

[0056] In this embodiment, sometimes, in order to make the word cloud more beautiful and to filter the word cloud segmentations, it is also necessary to adjust the segmentation combinations and display several combinations separately.

[0057] In a possible embodiment, the similar word segmentation grouping dictionary is shown in Table 2:

[0058]

[0059]

[0060] Table 2 Word segmentation and grouping dictionary

[0061] In this embodiment, in step S5, the word frequencies of the comment segmentations in the same group in the similar segmentation grouping dictionary are accumulated to obtain a word frequency list of the evaluation segmentations; based on the word frequency list, the word frequencies of the advantage key segmentations and the suggestion key segmentations are counted respectively to generate the advantage word cloud and the suggestion word cloud.

[0062] Specifically, based on the grouping dictionary, the word frequencies of all word segments in the same group are accumulated (such as "moisturizing for a long time" 20 times + "moisturizing for a long time" 25 times = "moisturizing effect lasting" 45 times), and the cumulative frequencies of advantage key words (such as "moisturizing effect lasting") and suggestion key words (such as "price is high") are counted separately for each product. The WordCloud library is used to render the font and color according to the word frequency size (green for the advantage word cloud and red for the suggestion word cloud) to generate a visual word cloud.

[0063] In a possible embodiment, the advantages word cloud data is shown in Table 3:

[0064]

[0065]

[0066] Table 3 Advantages Word Cloud Data In a possible embodiment, it is recommended that word cloud data be as shown in Table 4:

[0067]

[0068]

[0069] Table 4 Suggested word cloud data

[0070] In a possible example, based on the hundreds of product experience questionnaires distributed and collected, the answers to the two open questions of product advantages and product suggestions in the questionnaire are collected into the business database, and the present invention is used to generate the following word cloud: Figure 2 shown.

[0071] In summary, the present invention builds and trains a review effectiveness scoring model based on the BERT model and artificial rules; performs text cleaning and effectiveness screening on the target product review text through the review effectiveness scoring model to generate a standardized review text; segments the standardized review text according to punctuation to generate short review text sentences; removes stop words from the short review text sentences, and performs matching screening through a set vocabulary to generate a review word set to be collected; screens the review word set to be collected according to the set screening conditions, and adds the screened review word sets to the advantage review set and the suggestion review set; simplifies the standardized review text according to the simplified text rules to obtain a simplified review text; based on the advantage The review set and the suggested review set are used to identify an evaluation word list from the simplified review text through a matching algorithm based on text similarity; the evaluation word lists in the evaluation word list are converted into representative review word lists according to a set similar word grouping dictionary; the representative review word lists are grouped through a clustering algorithm based on text similarity and cosine similarity of embedding word vectors, and the grouping results are manually checked and corrected to generate a similar word grouping dictionary; the word frequencies of the review word lists in the same group in the similar word grouping dictionary are accumulated to obtain a word frequency list of the evaluation word lists; based on the word frequency list, the word frequencies of the advantage key word lists and the suggestion key word lists are respectively counted to generate an advantage word cloud and a suggestion word cloud. The present invention uses the segmentation of the combination of subject words and evaluation words as the basic unit, combined with the intelligent cleaning and dynamic new word extraction mechanism of the BERT model, to effectively solve the problems of semantic fragmentation and redundant dispersion of synonyms in traditional word cloud technology. When processing online review texts, it can automatically merge synonymous expressions such as "very full of bubbles" and "sufficient amount of bubbles" into unified key segmentations and accumulate word frequencies, significantly improving the information aggregation and expression integrity of the word cloud; at the same time, through the classification of advantage / suggestion review sets and manually corrected segmentation clustering strategies, it accurately distinguishes the positive and negative evaluation dimensions of users on products, taking into account semantic accuracy and natural language habits, and solves the parsing problems of typos, lack of punctuation and ambiguous sentiment in complex review texts, providing e-commerce platforms and market research scenarios with an efficient and intuitive multi-dimensional user demand analysis tool, facilitating product optimization decisions and precision marketing.

[0072] It should be noted that the method of the embodiments of the present disclosure can be performed by a single device, such as a computer or server. The method of the embodiments of the present disclosure can also be applied in a distributed scenario, where multiple devices cooperate to perform the method. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiments of the present disclosure, and the multiple devices will interact with each other to complete the method.

[0073] It should be noted that the above description is limited to some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0074] Example 2

[0075] See also Figure 3 Embodiment 2 of the present invention further provides a word cloud generation device based on product questionnaire review text, comprising:

[0076] Review text cleaning module 001 is used to build and train a review effectiveness scoring model based on the BERT model and manual rules; the review effectiveness scoring model is used to clean and screen the target product review text for effectiveness, generating standardized review text;

[0077] New word discovery module 002 is used to segment the standardized review text according to punctuation to generate review text short sentences; remove stop words from the review text short sentences, and match and filter them through a set vocabulary to generate a review word set to be collected; filter the review word set to be collected according to set filtering conditions, and add the filtered review word sets to the advantage review set and the suggestion review set;

[0078] The word segmentation matching module 003 is used to simplify the normalized review text according to the simplified text rules to obtain a simplified review text; based on the merit review set and the suggestion review set, identify an evaluation word list from the simplified review text using a text similarity-based matching algorithm; and convert the evaluation word list into representative review word according to a set similar word grouping dictionary;

[0079] The word segmentation grouping module 004 is used to group the representative review word segments using a clustering algorithm based on text similarity and cosine similarity of embedding word vectors, and manually check and correct the grouping results to generate a similar word segmentation grouping dictionary;

[0080] The word cloud generation module 005 is used to accumulate the word frequencies of the comment segmentations in the same group in the similar segmentation grouping dictionary to obtain a word frequency list of the evaluation segmentations; based on the word frequency list, the word frequencies of the advantage key segmentations and the suggestion key segmentations are respectively counted to generate the advantage word cloud and the suggestion word cloud.

[0081] In this embodiment, in the review text cleaning module 001, in the process of performing text cleaning and effectiveness screening on the target product review text through the review effectiveness scoring model, expressions, meaningless characters, and repeated texts in the target product review text are deleted, and the complete Chinese review is retained to obtain the effectiveness review text; for the text without punctuation in the effectiveness review text, punctuation prediction model is used to predict and supplement punctuation to generate the standardized review text.

[0082] In this embodiment, in the new word discovery module 002, in the process of filtering the review word set to be collected according to the set filtering conditions, the set filtering conditions are: the length of the review word is within the set length range, and the review word contains the review object and the corresponding review word.

[0083] In this embodiment, in the word segmentation matching module 003, in the process of simplifying the standardized comment text according to the simplified text rules, punctuation, repeated text and meaningless text are removed from the standardized comment text, and Chinese characters, key English characters and numeric characters are retained; the meaningless text is filtered through a stop word library.

[0084] In this embodiment, in the word segmentation matching module 003, in the process of identifying the evaluation word segmentation list from the simplified review text through the text similarity-based matching algorithm, the text similarity restriction condition in the text similarity-based matching algorithm is: the Levenstein distance is not less than the set threshold.

[0085] It should be noted that the information interaction, execution process, etc. between the modules of the above-mentioned system are based on the same concept as the method embodiment in Example 1 of the present application, and the technical effects they bring are the same as those of the method embodiment of the present application. For specific contents, please refer to the description in the method embodiment shown above in the present application, and no further details will be given here.

[0086] Example 3

[0087] Embodiment 3 of the present invention provides a non-transitory computer-readable storage medium, in which a program code for a word cloud generation method based on product questionnaire review text is stored. The program code includes instructions for executing embodiment 1 or any possible implementation thereof, a word cloud generation method based on product questionnaire review text.

[0088] Computer-readable storage media can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

[0089] Example 4

[0090] Embodiment 4 of the present invention provides an electronic device, including: a memory and a processor;

[0091] The processor and the memory communicate with each other via a bus; the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute a word cloud generation method based on product questionnaire review text in embodiment 1 or any possible implementation thereof.

[0092] Specifically, the processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor, which is implemented by reading software code stored in a memory. The memory can be integrated into the processor or located outside the processor and exist independently.

[0093] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable systems. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode.

[0094] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing system. They can be centralized on a single computing system or distributed across a network of multiple computing systems. Alternatively, they can be implemented using program code executable by a computing system, and thus, they can be stored in a storage system and executed by the computing system. In some cases, the steps shown or described herein can be performed in a different order than that shown, or they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.

[0095] Although the present invention has been described in detail above using general descriptions and specific embodiments, it will be apparent to those skilled in the art that modifications and improvements may be made thereto. Therefore, such modifications and improvements, without departing from the spirit of the present invention, are intended to be within the scope of protection claimed herein.

Claims

1. A word cloud generation method based on product questionnaire review text, characterized in that: include: Build and train a review effectiveness scoring model based on the BERT model and manual rules; Using the review effectiveness scoring model, the target product review text is cleaned and effectiveness screened to generate standardized review text; Segmenting the standardized review text by punctuation to generate short review text sentences; removing stop words from the short review text sentences, and performing matching screening through a set vocabulary to generate a review word set to be collected; screening the review word set to be collected according to set screening conditions, and adding the screened review word sets to the merit review set and the suggestion review set; Simplifying the standardized review text according to the simplified text rules to obtain a simplified review text; Based on the set of merit reviews and the set of suggestion reviews, an evaluation segmentation list is identified from the simplified review text using a text similarity-based matching algorithm; and the evaluation segmentation words in the evaluation segmentation list are converted into representative review segmentation words according to a set similar segmentation grouping dictionary; The representative review segmentation words are grouped by a clustering algorithm based on text similarity and cosine similarity of embedding word vectors, and the grouping results are manually checked and corrected to generate a similar segmentation grouping dictionary; Accumulate the word frequencies of the review word segments in the same group in the similar word grouping dictionary to obtain a word frequency list of the evaluation word segments; Based on the word frequency list, the word frequencies of the advantage key words and the suggestion key words are counted respectively to generate an advantage word cloud and a suggestion word cloud.

2. A word cloud generation method based on product questionnaire review text according to claim 1, characterized in that: In the process of performing text cleaning and effectiveness screening on the target product review text using the review effectiveness scoring model: The expressions, meaningless characters and repeated texts in the target product review text are deleted, and the complete Chinese review is retained to obtain the validity review text; for the text without punctuation in the validity review text, the punctuation prediction model is used to predict and supplement the punctuation to generate the standardized review text.

3. The word cloud generation method based on product questionnaire review text according to claim 2 is characterized in that: In the process of screening the review word set to be collected according to the set screening conditions: The set screening conditions are: the length of the comment segmentation is within the set length range, and the comment segmentation contains the comment object and the corresponding comment word.

4. The word cloud generation method based on product questionnaire review text according to claim 3 is characterized in that: In the process of simplifying the standardized review text according to the simplified text rules: Punctuation, repeated text and meaningless text are removed from the standardized review text, and Chinese characters, key English characters and numeric characters are retained; the meaningless text is filtered through a stop word library.

5. The word cloud generation method based on product questionnaire review text according to claim 4 is characterized in that: In the process of identifying the evaluation word list from the simplified review text by using the text similarity-based matching algorithm, the text similarity restriction condition in the text similarity-based matching algorithm is that the Levenstein distance is not less than a set threshold.

6. A word cloud generation device based on product questionnaire review text, using the word cloud generation method based on product questionnaire review text according to any one of claims 1 to 5, characterized in that: include: The review text cleaning module is used to build and train a review effectiveness scoring model based on the BERT model and manual rules; Using the review effectiveness scoring model, the target product review text is cleaned and effectiveness screened to generate standardized review text; The new word discovery module is used to segment the standardized review text according to punctuation to generate short review text sentences; remove stop words from the short review text sentences, and perform matching and screening through a set vocabulary to generate a set of review segmentation words to be collected; screen the set of review segmentation words to be collected according to set screening conditions, and add the screened review segmentation words to the advantage review set and the suggestion review set; A word segmentation matching module is used to simplify the normalized review text according to simplified text rules to obtain a simplified review text; Based on the set of merit reviews and the set of suggestion reviews, an evaluation segmentation list is identified from the simplified review text using a text similarity-based matching algorithm; and the evaluation segmentation words in the evaluation segmentation list are converted into representative review segmentation words according to a set similar segmentation grouping dictionary; A word segmentation grouping module is used to group the representative review words using a clustering algorithm based on text similarity and cosine similarity of embedding word vectors, and to manually check and correct the grouping results to generate a similar word segmentation grouping dictionary; A word cloud generation module is used to accumulate the word frequencies of the review word segments in the same group in the similar word segmentation grouping dictionary to obtain a word frequency list of the evaluation word segments; Based on the word frequency list, the word frequencies of the advantage key words and the suggestion key words are counted respectively to generate an advantage word cloud and a suggestion word cloud.

7. The word cloud generation device based on product questionnaire review text according to claim 6, characterized in that: In the review text cleaning module, expressions, meaningless characters, and repeated texts in the target product review text are deleted, and the complete Chinese reviews are retained to obtain the effective review text; for the text without punctuation in the effective review text, the comment punctuation prediction model is used to predict and supplement punctuation to generate the standardized review text.

8. The word cloud generation device based on product questionnaire review text according to claim 7, characterized in that: In the new word discovery module, the set screening conditions are: the length of the comment segmentation is within the set length range, and the comment segmentation contains the comment object and the corresponding comment word.

9. The word cloud generation device based on product questionnaire review text according to claim 8, characterized in that: In the word segmentation matching module, punctuation marks, repeated texts and meaningless texts are removed from the normalized review texts, and Chinese characters, key English characters and numeric characters are retained; the meaningless texts are screened through a stop word library.

10. The word cloud generation device based on product questionnaire review text according to claim 9, characterized in that: In the word segmentation matching module, the text similarity restriction condition in the text similarity-based matching algorithm is: the Levenstein distance is not less than a set threshold.