A user demand classification method for internet reviews
By extracting product features and user opinion words from Internet comment data, using Bi-LSTM and hierarchical clustering algorithm, and combining Kano model to classify user needs, the problem of traditional models being difficult to apply to Internet comment data is solved, and the automated analysis and classification of user needs is realized, and the targetedness of product development is improved.
Patent Information
- Application Number
- CN202310036663.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-10
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2043-01-10
AI Technical Summary
The traditional Kano model is difficult to effectively apply to Internet review data, and it is impossible to automatically extract and classify user needs, resulting in insufficient understanding of user needs in product development.
By extracting the name, user opinion words and negative words of product features from the Internet comment data, using the Bi-LSTM model for feature extraction, combining hierarchical clustering algorithm for feature clustering, and classifying user demands through the Kano model.
It realizes the automatic extraction of user demand information from Internet comments, improves the efficiency and accuracy of user demand analysis, and helps product developers better understand and adapt to user needs.
Smart Images

Figure CN115982470B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a user demand classification method oriented to Internet comments. Background Art
[0002] With the development of information technology and the popularization of mobile Internet, more and more users are posting feedback information related to product usage experience on Internet platforms such as shopping websites and social networks. This provides big data support for product demand analysis in the information age, which is conducive to improving the efficiency and intelligence level of user demand acquisition and analysis. However, the traditional survey-based demand analysis method is significantly different from the current demand analysis method based on Internet review data in terms of data input and data output. For example, the Kano model collects customer data through questionnaires, and divides customer needs into basic needs (Must-be requirements, MR), expected needs (One-dimensional requirements, OR), attractive requirements (AR), indifferent requirements (IR) and reverse requirements (RR) according to customer satisfaction when needs are met and unmet. Figure 1 However, considering the new characteristics of Internet review data, the traditional Kano model cannot be completely copied and applied to user demand analysis based on Internet review data. Instead, a demand analysis method that is adapted to the characteristics of Internet review data should be designed.
[0003] Therefore, it is necessary to provide a method for user demand analysis based on Internet review data, to automatically extract user demand information from Internet reviews, and to classify user needs in combination with the Kano model, thereby promoting product developers to have a deep understanding and grasp of user needs, and to make corresponding subsequent planning and improvements to product features to meet user needs. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to propose a user demand classification method for Internet reviews, which can be applied to Internet review data to conduct user demand analysis on product features, and use the Kano model to divide different user demand types according to product features, thereby promoting product developers to have an in-depth understanding and grasp of user needs, and to carry out corresponding subsequent planning and improvement of product features to quickly adapt to user needs.
[0005] Based on the above objectives, the present invention provides a user demand classification method for Internet comments, comprising:
[0006] Extract user demand information from Internet review data, including: product feature names, user opinion words, and negative words;
[0007] Calculate user attitudes and discussion popularity of product features based on user opinion words and negative words of the product features;
[0008] Based on user attitudes and discussion heat of product features, the user demand classification results of the product features are obtained through the Kano model.
[0009] Preferably, the user demand classification results of the product features are obtained by using a Kano model based on user attitudes and discussion heat of the product features, specifically including:
[0010] If the user attitude towards the product feature is negative and the discussion heat of the product feature is medium, the user demand classification result of the product feature obtained by the Kano model is basic demand;
[0011] If the user attitude towards the product feature is positive and the discussion heat of the product feature is high, the user demand classification result of the product feature obtained by the Kano model is expected demand;
[0012] If the user attitude towards the product feature is positive and the discussion heat of the product feature is low, then the user demand classification result of the product feature obtained by the Kano model is attractive demand;
[0013] If the user attitude towards the product feature is medium and there are more neutral comments, and the discussion heat of the product feature is low, then the user demand classification result of the product feature obtained by the Kano model is indifferent demand;
[0014] If the user attitude towards the product feature is medium and there are few neutral comments, and the discussion heat of the product feature is low, then the user demand classification result of the product feature obtained by the Kano model is reverse demand.
[0015] Preferably, the calculating of user attitudes toward product features based on user opinion words and negative words of product features specifically includes:
[0016] User i's user attitude towards product feature j css ij As shown in formula 3:
[0017]
[0018] In formula 3, N ij represents the number of opinions of user i on product feature j; represents the lth opinion of user i on product feature j; ss(*) represents the sentiment tendency of user opinion; user opinion It is composed of user i’s negative words and opinion words about product feature j, and its sentiment tendency Calculated by the support vector machine model.
[0019] Preferably, the discussion popularity of the product feature is calculated according to the following method:
[0020] User i's discussion popularity ca for product feature j of product p ij Calculate according to the following formula 6:
[0021]
[0022] In Equation 6, k is the total number of product features of product p.
[0023] Preferably, after extracting user demand information from Internet review data, the method further includes:
[0024] Cluster product features to obtain several product feature categories.
[0025] Preferably, after calculating the user attitude towards the product feature based on the user opinion words and negative words of the product feature, the method further includes:
[0026] The user attitude of product feature categories is calculated according to the following formula 5:
[0027]
[0028] In formula 5, css i represents the calculated average of user i's attitudes toward all k product features in the product feature category.
[0029] Preferably, after calculating the user attitude and discussion heat of the product features based on the user opinion words and negative words of the product features, the method further includes:
[0030] The discussion heat of the product feature category is calculated according to the following formula 7:
[0031]
[0032] Among them, ca i Represents the average value of user i’s discussion enthusiasm for all k product features in the product feature category.
[0033] The present invention also provides an electronic device comprising a central processing unit, a signal processing and storage unit, and a computer program stored in the signal processing and storage unit and executable on the central processing unit, wherein when the central processing unit executes the program, the user demand classification method for Internet comments as described above is implemented.
[0034] The present invention also provides a computer-readable storage medium, which stores a computer program. The computer program can be executed by at least one processor to enable the at least one processor to perform the steps of the user demand classification method for Internet reviews as described above.
[0035] The technical solution of the present invention extracts user demand information from internet review data, including the name of a product feature, user opinion words, and negative words. Based on the user opinion words and negative words for the product feature, user attitudes and discussion interest for the product feature are calculated. Based on the user attitudes and discussion interest for the product feature, a user demand classification result for the product feature is obtained using the Kano model. This allows user demand analysis of product features using internet review data, and the Kano model is used to classify different user demand types for product features. This facilitates product developers' in-depth understanding and grasp of user needs, enabling subsequent planning and processing of product features to quickly adapt to user needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0037] Figure 1 A flow chart of a method for classifying user needs for Internet reviews provided by an embodiment of the present invention;
[0038] Figure 2 A schematic diagram of a framework of a demand information extraction model based on Bi-LSTM provided in an embodiment of the present invention;
[0039] Figure 3 A schematic diagram showing demand information extracted by a Bi-LSTM-based demand information extraction model provided in an embodiment of the present invention;
[0040] Figure 4 A flowchart of a method for clustering product features based on a hierarchical clustering algorithm provided in an embodiment of the present invention;
[0041] Figure 5 A schematic diagram of a demand classification result tree diagram based on a hierarchical clustering algorithm provided in an embodiment of the present invention;
[0042] Figure 6 A schematic diagram showing the distribution of user attitudes and discussion interest for a product feature category provided by an embodiment of the present invention;
[0043] Figure 7 A schematic diagram of a user demand analysis model based on review data and the Kano model provided in an embodiment of the present invention;
[0044] Figure 8 A schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0045] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.
[0046] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present invention should have the usual meanings understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0047] The inventors of this invention consider that internet review data contains user feedback on purchased products or services. This useful information can reflect user demand for products or services, thereby providing direction and reference for product improvement. The technical solution of this invention can automatically extract user demand-related information from a large amount of unstructured internet review text, including product feature words, user opinion words, and negative words. Based on the Kano model, user needs are divided into basic needs, desired needs, attractive needs, indifferent needs, and reverse needs. This helps product developers gain a deeper understanding and grasp of user needs, allowing for subsequent planning and improvement of product features to quickly adapt to user needs.
[0048] The technical solutions of the embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0049] The embodiment of the present invention provides a user demand classification method for Internet comments, the specific process is as follows: Figure 1 As shown, the following steps are included:
[0050] Step S101: extracting user demand information from Internet review data;
[0051] Specifically, user demand information in Internet review data can be reflected by the names of product features, user opinions, and negative words.
[0052] In an exemplary embodiment, the user demand information extracted from Internet review data using a pre-trained Bi-LSTM (Bi-directional Long Short-Term Memory)-based demand information extraction model includes: F (name of product feature), E (user opinion word of product feature), N (negative word of product feature) and O (other words).
[0053] The demand information extraction model framework based on Bi-LSTM is as follows Figure 2 As shown, it includes embedding layer, Bi-LSTM layer, Softmax layer and output layer. Given an input sentence X={x1,x2,…,x n}, the model will predict the label Y corresponding to each word in the input sentence = {y1,y2,…,y n First, the input sentence is represented as a vector X through the embedding layer and input to the Bi-LSTM layer. In the Bi-LSTM layer, the forward LSTM calculates the sequence representation of each word from left to right. Backward LSTM calculates the sequence representation of each word from right to left Then the final feature representation of the word is obtained by connecting the context representation The technical solution of this invention uses three layers of Bi-LSTM. The output of the first Bi-LSTM layer serves as the input to the next Bi-LSTM layer. The feature representation of the output of the final Bi-LSTM layer serves as the input to the Softmax layer. The Softmax activation function calculates the probability distribution p of a set of labels {F, N, E, O}. Finally, the model predicts the label of each word in the input sentence based on the corresponding output of the Softmax layer.
[0054] The Bi-LSTM-based demand information extraction model is trained using input sentences consisting of words marked with labels {F, N, E, O}. Based on the trained demand information extraction model, the names of product features, user opinion words and negative words of product features are automatically extracted from N user review sentences, and the names are summarized and recorded as the product feature name set F = {w1, w2, ..., w nf}、User opinion word set E of product features = {w1,w2,…,w ne} and the negation word set N of product features = {w1,w2,…,w nn}.
[0055] For example, 2198 user reviews of the BYD Tang NEV, collected from the Autohome website, were used as the data source. The following data preprocessing steps were performed: 1) The 2198 reviews were split into 148,527 individual sentences based on punctuation marks such as ",.;!?"; 2) Each sentence in the review was segmented using the Python Jieba word segmentation tool; 3) Based on the corpus of 2198 reviews, a word vector representation for each sentence was trained using the word2vec model. Then, 2183 randomly selected review sentences were manually annotated, with each word in the sentence labeled as F (name of the product feature), N (negation of the product feature), E (opinion of the product feature), or O (other). Next, we used the Python Scikit-learn module to train the aforementioned Bi-LSTM-based demand information extraction model on 2,183 review sentences. The main model parameters are shown in Table 1. The number (proportion) of review sentences in the training set, validation set, and test set were 1,528 (70%), 327 (15%), and 328 (15%), respectively. Finally, we used the trained model to predict word labels for the remaining 146,344 review sentences.
[0056] Table 1
[0057]
[0058] Three commonly used machine learning model evaluation indicators are used for verification, namely precision, recall, and F1 value (F-measure), as shown in the following formula:
[0059]
[0060]
[0061]
[0062] In the above formula, TP represents the number of samples that the algorithm predicts as positive examples and are actually positive examples; FP represents the number of samples that the algorithm predicts as positive examples and are actually negative examples; FN represents the number of samples that the algorithm predicts as negative examples and are actually positive examples.
[0063] The results show (see Table 2) that the extraction accuracy of the demand information extraction model based on Bi-LSTM is 96%, among which the extraction accuracy of product feature names is 87% and the extraction accuracy of opinion words is 88%.
[0064] Table 2
[0065]
[0066] Using the trained Bi-LSTM-based demand information extraction model, we finally extracted 23,644 demand information from 2,198 reviews (148,527 review sentences in total), involving 657 unique product feature words, 17 unique negative words, and 1,843 unique opinion words, such as Figure 3 Among them, the most frequently used characteristic words include "power," "space," "interior," "seat," "steering wheel," and "cost-effectiveness," most of which are product feature labels recommended by Autohome for users to conduct targeted reviews (i.e., space, power, handling, energy consumption, comfort, appearance, interior, and cost-effectiveness); the most frequently used negative words include "no," "not," "no," and "not enough"; and the content of opinion words is rich, ranging from commonly used "good," "good," and "big" to more formal "precise," "light," and "powerful."
[0067] Step S102: clustering product features;
[0068] In this step, the names of the extracted product features are divided into different product feature categories with a hierarchical structure (i.e., clustered product features) using a hierarchical clustering algorithm. The same product feature category contains several similar or related product features.
[0069] The vector representation of the name of each extracted product feature is used as the input of the hierarchical clustering algorithm. The hierarchical clustering algorithm will output a tree-like hierarchy of the product feature names, which is usually represented by a tree diagram. The original data point is located at the bottom of the cluster tree, and the root node of the cluster is the vertex of the cluster tree. Assuming there are N data points to be clustered, hierarchical clustering will merge small categories into large categories "from bottom to top". First, each data point is regarded as a different cluster. By repeatedly merging the closest pair of clusters, until all data points belong to the same cluster. The basic process is as follows: Figure 4 As shown, it includes the following sub-steps:
[0070] Sub-step S401: Initialize the name of each extracted product feature (i.e., data point) as an independent class;
[0071] Sub-step S402: Calculate the distance between each class, merge the two classes with the closest distance into the same class, until all classes are classified into the same class;
[0072] In this substep, the distance between two data points a and b ||ab||2 is calculated by the Euclidean distance, as shown in Equation 1 below:
[0073]
[0074] In formula 1, a i is the i-th eigenvalue of data point a, bi is the i-th eigenvalue of data point b.
[0075] The distance Δ(A, B) between two classes A and B is measured using the Ward method, as shown in Equation 2:
[0076]
[0077] in, and Represent the centers of class A and class B respectively, n A and n B Represent the number of data points in class A and class B respectively, and Δ is the cost of merging class A and class B; Denotes the i-th data point. Given two clusters with equal center distances, Ward's method will tend to merge the smaller cluster.
[0078] Sub-step S403: Return the tree structure after all data points are classified.
[0079] Based on the tree hierarchy, several clustered product feature categories are selected.
[0080] For example, after sorting the names of the product features extracted above from high to low according to the frequency of occurrence, the names of the 80 product features with the highest frequency of occurrence are selected (see Table 3), and the hierarchical clustering algorithm is used to cluster them. The results are shown in Tables 4 and Figure 5 As shown, the 80 product feature words are divided into 4 major categories and 9 minor categories.
[0081] Table 3
[0082]
[0083]
[0084] Table 4
[0085]
[0086] Step S103: Calculate user attitudes towards product features.
[0087] Specifically, given a review r written by user i for a product p, the product p contains k product features. The user's opinion on each product feature (consisting of negative words and opinion words) can be obtained through the demand information extraction model based on Bi-LSTM mentioned above. Thus, the user attitude css of user i on product feature j (1≤j≤k) ij As shown in formula 3:
[0088]
[0089] In formula 3, Nij represents the number of opinions of user i on product feature j; represents the lth opinion of user i on product feature j; ss(*) represents the sentiment tendency of user opinion; user opinion It is composed of user i’s negative words and opinion words about product feature j, and its sentiment tendency Calculated by the Support Vector Machine (SVM) model.
[0090] For m labeled training samples in is the training sample set R n The i-th sample in y i For samples The SVM will generate a separation hyperplane so that the data on one side of the hyperplane belongs to one category and the data on the other side belongs to another category. Its decision function can be expressed as shown in Equation 4:
[0091]
[0092] Among them, α i and b are SVM model parameters, is a kernel function that implicitly maps samples to a high-dimensional space. As a linear classifier, SVM will try to find a hyperplane that maximizes the margin between the given positive and negative training samples. Ultimately, the sentiment tendency of positive opinions is marked as 1 by the SVM model, i.e. The sentiment tendency of negative opinions is marked as 0, that is, If user i does not evaluate product feature j, that is, N ij = 0, then the user's opinion on the product feature is considered neutral, i.e. css ij =0.5.
[0093] Based on Formula 3, user i's user attitude towards product p is the average of the user's user attitudes towards all k product features of the product, or the average of user i's user attitudes towards all k product features in a certain product feature category, which can be expressed as Formula 5:
[0094]
[0095] Therefore, the value range of user attitude is between 0 and 1. The closer the user attitude is to 0, the user has a negative attitude towards the product feature or the product feature category. The closer it is to 1, the user has a positive attitude towards the product feature or the product feature category. The closer it is to 0.5, the user has a neutral attitude towards the product feature or the product feature category.
[0096] Step S104: Calculate the discussion popularity of the product features.
[0097] Specifically, discussion heat reflects the level of interest in user discussions about product features. Longer user comments on a particular feature indicate a higher level of user attention to it. Therefore, we can conclude that a widely discussed feature is either a largely satisfied need (e.g., users rave about it) or a need that urgently needs improvement (e.g., users complain about it repeatedly).
[0098] Based on the above assumptions, given a review r written by user i for a product p, the product p contains k product features, define the discussion heat ca of user i on product feature j ij is the ratio of the number of opinions of the user on the product feature j to the total number of opinions of the user on all features of the product p, as shown in Formula 6:
[0099]
[0100] In formula 6, N ij represents the number of opinions of user i on product feature j; k is the total number of product features of product p.
[0101] Based on Formula 6, the discussion enthusiasm of user i on product p is the average of the discussion enthusiasm of the user on all k product features of the product, or the average of the discussion enthusiasm of user i on all k product features in a certain product feature category, as shown in Formula 7:
[0102]
[0103] Therefore, discussion heat reflects the extent to which a product feature is widely discussed by users, and its value range is between 0 and 1. The closer the discussion heat is to 1, the higher the users' attention to the product feature or the category of the product feature, which also indirectly reflects the importance the users attach to the product feature or the category of the product feature; the closer the discussion heat is to 0, the lower the users' attention to the product feature or the category of the product feature.
[0104] For example, for the nine product features obtained above, including space, power, control, energy consumption, comfort, appearance, interior, cost-effectiveness and body, the average user attitude and average discussion heat are calculated, and user comments are divided into positive comments, neutral comments and negative comments according to Table 5. Finally, the statistical results of user attitude and discussion heat for each product feature are obtained as shown in Tables 6 and Figure 6 shown.
[0105] Table 5
[0106]
[0107] Table 6
[0108]
[0109] Step S105: Based on the user attitude and discussion heat of the product feature, a user demand classification result of the product feature is obtained through the Kano model.
[0110] Specifically, the traditional Kano model collects customer data through questionnaires. On this basis, according to customer satisfaction when needs are met and not met, user needs are divided into basic needs, expected needs, attractive needs, indifferent needs and reverse needs, as shown in Table 7.
[0111] Table 7
[0112]
[0113] In order to obtain user needs from massive Internet review data and improve the efficiency of user demand acquisition and analysis, the present invention automatically extracts user demand-related information from Internet review data, including product feature words, user opinion words and negative words, and calculates user attitudes and discussion heat of product features. Among them, the name of the product feature represents the evaluation object of the user in the Internet review, such as product attributes, functions or services; user opinion words and negative words jointly reflect the user's attitude, opinion and evaluation of the product feature, thereby reflecting the user's demand for the product; user attitude css reflects the user's emotional attitude towards the product feature; discussion heat ca reflects the degree to which the product feature is widely discussed by users. On this basis, in order to further analyze the overall satisfaction of users with product features, the present invention proposes a user satisfaction index cs, which is the weighted sum of user attitude and discussion heat, and the calculation formula is as shown in Formula 8:
[0114] cs=w CSS css+w CA ca (Formula 8)
[0115] Among them, the weight w CSS and w CA The setting can not rely on expert experience, but on the average user attitude of the product and average discussion heat To determine, as shown in formula 9:
[0116]
[0117] In Formula 9, n represents the number of users who comment on the product, and k represents the number of product features of the product.
[0118] Therefore, combining the Kano model's connotations with the characteristics of online reviews, we can see that user attitudes indirectly reflect their satisfaction with product features. Positive, negative, and neutral user attitudes can be translated into "like," "dislike," and "indifferent" in the Kano model (Table 7). Discussion activity indirectly reflects the importance users place on product features, i.e., their importance. To this end, a systematic analysis of online review data yielded a method for categorizing user needs into different types, as shown in Table 8.
[0119] Table 8
[0120]
[0121] Finally, based on the user attitude, discussion heat and user satisfaction index of product features (or product feature categories), a user demand classification model based on Internet review data and Kano model is designed. Figure 7 The corresponding relationship between each indicator value and the user demand type is shown in Table 9, so that the demand classification result of the product feature (the product feature category) can be determined:
[0122] 1) Basic Needs (MR): The discussion of the product feature (the product feature category) is moderate, and users have a negative attitude towards the product feature (the product feature category). This type of need corresponds to a product problem that needs to be solved urgently, and the user comments are highly subjective and emotional.
[0123] 2) Expected demand (OR): The product feature (the product feature category) has a high level of discussion, users have a positive attitude towards the product feature (the product feature category), there are many non-negative comments about this type of demand, and user satisfaction is sensitive to the satisfaction status of this type of demand;
[0124] 3) Attractive Needs (AR): The discussion of the product feature (the product feature category) is low, and users have a positive attitude towards the product feature (the product feature category). This type of need has many positive comments, and users pay less attention to the satisfaction status of this type of need. However, when this type of need is satisfied, user attitude will significantly improve.
[0125] 4) Indifferent demand (IR): The discussion of the product feature (the product feature category) is low, the user attitude towards the product feature (the product feature category) is moderate, there are more neutral comments for this type of demand, and the user comments are not emotionally charged;
[0126] 5) Reverse demand (RR): The discussion heat of the product feature (the product feature category) is low, the user attitude towards the product feature (the product feature category) is medium, there are few neutral comments on this type of demand, and users are either highly satisfied or poorly satisfied with this type of demand.
[0127] Table 9
[0128]
[0129] For example, according to the calculation results in Table 5, and referring to Tables 8 and 9, the demand classification results based on Internet comment data are shown in Table 10. Among them, energy consumption and control are the basic needs of new energy vehicle users. Users of this type of demand have negative attitudes, high discussion enthusiasm, and more negative comments. Only when this type of demand is met will users be satisfied; power, comfort, and interior are the expected needs of new energy vehicle users. Users of this type of demand have positive attitudes, high discussion enthusiasm, and more non-negative comments. Whether this type of demand is met will significantly affect the user's satisfaction; appearance is the charm need of new energy vehicle users. Users of this type of demand have positive attitudes, high discussion enthusiasm, and more non-negative comments. Low popularity and many positive comments. If this type of demand is met, user satisfaction will be significantly improved; space is an undifferentiated demand of new energy vehicle users. Users have a medium attitude towards this type of demand, low discussion enthusiasm, and more neutral comments (that is, the number of neutral comments is greater than or equal to the number of positive comments or negative comments). Whether this type of demand is met has little effect on user satisfaction; cost-effectiveness and body are reverse demands of new energy vehicle users. Users have a medium attitude towards this type of demand, low discussion enthusiasm, and fewer neutral comments (that is, the number of neutral comments is less than half of the number of positive comments or negative comments). The higher the degree of satisfaction of this type of demand, the lower the user satisfaction.
[0130] Table 10
[0131]
[0132] Figure 8 10 is a schematic diagram showing a more specific hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other within the device via the bus 1050.
[0133] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the user demand classification method for Internet reviews provided in the embodiments of this specification.
[0134] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0135] The input / output interface 1030 is used to connect to an input / output module. It can be connected to a nonlinear receiver to receive information from the nonlinear receiver, enabling information input and output. The input / output module can be configured as a component within the device (not shown) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., while output devices may include a display, speaker, vibrator, indicator light, etc.
[0136] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).
[0137] The bus 1050 comprises a path for transmitting information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).
[0138] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0139] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the user demand classification method for Internet comments in the embodiment are implemented.
[0140] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk equipped with the computer device, a smart media card (SMC), a secure digital (SD) card, a flash memory card, etc. Of course, the computer-readable storage medium may also include both the internal storage unit of the computer device and its external storage device. In this embodiment, the computer-readable storage medium is generally used to store the operating system and various application software installed on the computer device, such as the program code of the user demand classification method for Internet comments in the embodiment. In addition, the computer-readable storage medium may also be used to temporarily store various types of data that have been output or are about to be output.
[0141] The technical solution of the present invention extracts user demand information from internet review data, including the name of a product feature, user opinion words, and negative words. Based on the user opinion words and negative words for the product feature, user attitudes and discussion interest for the product feature are calculated. Based on the user attitudes and discussion interest for the product feature, a user demand classification result for the product feature is obtained using the Kano model. This allows for demand analysis of product features using internet review data, and the Kano model is used to classify different user demand types for product features. This facilitates product developers' in-depth understanding and grasp of user needs, enabling subsequent planning and improvement of product features to quickly adapt to user needs.
[0142] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0143] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Within the scope of the present invention, the above embodiments or technical features in different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the present invention as described above, which are not provided in detail for the sake of simplicity.
[0144] In addition, to simplify the description and discussion, and in order not to obscure the present invention, known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, devices may be shown in the form of block diagrams to avoid obscuring the present invention, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the present invention is to be implemented (i.e., these details should be fully within the scope of understanding of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present invention, it will be apparent to those skilled in the art that the present invention may be implemented without these specific details or with variations in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0145] Although the present invention has been described in conjunction with specific embodiments thereof, many alternatives, modifications and variations of these embodiments will be apparent to those skilled in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may utilize the embodiments discussed.
[0146] The embodiments of the present invention are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A user demand classification method for Internet comments, comprising: Extract user demand information from Internet review data, including: product feature names, user opinion words, and negative words; Based on the user opinion words and negative words of product features, calculate the user attitude and discussion heat of the product features, specifically including: user Product Features User attitude As shown in formula 3: (Formula 3) In formula 3, Represents a user Product Features the number of comments; Represents a user Product Features No. Opinions; Indicates the sentiment tendency of user opinions; user opinions By user Product Features The negative words and opinion words are composed of Calculated by the support vector machine model; Based on user attitudes and discussion heat of product features, the user demand classification results of the product features are obtained through the Kano model.
2. The method according to claim 1, characterized in that The user demand classification results of the product features are obtained by using the Kano model based on the user attitudes and discussion heat of the product features, specifically including: If the user attitude towards the product feature is negative and the discussion heat of the product feature is medium, the user demand classification result of the product feature obtained by the Kano model is basic demand; If the user attitude towards the product feature is positive and the discussion heat of the product feature is high, the user demand classification result of the product feature obtained by the Kano model is expected demand; If the user attitude towards the product feature is positive and the discussion heat of the product feature is low, then the user demand classification result of the product feature obtained by the Kano model is attractive demand; If the user attitude towards the product feature is medium and there are more neutral comments, and the discussion heat of the product feature is low, then the user demand classification result of the product feature obtained by the Kano model is indifferent demand; If the user attitude towards the product feature is medium and there are few neutral comments, and the discussion heat of the product feature is low, then the user demand classification result of the product feature obtained by the Kano model is reverse demand.
3. The method according to claim 1, characterized in that The discussion popularity of the product features is calculated according to the following method: user About products Product Features Discussion heat Calculated according to the following formula 6: (Equation 6) In formula 6, For products The total number of product features.
4. The method according to any one of claims 1 to 3, characterized in that: After extracting user demand information from the Internet review data, the method further includes: Cluster product features to obtain several product feature categories.
5. The method according to claim 4, characterized in that After calculating the user attitude towards the product feature based on the user opinion words and negative words of the product feature, the method further includes: The user attitude towards product feature categories is calculated according to the following formula 5: (Formula 5) In formula 5, Indicates the calculated user For all product feature categories The average value of user attitudes towards product features.
6. The method according to claim 4, characterized in that After calculating the user attitude and discussion heat of the product features based on the user opinion words and negative words of the product features, the method further includes: The discussion heat of the product feature category is calculated according to the following formula 7: (Equation 7) in, Represents a user For all product feature categories The average discussion popularity of each product feature.
7. The method according to claim 6, characterized in that Based on the user attitude and discussion heat of the product features, the user demand classification results of the product features are obtained through the Kano model, specifically: Based on the user attitudes and discussion heat of the product feature category, the user demand classification results of the product feature category are obtained through the Kano model.
8. An electronic device comprising a central processing unit, a signal processing and storage unit, and a computer program stored in the signal processing and storage unit and executable on the central processing unit, characterized in that: When the central processing unit executes the program, the method according to any one of claims 1 to 7 is implemented.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which can be executed by at least one processor to enable the at least one processor to perform the steps of the user demand classification method for Internet reviews as described in any one of claims 1-7.
Citation Information
Patent Citations
User demand analysis method, device and equipment
CN114881677A
Product demand mining method and system fusing extension and online comments
CN115063171A