User demand comprehensive analysis method based on online user comment data

Through the LDA theme model and fine-tuned BERT model combined with the IGR-AHP model, the problems of low efficiency and high subjectivity of user demand analysis in the prior art are solved, and more accurate user demand identification and priority sorting are achieved.

CN120146699APending Publication Date: 2025-06-13CIVIL AVIATION UNIV OF CHINA
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510356025.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing user demand analysis method based on online comment data has problems such as high emotional dictionary labeling cost, high time consumption, and low analysis efficiency. It is difficult to deeply explore user demand information and it is difficult to accurately identify and extract important user demand points.

Method used

The LDA theme model is used to mine text, extract user needs, and calculate user needs emotional scores by fine-tuning the BERT model. The user demand weight value is calculated based on the IGR-AHP model, and finally the user's comprehensive evaluation index value is calculated through the weighted average method to prioritize user demands.

Benefits of technology

It improves the efficiency of user demand analysis, can more accurately identify and refine important user demand points, avoids the subjectivity of manual empowerment, and improves the scientific nature of the results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146699A_ABST
    Figure CN120146699A_ABST
Patent Text Reader

Abstract

The invention discloses a user demand comprehensive analysis method based on online user comment data. The method comprises: calculating a user demand intensity value; calculating a user demand emotion score; calculating a user demand weight value by adopting an IGR-AHP model based on the user demand emotion score; and respectively endowing the user demand intensity value and the user demand weight value with weight coefficients by using a weighted average method, thereby calculating a user comprehensive evaluation index value, and finally carrying out priority ranking on the user demands according to the user comprehensive evaluation index value. The method has the advantages that the LDA topic model is used for text mining, and the problem of how to effectively extract user demand information from a large number of user comments is solved; a user demand emotion score is calculated by fine tuning the BERT model, and the problem that a traditional emotion dictionary analysis method is tedious in process is solved; and finally, combining and considering the strength value and the weight value of the user demand through a weighted average method, and carrying out priority ranking on the user demand.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data analysis, and particularly relates to a comprehensive analysis method for user requirements based on online user review data. Background Art

[0002] Online reviews are the content posted by users on social platforms or e-commerce platforms, aiming to share their experiences and feelings about merchants, products or services, and are a form of communication and interaction by users through the network. At present, the analysis and research on online review data mainly focus on the field of e-commerce platforms, and mainly through analyzing users' shopping experiences to insight into user requirements, so as to promote product improvement and enhancement.

[0003] However, the existing user requirement analysis methods based on online review data have the following deficiencies: First, as a commonly used tool in sentiment analysis, the sentiment dictionary has problems such as high manual annotation cost, large time consumption, and low analysis efficiency. Second, the existing research usually relies on sentiment analysis as the main analysis means, but this method often fails to deeply excavate user requirement information and is difficult to accurately identify and refine important user requirement points. Therefore, how to improve the efficiency of data analysis and further refine the identification and refinement of user requirements is still a technical challenge to be solved in the field of online review data analysis. Summary of the Invention

[0004] In order to solve the above problems, the purpose of the present invention is to provide a comprehensive analysis method for user requirements based on online user review data.

[0005] In order to achieve the above purpose, the comprehensive analysis method for user requirements based on online user review data provided by the present invention includes the following steps carried out in sequence:

[0006] 1) Obtain user reviews and form an original data set and a training data set, then preprocess the original data set to obtain a user review word segmentation data set, then construct a weighted word bag model based on the user review word segmentation data set and input it into an LDA (Latent Dirichlet Allocation) topic model for training, extract user requirements in the original data set, label the user requirement categories for each user review, and finally calculate the user requirement intensity value;

[0007] 2) Fine-tune the BERT (Bidirectional Encoder Representations from Transformers) model using the above training data set, and then use the fine-tuned BERT model to predict the sentiment tendency of user reviews with different user requirements in the original data set with the category labels completed above, and calculate the user requirement sentiment score;

[0008] 3) Based on the above user demand sentiment scores, use the IGR-AHP model to calculate the user demand weight values;

[0009] 4) Use the weighted average method to assign weight coefficients to the user demand intensity values obtained in step 1) and the user demand weight values obtained in step 3) respectively, thereby calculating the user comprehensive evaluation index value, and finally prioritize the user demands according to the level of the user comprehensive evaluation index value.

[0010] In step 1), the method of obtaining user comments, forming the original dataset and the training dataset, then preprocessing the original dataset to obtain the user comment word segmentation dataset, then constructing a weighted word bag model based on the user comment word segmentation dataset and inputting it into the LDA (Latent Dirichlet Allocation) topic model for training, extracting the user demands in the original dataset, and annotating the user demand categories for each user comment, and finally calculating the user demand intensity value is as follows:

[0011] First, use the web crawler tool in Python to obtain multiple user comments of the product to be analyzed and other similar products from e-commerce websites respectively. After removing the duplicate comments and spam comments, obtain the original dataset for method verification and the training dataset for model fine-tuning respectively;

[0012] Then, combine the Baidu stop word library and the Harbin Institute of Technology stop word library to construct a stop word list; at the same time, collect professional terms and construct a custom word list; import the above stop word list and custom word list into the Jieba word segmentation tool in Pyhton, and use the Jieba word segmentation tool to preprocess the original dataset, mainly including: performing word segmentation on the user comments in the original dataset, thereby decomposing the sentences into independent words; removing the stop words and meaningless words in the user comments to obtain the user comment word segmentation dataset;

[0013] After that, use the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm to weight the words in the above user comment word segmentation dataset to construct a weighted word bag model; the calculation formula of TF-IDF is:

[0014]

[0015] where f(t, d) represents the number of occurrences of the word t in the user comment, n(d) represents the total number of words in the user comment word segmentation dataset, N represents the total number of user comments in the user comment word segmentation dataset, and n t represents the number of user comments containing the word t;

[0016] Import the above weighted bag-of-words model into the LDA topic model, and at the same time set the prior parameters α, β, the number of iterations passes, and multiple topics k to train the LDA topic model. Among them, the topic k is used to control the number of topics recognized by the model, the prior parameter α is used to control the sparsity of the document-topic distribution, the prior parameter β is used to control the sparsity of the topic-word distribution, and the number of iterations passes represents the number of times the model is trained. During the training of the LDA topic model, the LDA topic model is trained one by one with different topics k within the set topic range, and the perplexity evaluation method is used to select the optimal topic. Perplexity is an indicator to measure how well the model fits the data. The lower the perplexity, the stronger the model's ability to interpret the data. For each topic k, the calculation formula for perplexity perplexity is:

[0017]

[0018] where D represents the user comment word segmentation dataset, and N i represents the number of words in user comment i, and p(w i ) represents the occurrence probability of word t i in user comment i;

[0019] Use the results of the above multiple trainings to draw a topic-perplexity curve, and determine the optimal topic according to the elbow theory. The elbow theory is that when a significant inflection point appears in the perplexity curve, which is manifested as the curve changes from a sharp decline to a flat or rising trend, the corresponding topic is the optimal topic;

[0020] Rerun the LDA topic model with the above optimal topic to perform topic clustering on the user comment word segmentation dataset, and group the topic words with similar expressions into one category, thereby obtaining a topic-topic word matrix. This matrix shows the words with higher frequencies under each topic;

[0021] By manually summarizing and analyzing these topic words, extract the user demand categories, and label each user comment in the original dataset with the user demand categories;

[0022] Finally, according to the average topic strength of each topic in the LDA topic model (i.e., the mean of the probability distribution of the topic belonging to the document), calculate the user demand strength value;

[0023] The calculation formula for the average topic strength of each topic k is:

[0024]

[0025] where θ k,i represents the probability that topic k appears in user comment i, and N represents the total number of user comments in the user comment word segmentation dataset.

[0026] For each user requirement X, its strength value calculation formula is:

[0027]

[0028] Among them, K represents the total number of topics included, and S k represents the topic strength value of the included topics.

[0029] In step 2), the method of using the above training dataset to fine-tune the BERT (Bidirectional Encoder Representations from Transformers) model, and then using the fine-tuned BERT model to predict the sentiment tendency of user comments of different user requirements in the original dataset with category annotations completed, and calculating the user requirement sentiment score is as follows:

[0030] First, perform sentiment tendency annotation on user comments according to the content of user comments in the training dataset. Annotating 0 means that the user comment is a negative sentiment, and annotating 1 means that the user comment is a positive sentiment; to ensure the stability and generalization ability of model training, randomly divide the data in the training dataset with sentiment tendency annotation completed into a training set and a validation set according to a ratio of 8:2, which are used for parameter optimization and performance evaluation of the BERT model respectively;

[0031] Then use the pre-trained bert-base-chinese model and its tokenizer provided by the Hugging Face website to convert each user comment in the above training set and validation set into a high-dimensional vector that can capture the semantic meaning of the sentence text and input it into the BERT model for BERT model fine-tuning; during the fine-tuning process, the BERT model optimizes parameters through the AdamW optimizer on the training set to minimize the difference between the prediction result and the true label; adopt a multi-round optimization strategy, calculate the loss value through forward propagation, and update the weights using backpropagation; at the same time, perform performance evaluation on the validation set, mainly measuring the performance of the BERT model for sentiment classification through accuracy;

[0032] Finally, input the user comments in the original dataset with category annotations completed into the fine-tuned BERT model. Encode the comments through the tokenizer built into the BERT model, convert them into high-dimensional vectors, perform semantic analysis and sentiment tendency prediction, and calculate the sentiment score of the user requirements.

[0033] In step 3), the method of calculating the user requirement weight value using the IGR-AHP model based on the above user requirement sentiment score is as follows:

[0034] The weight of user requirements is closely related to the corresponding user experience of the product. The higher the degree of emotional impact, the higher the degree of attention of the consumer group to this attribute.

[0035] First, eliminate the uninformative comments in the original dataset with category labels completed, and then convert the user requirement sentiment scores obtained in step 2) into satisfaction scores to construct a requirement-satisfaction matrix.

[0036] Perform cross-validation processing on the above requirement-satisfaction matrix to form multiple sample combinations, calculate the information gain values of the requirement-satisfaction matrix under each sample respectively, then calculate the mean μ and variance σ of the normal distribution parameters using the information gain values of user requirements, and then calculate the relative distance λ between user requirements based on the mean μ and variance σ.

[0037] The formula for calculating the information gain value of user requirements is:

[0038] IG(X|Y) = Ent(X) - Ent(X|Y)

[0039] Among them, Ent(X) represents the information entropy of user requirement X, and Ent(X|Y) represents the uncertainty of user requirement X under the overall satisfaction Y of user comments.

[0040] Based on the above information gain value of user requirements, calculate the mean μ and variance σ, and the calculation formulas are:

[0041]

[0042] Calculate the relative distance λ between user requirements based on the mean μ and variance σ, and the calculation formula is:

[0043]

[0044] Among them, IG represents the information gain value of user requirements, t represents the number of sample combinations, and the relative distance λ between user requirements represents the similarity or difference in information gain between two requirements, μ a and μ b respectively represent the means of the information value-added of requirement a and requirement b, and σ a and σ b respectively represent the variances of the information gain values of requirement a and requirement b.

[0045] Based on the above relative distance λ between user requirements, mean μ and variance σ, perform hierarchical division on the relative distance λ between user requirements and set the scale c ab in the following judgment matrix C to construct the judgment matrix C; the specific rules are as follows:

[0046] When λ = 0, that is, μ a = μ b then cab = 1; when 0 < λ ≤ 1 and μ a > μ b , then c ab = 2; when 1 < λ ≤ 2.58 and μ a > μ b , then c ab = 3; when 2.58 < λ and μ a > μ b , then c ab = 4;

[0047] After the judgment matrix C is constructed, a consistency test is performed on it to ensure the rationality of the result;

[0048]

[0049] After the consistency test passes, the root method is used to calculate the weights of different user requirements and perform normalization processing, thereby determining the weight values w of different user requirements. The calculation formula of the root method is as follows:

[0050]

[0051] Among them, the scale c ab represents the relative importance degree of user requirements a and b, C represents the judgment matrix constructed by the scale c ab , and B represents the number of user requirements.

[0052] In step 4), the method of using the weighted average method to assign weight coefficients to the user requirement intensity value obtained in step 1) and the user requirement weight value obtained in step 3) respectively, thereby calculating the user comprehensive evaluation index value, and finally ranking the user requirements according to the level of the user comprehensive evaluation index value is:

[0053] After determining the different user requirement intensity values and weight values, these two indicators are used as important indicators to measure user requirements. The weight coefficients are assigned to them respectively through the weighted average method, thereby calculating the user comprehensive evaluation index W; finally, the user requirements are ranked according to the level of the user comprehensive evaluation index W.

[0054] Compared with the prior art, the present invention has the following beneficial effects:

[0055] 1. A user requirement analysis method is proposed: First, text mining is performed using the LDA topic model, which solves the problem of how to effectively extract user requirement information from a large number of user comments; Second, the BERT model is fine-tuned to calculate the user requirement sentiment score, which overcomes the problem of the cumbersome process of the traditional sentiment dictionary analysis method; Finally, the user requirement intensity value and weight value are considered together through the weighted average method, and the user requirements are ranked by priority;

[0056] 2. A method for calculating the weight of user requirements is proposed, which assigns weights to user requirements by combining the results of sentiment calculation and the information entropy theory, avoiding the subjectivity of manual weighting and improving the scientificity of the results. Brief Description of the Drawings

[0057] Figure 1 It is a flowchart of the user requirement analysis method based on online review data provided by the present invention.

[0058] Figure 2 It is a graph of the topic-perplexity curve in the LDA topic model. Detailed Description of the Preferred Embodiments

[0059] Next, taking the aerial photography drone product as an example, the present invention will be described in detail with reference to the accompanying drawings.

[0060] As Figure 1 shown, the comprehensive user requirement analysis method based on online user review data provided in this embodiment includes the following steps carried out in sequence:

[0061] 1) Obtain user reviews and form the original data set and the training data set, then preprocess the original data set to obtain the user review word segmentation data set, and then build a weighted word bag model based on the user review word segmentation data set and input it into the LDA (Latent Dirichlet Allocation) topic model for training, extract the user requirements in the original data set, and label the user requirement categories for each user review, and finally calculate the user requirement intensity value;

[0062] The attributes of the aerial photography drone in the user reviews are described by a series of topic words. For example, when evaluating the "flight performance", the term "flight performance" is not directly expressed, but their attitudes are expressed through words such as hovering and stability. Therefore, the LDA topic model is selected to mine the topic words of different user requirements for analysis.

[0063] Use the web crawler technology in Python to obtain 6649 user reviews of all available models of a certain brand of aerial photography drone online from JD.com. After removing the duplicate reviews and spam reviews, 6093 valid user reviews are obtained, which form the original data set for method verification;

[0064] To ensure the scientificity of the training data set, 42300 user reviews of digital products such as computers, mobile phones, and cameras are also obtained online from JD.com at the same time. After removing the duplicate reviews and spam reviews, 41707 valid user reviews remain, which form the training data set for model fine-tuning;

[0065] Combine the Baidu stop word library and the Harbin Institute of Technology stop word library to construct a stop word list. At the same time, collect the professional terms of drones to construct a custom drone word list. Then, import the above stop word list and custom drone word list into the Jieba word segmentation tool in Python, and use the Jieba word segmentation tool to preprocess the user evaluations in the above original dataset to construct a user comment word segmentation dataset;

[0066] Then, use the TF-IDF algorithm to weight the words in the user comment word segmentation dataset to construct a weighted word bag model; import the weighted word bag model into the LDA topic model, set the number of topics k to 2-15, the prior parameters α and β to 0.01 and 0.1 respectively, and the number of iterations passes to 1000 times, perform LDA topic model training and calculate the perplexity, and obtain the Figure 2 shown topic-perplexity curve. According to the elbow theory, initially determine that the optimal number of topics is 6 or 12.

[0067] After that, re-run the LDA topic model with 6 and 12 topics respectively to perform topic clustering on the user comment word segmentation dataset, group the topic words with similar expressions into one category, and after comparing the results, finally determine that the optimal number of topics is 6, and obtain the topic-topic word matrix as shown in Table 1:

[0068] Table 1 Topic-Topic Word Matrix

[0069]

[0070] Manually summarize and analyze the topic words, extract five categories of user needs, namely flight performance, aircraft control, logistics service, photography performance, and appearance design, and label each user comment in the original dataset with the user need category. The results are shown in Table 3.

[0071] Calculate the user need intensity value using the user need intensity value calculation formula. The results are shown in Table 2.

[0072] Table 2 User Need Intensity Value

[0073]

[0074]

[0075] Table 3 User Need Category Labeling Results

[0076]

[0077] 2) Use the above training dataset to fine-tune the BERT (Bidirectional Encoder Representations from Transformers) model, and then use the fine-tuned BERT model to predict the sentiment tendency of user comments with different user requirements in the above-mentioned original dataset with category annotations completed, and calculate the user requirement sentiment scores;

[0078] After completing the user requirement category annotation and intensity value calculation, sentiment tendency annotation is carried out item by item for 41,707 user comments in the training dataset. After the annotation is completed, these user comments are randomly divided into a training set and a validation set according to a ratio of 8:2, and then imported into the Python environment. The pre-trained bert-base-chinese model and its tokenizer are used to convert each user comment in the training set and the validation set into a high-dimensional vector and input it into the BERT model, and the AdamW optimizer is used for multiple rounds of optimization for fine-tuning. The fine-tuned BERT model is used to calculate the user requirement sentiment scores for the user comments in the original dataset with category annotations completed, and the results are shown in Table 4.

[0079] Table 4 Sentiment scores of each user requirement

[0080]

[0081] 3) Based on the above user requirement sentiment scores, use the IGR-AHP model to calculate the user requirement weight values;

[0082] After completing the sentiment analysis, the user requirement weight values are calculated based on the user requirement sentiment scores. After removing the comments in the original dataset with category annotations completed that do not contain requirement information, 5,870 valid user comments remain.

[0083] Based on the user requirement sentiment scores of the user comments in the original dataset, the satisfaction conversion of the user comments is carried out according to Table 5 and Table 6, and the specific conversion rules are as follows:

[0084] (1) For user comments that contain only a single requirement, the satisfaction conversion is carried out according to the user requirement sentiment score value. For requirements not mentioned, if the user requirement sentiment score of this user comment is greater than 0.2, it is marked as "basically satisfied", otherwise it is marked as "relatively dissatisfied";

[0085] (2) For the case where there are multiple requirements in a single user comment, according to its feature words, the multiple requirements are calculated separately by the satisfaction dictionary method;

[0086] (3) Use the lowest user requirement satisfaction as the total satisfaction of the user comment;

[0087] Table 5 Sentiment value and satisfaction conversion table

[0088]

[0089] Partial examples of the satisfaction dictionary in Table 6

[0090]

[0091]

[0092] After the transformation is completed, the requirement-satisfaction matrix is obtained as shown in Table 7, where X 0 represents the overall satisfaction of each user comment, and X a (a = 1, 2, 3, 4, 5) represents the satisfaction of different user requirements in each user comment.

[0093] Table 7 Requirement-satisfaction matrix

[0094]

[0095] To avoid the contingency of the calculation results, the requirement-satisfaction matrix is processed by 10-fold cross-validation. Randomly select 5 of them, and there are 252 sample combinations in total. Calculate the information gain values of the requirement-satisfaction matrix under each sample respectively, and construct the information gain value matrix of user requirements. The results are shown in Table 8.

[0096] Table 8 Information gain value matrix of user requirements

[0097]

[0098] Based on the information gain matrix of user requirements, calculate the normal distribution parameters (mean μ and variance σ) of the information gain values of user requirements. The results of the mean μ and variance σ are shown in Table 9.

[0099] Table 9 Distribution of user requirement parameters

[0100]

[0101] Calculate the relative distance λ between different user requirements using the normal distribution parameters, and construct the judgment matrix C from this:

[0102]

[0103] The result of the consistency test shows that the maximum eigenvalue of the judgment matrix C is 5.146, the consistency index CI value is 0.0365, and the RI value is 1.12. The CR value can be obtained as 0.0325, which is less than 0.1. Therefore, it is reasonable to calculate the weight values of the user requirements of aerial photography drones using this judgment matrix C.

[0104] From the judgment matrix C, the weights of user requirements are calculated by the root extraction method, and the weight values of user requirements are obtained after normalization. The results are shown in Table 10.

[0105] Table 10 Weight Values of User Requirements

[0106]

[0107] 4) Use the weighted average method to assign weight coefficients to the user requirement intensity values obtained in step 1) and the user requirement weight values obtained in step 3) respectively, thereby calculating the user comprehensive evaluation index value, and finally prioritize the user requirements according to the level of the user comprehensive evaluation index value;

[0108] After determining the user requirement intensity value and weight value, use them as important indicators to measure user requirements, and assign a weight coefficient of 0.5 to each by the weighted average method. The comprehensive evaluation value of user requirements is calculated and shown in Table 11.

[0109] Table 11 Comprehensive Evaluation Values of User Requirements

[0110]

[0111] According to the data in Table 11, sort the comprehensive evaluation values of user requirements from high to low, and the obtained priority order is X 4 >X 1 >X 5 >X 3 >X 2 . This sorting result indicates that in the development process of aerial photography drones, photographic performance and flight performance are the requirements that should be given priority attention and optimization.

Claims

1. A user demand comprehensive analysis method based on online user review data, characterized by: The user demand comprehensive analysis method based on online user review data comprises the following steps performed in sequence: 1) Obtain user comments and form an original data set and a training data set, then preprocess the original data set to obtain a user comment segmentation data set, then build a weighted bag-of-words model based on the user comment segmentation data set and input it into the LDA topic model for training, extract user needs from the original data set, and label each user comment with the user need category, and finally calculate the user need intensity value; 2) Fine-tune the BERT model using the above training dataset, and then use the fine-tuned BERT model to predict the sentiment tendency of user comments of different user needs in the original dataset with the above category annotations, and calculate the user need sentiment score; 3) Based on the above user demand sentiment scores, the IGR-AHP model is used to calculate the user demand weight value; 4) Use the weighted average method to assign weight coefficients to the user demand intensity value obtained in step 1) and the user demand weight value obtained in step 3), thereby calculating the user comprehensive evaluation index value, and finally prioritize the user demands according to the level of the user comprehensive evaluation index value.

2. The method for comprehensive analysis of user needs based on online user review data according to claim 1, characterized in that: In step 1), the user comments are obtained and composed into an original data set and a training data set, and then the original data set is preprocessed to obtain a user comment segmentation data set, and then a weighted bag of words model is constructed based on the user comment segmentation data set and input into the LDA topic model for training, user needs in the original data set are extracted, and each user comment is labeled with a user need category, and finally the method for calculating the user need strength value is: First, we use Python web crawler tools to obtain multiple user reviews of the product to be analyzed and other similar products from e-commerce websites online. After removing duplicate reviews and spam reviews, we obtain the original data set for method verification and the training data set for model fine-tuning. Then, we combined Baidu's stop word library and Harbin Institute of Technology's stop word library to build a stop word list; at the same time, we collected professional terms and built a custom word list; we imported the stop word list and the custom word list into Pyhton's Jieba word segmentation tool, and used the Jieba word segmentation tool to preprocess the original data set, mainly including: performing word segmentation on the user comments in the original data set, thereby decomposing the sentences into independent words; removing stop words and meaningless words in the user comments to obtain the user comment word segmentation data set; Then, the TF-IDF algorithm is used to weight the words in the above user comment word segmentation dataset to construct a weighted bag-of-words model; the calculation formula of TF-IDF is: Among them, f(t,d) represents the number of occurrences of word t in the user's comments, n(d) represents the total number of words in the user comment segmentation dataset, N represents the total number of user comments in the user comment segmentation dataset, and n t represents the number of user comments containing word t; The weighted bag-of-words model is imported into the LDA topic model, and the prior parameters α, β, the number of iterations passes and multiple topics k are set to train the LDA topic model, where topic k is used to control the number of topics recognized by the model, the prior parameter α is used to control the sparsity of the document-topic distribution, the prior parameter β is used to control the sparsity of the topic-word distribution, and the number of iterations passes represents the number of model training times; in the LDA topic model training process, different topics k are used one by one within the above-set topic range for LDA topic model training, and the perplexity evaluation method is used to select the optimal topic; for each topic k, the perplexity perplexity is calculated as: Where D represents the user review word segmentation dataset, N i represents the number of words in user comment i, p(w i ) represents word t in user comment i i The probability of occurrence; Use the results of the above multiple trainings to draw a topic-perplexity curve, and determine the optimal topic based on the elbow theory; Re-run the LDA topic model with the above optimal topics, perform topic clustering on the user comment word segmentation dataset, and group topic words with similar expressions into one category, thereby obtaining a topic-topic word matrix; this matrix shows the words with higher frequency under each topic; By manually summarizing and analyzing these keywords, we extract user demand categories and label each user comment in the original data set with the user demand category; Finally, the user demand intensity value is calculated based on the average topic intensity of each topic in the LDA topic model; The average topic strength for each topic k is calculated as: Among them, θ k,i represents the probability of topic k appearing in user comment i, and N represents the total number of user comments in the user comment segmentation dataset; For each user requirement X, the intensity value is calculated as follows: Among them, K represents the total number of topics included, S k Represents the topic strength values ​​of the included topics.

3. The method for comprehensive analysis of user needs based on online user review data according to claim 1, characterized in that: In step 2), the method of fine-tuning the BERT model using the above training data set, and then using the fine-tuned BERT model to predict the sentiment tendency of user comments of different user needs in the original data set with the above category annotations, and calculating the user need sentiment score is: First, the sentiment tendency of user comments in the training data set is annotated according to their content. The annotation 0 indicates that the user comment has negative sentiment, and 1 indicates that the user comment has positive sentiment. To ensure the stability and generalization ability of model training, the data in the training data set with sentiment tendency annotation is randomly divided into a training set and a validation set in a ratio of 8:2, which are used for parameter optimization and performance evaluation of the BERT model respectively. Then, using the pre-trained bert-base-chinese model and its word segmenter provided by the Hugging Face website, each user comment in the above training set and validation set is converted into a high-dimensional vector that can capture the semantics of the sentence text and input into the BERT model for BERT model fine-tuning. During the fine-tuning process, the BERT model is optimized on the training set through the AdamW optimizer to minimize the difference between the predicted results and the true labels. A multi-round optimization strategy is adopted to calculate the loss value through forward propagation and update the weights through backpropagation. At the same time, performance evaluation is performed on the validation set, mainly measuring the performance of the BERT model for sentiment classification through accuracy. Finally, input the user comments in the original dataset with completed category annotations into the fine-tuned BERT model; encode the comments through the tokenizer built into the BERT model, convert them into high-dimensional vectors, perform semantic analysis and sentiment tendency prediction, and calculate the sentiment scores of user needs.

4. The method for comprehensive analysis of user needs based on online user review data according to claim 1, characterized in that: In step 3), the method of calculating the user need weight value using the IGR-AHP model based on the above-mentioned user need sentiment scores is as follows: First, eliminate the uninformative comments in the original dataset with completed category annotations, and then convert the user need sentiment scores obtained in step 2) into satisfaction degrees to construct a need-satisfaction matrix; Perform cross-validation processing on the above-mentioned need-satisfaction matrix to form multiple sample combinations, calculate the information gain values of the need-satisfaction matrix under each sample respectively, then calculate the mean μ and variance σ of the parameters belonging to the normal distribution using the information gain values of user needs, and then calculate the relative distance λ between user needs based on the mean μ and variance σ; The calculation formula for the information gain value of user needs is: IG(XY) = Ent(X) - Ent(X|Y) where Ent(X) represents the information entropy of user need X, and Ent(X|Y) represents the uncertainty of user need X under the overall satisfaction degree Y of user comments; Based on the above-mentioned information gain values of user needs, calculate the mean μ and variance σ, and the calculation formula is: Calculate the relative distance λ between user needs based on the mean μ and variance σ, and the calculation formula is: Among them, IG represents the information gain value of user demand, t represents the number of sample combinations, the relative distance between user demands λ represents the similarity or difference between two demands in terms of information gain, μ a and μ b They represent the mean of the information value added of demand a and demand b, σ a and σ b Represents the variance of the information gain values ​​of demand a and demand b respectively; Based on the relative distance λ between user requirements, mean μ and variance σ, the relative distance λ between user requirements is divided into levels and the scale c in the following judgment matrix C is used. ab To construct the judgment matrix C, the specific rules are as follows: When λ=0, that is, μ a =μ b , then c ab =1; when 0<λ≤1 and μ a >μ b , then c ab =2; when 1<λ≤2.58 and μ a >μ b , then c ab =3; when 2.58<λ and μ a >μ b , then c ab =4; After the judgment matrix C is constructed, perform a consistency test on it to ensure the rationality of the results; After the consistency test passes, use the root extraction method to calculate the weights of different user needs and perform normalization processing, thereby determining the weight values w of different user needs. The calculation formula of the root extraction method is as follows: Among them, the scale c ab Indicates the relative importance of user needs a and b, and C represents the scale c ab The constructed judgment matrix, B represents the number of user requirements.

5. The method for comprehensive analysis of user needs based on online user review data according to claim 1, characterized in that: In step 4), the method of assigning weight coefficients to the user need intensity values obtained in step 1) and the user need weight values obtained in step 3) respectively using the weighted average method, thereby calculating the user comprehensive evaluation index value, and finally ranking the user needs according to the level of the user comprehensive evaluation index value is as follows: After determining the different user need intensity values and weight values, use these two indicators as important indicators to measure user needs, and assign weight coefficients to them respectively through the weighted average method, thereby calculating the user comprehensive evaluation index W; finally, rank the user needs according to the level of the user comprehensive evaluation index W.

Citation Information

Cited By

  • Jeans appearance emotion preference analysis method

    CN120561274A

  • A jeans appearance emotional preference analysis method

    CN120561274B

  • Comment classification method and device, equipment, storage medium and program product

    CN120687616A

  • Method and system for automatically generating customized format report based on AI

    CN120805866A