Park social media comment scoring and core phrase extraction method and device
By constructing a park perception scoring model and SHAP interpreter, the efficiency and accuracy of park social media comment analysis are solved, automated comment scoring and core phrase extraction are realized, and the intelligence level of park management is improved.
Patent Information
- Application Number
- CN202510230564.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-07-18
AI Technical Summary
The existing park social media comment analysis method relies on manual reading to be time-consuming and labor-intensive, and it is difficult to capture public feedback efficiently and comprehensively. The existing sentiment analysis model lacks refined analysis and intelligent keyword extraction, which cannot accurately reflect the emotional tendencies in the comments.
A park-aware scoring model is constructed, combined with SHAP value and word segmentation technology, and preprocessed and classified annotated social media comment data, trained using a pre-trained language model, generate a park-aware scoring model, and use a SHAP interpreter to calculate the contribution value of each character to extract the core phrases of the comment.
It realizes automated scoring and core phrase extraction of park social media comments, improves analysis efficiency and accuracy, reduces labor costs, and is suitable for social media comment analysis in park management and other places.
Smart Images

Figure CN120336699A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital landscape and intelligent data analysis, and particularly relates to a method and device for scoring park social media comments and extracting core phrases. Background Art
[0002] As an important part of the urban green space, urban parks not only provide residents with a wide range of ecosystem services, such as air purification, noise mitigation, and biodiversity conservation, but also offer places for citizens to relax, entertain, and socialize, greatly improving the quality of life and happiness of residents. In addition, urban parks also play an indispensable role in improving the urban ecological environment, promoting community cohesion, and enhancing the overall image of the city. With the rapid development of various social media platforms such as Dianping, Weibo, and Xiaohongshu, the public actively shares their experiences and evaluations of urban parks on these platforms. These comments not only cover the subjective feelings of users but also deeply reflect their cognitions and expectations regarding various aspects such as park environment, facilities, and services. Deeply understanding and analyzing the public's perception and evaluation of parks is of great practical significance and application value for optimizing park management, improving service quality, and formulating scientific plans and designs.
[0003] However, traditional methods for analyzing park social media comments mainly rely on manual reading and classification. This approach is not only time-consuming and laborious but also becomes overwhelmed when faced with a large and continuously growing volume of social media comments. With the explosive growth of the number of comments, it is difficult for manual analysis to efficiently and comprehensively capture and understand the true feedback of the public. In addition, manual analysis is easily affected by the subjective factors of analysts, making it difficult to ensure the objectivity and consistency of the analysis results. This subjective deviation problem is particularly prominent in cross-time or cross-region comparative studies. To improve the efficiency and accuracy of comment analysis, researchers have gradually turned their attention to automated natural language processing technologies, hoping to achieve rapid, accurate processing, and in-depth analysis of large amounts of text data through advanced algorithms and models.
[0004] Currently, sentiment analysis methods based on machine learning and deep learning have been widely applied to the sentiment classification of social media comments. These methods usually train sentiment classification models to convert the sentiment tendency in the comment text into a satisfaction score, thereby quantifying the user's emotional feedback and further exploring the relationship between environmental satisfaction and different features of the park. However, existing research mainly focuses on general sentiment classification models, which are usually trained on large-scale and highly general datasets and lack refined analysis of specific dimensions unique to park comments, such as accessibility, usability, and attractiveness.
[0005] In addition, existing sentiment analysis models have significant deficiencies in interpretability and lack intelligent keyword extraction methods. Traditional keyword extraction mainly relies on manual word frequency statistics after word segmentation, which ignores context, word meaning, grammatical structure, and the relationships between words. For example, "convenient transportation" and "inconvenient transportation" are completely opposite in sentiment, but word frequency statistics may only show that "transportation" is a high-frequency word and cannot accurately reflect its sentiment tendency. At the same time, when the comment content is long and contains many irrelevant words, the word frequency statistics method is more likely to be interfered by noise data, resulting in inaccurate keyword extraction results and requiring manual intervention, further increasing the subjectivity and complexity of the analysis. Therefore, there is an urgent need for a method that can efficiently and accurately process park social media comment texts, which can not only automatically extract the perceived scores of comments but also intelligently identify and extract the core phrases that affect the perceived scores. Summary of the Invention
[0006] To solve the technical problems existing in the prior art, the present invention provides a method and device for scoring park social media comments and extracting core phrases. By constructing a park perception scoring model and combining SHAP values with word segmentation technology, it can automatically extract the perceived scores of comments and further intelligently identify and extract the core phrases that affect the perceived scores.
[0007] The first object of the present invention is to provide a method for scoring park social media comments and extracting core phrases.
[0008] The second object of the present invention is to provide a computer device.
[0009] The first object of the present invention can be achieved by adopting the following technical solutions:
[0010] A method for intelligent scoring of park social media comments and extracting core phrases, characterized in that the method includes:
[0011] S1. Collect comment data of park social media, where the comment data includes comment dates and comment texts;
[0012] S2. Preprocess and classify and label the collected comment data to establish a park comment perception scoring data set;
[0013] S3. Train a pre-trained language model based on the established park comment perception scoring data set to generate a park perception scoring model for perceiving scores of park comments;
[0014] S4. Use the park perception scoring model and its corresponding checkpoint weights to construct a SHAP interpreter, and the SHAP interpreter is used to calculate the SHAP values of each character in each comment to quantify the contribution of each character to the park perception score;
[0015] S5. Input the comment data of all social media of the target park into the park perception scoring model, calculate the SHAP value of each character in each comment, perform word segmentation, filtering, and merging processing on the comment text to obtain merged phrases, extract the core phrases of the text according to the total SHAP value of the merged phrases, and output the perception score and core phrases of each comment.
[0016] Specifically, step S2 includes the following steps:
[0017] Clean the comment text, delete blank, meaningless, and duplicate data, and randomly select a preset number of comments from the cleaned comment data as the comment text data to be labeled.
[0018] Classify and label the comment text data to be labeled according to multiple dimensions to establish a park comment perception scoring data set.
[0019] Specifically, the classification and labeling of the comment text data to be labeled according to multiple dimensions includes: classifying and labeling the comment text data to be labeled according to 3 dimensions, and the 3 classification dimensions include: accessibility dimension, usability dimension, and attractiveness dimension. Each classification dimension contains three types of labels: 0, 1, and 2, where "0" indicates that the dimension is poor, "1" indicates good, and "2" indicates not involved.
[0020] Specifically, the pre-trained language model is the Chinese-RoBERTa-wwm-ext model;
[0021] Step S3 includes the following steps:
[0022] Divide the park comment perception scoring data set into a training set, a validation set, and a test set at a set ratio;
[0023] Use the BertAdam optimizer to update the parameters of the Chinese-RoBERTa-wwm-ext model;
[0024] Perform forward propagation, and use the cross-entropy loss function to calculate the loss between the model output and the true label;
[0025] Perform backpropagation and model parameter update, iterate training until the model converges, and use accuracy ACC, precision Precision, recall Recall, and F1 value to evaluate the effectiveness of the park perception scoring model;
[0026] Save the checkpoint weights corresponding to the training stage with the highest prediction accuracy to obtain the final park perception scoring model, and complete the training of the park perception scoring model.
[0027] Specifically, the calculation formula of the cross-entropy loss function is as follows:
[0028]
[0029] where N is the batch size, C is the number of categories, y i,c is the true label of the i-th sample, is the probability that the model predicts the i-th sample belongs to category c.
[0030] Specifically, the calculation formulas of the accuracy ACC, precision Precision, recall Recall, and F1 value are as follows:
[0031]
[0032] where n is the total number of all comments, y i is the actual label of the i-th comment, is the predicted label of the model, is the number of comments with correct classification results; TP is the true positive class, that is, the number of samples correctly predicted as the positive class by the model; TN is the true negative class, that is, the number of samples correctly predicted as the negative class by the model; FP is the false positive class, that is, the number of samples wrongly predicted as the positive class by the model; FN is the false negative class, that is, the number of samples wrongly predicted as the negative class by the model.
[0033] Specifically, step S4 includes the following steps:
[0034] Load the park perception scoring model through the transformers library, map the checkpoint weights of the park perception scoring model to the CPU, remove the keys related to the classifier in the weight dictionary, and remove the prefix in the weight keys to adapt to the model structure of the park perception scoring model;
[0035] Initialize the SHAP interpreter, and use the park perception scoring model and its corresponding checkpoint weights to construct the SHAP interpreter, which is used to calculate the SHAP values of each character in each comment.
[0036] Specifically, step S5 includes the following steps:
[0037] Input the comment data of all social media of the target park into the park perception scoring model, and use the SHAP interpreter to calculate the SHAP values of each character of the comment text;
[0038] Use the Jieba word segmentation tool to segment and label the parts of speech of the comments, use the HIT stop word dictionary to delete meaningless or redundant words, merge nouns and their modifiers to get merged phrases, and add the SHAP values of all characters in each merged phrase to get the total SHAP value of each merged phrase;
[0039] Sort the total SHAP values of the calculated combined phrases from high to low, and select several phrases with the top-ranked total SHAP values as the core phrases of the comment.
[0040] The second object of the present invention can be achieved by adopting the following technical solutions:
[0041] A computer device includes a processor and a memory for storing programs executable by the processor. It is characterized in that when the processor executes the programs stored in the memory, the above-mentioned intelligent scoring and core phrase extraction method for park social media comments is implemented.
[0042] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0043] The present invention provides a method and device for scoring and extracting core phrases from park social media comments. By constructing a park perception scoring model and combining SHAP value calculation and word segmentation technology, it can automatically extract the perception scores of comments, effectively extract the most representative core phrases in each comment, help managers quickly identify the key points of user attention, achieve efficient and accurate scoring of park comments, significantly improve the automation level of comment text analysis, and reduce labor costs. The present invention is not only applicable to park comment analysis, but also can be extended to the analysis of social media comments of other types of places and services, has wide applicability and promotion value, and provides a powerful tool for the intelligent management of the urban environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the structures shown in these drawings without creative efforts.
[0045] Figure 1 It is a flowchart of the intelligent scoring and core phrase extraction system for park social media comments in Embodiment 1 of the present invention;
[0046] Figure 2 It is a relationship diagram of three dimensions of park comment scoring involved in Embodiment 1 of the present invention;
[0047] Figure 3 It is a model architecture diagram of the intelligent scoring and core phrase extraction method for park social media comments in Embodiment 1 of the present invention;
[0048] Figure 4This is an example effect diagram of the perceived scoring results of accessibility, availability, and attractiveness of each park in a certain area in Embodiment 1 of the present invention, as well as negative keyword groups extracted from the comments. Detailed implementation manners
[0049] Next, the technical solution of the present invention will be further described in detail in combination with the drawings and embodiments. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. The implementation manners of the present invention are not limited thereto. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0050] Embodiment 1:
[0051] As Figure 1 shown, this embodiment provides a method for scoring park social media comments and extracting core phrases, and the method includes the following steps:
[0052] S1. Collect the comment data of the social media of the target park, where the comment data includes the comment date and the comment text.
[0053] In this embodiment, the API of social media platforms such as Dianping and Xiaohongshu is called to obtain and collect the social media comment data of the target park. Specifically, 62,106 comment data of 130 parks in Guangzhou, Guangdong Province are obtained, and the comment data includes the comment date and the comment text.
[0054] S2. Preprocess and classify and label the collected comment data to establish a park comment perception scoring data set.
[0055] S21. Clean the comment text, delete blank, meaningless, and duplicate data, and randomly select a preset number of comments from the cleaned comment data as the comment text data to be labeled.
[0056] Specifically, delete the blank, meaningless, and duplicate comments in the social media comment data to ensure the data quality, and randomly select 4,000 comments from the cleaned comments as the comment data to be manually classified and labeled.
[0057] S22. Classify and label the comment text data to be labeled according to multiple dimensions to establish a park comment perception scoring data set.
[0058] Magdalena Biernacka et al. divided the supply of urban parks into three levels: accessibility, usability, and attractiveness. First, the benefits that urban parks can provide can mainly be obtained when residents can reasonably enter and use this space. The existence and accessibility of parks are considered the most basic factors for park use. Accessibility refers to the ease with which an individual can overcome obstacles such as distance and time to reach and enter the park. It is an indicator to measure the relative opportunity for people to access or use the park. Second, the common view in the past regarding improving the quality of urban parks was to increase the quantity and scale of parks, that is, to improve the accessibility of parks. However, research on the accessibility of parks and people's well-being shows that the usability and quality of parks need more attention. In terms of providing mental health benefits, the two can play a more important role than the quantity of parks. Usability refers to whether people can freely reach and enter the park and use it safely for entertainment purposes at any time, reflecting the effective degree of the park to meet the functional needs of different user groups. Attractiveness is an important indicator for evaluating the quality of parks. It refers to whether the park meets the personal needs, expectations, and preferences of users when a person is willing to use and spend his or her time there. It will affect people's decisions to go to these spaces for activities and stay. Generally speaking, to fully realize the various benefits of parks, parks must have both accessibility and usability. Parks that perform well in these aspects are usually more attractive.
[0059] Specifically, the review text data to be labeled is classified and labeled according to three dimensions: the accessibility dimension, the usability dimension, and the attractiveness dimension. Each classification dimension contains three types of labels: 0, 1, and 2, where "0" indicates that the dimension is poor, "1" indicates good, and "2" indicates not applicable. For example, in the review example "Yuexiu Park has a good location and very convenient transportation. It can be directly reached by subway and bus. The air is fresh in the morning, and many elderly people exercise there. It has a large area with flowers and trees everywhere. The Five Rams Statue is very beautiful for taking pictures. It is a very good park. I recommend everyone to go and have a look!", the three dimensions of accessibility, usability, and attractiveness are all good. Therefore, the labels for all three dimensions are "1".
[0060] As Figure 2 shown, the relationship diagram of the three dimensions of accessibility, usability, and attractiveness. Review examples representing accessibility: Is there a green space around me? Is the green space within a certain distance from my place of residence? Review examples representing usability: Can I freely enter the green space? Am I welcome there? Review examples representing attractiveness: Do I want to stay there? Does the green space meet my preferences?
[0061] S3. Train the pre-trained language model based on the established park review perception scoring dataset to generate a park perception scoring model for perceiving and scoring park reviews.
[0062] In this embodiment, the pre-trained language model can be the Chinese-RoBERTa-wwm-ext model. The Chinese-RoBERTa-wwm-ext model is a pre-trained language model optimized for Chinese based on the RoBERTa model. The Chinese-RoBERTa-wwm-ext model adopts the Whole Word Masking (WWM) technology, that is, masking the whole word instead of a single character during training, so as to better capture the semantic relationships at the word level in Chinese. In addition, the model is extended and trained on a larger-scale Chinese corpus, further improving its understanding ability of the Chinese language. This model is widely used in natural language processing tasks such as Chinese text classification, named entity recognition, and sentiment analysis, and performs well in these tasks, especially having strong advantages in dealing with tasks of long texts or complex semantics.
[0063] Fine-tune the pre-trained Chinese-RoBERTa-wwm-ext model based on the park review perception scoring dataset, including the following steps:
[0064] S31. Divide the park review perception scoring dataset into a training set, a validation set, and a test set at a set ratio. In this example, the pre-processed review data is randomly divided into a training set, a validation set, and a test set at a ratio of 7:2:1.
[0065] S32. Use the BertAdam optimizer to update the parameters of the Chinese-RoBERTa-wwm-ext model.
[0066] Specifically, set the learning rate to 1e-5, the batch size to 10, and the sequence length to 150.
[0067] S33. Perform forward propagation, and use the cross-entropy loss function to calculate the loss between the model output and the true label. The calculation formula of the cross-entropy loss function is:
[0068]
[0069] where N is the batch size, C is the number of classes, y i,c is the true label of the i-th sample, is the probability that the model predicts that the i-th sample belongs to class c.
[0070] S33. Perform backpropagation and update the model parameters, and iteratively train until the model converges. Use accuracy, precision, recall, and F1 value to evaluate the effectiveness of the park perception scoring model.
[0071] Among them, the calculation formula for updating the model parameters is:
[0072]
[0073] Among them, θ is the model parameter, and α is the learning rate.
[0074] Specifically, the calculation formulas for accuracy ACC, precision Precision, recall Recall, and F1 value are respectively:
[0075]
[0076] Among them, n is the total number of all comments, y i is the actual label of the i-th comment, is the predicted label of the model, is the number of comments with correct classification results; TP is the true positive class, that is, the number of samples correctly predicted as the positive class by the model; TN is the true negative class, that is, the number of samples correctly predicted as the negative class by the model; FP is the false positive class, that is, the number of samples wrongly predicted as the positive class by the model; FN is the false negative class, that is, the number of samples wrongly predicted as the negative class by the model.
[0077] S35. Save the checkpoint weights corresponding to the training stage with the highest prediction accuracy to obtain the final park perception scoring model, and complete the training of the park perception scoring model.
[0078] Among them, the checkpoint weights are the model weight parameters regularly saved during the model training process, which can be used for model deployment or resuming training. The checkpoint weights can be understood as the parameters of the best-performing model, and it is equivalent to having it for the model to possess the ability of park perception scoring.
[0079] In this embodiment, a comparative experiment was conducted to further demonstrate the superiority of the model proposed in this embodiment. Table 1 is a comparison table of the accuracy, precision, recall rate, and F1 value of different methods. The experimental results are shown in Table 1. Among them, BERT is a pre-trained language model based on Transformer, which captures context information through a bidirectional encoder. Compared with traditional unidirectional language models, it can consider the context before and after words simultaneously, improving the performance of natural language understanding tasks; ERNIE is a pre-trained language model based on BERT proposed by Baidu. Its main innovation lies in enhancing the semantic understanding ability of the model by integrating external knowledge bases (such as entity information, concept maps, etc.). ERNIE can better capture the deep semantics at the lexical, sentence, and document levels by introducing knowledge enhancement learning, thus showing better performance than BERT in Chinese natural language processing tasks, especially in tasks involving common sense knowledge and multi-domain knowledge; LSTM is a variant of the recurrent neural network commonly used to process and predict time series data, widely used in natural language processing, speech recognition and other tasks, especially suitable for processing data with time order or context relationships, such as text generation and machine translation tasks; XGBoost is an efficient and scalable machine learning algorithm based on gradient boosting, widely used in classification and regression tasks of structured data.
[0080] In this embodiment, the park perception scoring model performs excellently in all evaluation indicators. Especially in terms of accuracy, precision, recall rate, and F1 value, it leads other models, showing its strong comprehensive ability in Chinese natural language processing tasks. It can reliably automatically score the perceived accessibility, usability, and attractiveness in social media comments. Especially in the dimensions of accessibility and usability, the park perception scoring model, with its high accuracy (0.8338 and 0.8568) and precision (0.8380 and 0.8570), is significantly better than BERT and ERNIE, and also maintains a good balance in terms of F1 value and recall rate, indicating that it can not only make accurate predictions but also effectively identify positive samples, and is suitable for scenarios with high requirements for comprehensive performance in various tasks. In contrast, the performance of LSTM and XGBoost is significantly weaker, especially in terms of F1 value and precision, far behind the park perception scoring model. The evaluation index data of the park perception scoring model compared with other models are shown in Table 1 specifically.
[0081] Table 1
[0082]
[0083]
[0084] S4. Use the park perception scoring model and its corresponding checkpoint weights to construct a SHAP interpreter, which is used to calculate the SHAP values of each character in each comment to quantify the contribution of each character to the park perception score.
[0085] S41. Load the park perception scoring model through the transformers library, map the checkpoint weights of the park perception scoring model to the CPU, remove the keys related to the classifier from the weight dictionary, and remove the prefix in the weight keys to adapt to the model structure of the park perception scoring model.
[0086] Use the BertTokenizer and BertForSequenceClassification libraries in the transformers library to load the park perception scoring model, map the weights to the CPU, remove the keys related to the classifier from the weight dictionary, and remove the prefix text_model. in the weight keys to adapt to the current model structure. Among them, the keys related to the classifier in the weight dictionary include: text_model.classifier.weight and text_model.classifier.bias. The BertForSequenceClassification library adds a linear layer and an activation function on the basis of BERT to convert the encoded representation of the input sequence into a class distribution, thereby realizing text classification. BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model based on the Transformer architecture. The BERT model is based on the encoder structure of Transformer and contains multiple encoder blocks (Encoder block), and each encoder block consists of a self-attention mechanism, a feed-forward neural network, layer normalization, and a residual connection.
[0087] S42. Initialize the SHAP interpreter. Use the park perception scoring model and its corresponding checkpoint weights to construct the SHAP interpreter, and the SHAP interpreter is used to calculate the SHAP values of each character in each comment.
[0088] Specifically, the shap.Explainer function is used to construct a SHAP explainer based on the park perception scoring model and its checkpoint weights. This SHAP explainer can quantify the contribution of each character to the park perception score, where the SHAP value reflects the degree of attention the model pays to each character when judging perception - the higher the SHAP value, the greater the degree of attention. The SHAP explainer quantifies the contribution of features to the model prediction based on the Shapley value. By calculating the Shapley value of each feature, the contribution of each feature to the model prediction can be measured, and by calculating the marginal contribution of each feature to the model prediction, the way the model makes decisions can be explained, which is used to interpret the prediction results of machine learning models.
[0089] S5. Input the comment data of all social media of the target park into the park perception scoring model, calculate the SHAP value of each character in each comment, perform word segmentation, filtering, and merging processing on the comment text to obtain merged phrases, extract the core phrases of the text according to the total SHAP value of the merged phrases, and output the perception score and core phrases of each comment.
[0090] As Figure 3 shown, it is the model architecture diagram of the intelligent scoring and core phrase extraction method for park social media comments. Step S5 specifically includes the following steps:
[0091] S51. Input the comment data of all social media of the target park into the park perception scoring model, and use the SHAP explainer to calculate the SHAP value of each character in the comment text. The SHAP value of each character is calculated by the following formula:
[0092]
[0093] where is the SHAP value of feature j, F is the set of all features, S is the subset of features that does not include feature j, |S| represents the number of elements in set S, |F| represents the number of elements in set F, f S (x S∪{j} ) is the model prediction using subset S and feature j, f S (x S ) is the model prediction using only subset S, and! represents the factorial operation. For example, n! = n × (n - 1) × (n - 2) × … × 2 × 1. (|F| - |S| - 1): This is an arithmetic expression representing the cardinality of set F minus the cardinality of set S minus 1. (|S|!(|F| - |S| - 1)!): This is the product of two factorials. The first is the factorial of the cardinality of set F, and the second is the factorial of the cardinality of set F minus the cardinality of set F minus 1.
[0094] After calculating the SHAP values of each character using the SHAP interpreter, the SHAP values corresponding to each character are output, which can measure the contribution of each character and realize the weight analysis of each word in the comment.
[0095] S52. Use the Jieba word segmentation tool to segment and label the part-of-speech of the comment, use the HIT stopword dictionary to delete meaningless or redundant words, merge nouns and their modifiers to obtain merged phrases, and add up the SHAP values of all characters in each merged phrase to obtain the total SHAP value of each merged phrase.
[0096] The steps of using the Jieba word segmentation tool for word segmentation and part-of-speech labeling and using the HIT stopword dictionary to delete redundant words specifically include:
[0097] Use the Jieba word segmentation tool to segment and label the part-of-speech of each comment, use the HIT stopword dictionary to filter out meaningless or redundant words; merge nouns and their modifiers (including adjectives and adverbs); for each remaining phrase, calculate its total SHAP value weight;
[0098] S53. Sort the total SHAP values of the calculated merged phrases from high to low, and select several phrases with the top-ranked total SHAP values as the core phrases of the comment.
[0099] In this embodiment, for example, the following comment on Huolushan Forest Park is processed:
[0100] "In Tianhe, Guangzhou, it's a great place for weekend hiking! The transportation is convenient. You can take the No. 6 subway and get off at Exit A of Longdong Station and you'll reach the foot of the mountain!"
[0101] The predicted label for its accessibility dimension is "1" and the probability is 0.9921.
[0102] Input the comment data of all social media of the target park into the park perception scoring model, use the SHAP interpreter to calculate the SHAP value of each character in the comment text, and the calculation results of the SHAP values of each character's accessibility dimension are as follows:
[0103] At:0.0017 Guang:-0.0042 Zhou:-0.0061 Tian:-0.0009 He:-0.0009,:-0.0003 Zhou:-0.0029 Mo:-2.8905 Tu:-0.0005 Bu:0.0009 of:0.0014 Hao:0.0021 Di:0.0012 Fang:0.0028! :0.0013 Jiao:0.0151 Tong:0.0123 Fang:0.0062 Bian:0.0239, :0.0062 Take:0.0134 No.:0.0052 Six:0.0093 Subway:0.0081 Tie:0.0147, :0.0022 Long:0.0023 Dong:0.0021 Station:0.0144 6A:0.0011
[0104] Exit:0.0088 Kou:0.0109 Immediately:0.0080 Reach:0.0087 Mountain:0.0078 Foot:0.0042! :-0.0009
[0105] Sort the calculated phrase weights, and select the top 4 phrases with the highest total SHAP value as the core phrases of this comment. The keyword phrase with the highest SHAP value is: Convenient transportation:0.0576, Subway Line 6:0.0373, Exit:0.0198, Nice place:0.0123.
[0106] Such as Figure 4As shown in the figure, it is an example effect diagram of the perceived score results of the accessibility, usability, and attractiveness of each park in a certain district, as well as the negative keyword groups extracted from the comments. The scores of each park in Tianhe District, Guangzhou City, in the three dimensions of accessibility, usability, and attractiveness, and the example negative keyword groups extracted from 429 comments of Huoluoshan Forest Park are drawn in ArcGIS Pro. The results show that in Tianhe District with convenient transportation and complete infrastructure, parks like Huoluoshan Forest Park have a relatively high score in terms of accessibility (score 0.84), but there is still much room for improvement in the two dimensions of usability (score 0.79) and attractiveness (score 0.64). The extraction of keyword groups explains the specific problems: in terms of accessibility, there are mainly problems such as "a bit far", "chaotic parking", and "difficult to park"; in terms of usability, it is prominently manifested as "few toilets", "a lot of garbage", and "gloomy", etc.; in terms of attractiveness, it mainly reflects problems such as "low playability", "ordinary scenery", and "nothing special", etc. These findings provide a clear direction for the improvement of park quality. It is recommended that park management departments can start to improve from the following aspects: optimizing the transportation connection system and improving parking management; increasing the configuration of sanitation facilities and strengthening daily maintenance; enriching entertainment projects and enhancing landscape features, etc. Through the intelligent scoring and core phrase extraction of social media comments in this embodiment, it not only accurately reflects the distribution characteristics of the perceived evaluation of each park, but also deeply reveals the specific problems, fully verifying the effectiveness and practicality of this method in park perception evaluation, and providing a scientific basis for the refined management and quality improvement of urban parks.
[0107] In summary, the present invention provides a method capable of efficiently and accurately processing park social media comment texts. By calling the application programming interface of the social media platform, the social media comment data of the target park is obtained, the collected comments are classified and labeled, and a park comment perception scoring data set is established. Based on the park comment perception scoring data set, the Chinese-RoBERTa-wwm-ext model is fine-tuned to obtain a park perception scoring model. Using the park perception scoring model and the checkpoint weights during its training process, a SHAP interpreter is defined and the SHAP values of each character in each comment in the park social media are calculated. Jieba is used to segment and label the parts of speech of the comments, the HIT stop word dictionary is used to delete redundant words, nouns and their modifiers (including adjectives and adverbs) are merged, and the SHAP values of all characters in each merged phrase are added up to calculate its total SHAP value. The multiple phrases with the highest SHAP values are selected as the core phrases of the comment. For all the social media comments collected for the target park, the trained park perception scoring model and the core phrase extraction method are applied to extract all the core phrases in the park, realizing the intelligent scoring of park social media comments and the effective extraction of core phrases, significantly improving the accuracy and efficiency of comment analysis, and providing strong data support for park management and service optimization. It can provide a scientific basis for the planning, design and management of parks, help managers optimize resource allocation and improve service quality, and ultimately achieve the sustainable development of urban parks and the comprehensive improvement of the quality of residents' lives.
[0108] Embodiment 2:
[0109] This embodiment provides a computer device, which can be a server, a computer, etc. It includes a processor, a memory, an input device, a display and a network interface connected through a system bus. The processor is used to provide computing and control capabilities. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. When the processor executes the computer program stored in the memory, it implements the method for intelligent scoring of park social media comments and extraction of core phrases in Embodiment 1 above, including:
[0110] S1. Collect the comment data of the park social media, where the comment data includes the comment date and the comment text;
[0111] S2. Preprocess and classify and label the collected comment data to establish a park comment perception scoring data set;
[0112] S3. Train the pre-trained language model based on the established park comment perception scoring data set to generate a park perception scoring model for perceiving the score of park comments;
[0113] S4. Use the park perception scoring model and its corresponding checkpoint weights to construct a SHAP interpreter, which is used to calculate the SHAP values of each character in each comment to quantify the contribution of each character to the park perception score;
[0114] S5. Input the comment data of all social media of the target park into the park perception scoring model, calculate the SHAP value of each character in each comment, perform word segmentation, filtering, and merging processing on the comment text to obtain merged phrases, extract the core phrases of the text according to the total SHAP value of the merged phrases, and output the perception score and core phrases of each comment.
[0115] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A method for intelligent scoring and core phrase extraction of park social media comments, characterized in that, The method includes: S1. Collect the comment data of the park's social media, where the comment data includes the comment date and the comment text; S2. Preprocess and classify and label the collected comment data to establish a park comment perception scoring dataset; S3. Train the pre-trained language model based on the established park comment perception scoring dataset to generate a park perception scoring model for perceiving and scoring park comments; S4. Use the park perception scoring model and its corresponding checkpoint weights to construct a SHAP interpreter, which is used to calculate the SHAP values of each character in each comment to quantify the contribution of each character to the park perception score; S5. Input the comment data of all social media of the target park into the park perception scoring model, calculate the SHAP values of each character in each comment, perform word segmentation, filtering, and merging processing on the comment text to obtain merged phrases, extract the core phrases of the text according to the total SHAP values of the merged phrases, and output the perception score and core phrases of each comment.
2. The method for intelligent scoring and core phrase extraction of park social media comments according to claim 1, characterized in that The step S2 includes the following steps: Clean the comment text, delete blank, meaningless, and duplicate data, and randomly select a preset number of comments from the cleaned comment data as the comment text data to be labeled; Classify and label the comment text data to be labeled according to multiple dimensions to establish a park comment perception scoring dataset.
3. A method for intelligent scoring and core phrase extraction of park social media comments according to claim 1, characterized in that, The classification and labeling of the comment text data to be labeled according to multiple dimensions includes: classifying and labeling the comment text data to be labeled according to 3 dimensions, and the 3 classification dimensions include: accessibility dimension, usability dimension, and attractiveness dimension. Each classification dimension contains three types of labels: 0, 1, 2, where "0" indicates that the dimension is poor, "1" indicates good, and "2" indicates not involved.
4. A method for intelligent scoring and core phrase extraction of park social media comments according to claim 1, characterized in that, The pre-trained language model is the Chinese-RoBERTa-wwm-ext model; The step S3 includes the following steps: Divide the park comment perception scoring dataset into a training set, a validation set, and a test set at a set ratio; Use the BertAdam optimizer to update the parameters of the Chinese-RoBERTa-wwm-ext model; Perform forward propagation and use the cross-entropy loss function to calculate the loss between the model output and the true label; Perform backpropagation and model parameter update, and iteratively train until the model converges. Use the accuracy ACC, precision Precision, recall Recall, and F1 value to evaluate the effectiveness of the park perception scoring model; Save the checkpoint weights corresponding to the training stage with the highest prediction accuracy to obtain the final park perception scoring model and complete the training of the park perception scoring model.
5. The method for intelligent scoring and core phrase extraction of park social media comments according to claim 4, characterized in that The calculation formula for using the cross-entropy loss function is: where N is the batch size, C is the number of classes, and y i,c is the true label of the i-th sample, is the probability that the model predicts the i-th sample belongs to class c.
6. A method for intelligent scoring and core phrase extraction of park social media comments according to claim 4, characterized in that, The calculation formulas for the accuracy ACC, precision Precision, recall Recall, and F1 value are respectively: where n is the total number of all comments, y i is the actual label of the i-th comment, is the predicted label of the model, is the number of comments with correct classification results; TP is the true positive class, that is, the number of samples correctly predicted as the positive class by the model; TN is the true negative class, that is, the number of samples correctly predicted as the negative class by the model; FP is the false positive class, that is, the number of samples wrongly predicted as the positive class by the model; FN is the false negative class, that is, the number of samples wrongly predicted as the negative class by the model.
7. A method for intelligent scoring and core phrase extraction of park social media comments according to claim 1, characterized in that The step S4 includes the following steps: Load the park perception scoring model through the transformers library, map the checkpoint weights of the park perception scoring model to the CPU, remove the keys related to the classifier from the weight dictionary, and remove the prefix in the weight keys to adapt to the model structure of the park perception scoring model; Initialize the SHAP interpreter, and use the park perception scoring model and its corresponding checkpoint weights to construct the SHAP interpreter, which is used to calculate the SHAP values of each character in each comment.
8. A method for intelligent scoring and core phrase extraction of park social media comments according to claim 7, characterized in that The step S5 includes the following steps: Input the comment data of all social media of the target park into the park perception scoring model, and use the SHAP interpreter to calculate the SHAP values of each character of the comment text; Use the Jieba tokenization tool to tokenize and label the parts of speech of the comment, use the HIT stopword dictionary to delete meaningless or redundant words, merge nouns and their modifiers to get merged phrases, and add the SHAP values of all characters in each merged phrase to get the total SHAP value of each merged phrase; Sort the total SHAP values of the calculated merged phrases from high to low, and select several phrases with the highest total SHAP values as the core phrases of the comment.
9. A method for intelligent scoring and core phrase extraction of park social media comments according to claim 8, characterized in that, The SHAP value of each character is calculated by the following formula: Among them, is the SHAP value of feature j, F is the set of all features, S is the feature subset that does not include feature j, |S| represents the number of elements in set S, |F| represents the number of elements in set F, and f S (x s∪{j} ) is the model prediction using subset S and feature j, and f S (x S ) is the model prediction using only subset S, and! represents the factorial operation.
10. A computer device, comprising a processor and a memory for storing processor-executable programs, characterized in that, When the processor executes the program stored in the memory, it implements the method for intelligent scoring and core phrase extraction of park social media comments according to any one of claims 1-9.