Multi-scale comprehensive sensing method for park city
By searching keywords and analyzing emotional tendency on the comment text data of park cities, the problem of insufficient analysis of multiple scale evaluation elements of user comments in the existing technology is solved, and the refined evaluation and optimization of multi-scale perception results of park cities is achieved.
Patent Information
- Application Number
- CN202411949733.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to conduct in-depth analysis of evaluation elements of multiple scales in user comments, resulting in the multi-scale perception results of urban parks not being fine and accurate enough.
By obtaining the review text data of the park city, the keyword search method is used to cut the data into phrases, and classification is performed based on multi-scale keyword groups and preset classification methods. Then, the emotional tendency and scores of phrases are determined through the emotional tendency analysis algorithm, the weights of each category are calculated, and the public's perception of different scales of the park is quantified.
The refined evaluation of multi-scale perception results of park cities is achieved, and more objective and accurate evaluation results are provided, providing strong data support for the optimization of park cities.
Smart Images

Figure CN120069976A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of smart cities, and particularly relates to a multi-scale comprehensive perception method for park cities. Background Art
[0002] With the acceleration of the urbanization process, the continuous expansion of cities, and the rapid increase in the global population, urban development across the globe faces various difficulties and challenges, such as resource shortages, environmental pollution, food safety, economy, transportation, and public safety. Along with the rapid development of science and technology in recent years, the emergence of technologies such as the Internet of Things, cloud computing, and big data has promoted the current cities towards the path of intelligent, inclusive, and sustainable development and formed a people-oriented social intelligent innovation and sustainable urban form. A smart city is the intelligence of a digital city, aiming to seamlessly connect the physical city with the digital city through Internet of Things technology, and use cloud computing and big data technology to process the perception data obtained in real time and provide intelligent services.
[0003] Urban perception is the foundation for building a smart city. With the increasing development of modern smart cities, urban residents have higher requirements for public spaces, especially urban parks. How to accurately evaluate urban public spaces and create a better-quality urban environment that urban residents are more satisfied with has become an important topic. In the existing methods, social media text data solves the problem that traditional methods cannot adapt to a large number of evaluation objects and heavy workloads. However, the evaluation research based on review texts does not dig deep enough into the text data, and the utilization of text data stays at relatively superficial levels such as the number of comments, word frequency, and overall sentiment analysis. However, user comments often contain evaluation elements at multiple scales of the evaluation object. The existing methods cannot analyze the multiple scales of user comments, resulting in the perception results of the multi-scales of parks in the city being not fine and accurate enough. Summary of the Invention
[0004] The present invention provides a multi-scale comprehensive perception method for park cities, which cuts the evaluation text data to obtain short sentences for scale division and sentiment tendency analysis, so as to evaluate the importance of relevant scales of parks, and objectively evaluate the park at different scales from the perspective of users to obtain more refined evaluation results.
[0005] The present invention provides a multi-scale comprehensive perception method for park cities, including:
[0006] Obtain the review text data of multiple parks in the park city; wherein, the review text data is network evaluation data or questionnaire survey data;
[0007] Use the keyword retrieval method to cut the review text data into short sentences, and classify the short sentences based on multi-scale keyword groups and a preset classification method;
[0008] Evaluate the short sentence through the sentiment analysis algorithm, determine the sentiment of the short sentence, and calculate the score of the short sentence;
[0009] Determine the text sentiment, the total text score, and the total park score on the same scale according to the sentiment and score of the short sentence;
[0010] Take the total park score and the total text scores at multiple scales as the dependent variable and independent variable, and obtain the weights of each category through the multivariate linear regression method;
[0011] Determine the proportion of the sentiment of each category, combine it with the weights of each category to quantify the public's perception of different scales of the park, and obtain the perception preference and satisfaction to optimize the corresponding scale of the park city.
[0012] Further, the step of using the keyword retrieval method to cut the review text data into short sentences and classify the short sentences based on the keyword groups at multiple scales and the preset classification method includes:
[0013] Clause the review text data F into single-content sentences S, and then tokenize the sentences S into multiple words W;
[0014] Preset the corresponding relationship between multiple words and scales to form a scale word list, and retrieve each word of the input tokenized sentence one by one. If the word W i (i = 1, 2, 3,.., n) exists in the category word list, return the category corresponding to the word and use it as the category of the sentence S i (i = 1, 2, 3,.., n);
[0015] If all words are not in the category word list, use the preset dual-channel feature fusion and adversarial training model for polar short sentence classification.
[0016] Further, in the step of using the preset dual-channel feature fusion and adversarial training model for polar short sentence classification if all words are not in the category word list, the dual-channel feature fusion and adversarial training model includes an input layer, a ChineseBERT layer, an FGM adversarial training layer, a dual-channel feature extraction layer, a feature fusion layer with multi-head attention, and an output layer;
[0017] The input layer preprocesses the short sentence, sets the maximum sequence length of the sentence to 32. For each sentence, if the length exceeds 32, it is truncated, and if it is less than 32, it is padded with 0;
[0018] The ChineseBERT layer uses ChineseBERT for word embedding representation. After the word vector matrix W passes through the ChineseBERT word embedding, the word embedding representation H = {h 1 ,h 2 ...h n} of each word is obtained, where h i is the word embedding vector of the i-th word;
[0019] The FGM adversarial training layer calculates the gradient g of the loss function of the model for each word embedding vector h i . According to the gradient, an adversarial perturbation is generated, that is, a perturbation is added to each word embedding, and the size of the perturbation value is set to 0.05;
[0020] The dual-channel feature extraction layer processes the adversarially trained word embedding representation through the DPCNN layer and the BiGRU layer to obtain a feature matrix;
[0021] The feature fusion layer of multi-head attention converts the feature matrices from different channels into Q, K, and V matrices corresponding to each head, then calculates the attention weight matrix of each head, and finally concatenates the outputs of all heads to obtain the representation form after feature fusion;
[0022] The output layer takes the fused features as the input of the model, and uses the activation function Soft-max to perform a fully connected operation on the concatenated result to obtain the final classification result.
[0023] Furthermore, in the dual-channel feature extraction layer,
[0024] After obtaining the adversarially trained word embedding representation, the DPCNN layer is sent to multiple convolutional blocks of the DPCNN. Through its convolutional operation and residual connection, its feature representation is obtained. As the depth of the model increases, higher-level features are obtained, and the obtained features are effectively integrated through the pooling operation. The specific formula is:
[0025]
[0026] D = MaxPooling(R)
[0027] where C is the convolutional feature obtained through one-dimensional convolution, R is the residual connection between the convolutional feature and the original sequence, D represents the feature sequence obtained through the pooling operation, and the above formula is calculated repeatedly to obtain the final feature matrix representation D;
[0028] The BiGRU layer takes the output after obtaining the adversarially trained word embedding representation as the input of the BiGRU. The formula is:
[0029]
[0030] B = h 1 ⊕ h 2 … ⊕ h n
[0031] Wherein, are the outputs of the forward and reverse GRUs at time t respectively, is the input at time t, and B is the concatenation result of all outputs.
[0032] Furthermore, the steps of evaluating the short sentence through the sentiment analysis algorithm, determining the sentiment of the short sentence, and calculating the score of the short sentence include:
[0033] Analyze the short sentence through the sentiment analysis API to obtain the score of the short sentence, and the calculation is as follows:
[0034] s = 5 × p
[0035] Wherein, p is the confidence level that the sentiment of the short sentence is positive, in the range [0, 1], and the score s of the short sentence is calculated, in the range [0, 5];
[0036] In order to analyze the positive and negative factors of the park evaluation, the short sentences are divided into positive and negative parts, and the calculation is as follows:
[0037]
[0038] Wherein, c is the sentiment of the short sentence, positive or negative; pos is positive; neg is negative; p is the confidence level that the sentiment of the sentence is positive.
[0039] Furthermore, the steps of determining the text sentiment, text total score, and park total score under the same scale according to the sentiment and score of the short sentence include:
[0040] For the scores s of the short sentences under the same scale i (i = 1, 2, 3,.., N) take the mean value to obtain the text total score under the same scale:
[0041]
[0042] For the text total scores m of multiple scales i (i = 1, 2, 3,.., M) take the mean value to obtain the park total score y:
[0043]
[0044] Count the number of different sentiment tendencies of multiple short sentences under the same scale, and use the ratio of the number of different sentiment tendencies as the text sentiment under this scale.
[0045] Further, the step of taking the overall park score and the overall text scores at multiple scales as the dependent variable and independent variables respectively, and obtaining the weights of various categories through the multivariate linear regression method includes:
[0046] Taking the scores m i (i = 1, 2, 3, 4, M) calculated for each scale and the overall park score y as the independent variable and the dependent variable respectively, and obtaining the weights w i (i = 1, 2, 3, 4, M) of various categories through the multivariate linear regression method:
[0047] y = w 1 x 1 + w 2 x 2 + … + w 5 x 5
[0048] The larger the weight w i , the larger the proportion of the corresponding scale score m i in the overall park evaluation, indicating that this category is a relatively important factor for visitors to evaluate the park; in addition, the fitting effect of the multivariate linear regression is evaluated by the coefficient of determination, and the calculation is as follows:
[0049]
[0050] where f i is the model predicted value; y i is the label value; is the average label value, and the range of the coefficient of determination is [0, 1]. The larger it is, the better the fitting effect.
[0051] The present invention also provides a multi-scale comprehensive perception device for a park city, including:
[0052] An acquisition module, configured to acquire review text data of multiple parks in the park city; wherein, the review text data is network evaluation data or questionnaire survey data;
[0053] A cutting module, configured to cut the review text data into short sentences by using a keyword retrieval method, and classify the short sentences based on a multi-scale keyword group and a preset classification method;
[0054] An evaluation module, configured to evaluate the short sentences through a sentiment analysis algorithm, determine the sentiment tendency of the short sentences, and calculate the scores of the short sentences;
[0055] A determination module, configured to determine the text sentiment tendency, the overall text score and the overall park score at the same scale according to the sentiment tendency and scores of the short sentences;
[0056] A calculation module, configured to use the total park score and the total text scores at multiple scales as the dependent variable and independent variables respectively, and obtain the weights of each category through a multivariate linear regression method;
[0057] An optimization module, configured to determine the proportion of the sentiment tendency of each category, combine it with the weights of each category to quantify the public's perception of different scales of the park, and obtain the perception preference and satisfaction degree, so as to optimize the corresponding scale of the park city.
[0058] The present invention also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0059] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0060] The beneficial effects of the present invention are as follows:
[0061] The present invention obtains the review text data of multiple parks in the park city, uses the keyword retrieval method to cut short sentences, and conducts multi-scale classification; through the sentiment tendency analysis algorithm, determines the sentiment tendency and score of the short sentences, and further determines the text sentiment tendency, the total text score and the total park score at the same scale; finally, obtains the weights of each category through a multivariate linear regression method; combines with determining the proportion of the sentiment tendency of each category to quantify the public's perception of different scales of the park, and obtains the perception preference and satisfaction degree, so as to optimize the corresponding scale of the park city. It makes an objective evaluation of the park at different scales from the perspective of users, obtains more refined evaluation results, and provides strong data support for the construction of the park city. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention.
[0063] Figure 2 It is a schematic structural diagram of the device according to an embodiment of the present invention.
[0064] Figure 3 It is a schematic internal structure diagram of the computer device according to an embodiment of the present invention.
[0065] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0066] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0067] Such as Figure 1As shown, the present invention provides a multi-scale comprehensive perception method for a park city, comprising:
[0068] S1. Obtaining textual review data of multiple parks in a park city; wherein the textual review data is online evaluation data or questionnaire survey data, and the number of obtained textual review data is more than 500 to avoid analysis errors caused by too few samples.
[0069] S2. Cut the review text data into short sentences using a keyword search method, and classify the short sentences based on multi-scale keyword groups and a preset classification method. Specifically, it includes:
[0070] S201, segmenting the review text data F into sentences S with a single content, and then segmenting the sentence S into multiple words W;
[0071] The obtained review text data F is divided into sentences S, that is, F = {S 1 ,S 2 ,…,S n}; Then, each sentence is divided into multiple words W, that is, S = {W 1 ,W 2 ,…,W n}.
[0072] S202, pre-set the correspondence between multiple words and scales to form a scale word list, search each word one by one for the input word segmented sentence, and if the word W i (i=1,2,3,..,n) exists in the category word list, then the category corresponding to the word is returned and used as the sentence S i (i=1,2,3,..,n) categories.
[0073] Combining the green space system evaluation indicators and the word frequency analysis of the evaluation data, the park evaluation is divided into five scales: transportation, aesthetics, maintenance and safety, market value, protection and inheritance, and the words that can represent these five scales are found from the high-frequency words. For example, the phrases for the scale of transportation are transportation, parking, parking lot, subway, bus, bus, far, near, convenient, and easy to find; the phrases for the scale of aesthetics are beautiful, beautiful, beautiful, beautiful, beautiful, creative, style, nostalgic, harmonious, and exquisite; the phrases for the scale of market value are charging, supporting facilities, tickets, snacks, prices, food, fees, shops, cost-effectiveness, and merchants; the phrases for the scale of maintenance and safety are facilities, pavement, toilets, management, pets, attitude, bamboo chairs, chairs, garbage, and toilets; the phrases for the scale of protection and inheritance are history, culture, folklore, architecture, museums, literature and art, memorial, cultural relics, gardens, and humanities.
[0074] S203. If all words are not in the category word list, then a preset dual-channel feature fusion and adversarial training model is used for polar short sentence classification. Among them, the dual-channel feature fusion and adversarial training model includes an input layer, a ChineseBERT layer, an FGM adversarial training layer, a dual-channel feature extraction layer, a feature fusion layer with multi-head attention, and an output layer;
[0075] Input layer:
[0076] The input layer preprocesses the short sentence, sets the maximum sequence length of the sentence to 32. For each sentence, if the length exceeds 32, it is truncated, and if it is less, it is padded with 0.
[0077] ChineseBERT layer:
[0078] The ChineseBERT layer is the encoding layer. ChineseBERT combines a hybrid encoding strategy of character and word masking. Among them, the character-level masking strategy randomly selects a certain proportion of individual characters for masking and predicts their original characters, and better understands semantic information by inferring the masked characters. Word masking randomly selects a certain proportion of words and masks them, and better understands context information by predicting the whole word. In short, it fuses information such as pinyin, glyph, and context from different perspectives, so as to better capture the structure and semantic information of Chinese texts and enhance the model's understanding ability of Chinese texts. Specifically: Use ChineseBERT for word embedding representation. When the word vector matrix W passes through ChineseBERT word embedding, the word embedding representation H = {h 1 , h 2 ...h n} of each word is obtained, where h i is the word embedding vector of the i-th word.
[0079] FGM adversarial training layer:
[0080] The Fast Gradient Method (FGM) adversarial training technology, as a defense measure, strengthens model training to ensure that the model can make correct decisions when encountering minor or deliberate interferences. FGM adds some perturbations to the word embedding vectors. These perturbations are not added randomly, but are calculated through the gradient information during word embedding, simulating the perturbations that may appear in adversarial attacks to strengthen the recognition of such attacks. After the Chinese-BERT performs word embedding representation, FGM is used to add adversarial training, creating adversarial samples through subtle perturbations to improve the resistance to adversarial attacks and input fluctuations. Specifically: For each word embedding vector h i, calculate the gradient g of its loss function with respect to the model, and generate an adversarial perturbation based on the gradient, that is, add a perturbation to each word embedding, and the size of the perturbation value is set to 0.05; the specific formula is:
[0081]
[0082] where L represents the loss function; H represents the word embedding matrix, that is, the input of the previous layer; θ is the model parameter; y is the corresponding label vector; represents the generated new adversarial sample vector; the perturbed word embedding matrix is denoted as
[0083] Dual-channel feature extraction layer:
[0084] The dual-channel feature extraction layer processes the word embedding representation after adversarial training through the DPCNN layer and the BiGRU layer to obtain a feature matrix.
[0085] 1) The DPCNN layer is a deep pyramid convolutional neural network. Through multiple convolutional techniques and residual connection techniques, by combining the pyramid pooling strategy, it can very effectively obtain the local features and global features of the text, thus achieving the classification effect. Specifically: after obtaining the word embedding representation after adversarial training, it is sent to multiple convolutional blocks of the DPCNN. Through its convolutional operation and residual connection, its feature representation is obtained. As the depth of the model increases, higher-level features are obtained, and the obtained features are effectively integrated through the pooling operation. The specific formula is:
[0086]
[0087] D = MaxPooling(R)
[0088] where C is the convolutional feature obtained through one-dimensional convolution, R is the residual connection between the convolutional feature and the original sequence, D represents the feature sequence obtained through the pooling operation, and the above formula is calculated repeatedly to obtain the final feature matrix representation D;
[0089] 2) The BiGRU layer is a neural network with bidirectional gated unit recurrence. By combining forward and backward GRU units, it can obtain effective information of the context. Specifically: the output after obtaining the word embedding representation after adversarial training is used as the input of the BiGRU. The formula is:
[0090]
[0091] B = h 1 ⊕ h 2 … ⊕ h n
[0092] where, They are the outputs of the forward and reverse GRUs at time t, is the input at time t, and B is the concatenation result of all outputs.
[0093] Multi-head attention feature fusion layer:
[0094] The multi-head attention feature fusion layer uses a two-layer neural network to extract sentence features from different perspectives to retain the diversity of features, and can more effectively capture local and global information of sentences. Since the feature vectors from different channels have different dimensions, if the features from different channels are simply concatenated and added, the advantages of each channel cannot be fully utilized. Fusing their features through the multi-head attention mechanism can learn feature information from different sources, not only retain their respective advantages, but also generate richer and more expressive feature representations, effectively improving the model performance. Different heads focus on different feature subspaces, and through multi-head parallel learning, the relationships between features can be better captured. Specifically: convert the feature matrices from different channels into Q, K, and V matrices corresponding to each head, then calculate the attention weight matrix for each head, and finally concatenate the outputs of all heads to obtain the fused feature representation; the calculation formula for feature fusion is:
[0095]
[0096] X = Concat(Head 1 ,…Head k )W
[0097] Head i refers to the i-th head in the multi-head attention, and X is the final attention weight matrix of the output.
[0098] Output layer:
[0099] The output layer takes the fused features as the input of the model, and uses the activation function Soft-max to perform a fully connected operation on the concatenation result to obtain the final classification result.
[0100] S3. Evaluate the short sentence through a sentiment analysis algorithm to determine the sentiment tendency of the short sentence and calculate the score of the short sentence. Specifically, it includes:
[0101] S301. Analyze the short sentence through a sentiment analysis API to obtain the score of the short sentence. The calculation is as follows:
[0102] s = 5 × p
[0103] where p is the confidence that the sentiment tendency of the short sentence is positive, in the range [0, 1], and the score s of the short sentence is calculated, in the range [0, 5];
[0104] S302. To analyze the positive and negative factors of park evaluation, the short sentences are divided into two parts: positive and negative, and the calculation is as follows:
[0105]
[0106] Among them, c is the sentiment tendency of the short sentence, positive or negative; pos is positive; neg is negative; p is the confidence level that the sentiment tendency of the sentence is positive.
[0107] S4. Determine the text sentiment tendency, text total score, and park total score on the same scale according to the sentiment tendency and score of the short sentence. Specifically, it includes:
[0108] S401. Calculate the mean value of the scores s of the short sentences on the same scale i (i = 1, 2, 3,.., N) to obtain the text total score on the same scale:
[0109]
[0110] S402. Calculate the mean value of the text total scores m of multiple scales i (i = 1, 2, 3,.., M) to obtain the park total score y:
[0111]
[0112] S403. Count the number of different sentiment tendencies of multiple short sentences on the same scale, and use the ratio of the number of different sentiment tendencies as the text sentiment tendency on this scale. For example, if there are 9 short sentences on the same scale, among which 5 short sentences have a negative sentiment tendency and 4 short sentences have a positive sentiment tendency, then the sentiment tendency of this scale is positive / negative = 4 / 5.
[0113] S5. Take the park total score and the text total scores of multiple scales as the dependent variable and independent variables, and obtain the weights of each category through the method of multiple linear regression.
[0114] Take the calculated scores m of each scale i (i = 1, 2, 3, 4, M) and the park overall score y as the independent variable and dependent variable respectively, and obtain the weights w of each category through the method of multiple linear regression i (i = 1, 2, 3, 4, M):
[0115] y = w 1 x 1 + w 2 x 2 +…+ w 5 x 5
[0116] The weight wi The larger it is, the corresponding scale score m i The larger the proportion in the overall evaluation of the park, indicating that this category is a relatively important factor for visitors to evaluate the park; in addition, the fitting effect of multiple linear regression is evaluated by the coefficient of determination, which is calculated as follows:
[0117]
[0118] where f i is the model predicted value; y i is the label value; is the average label value, and the range of the coefficient of determination is [0, 1]. The larger it is, the better the fitting effect.
[0119] S6. Determine the proportion of the emotional tendency of each category, and combine it with the weight of each category to quantify the public's perception of different scales of the park, and obtain the perception preference and satisfaction degree, so as to optimize the corresponding scale of the park city. For example, determine that the proportion of positive / negative emotional tendencies of several scales such as transportation, aesthetics, market value, maintenance and safety, protection and inheritance is 17 / 12, 20 / 25, 17 / 19, 4 / 5, 2 / 3 respectively. Among them, the weights of transportation, aesthetics and market value account for a high proportion. Therefore, according to the weights, the user's perception preference is determined to be transportation, aesthetics and market value. At the same time, the positive emotions of aesthetics and market value are lower than the negative emotions, while the positive emotion of transportation is greater than the negative emotion. Therefore, it is determined that the user is relatively satisfied with the transportation aspect, and it is necessary to optimize the aesthetics and market value scales of the park.
[0120] The present invention obtains the review text data of multiple parks in the park city, uses the keyword retrieval method to cut short sentences, and conducts multi-scale classification; through the emotional tendency analysis algorithm, determines the emotional tendency and score of the short sentences, and then determines the text emotional tendency, the total text score and the total park score under the same scale; finally, obtains the weights of each category by the multiple linear regression method; combines with determining the proportion of the emotional tendency of each category to quantify the public's perception of different scales of the park, and obtains the perception preference and satisfaction degree, so as to optimize the corresponding scale of the park city. Make an objective evaluation of the park at different scales from the user's perspective, obtain more refined evaluation results, and provide strong data support for the construction of the park city.
[0121] As Figure 2 shown, the present invention also provides a park city multi-scale comprehensive perception device, including:
[0122] An acquisition module 1, configured to acquire the review text data of multiple parks in the park city; wherein, the review text data is network evaluation data or questionnaire survey data;
[0123] The cutting module 2 is used to cut the review text data into short sentences by using the keyword retrieval method, and classify the short sentences based on the multi-scale keyword groups and the preset classification method;
[0124] The evaluation module 3 is used to evaluate the short sentences through the sentiment analysis algorithm, determine the sentiment tendency of the short sentences, and calculate the scores of the short sentences;
[0125] The determination module 4 is used to determine the text sentiment tendency, the total text score and the total park score at the same scale according to the sentiment tendency and score of the short sentences;
[0126] The calculation module 5 is used to take the total park score and the total text scores at multiple scales as the dependent variable and independent variables, and obtain the weights of each category through the multivariate linear regression method;
[0127] The optimization module 6 is used to determine the proportion of the sentiment tendency of each category, combine it with the weights of each category to quantify the public's perception of different scales of the park, and obtain the perception preference and satisfaction, so as to optimize the corresponding scale of the park city.
[0128] In one embodiment, the cutting module 2 includes:
[0129] The division unit is used to divide the review text data F into sentences, divide them into single-content sentences S, and then segment the sentences S into multiple words W;
[0130] The setting unit is used to preset the corresponding relationship between multiple words and scales to form a scale word list, and retrieve each word of the input segmented word sentence one by one. If the word W i (i = 1, 2, 3,.., n) exists in the category word list, the corresponding category of the word is returned and used as the category of the sentence S i (i = 1, 2, 3,.., n);
[0131] The classification unit is used to perform polar short sentence classification by using the preset dual-channel feature fusion and adversarial training model when all words are not in the category word list.
[0132] In one embodiment, in the classification unit, the dual-channel feature fusion and adversarial training model includes an input layer, a ChineseBERT layer, an FGM adversarial training layer, a dual-channel feature extraction layer, a feature fusion layer with multi-head attention, and an output layer;
[0133] The input layer preprocesses the short sentence, sets the maximum sequence length of the sentence to 32. For each sentence, if the length exceeds 32, it is cropped, and if it is less than 32, it is padded with 0;
[0134] The ChineseBERT layer uses ChineseBERT for word embedding representation. After the word vector matrix W passes through the ChineseBERT word embedding, the word embedding representation H = {h 1 , h 2 ... h n} of each word is obtained, where h i is the word embedding vector of the i-th word;
[0135] The FGM adversarial training layer calculates the gradient g of the loss function of the model for each word embedding vector h i , and generates an adversarial perturbation according to the gradient, that is, adds a perturbation to each word embedding, and the size of the perturbation value is set to 0.05;
[0136] The dual-channel feature extraction layer processes the word embedding representation after adversarial training through the DPCNN layer and the BiGRU layer to obtain a feature matrix;
[0137] The feature fusion layer of multi-head attention converts the feature matrices from different channels into Q, K, V matrices corresponding to each head, then calculates the attention weight matrix of each head, and finally splices the outputs of all heads to obtain the representation form after feature fusion;
[0138] The output layer takes the fused features as the input of the model, and uses the activation function Soft-max to perform a full connection on the splicing result to obtain the final classification result.
[0139] In one embodiment, in the dual-channel feature extraction layer of the classification unit,
[0140] After obtaining the word embedding representation after adversarial training, the DPCNN layer is sent to multiple convolutional blocks of the DPCNN, and through its convolutional operation and residual connection, its feature representation is obtained. As the depth of the model increases, higher-level features are obtained, and the obtained features are effectively integrated through the pooling operation. The specific formula is:
[0141]
[0142] D = MaxPooling(R)
[0143] where C is the convolutional feature obtained through one-dimensional convolution, R is the residual connection between the convolutional feature and the original sequence, D represents the feature sequence obtained through the pooling operation, and the above formula is calculated repeatedly to obtain the final feature matrix representation D;
[0144] The BiGRU layer takes the output after obtaining the word embedding representation after adversarial training as the input of the BiGRU. The formula is:
[0145]
[0146] B = h 1 ⊕ h 2 … ⊕ h n
[0147] Among them, are the outputs of the forward and backward GRUs at time t respectively, is the input at time t, and B is the concatenation result of all outputs.
[0148] In one embodiment, the evaluation module 3 includes:
[0149] An analysis unit, configured to analyze the short sentence through a sentiment analysis API to obtain a score of the short sentence, and the calculation is as follows:
[0150] s = 5 × p
[0151] Among them, p is the confidence that the sentiment of the short sentence is positive, in the range [0, 1], and the score s of the short sentence is calculated, in the range [0, 5];
[0152] A calculation unit, configured to divide the short sentence into positive and negative parts in order to analyze the positive and negative factors of the park evaluation, and the calculation is as follows:
[0153]
[0154] Among them, c is the sentiment of the short sentence, positive or negative; pos is positive; neg is negative; p is the confidence that the sentiment of the sentence is positive.
[0155] In one embodiment, the determination module 4 includes:
[0156] A mean calculation unit, configured to calculate the mean of the scores s of the short sentences on the same scale i (i = 1, 2, 3,.., N) to obtain the total text score on the same scale:
[0157]
[0158] A total score calculation unit, configured to calculate the mean of the total text scores m of multiple scales i (i = 1, 2, 3,.., M) to obtain the total park score y:
[0159]
[0160] A statistics unit, configured to count the number of different sentiment tendencies of multiple short sentences on the same scale, and use the ratio of the number of different sentiment tendencies as the text sentiment tendency on this scale.
[0161] In one embodiment, the calculation module 5 includes:
[0162] Take the calculated scale scores \(m_i\) i (\(i = 1, 2, 3, 4, \cdots, M\)) and the overall park score \(y\) as the independent variable and the dependent variable respectively, and obtain the weights \(w_i\) i (\(i = 1, 2, 3, 4, \cdots, M\)) of each category by means of multiple linear regression:
[0163] \(y = w_1x_1 + w_2x_2+\cdots+w_Mx_M\) 1 \(x_1\) 1 +\(w_2\) 2 \(x_2\) 2 +\(\cdots+\) \(w_M\) 5 \(x_M\) 5
[0164] The greater the weight \(w_i\) i , the greater the proportion of the corresponding scale score \(m_i\) i in the overall park evaluation, indicating that this category is a relatively important factor for visitors to evaluate the park; in addition, the fitting effect of multiple linear regression is evaluated by the coefficient of determination, which is calculated as follows:
[0165]
[0166] where \(f_i\) i is the model prediction value; \(y_i\) i is the label value; \(\overline{y}\) is the average label value, and the coefficient of determination ranges from \([0, 1]\). The larger it is, the better the fitting effect.
[0167] Each of the above modules and units is used to correspondingly execute each step in the above multi-scale comprehensive perception method for park cities. The specific implementation manner refers to the method embodiments described above and will not be elaborated here.
[0168] As Figure 3 shown, the present invention also provides a computer device, which can be a server, and its internal structure can be as Figure 3 shown. The computer device includes a processor, a memory, a network interface, and a database connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store all the data required for the process of the multi-scale comprehensive perception method for park cities. The network interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it realizes the multi-scale comprehensive perception method for park cities.
[0169] Those skilled in the art can understand, Figure 3The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied.
[0170] An embodiment of this application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements any one of the above-mentioned multi-scale integrated perception methods for park cities.
[0171] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned method embodiments. Among them, any reference to a memory, storage, database, or other medium provided in this application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be obtained in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0172] It should be noted that in this article, the terms "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, device, article, or method including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, device, article, or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, device, article, or method including that element.
[0173] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A multi-scale comprehensive perception method for park cities, characterized in that: include: Obtaining textual review data of multiple parks in a park city; wherein the textual review data is online evaluation data or questionnaire survey data; Using a keyword search method to cut the review text data into short sentences, and classifying the short sentences based on multi-scale keyword groups and a preset classification method; Evaluating the short sentence by using a sentiment tendency analysis algorithm, determining the sentiment tendency of the short sentence, and calculating the score of the short sentence; Determine the sentiment orientation of the text, the total score of the text and the total score of the park under the same scale according to the sentiment orientation and score of the short sentence; The total park score and the total text score at multiple scales were used as dependent and independent variables, and the weights of each category were obtained through multivariate linear regression method. Determine the proportion of emotional tendencies in each category and combine it with the weight of each category to quantify the public's perception of parks at different scales, and derive perceived preferences and satisfaction to optimize the corresponding scale of the park city.
2. The multi-scale comprehensive perception method of park city according to claim 1 is characterized in that: The step of cutting the review text data into short sentences by using a keyword search method, and classifying the short sentences based on multi-scale keyword groups and a preset classification method includes: Segment the review text data F into sentences S with a single content, and then segment the sentence S into multiple words W; Preset the corresponding relationship between multiple words and scales to form a scale word list, and search each word one by one for the input segmented sentence. If the word W i (i=1,2,3,..,n) exists in the category word list, then the category corresponding to the word is returned and used as the sentence S i (i=1,2,3,..,n) categories; If all words are not in the category vocabulary, the preset dual-channel feature fusion and adversarial training model polarity short sentence classification is used.
3. The multi-scale comprehensive perception method of park city according to claim 2 is characterized in that: If all words are not in the category vocabulary, a preset dual-channel feature fusion and adversarial training model is used in the step of polarity phrase sentence classification, wherein the dual-channel feature fusion and adversarial training model includes an input layer, a ChineseBERT layer, an FGM adversarial training layer, a dual-channel feature extraction layer, a multi-head attention feature fusion layer, and an output layer; The input layer preprocesses the short sentences and sets the maximum sequence length of the sentences to 32. For each sentence, if the length exceeds 32, it is pruned, and if it is less than 32, it is padded with 0; The ChineseBERT layer uses ChineseBERT for word embedding representation. After the word vector matrix W is embedded by ChineseBERT, the word embedding representation H = {h1, h2...h n }, where h i is the word embedding vector of the i-th word; The FGM adversarial training layer is used for each word embedding vector h i , calculate the gradient g of the loss function of the model, and generate an adversarial perturbation based on the gradient, that is, add perturbation to each word embedding, and the perturbation value is set to 0.05; The dual-channel feature extraction layer processes the word embedding representation after adversarial training through the DPCNN layer and the BiGRU layer to obtain the feature matrix; The feature fusion layer of multi-head attention converts the feature matrices from different channels into Q, K, and V matrices corresponding to each head, then calculates the attention weight matrix of each head, and finally concatenates the outputs of all heads to obtain the representation after feature fusion; The output layer uses the fused features as the input of the model and uses the activation function Soft-max to fully connect the splicing results to obtain the final classification results.
4. The multi-scale comprehensive perception method of park city according to claim 3 is characterized in that: In the dual-channel feature extraction layer, After obtaining the word embedding representation after adversarial training, the DPCNN layer is sent to multiple convolution blocks of DPCNN to obtain its feature representation through its convolution operation and residual connection. As the depth of the model increases, higher-level features are obtained, and the acquired features are effectively integrated through pooling operations. The specific formula is: D=MaxPooling(R) Among them, C is the convolution feature obtained by one-dimensional convolution, R is the residual connection between the convolution feature and the original sequence, and D represents the feature sequence obtained by the pooling operation. Repeat the above calculation to obtain the final feature matrix representation D; The BiGRU layer will output the word embedding representation after adversarial training as the input of BiGRU, and its formula is: in, are the outputs of the forward and reverse GRU at time t, is the input at time t, and B is the concatenation result of all outputs.
5. The multi-scale comprehensive perception method of park city according to claim 1 is characterized in that: The steps of evaluating the short sentence by using a sentiment tendency analysis algorithm, determining the sentiment tendency of the short sentence, and calculating the score of the short sentence include: The short sentence is analyzed through the sentiment analysis API to obtain the score of the short sentence, which is calculated as follows: s=5×p Where p is the confidence that the sentiment tendency of the short sentence is positive, ranging from [0, 1], and the score s of the short sentence is calculated, ranging from [0, 5]; In order to analyze the positive and negative factors of park evaluation, the short sentences are divided into positive and negative parts, and the calculation is as follows: Among them, c is the sentiment tendency of the short sentence, positive or negative; pos is positive; neg is negative; p is the confidence that the sentiment tendency of the sentence is positive.
6. The multi-scale comprehensive perception method of park city according to claim 1 is characterized in that: The step of determining the text sentiment orientation, the text total score and the park total score under the same scale according to the sentiment orientation and the score of the short sentence comprises: The score of the short sentence under the same scale is s i (i=1,2,3,..,N) calculate the average and get the total score of the text under the same scale: The total score m of the text at multiple scales i (i=1,2,3,..,M) calculate the mean and get the total score y of the park: The number of different sentiment tendencies of multiple short sentences under the same scale is counted, and the ratio of the number of different sentiment tendencies is used as the text sentiment tendency under this scale.
7. The multi-scale comprehensive perception method of park city according to claim 1 is characterized in that: The step of using the park total score and the text total score at multiple scales as dependent variables and independent variables, and obtaining the weights of each category by a multivariate linear regression method, comprises: The calculated scores m for each scale are i (i=1,2,3,4,M) and the overall park score y are used as independent variables and dependent variables respectively, and the weights w of each category are obtained by multivariate linear regression method. i (i=1,2,3,4,M): y=w1 x1+w2 x2+…+w5 x5 Weight w i The larger the value, the corresponding scale score m i The larger the proportion in the overall evaluation of the park, the greater the proportion of this category is in the overall evaluation of the park, indicating that this category is a relatively important factor in visitors' evaluation of the park; in addition, the fitting effect of the multivariate linear regression is evaluated by the determination coefficient, which is calculated as follows: Among them, f i is the model prediction value; y i is the label value; is the average label value, and the determination coefficient range is [0, 1]. The larger the value, the better the fitting effect.
8. A multi-scale integrated perception device for a park city, characterized in that: include: An acquisition module is used to acquire textual review data of multiple parks in a park city; wherein the textual review data is online evaluation data or questionnaire survey data; A cutting module, used to cut the review text data into short sentences by using a keyword search method, and classify the short sentences based on multi-scale keyword groups and a preset classification method; An evaluation module, used to evaluate the short sentence by using a sentiment tendency analysis algorithm, determine the sentiment tendency of the short sentence, and calculate the score of the short sentence; A determination module, used to determine the text sentiment orientation, the text total score and the park total score under the same scale according to the sentiment orientation and score of the short sentence; The calculation module is used to use the park total score and the text total score at multiple scales as dependent and independent variables, and obtain the weights of each category through multivariate linear regression method; The optimization module is used to determine the proportion of emotional tendencies in each category and combine it with the weights of each category to quantify the public's perception of parks at different scales, and to derive perceived preferences and satisfaction in order to optimize the corresponding scale of the park city.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.