Urban residential quarter livability evaluation method based on sentiment analysis and GA-BP

Through the method of combining sentiment analysis with GA-BP model, a livability evaluation system for urban residential communities was constructed, which solved the problem of distortion of evaluation results in the existing technology, and achieved scientific and explainable livability evaluation.

CN120579907AInactive Publication Date: 2025-09-02NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511087476.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-09-02
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the evaluation of livability of urban residential communities, there are problems in the existing technology that subjective empowerment laws are easily affected by humans, objective empowerment laws are difficult to consider the connotation of indicators, and machine learning models are affected by user emotional bias, resulting in distortion of evaluation results.

Method used

The evaluation index system was constructed through word frequency analysis by using emotion analysis and GA-BP method, combined with K-means clustering and quartile method to screen samples, train the GA-BP model, and use the SHAP method to explain the influencing factors.

Benefits of technology

Scientific evaluation from the perspective of residents has been realized, important influencing factors of livability have been revealed, differences in expert perspectives have been avoided, and credibility and interpretability of the evaluation have been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579907A_ABST
    Figure CN120579907A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of livability evaluation, and discloses an urban residential quarter livability evaluation method based on sentiment analysis and GA-BP, and the method comprises the following steps: constructing a residential quarter livability evaluation index system based on word frequency analysis guidance; performing sentiment analysis according to the evaluation data, performing k-means clustering according to objective condition attributes of the cells, and rejecting prejudice cell samples by using a quartile method to obtain training samples; a GA-BP model is trained, and cell livability evaluation is carried out; an SHAP method is used for explaining the model, and community livability influence factors concerned by residents are analyzed. According to the urban residential quarter livability evaluation method based on sentiment analysis and GA-BP, the BP neural network model fused with the improved genetic algorithm is constructed, the willingness of residents and index characteristics are fully considered, the view angle difference of experts is avoided, residential quarter livability evaluation is more scientific, and important influence factors of livability are revealed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of livability evaluation, and in particular to a livability evaluation method for urban residential areas based on sentiment analysis and GA-BP. Background Art

[0002] As the direct vehicle for urban life, objectively evaluating the livability of residential communities can help policymakers better understand the state of the human settlement environment and formulate relevant development strategies. Residential communities are the smallest units of urban population concentration, directly impacting residents' quality of life and well-being. Therefore, evaluating their livability has been a key research topic in recent years.

[0003] Livability evaluation often requires a comprehensive assessment of multiple indicators. Comprehensive indicator system evaluation models are the mainstream approach for multi-indicator evaluations both domestically and internationally. These methods construct an appropriate evaluation indicator system based on the subject being evaluated, use subjective and objective weighting methods to determine indicator weights, and finally select an aggregate model to produce a comprehensive evaluation result. However, subjective weighting methods are susceptible to human influence and distortion during operation, and they ignore the diverse characteristics of indicators. Objective weighting methods lack consideration of indicator connotations, fail to incorporate domain knowledge, and suffer from poor interpretability. While combined weighting methods combine the advantages of subjective and objective weighting methods to some extent, further research is needed to determine how to combine them to balance the subjective and objective advantages of each indicator. Most indicators still only reflect a single subjective or objective weight, and the process is cumbersome, with low operability and interpretability.

[0004] With the rapid development of big data and machine learning technologies, scholars at home and abroad have introduced machine learning into multi-metric evaluation, using machines to discover patterns and correlations from large amounts of data to address the challenges of metric weighting and aggregation. However, the application of machine learning in multi-metric evaluation is fraught with difficulties. For areas without precise label values, such as the livability evaluation of urban residential areas, methods such as volunteer scoring and classical evaluation models are used to derive sample evaluation results. These samples are then used to train machine learning models. However, these evaluation results are distorted by the subjective opinions of volunteers or by flaws in the comprehensive evaluation model. To address this training sample issue, user evaluation data from social media networks is used. However, user sentiment information is biased, and using biased data to train neural networks will undoubtedly reduce the credibility of the evaluation results.

[0005] Therefore, to address the above problems, the present invention conducts sentiment analysis on user evaluation data, clusters communities with similar characteristics using the k-means algorithm, adopts the quartile method to eliminate abnormal samples in each type of community, trains a GA-BP model for evaluating the livability of residential communities, and then evaluates the livability of all residential communities in City A. Finally, the SHAP method is used to identify the factors affecting the livability of residential communities. Summary of the Invention

[0006] The purpose of the present invention is to provide a livability evaluation method for urban residential areas based on sentiment analysis and GA-BP. Starting from the perspective of residents, based on user evaluation data, a training sample screening method of coupling sentiment analysis and K-means clustering is constructed to construct a BP neural network model that integrates an improved genetic algorithm. This method fully considers the residents' wishes and indicator characteristics, avoids the differences in experts' perspectives, makes the community livability evaluation more scientific, and reveals the important influencing factors of livability.

[0007] To achieve the above objectives, the present invention provides a method for evaluating the livability of urban residential areas based on sentiment analysis and GA-BP, comprising the following steps: Step S1: Based on the word frequency analysis, a residential area livability evaluation index system is constructed; Step S2: sentiment analysis is performed based on the evaluation data, k-means clustering is performed based on the objective conditions and attributes of the cells, and biased cell samples are eliminated using the quartile method to obtain training samples; Step S3: training the GA-BP model to evaluate the livability of the community; Step S4: Use the SHAP method to interpret the model and analyze the factors affecting the livability of the community that residents are concerned about.

[0008] Preferably, in step S1, traditional residential area livability evaluation indicators and real estate network evaluation data are integrated to extract residents' concerns. Under the guidance of hot words, an urban residential area livability evaluation index system is constructed from the perspective of residents. The specific process is as follows: Step S11: Acquire residential community research data, specifically including residential community data, infrastructure point of interest data, road network data, population data, and mobile phone signaling data; Step S12: Use word frequency analysis to assist in index system construction, and quantitatively analyze user evaluation information with the help of Jieba word segmentation algorithm to obtain text hot spots and their changing trends; Step S13: Under the guidance of hot words, a community livability evaluation index system is constructed from the perspective of residents from six aspects: living cost, residential safety, environmental health, living comfort, living convenience, and travel convenience.

[0009] Preferably, in step S2, based on the sentiment dictionary method, sentiment analysis is performed on the community evaluation data with the help of the ROST EA content mining system, characteristic vocabulary of sentiment words, degree words, and negative words are found from the text, and the sentiment value of each characteristic word is found in the sentiment dictionary, and sentiment classification is performed according to the accumulated sentiment value to judge the residents' emotional attitude towards the residential community.

[0010] Preferably, in step S2, the K-means algorithm is used to divide the residential communities into multiple clusters according to feature similarity, the elbow method is used to determine the initial number of clusters k, the silhouette coefficient S is selected as the cluster evaluation index, and the quartile method is used within the same cluster to remove the residential community samples with biased emotional attitudes that are less than the lower quartile and greater than the upper quartile, and the remaining residential communities constitute the training sample set.

[0011] Preferably, the silhouette coefficient calculation formula is as follows: (1); (2); in, For samples Silhouette coefficient; For samples The average distance to all other samples in the same cluster; For samples The average distance to all samples in the nearest cluster; is the overall silhouette coefficient of the cluster, and its value range is [-1, 1]; is the total number of samples in the dataset.

[0012] Preferably, in step S3, a BP neural network model integrating an improved genetic algorithm, i.e., a GA-BP model, is constructed and trained to evaluate the livability of the community; wherein the specific operation process of the GA-BP model includes the following steps: Step S31, determining the BP neural network topology; The BP neural network topology is a three-layer structure with 5 neurons in the hidden layer. The ReLu function is used as the activation function and the Adam algorithm is used for parameter optimization. Step S32: Encode the weight and bias parameters to form a chromosome, and use the genetic algorithm optimized by the elite selection strategy to determine the optimal individual; Step S33: Initialize using the optimal weights and bias parameters, and further optimize the model through training.

[0013] Preferably, in step S4, the SHAP method is used to interpret the GA-BP model and analyze the factors affecting the livability of the community that residents are concerned about; Among them, SHAP is used to analyze the contribution of each evaluation index to the livability of urban residential areas, as shown below: (3); in, Indicates the SHAP values ​​of input features; Indicates that except The set of all features except features; F represents the set of all features; Representation model; Indicates that features are included and collection Input of features in ; Indicates that only the collection is included The SHAP method calculates the features of the model through the data distribution and the explained variables. , original model Explanatory model , as shown below: (4); in, is the benchmark value, which represents the average predicted value of all samples; Indicates the Whether a feature participates in model prediction; Indicates the number of features in the model.

[0014] Therefore, the present invention adopts the above-mentioned urban residential area livability evaluation method based on sentiment analysis and GA-BP. Starting from the perspective of residents, based on user evaluation data, the training sample screening method of coupling sentiment analysis and K-means clustering is constructed to construct a BP neural network model integrating the improved genetic algorithm. It fully considers the residents' wishes and indicator characteristics, avoids the differences in experts' perspectives, makes the community livability evaluation more scientific, and reveals the important influencing factors of livability. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is the technical roadmap of the urban residential area livability evaluation method based on sentiment analysis and GA-BP of the present invention; Figure 2 It is the flow chart of GA-BP algorithm of the present invention; Figure 3 It is the clustering result diagram of the present invention; Figure 4 is the importance and contribution rate of the indicators of the present invention; where (a) is the contribution value of the SHAP analysis criterion layer indicator; (b) is the contribution value of each indicator of the SHAP analysis; (c) is the SHAP global bar chart; Figure 5 It is a bee colony diagram of SHAP values ​​of factors affecting habitability of the present invention; Figure 6 : This is a SHAP feature dependency graph in an embodiment of the present invention; wherein, (a) is the construction year; (b) is the park green space; (c) is the price; (d) is the greening rate; (e) is the cultural facilities; (f) is the property fee; (g) is the road network density; and (h) is the educational facilities. DETAILED DESCRIPTION

[0016] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.

[0017] like Figure 1 As shown, the urban residential area livability evaluation method based on sentiment analysis and GA-BP of the present invention includes the following steps: Step S1: Based on the word frequency analysis, a residential area livability evaluation index system is constructed; Step S2: sentiment analysis is performed based on the real estate website evaluation data, k-means clustering is performed based on the objective conditions and attributes of the community, and biased community samples are eliminated using the quartile method to obtain training samples; Step S3: training the GA-BP model and evaluating the livability of all communities in city A; Step S4: Use the SHAP method to interpret the model and analyze the factors affecting the livability of the community that residents are concerned about.

[0018] Example 1 Step S1: Integrate traditional residential area livability evaluation indicators and real estate network evaluation data to extract residents' concerns. Guided by hot words, construct an urban residential area livability evaluation indicator system from the perspective of residents.

[0019] Step S11: Acquire residential community research data.

[0020] The research data mainly includes residential community data of a certain city, infrastructure point of interest (POI) data, road network data, population data, and mobile phone signaling data.

[0021] (1) Residential community data: Residential communities with complete information were obtained through Baidu Maps AOI (area of ​​interest) screening. Residential community attribute data and evaluation data were then obtained from a real estate website. Residential community attribute data included community name, price, property fees, ownership period, construction year, greening rate, floor area ratio, property management company, and number of parking spaces. Evaluation data included review details, the community being reviewed, user ID, and posting time.

[0022] Here, we obtained complete information for 1,479 communities and 88,172 comments. After removing duplicate comments and deleting meaningless and short comments, 71,040 valid data items remained.

[0023] (2) Infrastructure POI data: obtained from the Baidu Map Open Platform, and screened out seven categories of data, including polluting enterprises, parks and green spaces, educational facilities, cultural facilities, medical facilities, commercial facilities, and transportation facilities, totaling more than 14,000 items, based on the Baidu Map POI industry classification rules and research needs.

[0024] (3) Mobile phone signaling data: Mobile phone signaling data is obtained from the operator and used for regression modeling to calculate the population density of each community.

[0025] Step S12: deriving text hot spots and their changing trends through word frequency statistical analysis.

[0026] The present invention adopts the precise mode of Jieba word segmentation algorithm to perform word segmentation, stop word removal, and part-of-speech tagging on the 71,040 text data obtained, and then counts the number of occurrences of each word. The word frequency statistical results are manually screened according to whether the words have analytical value, and the word frequency statistical results are cleaned. The word screening examples are shown in Table 1.

[0027] Table 1 Word screening table ;

[0028] The overall frequency of words in user reviews shows an exponential downward trend, indicating that residents' comments on residential communities are similar, with a focus on a specific area, demonstrating a significant trend of hot spots. We extracted 91 high-frequency words with a frequency greater than 1,000. The most frequent occurrences were related to amenities, price, transportation, surroundings, environment, and property management, corresponding to six key areas: living costs, residential safety, environmental health, residential comfort, living convenience, and travel convenience. The presence of high-frequency words such as price, property management fees, and property rights indicates that residents are particularly concerned about the cost of living, so we incorporated indicators related to living costs into the evaluation system.

[0029] Step S13: Under the guidance of hot words, construct an urban residential area livability evaluation index system from the perspective of residents.

[0030] This paper follows the principles of scientificity, typicality, operability, and differentiation, and comprehensively considers the actual situation of the study area, the availability of indicators, and the results of text word frequency analysis. It constructs a residential livability evaluation index system for a certain city with a total of 19 indicators from six aspects: living cost, residential safety, environmental health, living comfort, living convenience, and travel convenience, as shown in Table 2.

[0031] Table 2 Livability evaluation index system for a certain city residential area ;

[0032] For communities with incomplete real estate information on the Internet, the greening rate, volume ratio, and parking space ratio indicators were filled with the mean; the building age indicator was filled with the mode; the price indicator was closely related to the geographical location, so the multi-scale geographically weighted regression (MGWR) model was used to fill the null value; the property fee indicator was more affected by the geographical environment than the geographical location, so the Geographically Optimal Similarity (GOS) model was used to regress the property fee indicator. Finally, the goodness of fit of the MGWR model for the price indicator was R 2 The goodness of fit of the GOS model for property fee index is 0.87. 2 It is 0.71.

[0033] Step S2: sentiment analysis is performed based on the real estate website evaluation data, k-means clustering is performed based on the objective conditions and attributes of the community, and the community samples with biased sentiment attitudes are eliminated using the quartile method to obtain training samples.

[0034] Step S21: Based on the sentiment dictionary method, sentiment analysis is performed on the community evaluation data with the help of the ROST EA content mining system to determine the residents' emotional attitudes towards the residential community.

[0035] Step S22: Use K-means clustering to divide residential communities into different types, and use the quartile method to remove residential community samples with biased sentiment attitudes that are smaller than the lower quartile and larger than the upper quartile. The remaining residential communities constitute the training sample set.

[0036] Based on the indicators in Table 2, K-means clustering was used to classify residential communities. The initial number of clusters, k, was determined using the "elbow method." The silhouette coefficient (S) was selected as the clustering effectiveness evaluation indicator. Within the same category, the quartile method was used to remove residential community samples with sentiment attitudes that were less than the lower quartile and greater than the upper quartile. The silhouette coefficient calculation formula is as follows: (1); (2); in, For samples Silhouette coefficient; For samples The average distance to all other samples in the same cluster; For samples The average distance to all samples in the nearest cluster; is the overall silhouette coefficient of clustering, and its value range is [-1, 1]. The closer its value is to 1, the better the clustering effect is; is the total number of samples in the dataset.

[0037] Step S3: construct and train a BP neural network model (GA-BP model) that integrates an improved genetic algorithm to evaluate the livability of all communities in a city.

[0038] This paper constructs a BP neural network model (GA-BP model) that integrates an improved genetic algorithm to evaluate the livability of residential areas. This model uses a BP neural network as its main framework and employs a genetic algorithm optimized with an elite selection strategy to initialize the BP neural network's weights and bias parameters, thereby improving the model's convergence speed and prediction accuracy.

[0039] like Figure 2 As shown in the figure, the specific operation process of the GA-BP model includes the following steps: Step S31: Determine the BP neural network topology structure.

[0040] The network has a three-layer structure with five hidden layer neurons. The ReLu function was used as the activation function, and the Adam algorithm was used for parameter optimization. After normalizing each indicator, a multicollinearity test was conducted. After removing indicators with a variance inflation factor (VIF) greater than 4, namely population density, the remaining 18 indicators were input into the BP neural network.

[0041] Step S32: Encode the weights and bias parameters to form chromosomes, and use the genetic algorithm optimized by the "elite selection strategy" to determine the optimal individual.

[0042] Among them, the genetic algorithm optimized by the "elite selection strategy" directly copies the best individuals that appear in the population during the evolution process to the next generation without pairing and crossover, so as to preserve the genes of the best individuals and avoid the problem that the classical genetic algorithm is prone to losing the best individuals, resulting in a decrease in the algorithm's convergence efficiency or even failure to converge.

[0043] Step S33: Initialize using the optimal weights and bias parameters, and further optimize the model through training.

[0044] Step S4: Use the SHAP method to interpret the GA-BP model and analyze the factors affecting the livability of the community that residents are concerned about.

[0045] This paper uses SHAP to analyze the contribution of each evaluation index to the livability of urban residential areas, as shown below: (3); in, Indicates the SHAP values ​​of input features; Indicates that except The set of all features except features; F represents the set of all features; Representation model; Indicates that features are included and collection Input of features in ; Indicates that only the collection is included The input of the features.

[0046] The SHAP method calculates the benchmark value based on the data distribution and the model's explained variables. , original model Explanatory model , as shown below: (4); in, is the benchmark value, which represents the average predicted value of all samples; Indicates the Whether a feature participates in model prediction; Indicates the number of features in the model.

[0047] Example 2 1. Construct a residential community livability evaluation model.

[0048] This example used the ROST EA 1.9 content mining system to perform sentiment analysis on user review texts. 200 reviews were selected for manual judgment and compared with the sentiment analysis results. This method achieved an accuracy of 96%, making it suitable for subsequent analysis. Sentiment analysis values ​​were calculated for both individual reviews and neighborhoods, yielding the distribution of user review data and residents' sentiment toward the neighborhoods. Table 3 shows that the vast majority of reviews, totaling 93.87%, were positive, with generally positive sentiment being the largest, accounting for 37.55%. Analyzing neighborhoods with 50 or more reviews (294 in total), the majority of neighborhoods received generally positive or moderately positive reviews, accounting for 95.24%. Highly positive reviews accounted for only 2.04%, significantly less than the 26.08% for individual reviews, indicating a relatively dispersed distribution of highly positive reviews.

[0049] Table 3 Sentiment analysis results ;

[0050] In order to remove biased reviews and obtain more reliable training samples, this embodiment uses the sentiment analysis results to select 294 communities with 50 or more reviews as preliminary training samples, performs k-means clustering, and uses the quartile method to remove outliers according to category. According to the "elbow method", k=5 is the inflection point of the error square sum curve. At this time, the silhouette coefficient S value is 0.56, which can better distinguish each community and is the optimal point. Using k-means to perform five-category clustering, the results are as follows: Figure 3As shown in the figure, Category 1 and Category 2 communities are distributed in the innermost layer, with a relatively small number of communities, including 12 Category 1 communities and 29 Category 2 communities. Category 1 communities are clustered in the core area of ​​City A, and both categories are located within the main urban area. The middle layer is primarily composed of Category 3 and Category 4 communities, which are the largest, with a total of 163 communities, most of which are located on the edge of the main urban area. The outermost layer is Category 5 communities, with a total of 90 communities, the vast majority of which are located outside the main urban area. The distribution of communities by category shows a circular pattern, which is closely related to the stratification of location value, the timing of urban expansion, the distribution of infrastructure, and the development of transportation routes.

[0051] After using the quartile method to remove cells with abnormal sentiment scores within each cell type, a total of 281 training samples were obtained. The mean sentiment value of these training samples was 13.80, with a standard deviation of 3.98. 81% of the samples scored between 9 and 18, generally conforming to a normal distribution. The GA-BP model was trained using cell sentiment values ​​as labels and the quantized values ​​of the evaluation indicators as features. After multiple training runs, the genetic algorithm was trained with an initial population size of 1000, a crossover probability of 0.8, 80 iterations, a model learning rate of 0.8, a regularization parameter of 0.15, and a maximum number of 300 iterations.

[0052] This embodiment tests the goodness of fit of the model from three aspects: loss function iteration error, prediction accuracy, and PSI. As the number of iterations increases, the errors on both the training sample and the verification sample decrease exponentially, the model convergence effect and generalization ability are good, and the degree of fit is high. 10% of the data were used as verification samples for 10-fold cross-validation, and the final training sample error was 8.57, and the verification sample error was 8.77. According to the emotional value interval, the community was divided into seven categories: highly negative, moderately negative, generally negative, neutral, generally positive, moderately positive, and highly positive for classification accuracy testing. The PSI value on the training sample was 0.092, and the PSI value on the verification sample was 0.097. The model prediction accuracy was 0.75, and the prediction accuracy was relatively high.

[0053] 2. Distribution pattern of livability of residential communities.

[0054] The trained GA-BP model was used to evaluate the livability of all residential communities in City A. In terms of numerical distribution, the vast majority of residential communities in City A received sentiment scores greater than 5, indicating high overall scores, good livability, and a high sense of well-being among residents. However, the livability scores of residential communities in City A generally ranged from 5 to 25, accounting for 97%. Residents expressed highly positive sentiment in only a small number of communities, only 2.30%, indicating a relatively small disparity in livability between communities. Using the natural breakpoint method, communities were further categorized into five categories: most livable, livable, average, unlivable, and least livable. The number and proportion of communities in each category were calculated, as shown in Table 4. The average livability rating was the largest, accounting for 33.40%, followed by the unlivable and livable categories, at 29.68% and 20.15%, respectively. The most livable category accounted for the smallest proportion, at only 5.81%.

[0055] Table 4 Statistics of residential areas with different livability ;

[0056] In terms of spatial distribution, residential communities are clustered within the main urban area and distributed linearly along highways outside the city. The livability of residential communities in City A decreases gradually from the city center to the periphery. The most livable communities are primarily located near Xinjiekou in the city center. This area boasts excellent surrounding amenities, including a dense network of public services, medical and health facilities, and abundant educational, cultural, and commercial facilities. It is close to public transportation and has a dense network of subway lines, providing quick access to major business districts. Living and commuting are convenient, resulting in high resident satisfaction. As distance from the city center increases, surrounding amenities decrease, and the livability of residential communities decreases. Residential communities with lower livability are primarily located along the Yangtze River, east of the Qinhuai River, and outside the Shanghai-Chengdu Expressway. Exceptionally high livability values ​​are found south of the Qinhuai New River and west of the Qinhuai River. This is due to the presence of commercial clusters, proximity to the airport and railway station, and a rich road network. Furthermore, the livability of a community is not completely correlated with its location; the livability of neighboring communities varies, but is generally close.

[0057] 3. Analysis of factors affecting the livability of residential communities.

[0058] This example uses the SHAP method to analyze the impact of various indicators on the livability of residential areas and their contributions. The results are as follows: Figure 4As shown in the figure, building age, parks and green space, price, greening rate, cultural facilities, property management fees, road network density, and educational facilities are key indicators, accounting for 66% of all indicators' contribution. With the exception of residential safety, the other five factors have roughly equal impact on community livability. Environmental health indicators contribute the most, at 24%, followed by residential comfort indicators, at 23%. The impact of residential comfort indicators is primarily driven by building age, which far outweighs all other indicators. This is due to the high correlation between building age and multiple indicators. While there are fewer indicators related to living costs, each has a significant impact, demonstrating that while residents value environmental health and residential comfort, they also consider living costs. Residential safety indicators contribute very little to livability, at only 3%. Property management indicators also contribute the least, at 3%. This is due to the low variance within property management indicators and the fact that most communities have well-established property management systems.

[0059] The SHAP value bee colony diagram combines the importance and characteristic effect of each indicator to present the SHAP value distribution of each indicator. The indicators are sorted from top to bottom according to their importance. The positive or negative SHAP value of the characteristic point represents the positive or negative effect of the point. Its distribution indicates the size of the indicator's influence and its correlation with livability. Figure 5 As can be seen, building age is the most important factor influencing the livability of residential communities in City A. This indicator is negatively correlated with livability overall, with characteristic points concentrated in high-value areas and dispersed in low-value areas. The SHAP values ​​of most points range from -5 to 5. This indicates that newer buildings have a smaller positive effect on residential livability, while older buildings have a highly significant negative impact. Parks and green spaces do not have a significant correlation with residential livability. However, when park and green space is in high-value areas, it has a more pronounced impact on residential livability, with a more dispersed distribution and a greater influence. Prices, greening ratios, cultural facilities, and road network density are positively correlated with livability. The impact patterns of prices and cultural facilities on residential livability are similar. When the values ​​are small, the samples are clustered and the impact on livability is not significant. However, when the values ​​exceed a certain threshold, the samples become more dispersed, indicating a significant positive impact on livability. Property fees, educational facilities, and PM2.5 levels are negatively correlated with livability, with some high property fee values ​​having a positive effect on livability. Property ownership periods of 70 years have little impact on livability, but those below 70 years have a significant negative impact. Other indicators, such as bus stops, medical facilities, and floor area ratio, have no significant impact on livability.

[0060] The above eight key indicators are selected to further analyze their impact patterns. Figure 6The change of SHAP value as the characteristic value of each influencing factor changes is described.

[0061] (1) The relationship between building age and residential area livability is complex, first showing a linear decreasing relationship, then gradually diverging, such as Figure 6 (a) in the data. Before 2004, the age of a residential complex had a positive impact on livability, with this positive effect gradually weakening. Between 2004 and 2012, it had a negative impact on livability, with this negative effect gradually strengthening. After 2012, the impact of the age of a residential complex on livability showed a weaker pattern, with some samples experiencing a continuous decline and others experiencing a rebound. After 2017, a positive impact reappeared. This is closely related to the pattern of urban expansion. Older residential complexes tend to be located in the core urban area of ​​City A, offering significant advantages in terms of convenient living and transportation. With urban expansion, residential complexes are increasingly located further away from the city center. In recent years, with the advent of sustainable development and residents' pursuit of a higher quality of life, newly built communities have placed greater emphasis on livability. Therefore, the age of a residential complex exhibits a certain positive impact on livability in later periods.

[0062] (2) Park green space and property fees have a negative impact on most communities. They only have a positive effect when the index values ​​are extremely high or extremely low, such as Figure 6 (b) and (f) in the data. When there are no parks around a residential area, their impact on livability shows no clear pattern. As the number of parks and green spaces increases, they crowd out the building area of ​​surrounding facilities, negatively impacting livability in most communities. However, when the value is extremely high, the importance of the residential environment to residents exceeds the abundance of surrounding facilities, and parks and green spaces have a positive impact on the community. The correlation coefficient between property management fees and building age is as high as 0.51, indicating that communities with extremely low property management fees are mostly older communities with better locations and higher livability. Communities with extremely high property management fees are mostly villa communities with beautiful environments and extremely high livability.

[0063] (3) The impact patterns of price, greening rate, cultural facilities, and road network density on community livability are relatively similar, such as Figure 6 (c), (d), (e), and (g) all have approximately linear increasing relationships. When the indicator falls below a certain threshold, it has a negative impact on livability. As the indicator value increases, the negative impact gradually weakens, eventually becoming positive. The impact of price, cultural facilities, and road density on community livability is somewhat diffuse, while the impact of greening ratio is more concentrated.

[0064] (4) The number of educational facilities around the community and the livability of the community are approximately in a decreasing relationship, such as Figure 6(h) in the figure. When the total number of educational facilities around a residential area is less than 12, educational facilities have a positive impact on the residential livability. When the total number of educational facilities exceeds 12, educational facilities have a negative impact on the residential livability. This phenomenon indicates that when educational facilities meet the basic schooling needs of residents' children, they tend to be saturated, and further additions of educational facilities have no effect on improving livability.

[0065] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. The urban residential area livability evaluation method based on sentiment analysis and GA-BP is characterized by: The following steps are involved: Step S1: Based on the word frequency analysis, a residential area livability evaluation index system is constructed; Step S2: sentiment analysis is performed based on the evaluation data, k-means clustering is performed based on the objective conditions and attributes of the cells, and biased cell samples are eliminated using the quartile method to obtain training samples; Step S3: training the GA-BP model to evaluate the livability of the community; Step S4: Use the SHAP method to interpret the model and analyze the factors affecting the livability of the community that residents are concerned about.

2. The urban residential area livability evaluation method based on sentiment analysis and GA-BP according to claim 1 is characterized in that: In step S1, traditional residential area livability evaluation indicators and real estate network evaluation data are integrated to extract residents' hot topics. Guided by hot words, an urban residential area livability evaluation index system is constructed from the perspective of residents. The specific process is as follows: Step S11: Acquire residential community research data, specifically including residential community data, infrastructure point of interest data, road network data, population data, and mobile phone signaling data; Step S12: Use word frequency analysis to assist in index system construction, and quantitatively analyze user evaluation information with the help of Jieba word segmentation algorithm to obtain text hot spots and their changing trends; Step S13: Under the guidance of hot words, a community livability evaluation index system is constructed from the perspective of residents from six aspects: living cost, residential safety, environmental health, living comfort, living convenience, and travel convenience.

3. The urban residential area livability evaluation method based on sentiment analysis and GA-BP according to claim 1 is characterized in that: In step S2, based on the sentiment dictionary method, the ROST EA content mining system is used to perform sentiment analysis on the community evaluation data, find the characteristic vocabulary of sentiment words, degree words, and negation words from the text, and look up the sentiment value of each characteristic word in the sentiment dictionary. Sentiment classification is performed based on the accumulated sentiment values ​​to determine the residents' emotional attitudes towards the residential community.

4. The urban residential area livability evaluation method based on sentiment analysis and GA-BP according to claim 3 is characterized in that: In step S2, the K-means algorithm is used to divide residential communities into multiple clusters according to feature similarity. The elbow method is used to determine the initial number of clusters k. The silhouette coefficient S is selected as the cluster evaluation index. Within the same cluster, the quartile method is used to remove residential community samples with biased sentiment attitudes that are less than the lower quartile and greater than the upper quartile. The remaining residential communities constitute the training sample set.

5. The urban residential area livability evaluation method based on sentiment analysis and GA-BP according to claim 4 is characterized in that: The calculation formula of the silhouette coefficient is as follows: (1); (2); in, For samples Silhouette coefficient; For samples The average distance to all other samples in the same cluster; For samples The average distance to all samples in the nearest cluster; is the overall silhouette coefficient of the cluster, and its value range is [-1, 1]; is the total number of samples in the dataset.

6. The urban residential area livability evaluation method based on sentiment analysis and GA-BP according to claim 1 is characterized in that: In step S3, a BP neural network model integrated with an improved genetic algorithm, namely a GA-BP model, is constructed and trained to evaluate the livability of the community. The specific operation process of the GA-BP model includes the following steps: Step S31, determining the BP neural network topology; The BP neural network topology is a three-layer structure with 5 neurons in the hidden layer. The ReLu function is used as the activation function and the Adam algorithm is used for parameter optimization. Step S32: Encode the weight and bias parameters to form a chromosome, and use the genetic algorithm optimized by the elite selection strategy to determine the optimal individual; Step S33: Initialize using the optimal weights and bias parameters, and further optimize the model through training.

7. The urban residential area livability evaluation method based on sentiment analysis and GA-BP according to claim 1 is characterized in that: In step S4, the SHAP method is used to interpret the GA-BP model and analyze the factors affecting the livability of the community that residents are concerned about; Among them, SHAP is used to analyze the contribution of each evaluation index to the livability of urban residential areas, as shown below: (3); in, Indicates the SHAP values ​​of input features; Indicates that except The set of all features except features; F represents the set of all features; Representation model; Indicates the inclusion of features and collection Input of features in ; Indicates that only the collection is included The SHAP method calculates the features of the model through the data distribution and the explained variables. , original model Explanatory model , as shown below: (4); in, is the benchmark value, which represents the average predicted value of all samples; Indicates the Whether a feature participates in model prediction; Indicates the number of features in the model.