Street aesthetics evaluation and optimization method based on deep learning and multi-source data
By integrating multi-source data and deep learning technology, a street aesthetic evaluation model is built, and the problem of lack of data support in traditional methods is solved, precise evaluation and optimization of street aesthetics is achieved, and the scientificity and aesthetic quality of urban planning and design are improved.
Patent Information
- Application Number
- CN202510305529.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional street aesthetic assessment methods lack precise data support and scientific quantitative means, making them difficult to meet the needs of modern urban planning and design, and rarely consider multi-source data for comprehensive analysis.
By integrating multi-source data and deep learning technology, using street scene data and deep convolutional neural network model, combining user evaluation, a street aesthetic evaluation model is constructed, objective street indicators are calculated, and key factors and appropriate value ranges are derived, and optimization strategies are proposed.
It has achieved an accurate assessment of the aesthetics of the street, provided scientific basis, improved the aesthetic quality of the street and the urban image, and promoted the sustainable development of the city.
Smart Images

Figure CN120258207A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of urban planning, and particularly relates to a method for evaluating and optimizing street aesthetics based on deep learning and multi-source data. Background Art
[0002] In urban construction, street aesthetics is of great significance for improving the quality of the city, promoting the quality of life of residents, and the sustainable development of the city. However, traditional methods for evaluating street aesthetics are mostly qualitative analyses, lacking precise data support and scientific quantitative evaluation means, and are difficult to meet the needs of modern urban planning and design. At the same time, existing research has insufficient exploration of the relationship between street aesthetics and the physical environment, and rarely comprehensively considers multi-source data for comprehensive analysis. Summary of the Invention
[0003] Aiming at the deficiencies existing in the prior art, the purpose of the present invention is to provide a method for evaluating and optimizing street aesthetics based on deep learning and multi-source data. By integrating multi-source data and deep learning technology, it overcomes the limitations of traditional methods, accurately evaluates street aesthetics, provides a scientific basis for urban planning and design, effectively improves the aesthetic quality of streets and the urban image, and promotes the sustainable development of the city. To achieve the above object and other advantages according to the present invention, there is provided a method for evaluating and optimizing street aesthetics based on deep learning and multi-source data, including:
[0004] S1. Obtain research data and perform preprocessing;
[0005] S2. Calculate objective street indicators through an aesthetic evaluation module and comprehensively evaluate aesthetics. The objective street indicators include four types of elements: the physical space, environmental facilities, ground floor interface, and macroscopic form of the street;
[0006] S3. Deduce the relationship between objective street indicators and aesthetics through an analysis and optimization module, determine key factors and suitable value ranges, and propose optimization strategies.
[0007] Preferably, the comprehensive evaluation of aesthetics is a subjective street indicator, combined with street view data and a machine learning algorithm based on a deep convolutional neural network model. At the same time, the evaluation score of people for the visual environment is converted into an artificial intelligence scoring model. Specifically as follows:
[0008] Use street view data and a machine learning algorithm based on a deep convolutional neural network (DCNN: Dynamic Convolution Neural Network) model to conduct large-scale quality evaluations on street view photos.
[0009] Experimental users vote on two street view pictures randomly provided by the MIT crowdsourcing platform globally to select which picture looks more beautiful, more boring, more friendly, and more frustrating, thus forming a learning data set.
[0010] Compare each image sample i with other images i'. The calculation formula for the positive rate of image i along a certain perception index is as follows:
[0011]
[0012] The calculation formula for the negative rate is as follows:
[0013]
[0014] Where P i and N i represent the number of times image i is selected or not selected in the comparison, and e i is the number of times image i is considered equal to another image. The score of picture i is:
[0015]
[0016] Convert the evaluation score of a person's visual environment into an artificial intelligence scoring model. The evaluation model uses five-fold cross-validation to select the model, repeatedly uses randomly generated subsamples for training and validation, and the prediction accuracy of the finally generated evaluation model is 0.79.
[0017] Use the trained DCNN deep learning model to automatically score the site street view image data, and statistically calculate the average score of each street view picture by section.
[0018] Preferably, obtaining research data and performing preprocessing is to screen the grade attributes of road network data from OpenStreetMap. Specifically:
[0019] According to the OSM grade attribute screening, remove the following samples: viaducts, bridge roads, sections within residential communities; sections within 0.5 m from intersections; sections with no buildings within a buffer distance of 40 m on both sides of the center line; sections with a street segment length less than 50 m; sections with fewer than 4 street view sampling points and sections where reasonable street view-related data cannot be obtained;
[0020] Through manual comparison, it is found that if the distance between street view image sampling points is too far or too close to 25 m, the landscape covered by the intercepted images will have too low or too high homogeneity, and it is easy to miss some street view features. Therefore, it is necessary to determine an appropriate distance.
[0021] Compared with the prior art, the beneficial effects of the present invention are: by integrating various data resources and advanced technologies, realizing the accurate evaluation of street aesthetics, and providing optimization strategies to improve the aesthetic quality of streets. Description of the Drawings
[0022] Figure 1Schematic diagram of the street width acquisition method for the street aesthetic evaluation and optimization method based on deep learning and multi-source data according to the present invention;
[0023] Figure 2 Schematic diagram of the target recognition technology analysis framework for the street aesthetic evaluation and optimization method based on deep learning and multi-source data according to the present invention;
[0024] Figure 3 Schematic diagram of identifying store signs based on the Maskrcnn architecture for the street aesthetic evaluation and optimization method based on deep learning and multi-source data according to the present invention;
[0025] Figure 4 Schematic diagram of the analysis framework of the semantic segmentation model based on ResNet for the street aesthetic evaluation and optimization method based on deep learning and multi-source data according to the present invention;
[0026] Figure 5 Schematic diagram of segmenting greenery through the street view semantic model of ResNet for the street aesthetic evaluation and optimization method based on deep learning and multi-source data according to the present invention;
[0027] Figure 6 Schematic diagram of sampling building texture and open space of a 500m×500m sampling unit for the street aesthetic evaluation and optimization method based on deep learning and multi-source data according to the present invention;
[0028] Figure 7 Schematic diagram of the method for predicting human perception of street images for the street aesthetic evaluation and optimization method based on deep learning and multi-source data according to the present invention;
[0029] Figure 8 Flow chart of the street aesthetic evaluation and optimization method based on deep learning and multi-source data according to the present invention. Detailed implementation manners
[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0031] Refer to Figure 8 , a street aesthetic evaluation and optimization method based on deep learning and multi-source data, including: S1, obtaining research data and performing preprocessing;
[0032] S2. Calculate the objective indicators of the street through the aesthetic evaluation module and comprehensively evaluate the aesthetics. The objective indicators of the street include four types of elements: the physical space, environmental facilities, ground floor interface, and macroscopic form of the street.
[0033] S3. Deduce the relationship between the objective indicators of the street and aesthetics through the analysis and optimization module, determine the key factors and the appropriate value range, and propose optimization strategies.
[0034] Furthermore, the physical space of the street measures the physical space elements, including street width, street aspect ratio, and discount rate.
[0035] The environmental facilities measure the environmental facility elements, including the number of storefront signs and green view rate.
[0036] The ground floor interface measures the ground floor interface elements, including interface permeability and function density.
[0037] The four types of elements of the macroscopic form measure the macroscopic form elements, including building density, floor area ratio, and open space ratio.
[0038] Furthermore, the comprehensive evaluation of aesthetics is a subjective indicator of the street. Combining street view data and a machine learning algorithm based on a deep convolutional neural network model, the evaluation score of people for the visual environment is converted into an artificial intelligence scoring model. Specifically:
[0039] Use street view data and a machine learning algorithm based on a deep convolutional neural network model to conduct a large-scale quality evaluation of street view photos.
[0040] The experimental users vote on two street view pictures randomly provided by the MIT crowdsourcing platform globally, and choose which picture looks more beautiful, more boring, more friendly, and more frustrating, thus forming a learning dataset.
[0041] Compare each image sample with other images and calculate the positive rate of the image along a certain perception index.
[0042] Convert the evaluation score of people for the visual environment into an artificial intelligence scoring model. The evaluation model uses five-fold cross-validation for model selection, repeatedly uses randomly generated subsamples for training and validation, and the prediction accuracy of the finally generated evaluation model is 0.79.
[0043] Use the trained DCNN deep learning model to automatically score the site street view image data, and statistically calculate the average score of each street view picture by section.
[0044] Example 1
[0045] I. Data collection
[0046] 1. Section screening
[0047] (1) Obtain road network data from OpenStreetMap (OSM), and define streets as roads that carry people's daily social life. Screen according to the OSM level attributes, and remove elevated roads, bridge roads, and road sections within residential communities, as these road sections usually do not conform to the research definition of streets or are not representative.
[0048] (2) Exclude the section within 0.5 m from the intersection, because the traffic in the intersection area is complex, and the environmental characteristics are quite different from those of general street sections, which may affect the accuracy and consistency of the data.
[0049] (3) Delete the road sections without buildings within the buffer distance of 40 m on both sides of the center line to ensure that the research road sections have a certain built environment to meet the actual scene requirements of street aesthetic research.
[0050] (4) Screen out the road sections with a length less than 50 m. Too short road sections are difficult to comprehensively reflect the characteristics and functions of the streets. At the same time, remove the road sections with less than 4 street view sampling points and those that cannot obtain reasonable street view-related data, and finally determine the effective road sections.
[0051] 2. Street view image acquisition
[0052] (1) Through manual comparative analysis, determine that the spacing of street view image sampling points is 25 m. If the spacing is farther than 25 m, the homogeneity of the intercepted images covering the landscape is too low, and important street view features may be missed; if it is closer than 25 m, the homogeneity is too high and the diversity of the streets cannot be effectively reflected.
[0053] (2) For each positioning point, cut the panoramic photo into 6 pieces, and then select 4 pictures from them. Such an operation can obtain street view information from different angles and improve the richness and comprehensiveness of the data.
[0054] II. Data processing
[0055] 1. Street view image preprocessing
[0056] (1) Clean the collected street view images, remove the noise, blurred parts, and interference elements irrelevant to the perception of street beauty in the images, such as flying birds in the sky, irrelevant buildings in the distance, etc., to improve the quality of the image data and make the subsequent analysis focus more on the characteristics of the streets themselves.
[0057] (2) Standardize the images, unify parameters such as brightness, contrast, and color saturation of the images, so that the images collected under different times and weather conditions are comparable, facilitating subsequent analysis and processing. For example, adjust the brightness of all images to the same standard range to ensure that the analysis of street elements (such as building colors, greening, etc.) is not affected by differences in lighting conditions.
[0058] 2. 3D Building Map Data Processing
[0059] (1) Geometric correction is performed on the 3D building map data to correct problems such as building position deviation and shape deformation that may occur during data collection, ensuring the accuracy of building position and shape. For example, for the building position offset caused by measurement errors, it is corrected to match the actual geographical location.
[0060] (2) Spatial registration is carried out to precisely match the 3D building map data with street view image data, road network data, etc., so that the data from different data sources is consistent in space, providing a basis for comprehensive analysis. By establishing a unified spatial coordinate system, the 3D model of the building is accurately corresponded to the corresponding street view image and road position, in order to accurately analyze the relationship between the building and the street environment.
[0061] III. Obtaining Objective Street Indicators
[0062] 1. Calculation of Physical Space Element Indicators
[0063] Specifically as follows in the table:
[0064]
[0065] Note: "-" indicates no unit.
[0066] (1) Street Width Calculation
[0067] Using the centerline and building base vector information collected by GIS, a series of buffers with a 1m interval are made outward from the centerline (the maximum width is 40m). For example, for the centerline of a street, buffers are generated on both sides at intervals of 1m until the maximum width of 40m is reached.
[0068] Calculate the area ratio of the unilateral buffer area and the area after removing the building base. By calculating the total buffer area and the area occupied by the building, the proportional relationship between the two is obtained.
[0069] Query the mutation point with the largest difference in the ratio of adjacent two buffers, and define the outer wall line of the smaller buffer in this case as the side interface line of this side. This step can accurately determine the boundary of the street. For example, at a certain mutation point, the area ratio of the buffer changes significantly, and the outer wall line of the smaller buffer corresponding to this point is the side interface line of one side of the street.
[0070] Calculate the distance from the side interface lines on both sides of the road section to the centerline, and add the distances on both sides to get the street width. Through precise measurement and calculation, the actual width of the street is obtained, providing accurate data for subsequent analysis of the relationship between street width and aesthetics.
[0071] (2) Street Aspect Ratio Calculation
[0072] First, obtain the geospatial street vector data of Shanghai from the OSM website. These data contain the basic geometric information of the streets.
[0073] Based on the deep multi-task learning of the densely connected convolutional network, identify the aspect ratio types of each street view image on each road section. Use the deep learning model to analyze the street view images, and divide the aspect ratio into four categories: 0 < H / W < 1, 1 < H / W < 2, 2 < H / W < 4, and H / W > 4. For example, for a street view image, the model determines the aspect ratio type it belongs to by identifying the building height and street width.
[0074] Count the number of street view images of the four aspect ratio types on the street by road section, and take the category with the largest number as the aspect ratio type of this road section. Through the statistical analysis of all street view images within the road section, determine the most representative aspect ratio type of this road section to study the relationship between the aspect ratio and the street aesthetics.
[0075] (3) Calculation of the building line ratio
[0076] According to the calculation method of Harvey et al., since the street has at least one lane, and the unilateral width of the crosswalk is at least 2 - 3m, the results with a unilateral buffer distance less than 3m are excluded. This step is to ensure that the building part of the street considered when calculating the building line ratio is of practical significance.
[0077] Then take the second highest d value or take the nearby value. When all area ratios and the differences of the ratios are 0, it means there is no building on this side, and the width is set to 40m. After determining the effective buffer distance, further determine the key parameter d value used for calculating the building line ratio.
[0078] After determining the interface line, screen the buildings that intersect with the street boundary line, and obtain the ratio of the intersection length with the interface line to the total length of the street segment, which is the building line ratio. By accurately measuring the intersection situation between the building and the street boundary, calculate the building line ratio to evaluate the impact of the fitting degree between the building and the street on aesthetics.
[0079] 2. Calculation of environmental facility element indicators
[0080] (1) Counting the number of storefront signs
[0081] Adopt the object recognition technology. The analysis steps of this technology include four stages: large-scale image acquisition, selection and calibration of representative images, training and verification of the evaluation model, and large-scale index calculation.
[0082] First, manually mark the storefront sign features in 5000 site street view images through the open-source data annotation platform MakeSense to obtain the training set data. During the marking process, the staff carefully mark the position, shape and other features of the storefront signs in each image to provide accurate basic data for model training.
[0083] Secondly, based on the training set, taking the Residual Neural Network (ResNet) as the basis, the architecture of the instance segmentation algorithm Mask R-CNN is adopted for training and recognition to establish a deep learning image recognition model. By leveraging the powerful learning ability of the deep learning algorithm, the model can accurately identify storefront signs. After testing, this algorithm can effectively identify storefront signs in images, with an average Intersection over Union index of 0.87, demonstrating a high accuracy rate.
[0084] After the model training is completed, the image data of all points are input, the average number of storefront signs at all sampling points in the street section is calculated, and this value is used as the average number of storefront signs for the section to avoid differences caused by different sampling lengths and total number of samples. Through the processing of a large amount of image data, the average number of storefront signs for each section is accurately counted, providing reliable data for analyzing the relationship between the number of storefront signs and the street aesthetic.
[0085] (2) Green view rate measurement
[0086] This study uses the ADE20K scene parsing data and conducts partial segmentation data training. Taking the ResNet architecture as the encoder and PPM + deep supervision trick as the decoder, it obtains the features of the environment, surrounding elements, and the graphic itself, realizes the intelligent segmentation of objective scene elements, predicts a class label for each pixel, and the model accuracy is 86%.
[0087] With the above technical support, the elements of the Shanghai street space are interpreted. The image data of all points are input, the plants in the street view images are identified and located, the pixel ratios of the four directions corresponding to each street point are statistically calculated, and the average value is calculated. Through the accurate identification and statistics of the plant pixels in the street view images, the green view rate is obtained to evaluate the impact of the greening degree on the street aesthetic.
[0088] 3. Calculation of the underlying interface element indicators
[0089] (1) Interface permeability calculation
[0090] The features of door openings and windows in 5000 site street view images are manually marked through the Make Sense platform. Windows, residential windows, overhanging corridors, fences, and block entrances / gates are classified as transparent classes to obtain training set data. During the marking process, the transparent and non-transparent interface features are accurately distinguished to provide accurate annotation data for model training.
[0091] Taking the ResNet network architecture as the encoder and PPM + deep supervision trick as the decoder, an element interpretation is performed on larger-scale Shanghai street view images, a deep learning model is established, and the model accuracy is 91%. The deep learning model is used to analyze large-scale street view images to identify transparent interface elements.
[0092] After the model training is completed, the image data of all points are input, and the transparent interface elements in 4 directions at each point are identified, such as the openings of streets, glass interfaces, etc., and the ratio of the transparent interface to the building interface in each street view picture at this point is obtained. The calculation formula is:
[0093]
[0094] where M is the permeability of a certain street, F i is the number of pixels of the transparent interface at the bottom layer of the street corresponding to a certain sampling point, F′ is the number of pixels of the street bottom layer interface corresponding to this sampling point, and n is the total number of sampling points.
[0095] Specifically, it is shown in the following table:
[0096]
[0097] (2) Calculation of street function density or mixing degree
[0098] According to the POI data in the Gaode Map in 2019, the street functions are specifically divided into 9 categories such as residential, commercial, transportation, and catering and entertainment. These classifications cover the main functional types of the streets and can comprehensively reflect the vitality and diversity of the streets.
[0099] The calculation formula for the function mixing degree of the street is:
[0100]
[0101] where n represents the number of categories, P i is the proportion of the i-th type of POI in the POI. To avoid the influence of data errors and accidental factors, the POI categories with a total number of less than 5 are not included in the function mixing degree. The higher the information entropy, the greater the mixing degree and the higher the diversity. For those with only 1 effective category, the information entropy is recorded as 0.
[0102] 4. Calculation of macroscopic morphological element indicators
[0103] (1) The block scale is generally not greater than 500m. Therefore, a 500m × 500m unit sampling is performed on the research scope from the OSM website, and fine data of streets, plots, and buildings are obtained according to the street form and function. This sampling method can effectively obtain relevant data on the street surrounding environment at the macroscopic level while ensuring the accuracy and representativeness of the data.
[0104] (2) Statistically analyze the ground floor area of buildings within the blocks surrounding each road section, calculate its ratio to the sampled block, and obtain the building density value. By calculating the proportion of the ground floor area of buildings in the block area, accurately evaluate the building density to study its impact on the street aesthetic.
[0105] (3) Statistically analyze the ground floor area of buildings within the square areas surrounding each road section, multiply it by the number of building floors to obtain the total building area, calculate its ratio to the square area, and obtain the estimated floor area ratio. By comprehensively considering the ground floor area of buildings, the number of floors, and the area of the region, calculate the floor area ratio to analyze its relationship with the street aesthetic.
[0106] (4) Statistically analyze the open space area within the square areas of each road section, divide it by the square area, and obtain the open space ratio within the research scope. By calculating the open space ratio, evaluate its contribution to the street aesthetic.
[0107] IV. Street Aesthetic Evaluation
[0108] 1. Obtaining Subjective Indicators
[0109] (1) Drawing on the method of Hu Chunbo et al. for artificial intelligence training scoring based on the public's perception, using street view data and machine learning algorithms based on the deep convolutional neural network (DCNN) model, conduct large-scale quality evaluations on street view photos. The machine learning training samples are sourced from the experimental data of the "Urban Pulse" project.
[0110] (2) The experimental users vote on two street view pictures randomly provided by the crowdsourcing platform of the Massachusetts Institute of Technology, selecting which picture looks more beautiful, more boring, more friendly, and more frustrating, thus forming a learning data set. During the voting process, the users evaluate the street view pictures based on their subjective feelings, providing diverse subjective data for subsequent model training.
[0111] (3) Compare each image sample i with other images i'. The calculation formula for the positive rate of image i along a certain perception index is as follows:
[0112]
[0113] The calculation formula for the negative rate is as follows:
[0114]
[0115] where P i and N i represent the number of times image i is selected or not selected in the comparison, and e i is the number of times image i is considered equal to another image. The score of picture i is:
[0116]
[0117] (4) Convert the human evaluation score of the visual environment into an artificial intelligence scoring model. The evaluation model uses five-fold cross-validation to perform model selection, repeatedly uses randomly generated subsamples for training and validation, and the prediction accuracy of the finally generated evaluation model is 0.79. Through cross-validation and continuous optimization, the accuracy and reliability of the model are improved.
[0118] (5) Use the trained DCNN deep learning model to automatically score all the street view image data of the site, count the average score of each street view picture by section, and realize the quantitative evaluation of the street beauty. Through the automatic scoring of a large number of street view images, the beauty degree of the street can be comprehensively and objectively evaluated, providing a data basis for subsequent analysis.
[0119] Construction and Analysis of the Beauty Evaluation Model
[0120] (1) Regard the 10 indicators of 3 categories of the selected site (street width, aspect ratio, building line adherence rate, number of storefront signs, green view rate, permeability, function density, building density, plot ratio, open space ratio) as independent variables, and the natural logarithm of the street beauty evaluation as the dependent variable. Use multiple linear regression to analyze the influence of street indicators on the street beauty evaluation. The calculation formula is:
[0121]
[0122] Among them, β1~β 10 are regression coefficients, β0 is a constant term, and e is an error item.
[0123] (2) If the VIF value after regression is less than 5, the regression model passes the hypothesis test. In addition, after testing, if the correlation coefficients between independent variables are all less than 0.7, there is no collinearity.
[0124] V. Relationship Analysis
[0125] (1) According to the results of multiple linear regression analysis, it is found that the beauty evaluation is significantly correlated with two factors of the street physical environment (street width and street aspect ratio) with the street beauty evaluation at the 0.001 level, and the street width is positively correlated and the aspect ratio is negatively correlated. This indicates that the larger the street width, the stronger the beauty evaluation; the larger the aspect ratio, the lower the beauty score. For example, a wide street can bring an open and comfortable feeling, while a too high building-to-street width ratio may make people feel depressed.
[0126] (2) Among the environmental facility elements, the green view rate has a strong positive correlation with the street aesthetic evaluation, and the average number of storefront signs also has a significantly positive correlation with the street aesthetic evaluation at the 0.001 level. This means that increasing the greening and the number of storefront signs can enhance the street aesthetic. Urban greening can create a quiet and comfortable atmosphere, while rich storefront signs can increase the vitality and attractiveness of the street.
[0127] (3) The correlation between the interface permeability and the street aesthetic evaluation is weak and negative. This shows that the more permeable the street interface is not necessarily better. An appropriate permeable interface allows individuals to observe the internal elements of the building and experience the rich street characteristics, but too high a permeability may affect the aesthetic. For example, a completely transparent street interface may make the street lack a sense of hierarchy and privacy.
[0128] (4) At the macroscopic level, the building density has the most obvious negative correlation with the aesthetic perception of the street. Too high a building density will make individuals feel crowded, thus affecting the aesthetic perception of the street. This is mutually confirmed with the fact that the green view rate is the most important factor guiding people's pleasant perception. Appropriate combination of building density and greening can enhance the street aesthetic.
[0129] (5) The expected line - adhering rate and the proportion of open space are related to the street aesthetic perception, but the statistical analysis results show that the two only have a weak correlation with the street aesthetic perception, and these two factors can be appropriately ignored in the subsequent discussion.
[0130] The equipment quantity and processing scale described here are used to simplify the description of the present invention. Applications, modifications, and variations of the present invention are obvious to those skilled in the art.
[0131] Although the embodiments of the present invention have been disclosed as above, it is not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily achieved. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to specific details and the examples shown and described here.
Claims
1. A method for street aesthetic evaluation and optimization based on deep learning and multi-source data, characterized in that, Including: S1. Obtain research data and perform preprocessing; S2. Calculate the objective indicators of the street through the aesthetic evaluation module and comprehensively evaluate the aesthetics. The objective indicators of the street include four types of elements: physical space, environmental facilities, ground floor interface, and macroscopic form of the street; S3. Deduce the relationship between the objective indicators of the street and aesthetics through the analysis and optimization module, determine the key factors and the appropriate value range, and propose optimization strategies.
2. The street aesthetic evaluation and optimization method based on deep learning and multi-source data according to claim 1, wherein In step S1, the research data obtained is to screen the hierarchical attributes of the road network data from the OpenStreetMap. The specific steps are as follows: S11. According to the OSM hierarchical attribute screening, remove the following samples: viaducts, bridge roads, sections within residential communities; sections within 0.5 m from intersections; sections with no buildings within a buffer distance of 40 m on both sides of the center line; sections with a street segment length less than 50 m; Sections with fewer than 4 street view sampling points and sections where reasonable street view-related data cannot be obtained; S12. Set the spacing of street view image sampling points to 25 m.
3. The method for evaluating and optimizing the street aesthetic feeling based on deep learning and multi-source data according to claim 1, wherein In step S2, the physical space of the street is to measure the physical space elements, including street width, street height-width ratio, and building frontage ratio; The environmental facilities are to measure the environmental facility elements, including the number of storefront signs and green view rate; The ground floor interface is to measure the ground floor interface elements, including interface permeability and function density; The four types of macroscopic form elements are to measure the macroscopic form elements, including building density, floor area ratio, and open space ratio.
4. The method for evaluating and optimizing the street aesthetic feeling based on deep learning and multi-source data according to claim 3, wherein The calculation of the physical space of the street is specifically as follows: The street width is obtained by using the center line and building base vector information collected by GIS. A series of buffers with an interval of 1 m are made outward from the center line, and the ratio of the area of the unilateral buffer and the area after removing the building base is calculated. Query the mutation point with the largest difference in the ratio between two adjacent buffers. Define the outer wall line of the smaller buffer in this case as the interface line on this side, and calculate the distance from both sides of the section to the center line; The street height-width ratio is obtained by obtaining the site geospatial street vector data from the OSM website. Based on the deep multi-task learning of the densely connected convolutional network, identify the height-width ratio types of each street view picture on each section and classify them into 4 types: 0 < H / W < 1, 1 < H / W < 2, 2 < H / W < 4, H / W > 4; finally, count the number of street view pictures of the 4 height-width ratio types on the street by section, and take the type with the largest number as the height-width ratio type of this section; The building frontage ratio is obtained by first excluding the results with a unilateral buffer distance less than 3 m, and then taking the second highest d value or taking a nearby value; when all area ratios and the differences in ratios are 0, it means there are no buildings on this side, and the width is set to 40 m; after determining the interface line, screen the buildings intersecting the street boundary line, and obtain the ratio of the intersection length with the interface line to the total length of the street section.
5. The street aesthetic evaluation and optimization method based on deep learning and multi-source data according to claim 3, characterized in that The specific calculation in the environmental facilities is as follows: Number of storefront signs: The target recognition technology is used to identify storefront signs. Specifically, the characteristics of storefront signs in 5000 site street view images are manually marked through the open source platform MakeSense to obtain training set data. Then, based on this, a deep learning model is trained and established using the Mask R-CNN architecture based on ResNet. The average intersection over union index of this model test reaches 0.87 and the accuracy is high. Finally, all point image data is input, and the average value of the number of storefront signs at the sampling points of the street section is calculated as the average number of storefront signs per section, avoiding the influence of sampling differences; Green view rate: Using semantic scene parsing technology, trained with ADE20K data, encoded with ResNet architecture and decoded with PPM+deepsupervisiontrick to achieve intelligent segmentation of scene elements; Based on this, the spatial elements of Shanghai streets are interpreted, plants are identified by inputting point image data, and the average pixel ratio of the four directions of each point is statistically calculated to provide data support for the research.
6. The street aesthetic evaluation and optimization method based on deep learning and multi-source data according to claim 3, characterized in that, The specific calculation of the bottom layer interface is as follows: Interface permeability: The training set is obtained by marking the characteristics of 5000 site street view images through the Make Sense platform. A deep learning model is constructed using the ResNet network architecture combined with PPM+deep supervision trick, and its accuracy reaches 91%. It is used to interpret the elements of large-scale street view images. After inputting all point image data, the proportion of the permeable interface at each point is identified and calculated; Street function density and mixing degree: According to the POI data in Gaode Map in 2019, the street functions are specifically divided into 9 categories such as residential, commercial, transportation, and catering and entertainment, and then the function mixing degree of the street is calculated.
7. The method for street aesthetic evaluation and optimization based on deep learning and multi-source data according to claim 3, wherein The specific calculation of the four types of elements of the macroscopic form is as follows: (1) Conduct unit sampling of 500m×500m for the research scope from the OSM website, and obtain detailed data of streets, plots, and buildings according to street form and function. (2) Statistically calculate the ground floor area of buildings within the surrounding blocks of each section, calculate its ratio to the sampled block, and obtain the building density value; (3) Statistically calculate the ground floor area of buildings within the surrounding square area of each section, multiply by the number of building floors to get the total building area, and calculate its ratio to the area of the square area to obtain the estimated floor area ratio; (4) Statistically calculate the open space area within the square area of each section, divide by the square area, and obtain the open space ratio within the research scope.
8. The method for evaluating and optimizing the street aesthetic feeling based on deep learning and multi-source data according to claim 3, characterized in that, The comprehensive evaluation of aesthetics is a subjective index of the street. Combining street view data and a machine learning algorithm based on a deep convolutional neural network model, the evaluation score of people for the visual environment is converted into an artificial intelligence scoring model. Specifically: Using street view data and a machine learning algorithm based on a deep convolutional neural network model, conduct a large-scale quality evaluation of street view photos; The experimental users vote on two street view pictures randomly provided by the MIT crowdsourcing platform globally, and choose which picture looks more beautiful, more boring, more friendly, and more frustrating, thus forming a learning data set; Compare each image sample with other images, and calculate the positive rate of the image along a certain perception index; Convert the human's evaluation score of the visual environment into an artificial intelligence scoring model. The evaluation model uses five-fold cross-validation for model selection, repeatedly uses randomly generated subsamples for training and validation, and the prediction accuracy of the finally generated evaluation model is 0.79; Use the trained DCNN deep learning model to automatically score the site street view image data, and count the average scores of each street view picture by section.
9. A street aesthetic evaluation and optimization device based on deep learning and multi-source data, characterized in that, Including: A data acquisition module for obtaining multi-source data of OSM road networks, street view images, and POIs; A data processing module for preprocessing the data; A beauty evaluation module for using a deep learning model to calculate element indicators such as street physical space, environmental facilities, ground floor interface, and macroscopic form, and comprehensively evaluate beauty; An analysis and optimization module for deriving the relationship between indicators and beauty, determining key factors and suitable value ranges, and proposing optimization strategies.
10. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 8.
Citation Information
Cited By
Street environment quality evaluation method and system based on multi-source spatial data
CN121936999A