A city park perception value evaluation method based on multi-modal data fusion
By using multimodal data fusion and multi-task learning, the problems of long data collection cycles and limited spatial coverage in traditional park evaluation methods have been solved. This has enabled an integrated subjective and objective park evaluation, providing standardized and quantitative evaluation results and improving the efficiency and accuracy of the evaluation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-22
- Publication Date
- 2026-04-10
AI Technical Summary
Traditional park evaluation methods rely on questionnaires and on-site surveys, which have long data collection cycles, high costs, and limited spatial coverage. Remote sensing monitoring is difficult to identify the status of internal facilities and experience details. Existing research lacks a standardized evaluation framework that integrates subjective and objective perspectives, making it difficult to achieve unified quantification at the national level.
By fusing multimodal data, including social media comments, remote sensing data, POI data, and road network information, and using pre-trained language models for multi-task learning, combined with semantic segmentation and object detection models, subjective perception and objective environmental indices are constructed to achieve quantitative assessment.
It has achieved automated, integrated subjective and objective park assessments at the national level, providing standardized and quantitative results, ensuring comparability across cities and time periods, and improving the coverage and representativeness of the assessments.
Smart Images

Figure CN121545160B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the cross field of remote sensing and geographic information science, computer vision and natural language processing, and particularly relates to a city park perception value evaluation method based on multi-modal data fusion. BACKGROUND
[0002] City parks are important places for residents to relax, socialize and engage in healthy activities, and their environmental quality directly affects the livability of the city. Traditional park evaluation relies on questionnaires and field surveys, which have long data collection periods, high costs and limited spatial coverage. Remote sensing monitoring has macro advantages, but it is difficult to identify the status of internal facilities and experience details in parks. With the popularity of user-generated content, social media comments and photos have become an important source of public perception. Existing research is mostly single-modal and weakly integrated, lacking a standardized evaluation framework that integrates subjective and objective factors, making it difficult to achieve unified quantification at the national scale. SUMMARY
[0003] The purpose of the present application is to achieve automated and integrated park evaluation at the national scale, and to form a standardized and quantitative result by unifying data standards and fusing multiple sources.
[0004] The above technical purpose of the present application is achieved by the following technical solutions:
[0005] A city park perception value evaluation method based on multi-modal data fusion, the method comprising:
[0006] Obtaining comment data of park social media, park boundary data, POI data, remote sensing land class proportion and road network information;
[0007] Using a pre-trained language model to perform multi-task learning reasoning on the text comment data, outputting multi-dimensional perception labels and sentiment intensity scores, and aggregating to generate subjective perception results;
[0008] Using a semantic segmentation model, an object detection model and a visual language model to identify scenes, characters / facility elements from image comment data, respectively, to obtain different scene element coverage, character / facility counts, and visual element language representations;
[0009] Constructing a plurality of objective environmental indices containing variable direction symbols to describe different environmental states, extracting variable information from the scene element coverage, character / facility counts, visual element language representations, POI data, remote sensing land class proportion and road network information, and constructing a mapping between the extracted variable information and the corresponding objective environmental indices to quantify each objective environmental index; the variable direction symbol is used to describe the positive or negative impact of the variable on the corresponding objective environment, and the objective environmental indices are constructed by weighted summation according to the positive or negative contribution direction of the variable;
[0010] The subjective perception result is aligned with the objective environment index in the park and period dimensions to output.
[0011] In some embodiments of the present application, the POI data includes data within the park boundary and within a buffer zone outside the park; the buffer zone is an area formed by extending a preset distance outward along the park boundary.
[0012] The remote sensing land class proportion and road network information are data within the park boundary.
[0013] In some embodiments of the present application, the park boundary data is derived from AOI data of an online map; missing boundary data is completed by other open source map data.
[0014] In some embodiments of the present application, the language model includes a shared encoder and a multi-label classification head and a multi-target regression head connected in parallel on the shared encoder.
[0015] The multi-label classification head outputs a multi-dimensional binary classification result as a multi-dimensional perception label.
[0016] The multi-target regression head outputs a multi-dimensional sentiment intensity score value.
[0017] In the model inference stage, the inference result of the comment is mapped to the corresponding park-period unit; in each unit, aggregation calculation is performed: for the perception label, the frequency of its mention is counted respectively, and the arithmetic mean of the sentiment intensity score of the activated sample is calculated as the subjective perception quantitative result of the park in the period.
[0018] In some embodiments of the present application, the objective environment index is calculated based on the following formula:
[0019]
[0020] In the formula, j is the serial number of the objective environment index; is the direction symbol of the variable, +1 represents positive promotion, and -1 represents negative inhibition; is the value of the jth original variable constituting the index; and are the mean and standard deviation of the variable in the whole sample, respectively.
[0021] In some embodiments of the present application, the variables extracted from the road network information include at least one of the total length of roads and road density.
[0022] In some embodiments of the present application, the objective environment index includes at least one of vegetation quality, water quality, original ecological degree, recreational interaction, road facility, service facility, public transportation, and commercial activity.
[0023] In some embodiments of the present application, the acquired data is divided into multiple periods according to timestamps, and indexed based on park number-period as the primary key;
[0024] The subjective perception results and the objective environment indexes are mapped to the index based on the primary key, and the output in the form of a structured data table is acquired.
[0025] In some embodiments of the present application, the output results at least include park identification, park name, park location, subjective perception results and objective environment indexes.
[0026] In some embodiments of the present application, the output results further include park land class proportions based on remote sensing, and different scene element coverages, person / facility counts and visual element language representations acquired by processing image comment data.
[0027] Compared with existing methods relying on single modalities or empirical weight design, the present application realizes tight coupling fusion and standardized quantification of "text-image-remote sensing / POI / road network" on a national scale: integrated data caliber, explicit variable-index mapping and signed Z-score aggregation ensure comparability across cities and periods; multi-task text modeling and visual three-branch parallel design significantly improve the coverage of subjective perception and the representation of objective environment; unified quality control and missing data processing strategy, (park x period) structured output and reproducible training-inference parameter configuration provide an engineered and sustainable technical path for dynamic monitoring and policy evaluation. BRIEF DESCRIPTION OF DRAWINGS
[0028] Figure 1 The method flowchart in some embodiments of the present application.
[0029] Figure 2 The text multi-task learning model structure diagram in some embodiments of the present application.
[0030] Figure 3 The visual element extraction model structure diagram in some embodiments of the present application.
[0031] Figure 4 The objective environment index construction flowchart and variable mapping relationship diagram in some embodiments of the present application.
[0032] Figure 5 The evaluation result structured output diagram in some embodiments of the present application. DETAILED DESCRIPTION
[0033] The technical solutions of the present application will be further described below in combination with the drawings and embodiments.
[0034] Embodiment 1
[0035] The embodiment takes 6184 open parks in cities nationwide as the evaluation object to specifically illustrate the implementation process of the scheme.
[0036] The data sources for the park perceived value evaluation include review data (including text and images) of social media, remote sensing land class proportion data, point of interest (POI) data and road network information data, and in addition, the park boundary (AOI) data is used to delimit the data range. The obtained data is divided into three periods (TIME_PERIODS) according to the timestamp: 2006-2014, 2015-2019 and 2020-2025.
[0037] The embodiment selects the "Dazhongping" platform as the social media data source, and takes the point of interest (POI) ID under the "park" category as the index object to collect the user review details (including review ID, publishing time, review text, associated image file) and the basic attribute information (name, number / ID, latitude and longitude coordinates, administrative division) of the parks in the nationwide range in a targeted manner. After cleaning, the original database is established, and the data contains 1.56 million reviews and 1.23 million review images.
[0038] The AOI data comes from online map platforms (in this embodiment, the data comes from the Baidu Map Open Platform), and the missing part is supplemented by open source map data (in this embodiment, it is supplemented from OpenStreetMap (OSM)). The POI index is counted within the park boundary and the 500m buffer zone outside it, and the remote sensing land class proportion data and the road network information data are limited within the park boundary. The remote sensing land class proportion data comes from the Dynamic World 10-meter resolution real-time land cover dataset on the Google Earth Engine platform; the road network data is obtained from the OpenStreetMap (OSM) open source map database.
[0039] The obtained data is used to evaluate the perceived value of the park, and the process is as shown in Figure 1 , including the following steps:
[0040] Step 1, data cleaning of review text data;
[0041] The review text is sequentially subjected to emoticon escaping, HTML tag deletion, URL removal, illegal character filtering, Unicode normalization and space normalization processing to form normalized text, and the text data cleaning is completed.
[0042] Each record (normalized text) is mapped to a three-period label according to the timestamp (review time) and associated with the park number to form a "park-period" primary key index.
[0043] Step 2, subjective perception extraction of review text (normalized text);
[0044] From the aforementioned 1.56 million comments, 20,000 comments were randomly selected for manual annotation to build a training data set. To ensure data quality, three field experts independently annotated and cross-validated the constructed data set. After testing, the manual annotation consistency of each perception dimension reached a high standard: accessibility κ = 0.82, α = 0.84; order κ = 0.79, α = 0.81; comfort κ = 0.80, α = 0.82; cultural aesthetics κ = 0.78, α = 0.80; interactive openness κ = 0.76, α = 0.78, providing reliable supervision signals for model training.
[0045] As shown in Figure 2 The embodiment adopts Chinese RoBERTa-wwm-ext as a shared encoder (the maximum sequence length MAX_LEN = 512 of word segmentation, the token sequence obtained by word segmentation of the comment text is input into the shared encoder), and parallel multi-label classification head and multi-target regression head are connected on it. The classification head and the regression head are both full connection layers (Linear Layer), the multi-label classification head is connected with the Sigmoid function, and the multi-target regression head is connected with Clamp(0,5).
[0046] The normalized comment text (Normalized Text) obtained in step one is input into the shared encoder, and the representation output by the shared encoder is fed into the two types of task heads. The classification head produces six-dimensional binary labels with a Sigmoid threshold of 0.5; the regression head outputs six-dimensional intensity scores. The sixth dimension is "ambiguous / non-specific perception", which uses the masking mechanism (-100 mask) to only absorb model noise and not output results, so it actually contains five-dimensional perception classification. Missing labels are removed with a -100 mask. The joint loss uses uncertainty weighting:
[0047]
[0048] wherein, is the classification task loss, and the calculation method is to take the average of the six-dimensional classification loss first, and then take the average of the valid samples, is the regression task loss, and only the elements of are calculated to take the average of the mean square error and all valid elements; , respectively represent the uncertainty of the classification task and the regression task, which are learnable parameters; during training, weighted random sampling is used to make the expected sampling ratio of park samples to general Chinese sentiment analysis data set (using the open-source Weibo sentiment sample WeiboSenti_100K) about 5:1.
[0049] where "valid sample" for classification task refers to the whole review, as long as the review contains at least one valid class label (i.e. class_labels vector is not all sentinel value -100), the review is considered as a "valid sample". "Valid element" for regression task refers to the specific value in the matrix. Because a review might only rate on partial dimensions (e.g. only "comfort" is evaluated), other dimensions are marked as -100. Regression loss only calculates the error for those non -100 specific dimensions, ignoring the missing values.
[0050] Model inference process and output results are as follows:
[0051] Multi-label classification head (output 6-dimension Logits -> Sigmoid -> probabilities):
[0052] Output vector (probabilities): [0.02, 0.05, 0.98, 0.10, 0.95, 0.01]; corresponding dimensions: [accessibility, orderliness, comfort / greening experience, cultural aesthetics, interactive openness / service facilities / hygiene (or 5th dimension), ambiguous / other].
[0053] Six-dimension binary labels are generated with a threshold of 0.5 after Sigmoid:
[0054] 3rd dimension (comfort / greening experience): 0.98 > 0.5 -> activated (Label=1).
[0055] 5th dimension (service facilities / hygiene): 0.95 > 0.5 -> activated (Label=1).
[0056] Other dimensions: not activated (Label=0).
[0057] The model judges that the review discusses two topics of "comfort" and "service facilities".
[0058] Multi-objective regression head (output 6-dimension Score -> Clamp -> intensity):
[0059] Output vector (scores): [2.5, 3.0, 4.8, 2.9, 1.2, 3.0]; corresponding dimensions are the same as above.
[0060] 3rd dimension (comfort): score 4.8 (close to 5 points), corresponding to the positive evaluation of "good greening, comfortable" in the review.
[0061] 5th dimension (service facilities): score 1.2 (close to 1 point), corresponding to the negative evaluation of "dirty toilet, poor experience" in the review.
[0062] Other dimensions: Although the model outputs numerical values (such as 2.5), these scores will be ignored / masked in the subsequent aggregation stage due to the classification label being 0.
[0063] The inference stage maps the inference results of individual reviews to the corresponding "Park-Period" unit according to the timestamp and park ID of the review; then performs aggregation calculation within each unit: for the five-dimensional perception indicators, the frequency of their mentions is counted to reflect the attention, and the arithmetic mean of the sentiment intensity scores of the activated samples (classification label 1) is calculated as the subjective perception quantification result of the park in that period.
[0064] Step three: image (review image) visual element extraction (text expression)
[0065] As shown in Figure 3 , this embodiment uses SegFormer semantic segmentation model, YOLOv8 model, and Chinese-CLIP to extract visual elements from the review image, including:
[0066] (1) Input the image into the semantic segmentation model, aggregate different elements and input the coverage rate of each element;
[0067] This embodiment uses the SegFormer semantic segmentation model to realize image element classification, aggregates the image into building (building), sky (sky), vegetation (tree, grass, plant_flower), water body (water), road (road_path), bare soil (bare_soil), hardened ground (hardened_ground) and other elements on the ADE20K system, and outputs the element coverage rate. In this embodiment, it is derived as green coverage rate ("green_coverage_percent", calculated by combining tree, grass, and flower pixels), water area coverage rate ("water_area_percent"), hardened ground coverage rate ("hardened_ground_percent"), bare soil coverage rate ("bare_soil_percent"), and building coverage rate ("building_percent");
[0068] (2) Identify specific targets in the image and count them through the target detection model;
[0069] In this embodiment, the YOLOv8 model is used for target recognition, including recognizing and counting human or specific object targets such as persons, benches, chairs, tables, dogs, birds, signboards, and lights, and using the minimum target size filtering rule to reduce small target noise.
[0070] (3) Use a multi-modal learning model for visual-linguistic representation;
[0071] Use Chinese-CLIP to set multiple prompts for each label. In this embodiment, 3 Chinese prompt phrases are constructed for each label, covering lighting, neatness, path structure, activity scene, and vegetation landscape. The maximum probability is taken as the confidence, and the default threshold is 0.15. The threshold for wide paths (>1.5m) and open view is 0.6.
[0072] For example, a pre-defined CLIP hierarchical visual label library is used, including lush vegetation, blooming flowers, and turbid water with floating objects. The Chinese-CLIP model is used to match the image with the pre-defined visual semantic label to obtain certain physical scene attributes. For each label, the text encoder of Chinese-CLIP generates multiple prompt phrases (Prompt Ensembling, such as "a park photo containing dense vegetation"), and Chinese-CLIP calculates the cosine similarity between the image and these prompt vectors to realize zero-shot classification of image content.
[0073] (4) Aggregate the results extracted in (1)-(3) into coverage (proportion), number of facilities (count / density), and scene concept (proportion / indicator variable) in "Park x Period".
[0074] Step four: Multi-source geographic data processing
[0075] Variable interpretation is performed on remote sensing data, including bare ground area proportion ("bare_ratio percent"), building area proportion ("built_ratio percent"), grassland area proportion ("grass_ratio percent"), tree area proportion ("trees_ratio percent"), and water area proportion ("water_ratio percent"). Align the interpreted variables by "Park x Period";
[0076] Points of interest map "Life Services" to service facility density, "Transportation Facilities" to public transportation, and "Shopping / Dining / Hotels" to commercial activities (unit: points / km). 2 );
[0077] The road network is calculated by dividing the total road length (km) by the park area (km²). 2 Calculate road density (km / km) 2 ).
[0078] Step 5: Construction of Eight Objective Environmental Indices
[0079] like Figure 4 As shown, the objective environment index is quantified based on the information extracted from images, remote sensing land cover data, POI data, and road network data in steps three and four.
[0080] This embodiment constructs eight objective environmental indices, including vegetation quality, water quality, degree of pristine environment, recreational interaction, road facilities, service facilities, public transportation, and commercial activities. Let the first... Item Index ( =1,2,...,8) contains k j The normalized variable. The direction sign of each variable is (+1 represents positive promotion, -1 represents negative inhibition), then the j-th objective environment index I j The calculation formula is:
[0081]
[0082] In the formula For the first part of the index Original variable values (such as green coverage rate, road density, etc.); and These are the mean and standard deviation of the variable in the entire sample (i.e., Z-score standardization). For example, when calculating the "vegetation quality index", the variable set is selected as {green coverage rate, grassland area ratio, tree area ratio, frequency of dense vegetation label, frequency of blooming flower label}, and the direction sign of all variables is set to +1.
[0083] The eight indexes include: vegetation quality (such as green coverage, lush vegetation, blooming flowers, grass area ratio, tree area ratio, all positive), water quality (water area coverage, water cleanliness, water turbidity / floatable label frequency (negative), water area ratio, positive), original ecological degree (bare soil coverage, bare soil area ratio is positive; hardening ground coverage, building coverage, building area ratio is negative), recreational interaction (people, children playing, picnic / meeting, square dance, fitness exercise, chess / card playing, dog, seesaw, bird, clean and tidy, positive), road facilities (road, lamp, fence, road density, wide path is positive; narrow path can be included as negative according to needs), service facilities (chair, table, public toilet, gallery, service facility density, positive), public transportation (traffic facility density, positive), commercial activities (commercial POI density, positive).
[0084] The variable-index mapping and variable positive-negative marking of the eight objective environment indexes are fixed and standardized by Z-score in the full sample range. The aggregation is performed according to the above signed mean formula.
[0085] Step six: result alignment and output
[0086] As shown in Figure 5 , the subjective and objective results (subjective perception and objective environment index) are connected by taking (park number ParkID, period) as the primary key. The output (core field) includes park ParkID, name, administrative division (province / city / district), longitude and latitude, text scores (six dimensions) of each period, visual elements and ratio indicators, remote sensing land area proportion (2010 / 2015 / 2020) and eight indexes (by period) of structured data table, so as to subsequent batch calculation and full-process traceability. The output data supports CSV and database two kinds of landing forms.
[0087] Among them, the remote sensing land area proportion refers to the area proportion of five types of land (trees, grass, water, bare land, buildings) obtained based on remote sensing image interpretation. The park social media data spans 2006-2025, but high-precision land use data is limited. Considering the relatively stable surface coverage characteristics of urban parks, in order to match the three urban development stages, when outputting data, the remote sensing data selects the remote sensing ecological interpretation results of three typical years for period mapping: 2010 remote sensing data represents the 2006-2014 period; 2015 remote sensing data represents the 2015-2019 period; 2020 remote sensing data represents the 2020-2025 period.
[0088] Step seven: data quality control and missing data processing
[0089] Columns with standard deviation of 0 are set to 0 before standardization; road density is imputed with the median of the full sample; service facility density, public transportation, commercial activity, and missing visual / remote sensing variables are set to 0; missing park area is imputed with 10 -6 Instead of dividing by zero to avoid; no Winsorize or truncation processing.
[0090] This example evaluates the performance of the core text multi-task learning model (RoBERTa-wwm-ext). The training configuration is as follows: learning rate 2×10 -5 , AdamW optimizer, batch size 32, 10 rounds of training, Warmup ratio 0.1, Dropout 0.1. The model test set performance is as follows:
[0091] 1. Multi-label classification task: In five-dimensional perception classification, the macro average F1 value (Macro F1) reaches 0.805, the micro average F1 value (Micro F1) is 0.887, the exact match rate (Exact Match Ratio) is 0.609, and the Hamming Loss is only 0.089, indicating that the model can accurately identify complex review semantic categories.
[0092] 2. Sentiment intensity regression task: Mean absolute error (MAE) is 0.323, and coefficient of determination (R ) is 0.774, indicating that the model's predicted sentiment intensity is highly fitted to the human score.
[0093] 3. Auxiliary task: In the sentiment binary classification task, the F1 value reaches 0.981, verifying the effectiveness of the transfer learning strategy.
Claims
1. A city park perception value evaluation method based on multi-modal data fusion, characterized in that, The method comprises: acquiring comment data of park social media, park boundary data, POI data, remote sensing land class proportion and road network information; performing multi-task learning reasoning on the text comment data by using a pre-trained language model, outputting multi-dimensional perception labels and sentiment intensity scores, and aggregating to generate subjective perception results; respectively using a semantic segmentation model, an object detection model and a visual language model to identify scenes, characters / facility elements from image comment data, and acquiring different scene element coverage, character / facility counts and language representations of visual elements; the language model comprises a shared encoder and a multi-label classification head and a multi-target regression head connected in parallel on the shared encoder; the multi-label classification head outputs multi-dimensional binary classification results as multi-dimensional perception labels; the multi-target regression head outputs multi-dimensional sentiment intensity score values; in the model reasoning stage, the reasoning results of the comments are mapped to corresponding park-period units; in each unit, the frequency of each perception label being mentioned is counted, and the arithmetic mean of the sentiment intensity scores of the activated samples is calculated as the subjective perception quantification result of the park in the period; a plurality of objective environment indexes containing variable direction symbols are constructed to describe different environment states, variable information is extracted from the scene element coverage, character / facility count, language representation of visual elements, POI data, remote sensing land class proportion and road network information, and the extracted variable information is mapped to the corresponding objective environment index to quantify each objective environment index; the variable direction symbol is used to describe the positive or negative impact of the variable on the corresponding objective environment, and the objective environment index is constructed by weighted summation according to the positive and negative contribution directions of the variable; the objective environment index is calculated based on the following formula: ; where j is the serial number of the objective environmental index; is the direction symbol of the variable, +1 represents positive promotion, and -1 represents negative inhibition; is the jth original variable value constituting the index; is the jth original variable value constituting the index; and are the mean and standard deviation of the variable in the whole sample, respectively. align the subjective perception results and the objective environment indexes in the park and period dimensions to output.
2. The method of claim 1, wherein, The POI data includes data within the park boundary and within a buffer zone outside the park; the buffer zone is an area formed by extending a predetermined distance outward along the park boundary; The remote sensing land class proportion and road network information are data within the park boundary.
3. The method of claim 1, wherein, The park boundary data is derived from AOI data of an online map; missing boundary data is completed by other open source map data.
4. The method of claim 1, wherein, The variables extracted from the road network information include at least one of total road length and road density.
5. The method of claim 1, wherein, The objective environment index includes at least one of vegetation quality, water quality, original ecological degree, recreational interaction, road facility, service facility, public transportation and commercial activity.
6. The method of claim 1, wherein, The acquired data is divided into multiple periods according to the timestamp, and indexed by park number-period as the primary key; map the subjective perception results and the objective environment indexes to the index based on the primary key to obtain output in the form of a structured data table.
7. The method of claim 1, wherein, The output result at least contains park identification, park name, park location, subjective perception result and objective environment index.
8. The method of claim 7, wherein, The output result also includes park land class proportion obtained based on remote sensing; and different scene element coverage, character / facility count and language representation of visual elements obtained by processing image comment data.
Citation Information
Patent Citations
Street space quality intelligent evaluation and image generation method and system based on deep learning
CN120724519A
Personalized multimodal shopping assistant system using adaptive hybrid processing and method thereof
KR102865941B1