Ecological travel index estimation method based on panoramic semantic segmentation and adaptive model

CN117910698BActive Publication Date: 2026-09-25YUNNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410023142.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-08
Publication Date
2026-09-25
Estimated Expiration
2044-01-08

AI Technical Summary

Technical Problem

[0006]本发明实施例的目的在于提供基于全景语义分割和自适应模型的生态出行指数估算方法,解决了现有技术中全景图像中存在的图像畸变和对象变形、标注数据的缺乏等问题,同时为满足行人感知性需求并降低路径选择的复杂性,提出并开发了一个可步行性框架作为一种有效的解决方案

Benefits of technology

[0068]本方法相较于现有方法具有诸多优点与积极效果。首先,提出了一种基于深度学习技术与街道全景图像测量生态出行综合指数的研究框架,这是一种综合考虑行人感知与实际物理情况,并结合街道全景图像大数据的特征和主观视觉感知的理解街道环境对生态出行影响的新范式,这种方法能够提高城市规划和设计的效率和准确性,节省成本和时间。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117910698B_ABST
    Figure CN117910698B_ABST
Patent Text Reader

Abstract

The application discloses an ecological travel index estimation method based on panoramic semantic segmentation and adaptive model, including the following steps: step one, collecting map data; step two, constructing a panoramic image semantic segmentation model and performing image processing; step three, training a panoramic semantic segmentation and adaptive model; step four, evaluating street ecological quality based on the panoramic image semantic segmentation model; step five, evaluating physical walkability based on GIS; step six, evaluating perceptual walkability based on a man-machine confrontation model; and step seven, obtaining an ecological travel index. The method can better understand the influence of street environment on ecological travel, improve the efficiency and accuracy of urban planning and design, intelligently analyze and evaluate the urban form and the travel conditions of each region, and help better grasp the sustainability of urban development and optimize the urban space structure and layout.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of computer vision and urban planning in smart cities, and in particular relates to a computer-aided auditing method for assessing the comprehensive hopability of a city based on natural street view panoramic images. Background Technology

[0002] Accelerated urbanization and increased reliance on automobiles have led to increasingly severe urban traffic congestion and environmental pollution. Therefore, constructing walkable city streets and areas has become an indispensable factor in urban planning and design. Walking, as a green mode of transportation, not only reduces traffic congestion and environmental pollution but also promotes people's physical and mental health. Walkable city streets and areas also provide residents with spaces for socializing, exercising, and engaging in other daily activities, while simultaneously promoting urban economic and social development. Therefore, urban planning and design need to measure individual accessibility and exposure to factors such as urban green spaces and open sky from a pedestrian perspective to formulate corresponding urban planning and design recommendations, which is of great significance for promoting ecological mobility and urban livability.

[0003] In early research, residents' social relationships were used to describe community life. However, most street studies were conducted through face-to-face interviews, questionnaires, or a combination of expert panel ratings. Some of these methods measured the street environment and tested its association with walking behavior; others assessed indicators such as land use diversity and street green visibility to develop and apply the Pedestrian Environment Index (PEI). However, these methods could only be used for small areas of interest because they were time-consuming, had small sample sizes, were costly, and inefficient.

[0004] In recent years, with the rise of map services and the continuous updating of crowdsourced street view data, large-scale quantitative measurement of street view images (SVIs) based on computer vision technology has become a trend in the era of big data. Using SVIs for visual ecological quality and visual walkability assessment can evaluate pedestrians' psychological feelings and subjective comfort levels regarding the environment. By measuring the physical quality of street visuals, such as greenery coverage, visual congestion, outdoor fencing, and visual pavement width, the physical characteristics of these street view images can be quantified. However, calculating these physical characteristics is difficult to represent the subjective on-site perception of ecological travel.

[0005] Currently, numerous urban studies are underway, utilizing computer vision techniques to further process panoramic street view images (PSVIs) to analyze various indicators. However, due to the inherent rectangular projection of panoramic images, significant image distortion and object deformation often exist. This makes many convolutional neural networks and learning methods unsuitable for panoramic image processing. Besides the distortion problem, the lack of labeled data is another key challenge hindering progress in panoramic image visual processing. Therefore, it is necessary to train models specifically and employ appropriate methods to measure comprehensive indices of ecological travel and explore their relationship with visual elements in the street-based built environment. Summary of the Invention

[0006] The purpose of this invention is to provide an ecological travel index estimation method based on panoramic semantic segmentation and adaptive models, which solves the problems of image distortion and object deformation and lack of labeled data in panoramic images in the prior art. At the same time, in order to meet the pedestrian perception needs and reduce the complexity of path selection, a walkability framework is proposed and developed as an effective solution.

[0007] To solve the above-mentioned technical problems, the technical solution adopted by this invention is an ecological travel index estimation method based on panoramic semantic segmentation and adaptive models, comprising the following steps:

[0008] Step S1: Collect map data;

[0009] Step S2: Construct a panoramic image semantic segmentation model and perform image processing;

[0010] Step S3: Train the panoramic semantic segmentation and adaptive model;

[0011] Step S4: Assess the ecological quality of the streets based on the panoramic image semantic segmentation model;

[0012] Step S5: GIS-based physical hopability assessment;

[0013] Step S6: Perceptual hop assessment based on human-machine adversarial model;

[0014] Step S7: Obtain the Eco-Travel Index.

[0015] Furthermore, step S1 includes:

[0016] Step S11: Obtain road network data of the target study area from the Open Street Map (OSM), perform rasterization and vectorization processing on the road network data image, and perform georegistration operations to convert the latitude and longitude coordinates in the OSM into the local coordinate system;

[0017] Step S12: Sample the vectorized road network data image and set the parameters pitch and field of view width;

[0018] Step S13: Write a program to access the Street View Map Application Programming Interface (API) and collect panoramic street image data (PSVIs) and corresponding Points of Interest (POI) data for the target study area;

[0019] Step S14: Obtain the basic land use type EULUC data of the target study area from EULUC-China and input it into ArcGIS to obtain the land use type data of the target study area;

[0020] Step S15: Perform data cleaning, remove invalid access data, and integrate and classify the collected PSVIs, POIs, and EULUC data.

[0021] Furthermore, step S2 includes:

[0022] Step S21: Construct a panoramic image semantic segmentation model;

[0023] The panoramic image semantic segmentation model includes an input embedding layer, a location encoding layer, an encoder layer, a multi-scale feature extraction layer, a decoder layer, and a prediction layer;

[0024] The input embedding layer is used to divide the input image into a series of image blocks and convert each image block into a vector representation;

[0025] The positional encoding layer is used to add positional information to the input image patch, helping the model understand the relative relationships between different positions;

[0026] The encoder layer consists of multiple stacked encoders. Each encoder contains two sub-layers. It uses a multi-head self-attention mechanism to extract contextual information by calculating the correlation between different locations in the input image patch and performing non-linear transformation and mapping on the features processed by the self-attention mechanism.

[0027] The multi-scale feature extraction layer is used to map input features to different scales and extract richer feature representations through different pooling operations and projection transformation modules. Finally, the features at different scales are fused into a multi-scale feature representation.

[0028] The decoder layer, composed of multiple stacked decoders, is used to associate the features generated by the multi-scale feature extraction layer with the features at each location to obtain the correspondence between the input image patch and the output image patch;

[0029] The prediction layer is used to map the final semantic segmentation output of the image based on the correspondence between the input image patches and the output image patches using a function.

[0030] Step S22: Input an image with a given shape of H×W×D into the encoder and perform patterning processing on the image; H, W, and D are the height, width, and depth of the image, respectively;

[0031] Step S23: The image is processed through three layers of MaxPools of different sizes to construct a spatial feature map and feature mappings f between different scale resolutions. i Where i∈(1,2,3) represents different step sizes (8,16,32). After downsampling with different step sizes, the channel size gradually increases, and the channel size D corresponding to different step sizes is... i ∈(128,256,512);

[0032] Step S24: In each scale space, three nested convolutional layers with the same filter size and stride are used to explore the depth space feature map at each scale, mapping the features at different scale resolutions to f. i Parsed as a uniform shape And input it into the decoder, where the number of embedded channels T is set to 128;

[0033] Step S25: Apply the projection transformation method in each scale space. The specific steps are as follows:

[0034] Step S251: The coordinate system of the original panoramic image is defined by latitude θ∈[0,2π), longitude θ∈[0,2π]. Convert it to an x,y Cartesian product coordinate system and perform a rectangular projection transformation. The specific formula is as follows: Where (θ0, Φ0) are the original spherical coordinates, (θ1, φ1) = (0, 0), θ1 is the center latitude and φ1 is the center longitude;

[0035] Step S252: Using the distortion transformation formula, the area covered by the panoramic image is represented as a set of coordinates (θ). m ,φ m The specific formula is:

[0036] Where, θ m , Φ m It is the angular matrix of the panoramic image of size H×W, composed of a two-dimensional Euler angle series, where m is the panoramic image number; the model divides the image into k slices, θ k , Φ k Let N be the angle matrix numbered k for the slices scanned by the model, and N be the number of pixels in the region; x, y refer to the planar coordinates of a specific location in the panoramic image. The reliability of the orientation information at that location is evaluated by calculating the average value of the (x, y) orientation information at that specific location in the panoramic image and using the ratio of the derivatives.

[0037] Step S253: Convert the formula in S252 to... For any pixel position in the image starting from φ=0, corresponding distortion and warping transformations can be performed;

[0038] Step S26: After the decoder completes the association and decoding of feature information of feature images at different scales, it is input into the prediction layer. The prediction layer will output the final panoramic image semantic segmentation result that matches the size of the input image according to the number of semantic categories of each task.

[0039] Furthermore, step S3 includes:

[0040] Step S31: In the panoramic image semantic segmentation results, use the pinhole image dataset Cityscapes with its set of labeled images I s ={(x s ,y s )∣x s ∈R H×W×3 ,y s ∈{0,1} H×W×n The source domain is} and the target domain is the SynPass synthetic panoramic image dataset I. t ={(x t )∣x t ∈R H×W×3}, The label information from the SynPass dataset is not used here, where I s Source domain sample image group, x s The original RGB three-channel image of the training set, y s Refers to the corresponding sample training labels, R H×W×3 Let {0,1} represent a real matrix of size H×W×3, where H and W are the height and width of the matrix, respectively, and 3 represents the dimension of each element in the matrix, corresponding to the RGB color channels; H×W×n Let {0,1} represent a binary matrix of size H×W×n, where {0,1} represents the set of elements that can only take the values ​​0 or 1, H and W represent the height and width of the matrix, respectively, and n is the number of shared classes.

[0041] Step S32: Apply the panoptic segmentation loss function In source domain I s The model is then fitted and trained, where... Used to label the category to which each pixel belongs. When a pixel (i,j) belongs to the nth category, otherwise For pixels The probability of being predicted as the nth class; p s (y s |x sThe probability of the predicted class obtained by the model during training and validation in the source domain includes the predicted class (1-n), the slice height of iH, the slice width of jW, and the source domain sample image x. s With tag y s Conduct training and verification;

[0042] Step S33: During the source domain fitting process, use the loss calculation formula. in, For target domain pixels The pseudo-labels are used to optimize the model; they are given by the class most likely predicted by the model. For target domain pixels The probability of being predicted as the nth class; y t To fit the selected sample labels from the source domain to the target domain. This refers to the degree of fit between the pseudo-labels obtained by the model in the target domain and the predicted categories;

[0043] Step S34: Using the result obtained in S33 as a prototype, further training is performed on the original target domain SynPass synthetic panoramic image dataset using the loss calculation method in step S33. The correlation in different feature spaces is continuously learned. After adaptation in the original domain, the model is trained with label information as an intermediate domain and saved as the prototype of the intermediate stage.

[0044] Step S35: Using unlabeled panoramic street image data PSVIs as the final target domain, extract features through multiple stages and scales and embed them to construct a prototype;

[0045] Step S36: Match the target data with the source data, and classify the target data according to the transformed labeled source data to obtain a usable panoramic dataset with coarse-grained labels. At the same time, obtain the final optimal model for unsupervised domain adaptation from pinhole street view images to panoramic images.

[0046] Furthermore, step S4 includes:

[0047] Step S41: Input the obtained PSVIs data with geographical location information of the study area into the trained panoramic image semantic segmentation model to perform semantic segmentation prediction of the panoramic image, obtain the position and proportion information of different elements, and output multi-dimensional element segmentation results.

[0048] Step S42: Calculate the street ecological index, which is to calculate the proportion of green vegetation, sky, and road elements in the panoramic street image, represented as: Greenery Index (GVI), Sky View Factor (SVF), and Relative Road Width (RRW), respectively. The calculation formula is:

[0049]

[0050] Among them, PV index This represents the street ecological index for each location, Area p_i The area represents the total number of index pixels captured along direction i in the panoramic image at each sampling point, given by the semantic segmentation model; t_i This represents the total number of pixels of the visual element in the PSVI at a given sampling point.

[0051] Furthermore, step S5 includes:

[0052] Step S51: Construct a GIS network dataset from POI, OSM, and EULUC data;

[0053] Step S52: Calculate the number of facility types within multiple different walking ranges from each sampling point on the street, expand and classify the POI facility data into different weights, and assign gradually decreasing distance decay coefficients to different walking ranges according to different EULUC and road line data types of road network data.

[0054] Step S53: Obtain the physical walk score for each sample point by multiplying the distance from the sampling point to the POI in different ranges by the corresponding attenuation coefficient and overlapping the weights.

[0055] Furthermore, step S6 includes:

[0056] Step S61: Train volunteers who understand the regional socio-economic background or have experience in urban street design research to be responsible for conducting urban perceived hopability scoring.

[0057] Step S62: The system will collect panoramic street image datasets of the study area from the city data collection module and distribute them to each volunteer in a non-completely random order for the volunteers to rate. The volunteers will use numbers to quantify the four criteria of perceived hopability, namely, to rate the perceived hopability in four dimensions: safety, convenience, comfort, and attractiveness. The score ranges from 0 to 100, and the higher the score, the stronger the consistency with perception.

[0058] Step S63: Construct a human-machine adversarial scoring architecture, including a random forest model, to analyze the correlation between the four-dimensional perception dimension and multi-dimensional visual elements on the evaluation results. This requires using the Pearson correlation coefficient ρ for each pair of corresponding variables. x,y The specific formula is as follows: Where Cov(X,Y) represents the covariance of variables X and Y, and σ x and σ y These represent the standard deviations of variables X and Y, respectively. i and yi Let x and y represent the feature values ​​of the i-th sample, respectively, and n represent the number of samples, μ x and μ y y and y represent the mean values ​​of the feature values, respectively. In the random forest model, two-thirds of the samples are used for classification as subsample datasets or model fitting, and the remaining one-third are used as out-of-bag (OBB) data to evaluate the overall error of the model and identify the importance of the variables.

[0059] Step S64: After the volunteers have completed the rating of the first 50 images provided by the system, the system will build a random forest set to adapt to the rating process, and combine the relationship between different element features and the four-dimensional perception score to give the volunteers the predicted score for the 51st image and thereafter.

[0060] Step S65: If, in a certain perceptual dimension, the difference between the volunteer's subjective rating and the model's predicted score is greater than 10 points, and this situation occurs in more than 5 images, the model will re-collect the volunteer's rating features and continue learning based on the previous model until this situation occurs.

[0061] Step S66: If a photo is rated multiple times by multiple volunteers, set the final score of the photo as the median.

[0062] Step S67: If, during the continuous scoring of 200 images of the same perceptual dimension, the difference between the machine prediction score and the human annotation score is within ±5 points, and there are no five consecutive instances where the score is greater than 10 points, then it is considered a compromise between human and machine.

[0063] Step S68: During the process of a volunteer evaluating 500-1000 images, if the OOB validation error of the fitted model is less than 5 points, the user rating process will stop and a human-computer adversarial perceptual hop score dataset will be output.

[0064] Furthermore, step S7 includes:

[0065] Step S71: Conduct a comprehensive evaluation of the street ecological index, physical walkability, and perceived walkability indicators, and standardize the results to a range of 1 to 10;

[0066] Step S72: By comparing expert questionnaires, the weight of each indicator was calculated using the Analytic Hierarchy Process (AHP).

[0067] Step S73: Add all weights together to get 1. Multiply each variable by its corresponding weight and add them together to get the comprehensive Ecological Mobility Index (ETI).

[0068] This method offers numerous advantages and positive effects compared to existing approaches. Firstly, it proposes a research framework based on deep learning technology and panoramic street images to measure a comprehensive ecological travel index. This new paradigm comprehensively considers pedestrian perception and actual physical conditions, combining the characteristics of large-scale panoramic street images with subjective visual perception to understand the impact of the street environment on ecological travel. This approach can improve the efficiency and accuracy of urban planning and design, while saving costs and time.

[0069] Secondly, this method uses deep learning technology to intelligently analyze and evaluate urban morphology and mobility in various regions, which can better grasp the sustainability of urban development, help promote ecological mobility and urban livability, optimize urban spatial structure and layout, improve resource utilization efficiency and ecological protection level, and better achieve coordinated development of urban economy, society, culture and environment. Attached Figure Description

[0070] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0071] Figure 1 This is a flowchart of the overall research method of the present invention.

[0072] Figure 2 This is a perceptual hop scoring architecture based on a human-machine adversarial model in an embodiment of the present invention. Detailed Implementation

[0073] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0074] This invention provides a method for estimating the ecological travel index based on panoramic semantic segmentation and an adaptive model, such as... Figure 1 As shown, the process is divided into three main stages: data collection and processing, panoramic image semantic segmentation and domain adaptive model training, and human-machine adversarial model and physical hop count analysis. The specific implementation steps are as follows:

[0075] S1: Data collection, the specific steps are as follows:

[0076] S11: Obtain road network data of the target study area from the Open Street Map (OSM), perform rasterization and vectorization on the road network data image, and perform georegistration operations to convert the latitude and longitude coordinates in the OSM into the local coordinate system.

[0077] S12: Sample street viewpoints every 30 meters from the vectorized road network data image and set corresponding parameters such as pitch and field of view width to ensure the quality and accuracy of image acquisition.

[0078] S13: Write a program to access the Street View map application programming interface and collect panoramic street view images (PSVIs) and corresponding points of interest (POIs) data for the target study area.

[0079] S14: Obtain the Essential Urban Land Use Classification (EULUC) data for the target study area from EULUC-China and input it into ArcGIS to obtain the land use type data for the target study area.

[0080] S15: Perform data cleaning, remove invalid access data to ensure the reliability of subsequent urban planning analysis and deep learning model training, and integrate and classify the collected PSVIs, POIs and EULUC data.

[0081] S2: Construct a panoramic image semantic segmentation model and perform image processing. The specific steps are as follows:

[0082] S21: Construct a panoramic image semantic segmentation model, which includes an input embedding layer, a location encoding layer, an encoder layer, a multi-scale feature extraction layer, a decoder layer, and a prediction layer.

[0083] (1) Input embedding layer: Divide the input image into a series of image blocks and convert each image block into a vector representation.

[0084] (2) Position encoding layer: Adds position information to the input image patch to help the model understand the relative relationship between different positions.

[0085] (3) Encoder layer: It consists of multiple encoders stacked together. Each encoder contains two sub-layers. It uses a multi-head self-attention mechanism to extract contextual information by calculating the correlation between different positions in the input image patch and to perform nonlinear transformation and mapping on the features processed by the self-attention mechanism.

[0086] (4) Multi-scale feature extraction layer: The input features are mapped to different scales, and richer feature representations are extracted through different pooling operations and projection transformation methods. Finally, the features at different scales are fused into a multi-scale feature representation.

[0087] (5) Decoder layer, which consists of multiple decoders stacked together, associates the features generated by the multi-scale feature extraction layer with the features at each location to obtain the correspondence between the input image patch and the output image patch.

[0088] (6) Prediction layer: Based on the correspondence between the input image patch and the output image patch, the softmax activation function is used to map the image to the final semantic segmentation output result.

[0089] S22: Input an image of a given shape H×W×D into the encoder. The model will pattern the image, where H, W, and D are the height, width, and depth of the image, respectively.

[0090] S23: The processed image passes through three MaxPools of different sizes, mapping the spatial feature map to different resolution scales. i Where i∈(1,2,3) represents different step sizes (8,16,32). After downsampling with different step sizes, the channel size gradually increases, and the channel size D corresponding to different step sizes is... i ∈(128,256,512);

[0091] S24: In each scale space, three nested convolutional layers with the same filter size and stride are used to explore the deep space feature map at each scale, mapping the multi-scale feature map f. i Parsed as a uniform shape And input it into the decoder, where the number of embedded channels is set to 128.

[0092] S25: To address the image distortion and object deformation issues present in panoramic image data, a projection transformation method is proposed, with the following specific steps:

[0093] S251: The coordinate system of the original panoramic image is latitude θ∈[0,2π), longitude θ∈[0,2π]. Convert it to an x,y Cartesian product coordinate system and perform a rectangular projection transformation. The specific formula is as follows: Where (θ0, Φ0) are the original spherical coordinates, (θ1, φ1) = (0, 0), θ1 is the center latitude, and φ1 is the center longitude.

[0094] S252: This type of rectangular projection will have considerable distortion, so a distortion transformation formula is needed to represent the area covered by the panoramic image as a set of coordinates (θ). m ,φ mThe specific formula is:

[0095] Where, θ m Φ m It is an angular matrix of size H×W panoramic images, composed of a series of equally variable two-dimensional Euler angles, where m is the panoramic image number. The model will split the image into k slices, θ k Φ k Let N be the angle matrix numbered k for the slices scanned by the model, where N is the number of pixels in the region. The reliability of the orientation information at a specific location in the panoramic image is evaluated by calculating the average value of the (x, y) orientation information and using the ratio of the derivatives.

[0096] S253: Convert the above formula to It can be observed that there is a strong correlation between image distortion and cos(φ), that is, any pixel position in the image starting from φ=0 can be transformed with corresponding distortion.

[0097] S254: Based on the above three equations, feature maps are extracted at multiple scales. On this basis,

[0098] An embedded projection transformation method helps preserve the shape and spatial layout of objects, making the 2D reference index of feature maps and panoramic representation more accurate and consistent.

[0099] S26: After the decoder completes the association and decoding of feature information of feature images of different scales, it is input into the last softmax prediction layer in the panoramic semantic segmentation model. This layer will output the final semantic segmentation result that matches the size of the input image according to the number of semantic categories of each task.

[0100] S3: Training the panoramic semantic segmentation and adaptive model, the specific process includes:

[0101] S31: In the proposed panoramic image semantic segmentation model, the Cityscapes pinhole image dataset and its set of labeled images are first used. s ={(x s ,y s )∣x s ∈R H×W×3 ,y s ∈{0,1} H×W×n The source domain is} and the target domain is the SynPass synthetic panoramic image dataset I. t ={(x t )∣x t ∈R H×W×3The label information from the SynPass dataset is not used here. Where I s Source domain sample image group, x s The R refers to the original RGB three-channel image of the training set. H×W×3 This represents a real matrix of size H×W×3, where H and W are the height and width of the matrix, respectively, and 3 represents the dimension of each element in the matrix, corresponding to the RGB color channels; y s Refers to the corresponding sample training labels, {0,1} H×W×n Let {0,1} represent a binary matrix of size H×W×n, where {0,1} represents the set of elements that can only take the values ​​0 or 1, H and W represent the height and width of the matrix, respectively, and n is the number of shared classes.

[0102] The purpose of this step is to adapt the model from the source domain to the target domain using n shared classes. The model will adjust from the source domain to the target domain using n shared classes without a target domain label.

[0103] S32: Using the panoptic segmentation loss function In source domain I s The model is trained by fitting the source domain sample image x. s With tag y s Training and validation are conducted. Among them... Used to label the category to which each pixel belongs. When a pixel (i,j) belongs to the nth category, otherwise For pixels The probability p of being predicted as the nth class s (y s |x s The probability of the predicted class obtained by training and validating the model in the source domain includes the predicted class from 1 to n, the slice height of iH, and the slice width of jW.

[0104] S33: In order to apply the pre-trained model to the target data, the loss calculation formula is used during the original domain fitting process. in, For target domain pixels The pseudo-labels are given by the model and used to optimize the model. y t To fit the selected sample labels from the source domain to the target domain. For target domain pixels The probability of being predicted as the nth class. This refers to the degree of fit between the pseudo-labels derived by the model and the predicted categories in the target domain.

[0105] S34: Then, using the results of this stage as a prototype, the model is further trained on the original target domain SynPass synthetic panoramic image dataset using the loss function calculation formula in S33. The model continuously learns the correlation in different feature spaces to fit the difference between sample labels and pseudo-labels between the source domain and the target domain. After adaptation in the original domain, the model is trained with the label information as an intermediate domain and saved as the prototype of the intermediate stage.

[0106] S35: Using the unlabeled panoptic dataset PSVIs as the final target domain, we extract features and embed them into prototypes through multi-stage and multi-scale extraction, making the labels of the final target domain more robust and expressive.

[0107] S36: Match the target data with the source data, and classify the target data according to the transformed labeled source data to obtain a usable panoramic dataset with coarse-grained labels. At the same time, obtain the final optimal model for unsupervised domain adaptation from pinhole street view images to panoramic images.

[0108] The human-machine adversarial model and indicator evaluation module includes street ecological quality assessment based on panoramic semantic segmentation, physical hopability assessment based on GIS, and perceptual hopability assessment based on the human-machine adversarial model.

[0109] S4: Street ecological quality assessment based on panoramic semantic segmentation model, including the following process:

[0110] S41: Input the obtained PSVIs data with geographical location information of the study area into the trained panoramic image semantic segmentation model to perform semantic segmentation prediction of the panoramic image, so as to obtain information such as the position and proportion of different elements. The model will finally output the segmentation results of multi-dimensional elements.

[0111] S42: Calculate the street ecological index, which requires calculating the proportions of green vegetation, sky, and road elements in the panoramic street image. These are represented as: Greenery Index (GVI), Sky View Factor (SVF), and Relative Road Width (RRW), respectively. The calculation formula is:

[0112]

[0113] Among them, PV index This represents the street ecological index for each location, Area p_i The area represents the total number of index pixels captured in direction i of the panoramic image at each sampling point, given by the semantic segmentation model; t_i This represents the total number of pixels of the visual element in the PSVI at a given sampling point.

[0114] S5: GIS-based physical walkability assessment, including the following process:

[0115] S51: Construct a GIS network dataset from POI data, OSM road network data, and EULUC data.

[0116] S52: Calculate the number of facility types within a walking range of 400, 800, 1200, 1600 and 2000 meters from each sampling point on the street, expand the POI facility data into different weights, and assign gradually decreasing distance decay coefficients to different walking ranges according to different EULUC and road line data types of road network data, in order to measure gait speed and walking ability to reach different walking ranges.

[0117] S53: The physical walk score for each sample point is obtained by multiplying the distance from the sampling point to the POIs in different ranges by the corresponding attenuation coefficient and overlapping the weights.

[0118] S6: Perceptual hop assessment based on human-machine adversarial model, including the following process:

[0119] S61: Train 40 volunteers who understand the regional socio-economic background or have experience in urban street design research. They will be responsible for conducting urban perceived hopability assessments.

[0120] S62: The system will collect panoramic street image datasets of the study area from the city data collection module and assign them to each volunteer for scoring in a non-completely random order to reduce dataset error. Four criteria for quantifying perceived walkability are used: safety, convenience, comfort, and attractiveness. Perceived walkability is scored on a scale of 0 to 100, with higher scores indicating stronger consistency with perceived values.

[0121] S63: The human-machine adversarial scoring architecture includes a Random Forest (RF) model to fit the relationship between visual element target features and user ratings across different perceptual dimensions. Visual element target features are the proportions of each element in the multi-dimensional feature vector obtained after semantic segmentation of each sample image using a panoramic semantic segmentation model.

[0122] To analyze the correlation between the four-dimensional perception dimension and multi-dimensional visual elements on the evaluation results, Pearson correlation coefficients were used for pairwise corresponding variables. In the equation... In this context, Cov(X,Y) represents the covariance of variables X and Y, and σ x and σ y These represent the standard deviations of variables X and Y, respectively. i and y i Let x and y represent the feature values ​​of the i-th sample, respectively, and n represent the number of samples, μ x and μy ...

[0123] S64: After the volunteers complete the rating of the first 50 images provided by the system, the system will build a random forest set to adapt to the rating process, and combine the relationship between different element features and the four-dimensional perception score to give the volunteers the predicted rating for the 51st image and subsequent images.

[0124] S65: If, in a certain perceptual dimension, the difference between the volunteer's subjective rating and the model's predicted score is greater than 10 points, and this situation occurs in more than 5 images, the model will re-collect the volunteer's rating features and continue learning based on the previous model.

[0125] S66: If a photo is rated multiple times by multiple volunteers, the final score of the photo is set as the median to ensure the reliability of subjective scoring.

[0126] S67: The human-machine adversarial compromise is defined as the difference between the machine's predicted score and the human's annotated score within ±5 points during the continuous scoring of 200 images of the same perceptual dimension, and without any five consecutive instances where the difference is greater than 10 points.

[0127] S68: Finally, during the process of a volunteer evaluating approximately 500-1000 images, if the OOB validation error of the fitted model is less than 5 points, the user rating process will stop, and a human-computer adversarial perception hopability score dataset will be output.

[0128] S7: Using the Analytic Hierarchy Process (AHP), the weights of each indicator and different variables are obtained to comprehensively evaluate indicators such as street ecological quality (including GVI, SVF, and RRW), physical walkability, and perceived walkability (including safety, comfort, convenience, and attractiveness), and summarize them into a comprehensive Ecological Travel Index (ETI) for urban planners to conduct planning studies in the corresponding research areas. The specific steps are as follows:

[0129] S71: Conduct a comprehensive assessment of indicators such as street ecological quality, physical walkability, and perceived walkability, and standardize the results of these indicators to a range of 1 to 10.

[0130] S72: Compare the results using expert questionnaires and calculate the weight of each indicator using the Analytic Hierarchy Process (AHP), where all weights are summed to 1.

[0131] S73: Each variable is multiplied by its corresponding weight and then summed to obtain the comprehensive Ecological Mobility Index (ETI).

[0132] The key features of this invention are as follows: First, it proposes a research framework based on deep learning technology and panoramic street images for measuring the comprehensive ecological mobility index. This method comprehensively considers pedestrian perception and actual physical conditions, combining the characteristics of panoramic PSVIs big data and subjective visual perception to better understand the impact of the street environment on ecological mobility. This approach can improve the efficiency and accuracy of urban planning and design while saving costs and time. Furthermore, this method contributes to promoting the development of ecological mobility and urban livability. Utilizing deep learning technology for intelligent analysis and evaluation of urban morphology and mobility in various areas helps to better grasp the sustainability of urban development, optimize urban spatial structure and layout, improve resource utilization efficiency and ecological protection levels, and ultimately better achieve the coordinated development of urban economy, society, culture, and environment.

[0133] Secondly, in the field of urban planning research based on PSVIs, this invention is the first to consider the problems of image distortion and scarce annotations in panoramic images. It proposes a novel panoramic image semantic segmentation model to address these two issues, and further proposes a transferable, unsupervised, adaptive, multi-stage prototype strategy for outdoor street scenes to address the scarcity of panoramic image samples. This strategy is integrated into the method of this invention for the analysis of ecological travel comprehensive indices. For example, when planning pedestrian walkways, the proposed model, with its higher accuracy, analyzes elements such as green vegetation and walking width in panoramic images. This allows for a more scientific determination of the layout and width of pedestrian walkways, reducing traffic conflicts between pedestrians and vehicles and improving pedestrian safety and comfort.

[0134] Finally, the unsupervised domain adaptive multi-stage prototyping strategy proposed in this invention can be used for self-labeling of panoramic image data of the research area, enabling the dataset to be used as a training or validation set when subsequent segmentation tasks for similar tasks are required. Furthermore, the methods for collecting and processing this dataset can be extended to other cities or regions to provide reference for addressing related issues.

[0135] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0136] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. An ecological travel index estimation method based on panoramic semantic segmentation and adaptive models, characterized in that, Includes the following steps: Step S1: Collect map data; Step S2: Construct a panoramic image semantic segmentation model and perform image processing; Among them, S21: The constructed panoramic image semantic segmentation model includes an input embedding layer, a location encoding layer, an encoder layer, a multi-scale feature extraction layer, a decoder layer, and a prediction layer; Step S22: Input a given shape The image is fed into the encoder and patterned; H, W, and D are the height, width, and depth of the image, respectively. Step S23: The image is processed through three layers of MaxPools of different sizes to construct a spatial feature map and feature mapping between different scale resolutions. ,in Indicate different step sizes After downsampling at different lengths, the channel size gradually increases, and the channel size corresponding to different lengths... ; Step S24: In each scale space, three nested convolutional layers with the same filter size and stride are used to explore the depth space feature map at each scale, mapping the features at different scale resolutions. Parsed as a uniform shape And input it into the decoder, where the number of embedded channels T is set to 128; Step S25: Apply the projection transformation method in each scale space. The specific steps are as follows: S251: The original panoramic image is based on a spherical latitude and longitude coordinate system. It is converted into a Cartesian product coordinate system through equidistant projection transformation, and the coordinates of the center point of the sphere are set. S252: The area covered by the panoramic image is mapped to a set of coordinates, the image is split into multiple slices, and the reliability of the orientation information is evaluated by calculating the average value of the orientation information at a specific location in the image and using the ratio of the derivatives. S253: Formula for applying the corresponding distortion and warping transformation to any pixel position in the image; Step S3: Train the panoramic semantic segmentation and adaptive model; Step S31: In the panoramic image semantic segmentation results, use the Cityscapes pinhole image dataset with a set of its labeled images. The source domain is the SynPass synthetic panoramic image dataset, and the target domain is the SynPass synthetic panoramic image dataset. The label information from the SynPass dataset is not used here. Source domain sample image group, Refers to the original RGB three-channel images of the training set. Refers to the corresponding sample training labels. This represents a real matrix of size H×W×3, where H and W are the height and width of the matrix, respectively, and 3 represents the dimension of each element in the matrix, corresponding to the RGB color channels. Let {0,1} represent a binary matrix of size H×W×n, where {0,1} represents the set of elements that can only take the values ​​0 or 1, H and W represent the height and width of the matrix, respectively, and n is the number of shared classes. Step S32: Apply the panoptic segmentation loss function In the source domain The model is then fitted and trained, where... Used to label the category to which each pixel belongs. When a pixel (i,j) belongs to the nth category, ,otherwise For pixels The probability of being predicted as the nth class; This refers to the probability of the predicted class obtained by the model during training and validation in the source domain, including the predicted class from 1 to n, the slice height of iH, the slice width of jW, and the source domain sample images required. With tags Conduct training and verification; Step S33: During the source domain fitting process, use the loss calculation formula. ,in, For target domain pixels The pseudo-labels are used to optimize the model; they are given by the class most likely predicted by the model. For target domain pixels The probability of being predicted as the nth class; To fit the selected sample labels from the source domain to the target domain. This refers to the degree of fit between the pseudo-labels obtained by the model in the target domain and the predicted categories; Step S34: Using the result obtained in S33 as a prototype, further training is performed on the original target domain SynPass synthetic panoramic image dataset using the loss calculation method in step S33. The correlation in different feature spaces is continuously learned. After adaptation in the original domain, the model is trained with label information as an intermediate domain and saved as the prototype of the intermediate stage. Step S35: Using unlabeled panoramic street image data PSVIs as the final target domain, extract features through multiple stages and scales and embed them to construct a prototype; Step S36: Match the target data with the source data, and classify the target data according to the transformed labeled source data to obtain a usable panoramic dataset with coarse-grained labels, and at the same time obtain the final optimal model for unsupervised domain adaptation from pinhole street view images to panoramic images. Step S4: Assess the ecological quality of the streets based on the panoramic image semantic segmentation model; Step S5: GIS-based physical hopability assessment; Step S6: Perceptual hop assessment based on human-machine adversarial model; Step S7: Obtain the Eco-Travel Index.

2. The method for estimating the ecological travel index based on panoramic semantic segmentation and an adaptive model according to claim 1, characterized in that, Step S1 includes: Step S11: Obtain road network data of the target study area from the Open Street Map (OSM), perform rasterization and vectorization processing on the road network data image, and perform georegistration operations to convert the latitude and longitude coordinates in the OSM into the local coordinate system; Step S12: Sample the vectorized road network data image and set the parameters pitch and field of view width; Step S13: Write a program to access the Street View Map Application Programming Interface (API) and collect panoramic street image data (PSVIs) and corresponding Points of Interest (POI) data for the target study area; Step S14: Obtain the basic land use type EULUC data of the target study area from EULUC-China and input it into ArcGIS to obtain the land use type data of the target study area; Step S15: Perform data cleaning, remove invalid access data, and integrate and classify the collected PSVIs, POIs, and EULUC data.

3. The method for estimating the ecological travel index based on panoramic semantic segmentation and an adaptive model according to claim 1, characterized in that, In step S21: The input embedding layer is used to divide the input image into a series of image patches and convert each image patch into a vector representation; The positional encoding layer is used to add positional information to the input image patch, helping the model understand the relative relationships between different positions; The encoder layer consists of multiple stacked encoders. Each encoder contains two sub-layers. It uses a multi-head self-attention mechanism to extract contextual information by calculating the correlation between different locations in the input image patch and performing non-linear transformation and mapping on the features processed by the self-attention mechanism. The multi-scale feature extraction layer is used to map input features to different scales and extract richer feature representations through different pooling operations and projection transformation modules. Finally, the features at different scales are fused into a multi-scale feature representation. The decoder layer, composed of multiple stacked decoders, is used to associate the features generated by the multi-scale feature extraction layer with the features at each location to obtain the correspondence between the input image patch and the output image patch; The prediction layer is used to map the final semantic segmentation output of the image based on the correspondence between the input image patches and the output image patches using a function. Step S25 includes: employing a projection transformation method in each scale space, with the specific steps as follows: Step S251: Latitude of the coordinate system to which the original panoramic image belongs ,longitude Convert it to The specific formula for performing a rectangular projection transformation on a Cartesian product coordinate system is as follows: in, The coordinates are in the original spherical coordinate system. With center latitude, Central longitude; Step S252: Use the distortion transformation formula to represent the area covered by the panoramic image as a set of coordinates. The specific formula is as follows: ,in, It is the angular matrix of the panoramic image of size H×W, composed of a two-dimensional Euler angle series, where m is the panoramic image number; the model divides the image into k slices. Let k be the angle matrix representing the slices scanned by the model. It is the number of pixels in the region; x, y refers to the planar coordinates of a specific location in the panoramic image. The reliability of the orientation information at that location is evaluated by calculating the average value of the (x, y) orientation information at that specific location in the panoramic image and using the ratio of the derivatives. Step S253: Convert the formula in S252 to... , for from Any pixel in the starting image can be subjected to corresponding distortion and warping transformations; It also includes step S26: After the decoder completes the association and decoding of feature information of feature images of different scales, it is input into the prediction layer. The prediction layer will output the final panoramic image semantic segmentation result that matches the size of the input image according to the number of semantic categories of each task.

4. The method for estimating the ecological travel index based on panoramic semantic segmentation and an adaptive model according to claim 1, characterized in that, Step S4 includes: Step S41: Input the obtained PSVIs data with geographical location information of the study area into the trained panoramic image semantic segmentation model to perform semantic segmentation prediction of the panoramic image, obtain the position and proportion information of different elements, and output multi-dimensional element segmentation results. Step S42: Calculate the street ecological index, which is to calculate the proportion of green vegetation, sky, and road elements in the panoramic street image, represented as: Greenery Index (GVI), Sky View Factor (SVF), and Relative Road Width (RRW), respectively. The calculation formula is: , in, This represents the street ecological index for each location. This represents the total number of index pixels captured in direction i of the panoramic image at each sampling point, given by the semantic segmentation model; This represents the total number of pixels of the visual element in the PSVI at a given sampling point.

5. The method for estimating the ecological travel index based on panoramic semantic segmentation and an adaptive model according to claim 1, characterized in that, Step S5 includes: Step S51: Construct a GIS network dataset from POI, OSM, and EULUC data; Step S52: Calculate the number of facility types within multiple different walking ranges from each sampling point on the street, expand and classify the POI facility data into different weights, and assign gradually decreasing distance decay coefficients to different walking ranges according to different EULUC and road line data types of road network data. Step S53: Obtain the physical walk score for each sample point by multiplying the distance from the sampling point to the POI in different ranges by the corresponding attenuation coefficient and overlapping the weights.

6. The method for estimating the ecological travel index based on panoramic semantic segmentation and an adaptive model according to claim 1, characterized in that, Step S6 includes: Step S61: Train volunteers who understand the regional socio-economic background or have experience in urban street design research to be responsible for conducting urban perceived hopability scoring. Step S62: The system will collect panoramic street image datasets of the study area from the city data collection module and distribute them to each volunteer in a non-completely random order for the volunteers to rate. The volunteers will use numbers to quantify the four criteria of perceived hopability, namely, to rate the perceived hopability in four dimensions: safety, convenience, comfort, and attractiveness. The score ranges from 0 to 100, and the higher the score, the stronger the consistency with perception. Step S63: Construct a human-machine adversarial scoring architecture, including a random forest model, to analyze the correlation between the four-dimensional perception dimension and multi-dimensional visual elements on the evaluation results. This requires using the Pearson correlation coefficient on each pair of corresponding variables. The specific formula is as follows: ,in, Representing variables and covariance, and Representing variables respectively and standard deviation and Let x and y represent the feature values ​​of the i-th sample, respectively, and n represent the number of samples. and y and y represent the mean values ​​of the feature values, respectively. In the random forest model, two-thirds of the samples are used for classification as subsample datasets or model fitting, and the remaining one-third are used as out-of-bag (OOB) data to evaluate the overall error of the model and identify the importance of the variables. Step S64: After the volunteers have completed the rating of the first 50 images provided by the system, the system will build a random forest set to adapt to the rating process, and combine the relationship between different element features and the four-dimensional perception score to give the volunteers the predicted score for the 51st image and thereafter. Step S65: If, in a certain perceptual dimension, the difference between the volunteer's subjective rating and the model's predicted score is greater than 10 points, and this situation occurs in more than 5 images, the model will re-collect the volunteer's rating features and continue learning based on the previous model until this situation occurs. Step S66: If a photo is rated multiple times by multiple volunteers, set the final score of the photo as the median. Step S67: If, during the continuous scoring of 200 images of the same perceptual dimension, the difference between the machine prediction score and the human annotation score is within ±5 points, and there are no five consecutive instances where the score is greater than 10 points, then it is considered a compromise between human and machine. Step S68: During the process of a volunteer evaluating 500-1000 images, if the OOB validation error of the fitted model is less than 5 points, the user rating process will stop and a human-computer adversarial perceptual hop score dataset will be output.

7. The method for estimating the ecological travel index based on panoramic semantic segmentation and an adaptive model according to any one of claims 1 to 6, characterized in that, Step S7 includes: Step S71: Conduct a comprehensive evaluation of the street ecological index, physical walkability, and perceived walkability indicators, and standardize the results to a range of 1 to 10; Step S72: By comparing expert questionnaires, the weight of each indicator was calculated using the Analytic Hierarchy Process (AHP). Step S73: Add all weights together to get 1. Multiply each variable by its corresponding weight and add them together to get the comprehensive Ecological Mobility Index (ETI).

Citation Information

Patent Citations

  • Street quality evaluation method based on man-machine confrontation scoring

    CN111126864A

  • Semi-supervised semantic segmentation method based on self-adaptive false label correction

    CN115761735A