Method for recognizing spaces to be optimized in built environment on basis of morphological hierarchy model
Patent Information
- Application Number
- PCT/CN2025/134545
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2025-11-13
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025134545_01102026_PF_FP_ABST
Abstract
Description
A method for identifying the built environment space to be optimized based on a morphological hierarchical model Technical Field
[0001] This invention relates to the field of urban planning, and in particular to a method for identifying spaces in the built environment that need optimization. Background Technology
[0002] The urban built environment, comprising buildings, green spaces, and roads, is a vital carrier of human social development and daily life. The rationality and quality of its spatial layout directly impact residents' comfort, urban efficiency, and sustainable development capabilities. In today's increasingly dense and complex urban landscape, coupled with dwindling land resources and diverse resident demands, numerous spaces within the built environment require optimization. These spaces manifest as incompatible aesthetics, low space utilization, and poor environmental quality. Their existence not only affects the overall image and operational efficiency of the city but also lowers residents' quality of life. Therefore, accurate identification of spaces requiring optimization within the built environment is a crucial prerequisite for developing scientific and rational optimization strategies to enhance urban competitiveness and sustainable development. Its significance lies in enabling urban planners and decision-makers to proactively promote urban renewal initiatives, optimize resource allocation, and effectively improve the quality of the urban built environment.
[0003] Current methods for identifying areas of built environment to be optimized have certain limitations. On the one hand, traditional methods relying on urban planners' experience for identification suffer from low efficiency, strong subjectivity, and difficulty in fully considering the complexity and dynamism of the built environment. On the other hand, identification methods that treat the research object as a homogeneous whole and divide it horizontally lack hierarchical delineation of spatial structure and ignore the interaction and synergy of morphological characteristics of different levels of the built environment. This results in a lack of systematicity, making it impossible to accurately locate areas to be optimized. Consequently, the accuracy and reliability of the identification results are greatly reduced, making it difficult to provide effective support for precise policy implementation. Summary of the Invention
[0004] To address the shortcomings mentioned in the background art, the present invention aims to provide a method for identifying the built environment space to be optimized based on a morphological hierarchical model, which can automatically identify the built environment space to be optimized with high precision.
[0005] The objective of this invention can be achieved through the following technical solutions:
[0006] A method for identifying the built environment space to be optimized based on a morphological hierarchical model includes the following steps:
[0007] Step 1: Construction Environment Data Collection and Preprocessing
[0008] Using drones with storage capacity of 5T or more, equipped with multispectral cameras and LiDAR systems with an accuracy of 100mm or more, three-dimensional morphological data of the built environment of the target are collected and identified, specifically including building morphology data, green space morphology data, and road morphology data. The data is cleaned to remove outliers. The data is standardized into a voxel network, with a voxel volume of 1 cubic decimeter. Each voxel stores the three-dimensional morphological information of the building, green space, or road at the corresponding spatial location, generating a voxel network dataset.
[0009] Step Two: Preliminary Classification of Built Environment Morphology
[0010] Collect street boundary vector data of the target built environment and obtain a voxel network dataset within its coordinate range; based on the voxel network dataset, use multi-task Bayesian federated learning (BFL) to divide the street boundary vector data of the target built environment into a central area, a transition area and an edge area.
[0011] Specifically, three sub-tasks are first defined: identifying the central area, transition area, and edge area of the target built environment; on local devices, the three sub-tasks are jointly modeled through a multi-output Gaussian process (MOGP), the training set is used to optimize the model parameters, a prior distribution is introduced, and the three-dimensional combination features of building voxels, green space voxels, and road voxels in the voxel network data are captured. Then, the posterior distributions on different devices are uploaded to the global processor for aggregation to form a global MOGP prior, which is then distributed back to each local device for the next round of training. After no less than 50 iterations, the model converges, and the street boundary vector data of the target built environment is classified into the central area, transition area, and edge area.
[0012] Step 3: Revise the morphological gradation of the built environment
[0013] Based on the hierarchical results of the street boundary vector data of the target built environment, blocks belonging to the same partition and located in the spatial neighborhood are connected to form several street sets. If the total area of the blocks in the street set is not less than 1 square kilometer, the verification is passed; if it is less than 1 square kilometer, the partition to which it belongs is corrected. During the correction process, if the adjacent street sets of this street set all belong to the same partition, then this street set is also assigned to this partition; if the blocks surrounding this street set belong to different partitions, the three-dimensional combination features of the voxel network within the target street set are compared to which adjacent street set is most similar, and this street set is assigned to the partition to which the most similar street set belongs.
[0014] Step 4: Preliminary identification of areas to be optimized
[0015] Using a near-Earth satellite with over 100T of storage space, equipped with a multispectral camera and a lidar system with an accuracy of over 10m, nighttime light data of the target built environment was collected over one calendar month. Based on land use attributes, commercial, residential, and industrial blocks within the target built environment of the central, transition, and edge zones were extracted, and their corresponding nighttime light data averages were matched according to coordinates. For blocks at the same tier and with the same land use attribute, the XGBoost model was used to divide each tier and function block into three intervals based on the average nighttime light data. Blocks in the lowest interval were identified as the space to be optimized. Based on land use status information, blocks with a land use status of "under construction" or "completed but not yet operational" were removed, resulting in a preliminary dataset of the commercial, residential, and industrial blocks to be optimized within the central, transition, and edge zones.
[0016] Step 5: Spatial hierarchy recognition to be optimized
[0017] For the initial optimization space in the central area, a drone equipped with a multispectral camera was used to collect image data of the corresponding coordinate blocks. A circling acquisition was performed every 10 meters in height. Based on the central area case database, the Vision Transformer (ViT) was used to determine whether the streetscape of various blocks in the central area needed optimization. For the initial optimization space in the transition area, a street view acquisition vehicle equipped with a 360° panoramic camera was used to collect streetscape image data of the corresponding coordinate blocks. Based on the transition area case database, the Convolutional Neural Network (CNN) was used to determine whether the streetscape of various blocks in the transition area needed optimization. For the initial optimization space in the edge area, satellite image data of the corresponding coordinate blocks was acquired via near-Earth satellites. Stereo image pairs were used to capture images of the target at 60° intervals to obtain overlapping stereo images. Matching and calculation were performed to extract the 3D information of the graphic elements of the target area. Based on the edge area case database, the CNN was used to determine whether the appearance of various blocks in the edge area needed optimization. This resulted in a hierarchical recognition dataset of optimization spaces for commercial, residential, and industrial blocks within the central, transition, and edge areas.
[0018] Step Six: Interaction and Feedback on Recognition Results
[0019] The target built environment's hierarchical recognition dataset of spaces to be optimized, output from step five, is imported into a holographic sand table display device equipped with a wearable 3D motion capture system and a large display screen with a resolution of 8K or higher. This device displays the distribution of commercial, residential, and industrial blocks belonging to the central, transition, and edge areas of the target built environment, allowing for virtual human-computer interaction and recording feedback information. Information on blocks deemed not to be spaces to be optimized is fed back to the corresponding hierarchical case database, updating the corresponding hierarchical recognition results of spaces to be optimized.
[0020] Furthermore, in step two, the three-dimensional combined features of building voxels, green space voxels, and road voxels in the voxel network data are captured. Specifically, this is achieved by employing three sub-tasks of joint modeling using a multi-output Gaussian process (MOGP) to capture the three-dimensional combined features. First, based on the voxel network dataset, a prior distribution is defined, and the spatial distribution patterns of building, green space, and road voxels are embedded into the covariance function. The function quantifies the spatial correlation of different voxel types, covering average building height, building height variation, green space coverage, green space height variation, green space plant morphology richness, road network density, and road network connectivity, as shown in the table below.
[0021] A multidimensional covariance matrix is constructed to reflect the statistical regularity of the three-dimensional combination features. During the local device training phase, when optimizing the MOGP model parameters using the training set, a class weight factor is introduced to assign dynamic weights to building, green space, and road voxels respectively. By maximizing the marginal likelihood function, the weights are adaptively adjusted to strengthen the dominant features. The building and road voxel densities are higher in the central area, and their weights account for a larger proportion in the covariance calculation, decreasing sequentially in the transition and edge areas. The green space voxel density is higher in the edge area, and its weights account for a larger proportion in the covariance calculation, decreasing sequentially in the edge and central areas. During training, the posterior distribution of each subtask is updated through variational inference, and the parameters of each device are fused using a weighted average method to generate a global prior distribution. After at least 50 iterations, the model converges, and the global prior is used to represent the differences in the three-dimensional combination features of different zones.
[0022] Furthermore, in step three, the three-dimensional combination features of the voxel network within the target block set are compared to which adjacent block set is closest. Specifically, the following indicators of the voxel network within this block set are extracted: three-dimensional fractal dimension, sky openness, plot ratio, building density, green patch density, and road intersection density, as shown in the table below.
[0023] A feature matching algorithm is used to compare the extracted three-dimensional combined features with the features of adjacent street sets one by one to obtain the similarity score between each adjacent street set and the target street set; the tier type to which the street set with the highest score belongs is selected as the tier type to which the target street set belongs.
[0024] Further, in step four, each functional block is divided into three intervals based on the average nighttime light data. This is done independently for different functional blocks at different levels. The average nighttime light data of each functional block at each level is normalized after removing outliers. Rules are set for model training: the average nighttime light data of blocks of the same functional type in the central area, transition area, and edge area decreases sequentially. Nighttime light data from 6 PM to 10 PM is collected for commercial land, from 6 PM to midnight for residential land, and from 6 PM to 6 AM for industrial land. The XGBoost model is trained using the preprocessed dataset to learn how to map blocks to three nighttime activity intervals: high, medium, and low. After model training, the model is applied to the test dataset to verify classification accuracy. The average nighttime light data, the corresponding level, and the land use function attribute are extracted from the blocks to be classified, input into the XGBoost model, and the nighttime activity interval to which the block belongs is output.
[0025] Further, in step five, it is determined whether the streetscape of various blocks in the central area needs optimization. Specifically, this involves constructing a central area case database containing image data before and after optimization for no less than 500 cases of streetscape optimization in the central area; preprocessing the image data of the blocks in the initial central area to be optimized by using a Vision Transformer to segment the image into 10*10 pixel blocks, which are linearly embedded in a high-dimensional space to form a set of feature vectors, which are input into the Transformer encoder for information interaction and feature fusion to obtain the image feature codes of the blocks; calculating the similarity between the feature codes, and determining the case image data most similar to the block to be judged. If the case image data belongs to the pre-optimization data, the block to be judged is determined to belong to the space to be optimized; if the case image data belongs to the post-optimization data, the block to be judged is determined not to belong to the space to be optimized.
[0026] The beneficial effects of this invention are:
[0027] 1. This invention is based on multi-output Gaussian process joint modeling, dynamically adjusting the category weight factors of building, green space, and road voxels to enhance the differences in 3D combined features of different zones. In the central zone, the density of building and road voxels is higher, and their weights account for a larger proportion in the covariance calculation, ensuring that high-density development features are accurately captured; in the edge zone, the density of green space voxels is higher, and their weights account for a larger proportion, highlighting ecological features. Through no less than 50 iterations of training, the model adapts to the differences in zone features, significantly improving the segmentation accuracy of the central zone, transition zone, and edge zone, avoiding the static bias of traditional methods.
[0028] 2. This invention extracts multi-dimensional indicators such as sky openness, floor area ratio, and building density, and combines them with a feature matching algorithm to quantify the similarity of adjacent street blocks, ensuring the objectivity and accuracy of the zoning correction. For street blocks with an area of less than 1 square kilometer, the similarity score between their three-dimensional combined features and adjacent street blocks is compared, and they are assigned to the most similar zoning. This mechanism effectively solves the problem of ambiguous classification of streets with insufficient area, improving the scientific nature and reliability of the zoning results.
[0029] 3. This invention collects nighttime light data for commercial, residential, and industrial land at different times, and uses the XGBoost model to divide the area into high, medium, and low activity zones, scientifically identifying low-activity blocks as spaces to be optimized. By independently dividing the nighttime light data for different functional zones, the functional adaptability of the identification is improved. This method not only improves the accuracy of identifying spaces to be optimized but also provides data support for urban planning.
[0030] 4. This invention, based on a database of case studies in the central, transition, and edge areas, extracts image feature codes using Vision Transformer and convolutional neural networks. It then matches and optimizes case data before and after optimization to achieve automated determination of streetscape improvements. This method reduces reliance on human experience and ensures the objectivity and operability of the streetscape assessment.
[0031] 5. This invention enables real-time interaction between the user and a holographic sand table display device, and updates model parameters based on a revised case database, forming a closed loop of "identification-feedback-optimization". This mechanism enhances the dynamic iteration capability of spatial identification results, providing strong support for urban planning decisions. Attached Figure Description
[0032] Figure 1 is a flowchart of the method of the present invention;
[0033] Figure 2 shows the results of nighttime light recognition in Shanghai.
[0034] Figure 3 shows the results of the tiered identification of inefficient land use in Shanghai. Detailed Implementation
[0035] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0036] A method for identifying the built environment space to be optimized based on a morphological hierarchical model, as shown in Figure 1, includes the following steps:
[0037] I. Built Environment Data Acquisition and Preprocessing. Using drones with over 5TB of storage, equipped with multispectral cameras and LiDAR systems with an accuracy of over 100mm, three-dimensional morphological data of the target built environment are acquired and identified. This includes building morphology data, green space morphology data, and road morphology data. The data is cleaned to remove outliers. The data is standardized to a voxel network format, with each voxel having a volume of 1 cubic decimeter. Each voxel stores the three-dimensional morphological information of the corresponding building, green space, or road in its spatial location, generating a voxel network dataset.
[0038] II. Preliminary Classification of Built Environment Morphology. The block boundary vector data of the target built environment is collected, and a voxel network dataset within its coordinate range is obtained. Based on the voxel network dataset, multi-task Bayesian federated learning (BFL) is used to divide the block boundary vector data of the target built environment into central, transition, and edge regions. Specifically, three sub-tasks are defined: identifying the central, transition, and edge regions of the target built environment. At the local devices, the three sub-tasks are jointly modeled using a multi-output Gaussian process (MOGP). The training set is used to optimize the model parameters, and a prior distribution is introduced to capture the three-dimensional combination features of building voxels, green space voxels, and road voxels in the voxel network data. The posterior distributions from different devices are then uploaded to a global processor for aggregation, forming a global MOGP prior, which is distributed back to each local device for the next round of training. After at least 50 iterations, the model converges, and the block boundary vector data of the target built environment is classified into central, transition, and edge regions.
[0039] The capture of the three-dimensional combined features of building voxels, green space voxels, and road voxels in the voxel network data is specifically achieved through three sub-tasks: using a multi-output Gaussian process (MOGP) joint modeling to capture the three-dimensional combined features. First, based on the voxel network dataset, a prior distribution is defined, and the spatial distribution patterns of building, green space, and road voxels are embedded in the covariance function. The function quantifies the spatial correlation of different voxel types, covering average building height, building height variation, green space coverage, green space height variation, green space plant morphology richness, road network density, and road network connectivity (see table below).
[0040] A multidimensional covariance matrix is constructed to reflect the statistical regularity of the three-dimensional combination features. During the local device training phase, when optimizing the MOGP model parameters using the training set, a class weight factor is introduced to assign dynamic weights to building, green space, and road voxels respectively. By maximizing the marginal likelihood function, the weights are adaptively adjusted to strengthen the dominant features. The building and road voxel densities are higher in the central area, and their weights account for a larger proportion in the covariance calculation, decreasing sequentially in the transition and edge areas. The green space voxel density is higher in the edge area, and its weights account for a larger proportion in the covariance calculation, decreasing sequentially in the edge and central areas. During training, the posterior distribution of each subtask is updated through variational inference, and the parameters of each device are fused using a weighted average method to generate a global prior distribution. After at least 50 iterations, once the model converges, the global prior is used to characterize the differences in the three-dimensional combination features of different zones.
[0041] 3. Correct the morphological classification of the built environment; based on the classification results of the block boundary vector data of the target built environment, connect blocks belonging to the same partition and located in the spatial neighborhood to form several block sets; if the total area of the blocks in the block set is not less than 1 square kilometer, the verification is passed; if it is less than 1 square kilometer, the partition to which it belongs is corrected; during the correction process, if the block sets adjacent to this block set all belong to the same partition, then this block set is also assigned to this partition; if the blocks surrounding this block set belong to different partitions, compare and select which adjacent block set the three-dimensional combination feature of the voxel network in the target block set is most similar to, and assign this block set to the partition to which the most similar block set belongs;
[0042] The comparison determines which adjacent block set has the closest three-dimensional combination features of the voxel network within the target block set to the target block set. Specifically, this is achieved by extracting the following indicators from the voxel network within this block set: three-dimensional fractal dimension, sky openness, floor area ratio, building density, green patch density, and road intersection density (see table below).
[0043] A feature matching algorithm is used to compare the extracted three-dimensional combined features with the features of adjacent street sets one by one to obtain the similarity score between each adjacent street set and the target street set; the classification type of the street set with the highest score is selected as the classification type of the target street set.
[0044] IV. Preliminary Identification of Spaces to be Optimized: Using near-Earth satellites with storage capacity of over 100T equipped with multispectral cameras and lidar systems with an accuracy of over 10m, nighttime light data of the target built environment was collected over one calendar month. Based on the land use function attributes within the spaces to be optimized, commercial blocks, residential blocks, and industrial blocks in the target built environment of the central area, transition area, and edge area were extracted, and the corresponding average nighttime light data was matched according to the coordinates. For blocks in the same tier and with the same land use function attribute, the blocks of each tier and function were divided into three intervals based on the average nighttime light data according to the XGBoost model. The blocks in the lowest interval were identified as spaces to be optimized. Based on the land use status information, blocks with the land use status of "under construction" or "completed but not yet operational" were removed, resulting in a preliminary dataset of commercial blocks, residential blocks, and industrial blocks in the central area, transition area, and edge area to be optimized.
[0045] The step of dividing each functional block into three intervals based on the average nighttime light data means independently dividing the blocks into different functional zones within the space to be optimized. Outliers are removed and the data is normalized for the average nighttime light data collected from each functional block. Rules are set for model training: the average nighttime light data of blocks belonging to the same functional type in the central area, transition area, and edge area decrease sequentially. Nighttime light data from 6 PM to 10 PM is collected for commercial land, from 6 PM to midnight for residential land, and from 6 PM to 6 AM for industrial land. The XGBoost model is trained using the preprocessed dataset to learn how to map blocks to high, medium, and low nighttime activity intervals. After model training, it is applied to the test dataset to verify classification accuracy. The average nighttime light data, functional zone, and land use attribute of the block to be classified are extracted, input into the XGBoost model, and the nighttime activity interval to which the block belongs is output.
[0046] V. Hierarchical Recognition of Spaces to be Optimized; For the initial space to be optimized in the central area, a drone equipped with a multispectral camera was used to collect image data of the corresponding coordinate blocks, performing a surround acquisition every 10 meters in height. The VisionTransformer (ViT) was used based on the central area case database to determine whether the streetscape of various blocks in the central area needed optimization. For the initial space to be optimized in the transition area, a street view acquisition vehicle equipped with a 360° panoramic camera was used to collect streetscape image data of the corresponding coordinate blocks. The Convolutional Neural Network (CNN) was used based on the transition area case database to determine whether the streetscape of various blocks in the transition area needed optimization. For the initial space to be optimized in the edge area, satellite image data of the corresponding coordinate blocks was acquired via near-Earth satellites. Stereo image pairs were used to capture images of the target at 60° intervals to obtain overlapping stereo images. Matching and calculation were performed to extract the 3D information of the graphic elements of the target area. The CNN was used based on the edge area case database to determine whether the appearance of various blocks in the edge area needed optimization. This resulted in a hierarchical recognition dataset of spaces to be optimized for commercial, residential, and industrial blocks within the central, transition, and edge areas.
[0047] The process of determining whether the streetscape of various blocks in the central area needs optimization involves constructing a central area case database containing image data before and after optimization for no fewer than 500 cases of streetscape optimization in the central area; preprocessing the image data of the blocks in the initial central area to be optimized by using a Vision Transformer to segment the image into 10*10 pixel blocks, which are linearly embedded in a high-dimensional space to form a set of feature vectors, which are then input into the Transformer encoder for information interaction and feature fusion to obtain the image feature codes of the blocks; calculating the similarity between the feature codes, and determining the case image data most similar to the block to be determined. If the case image data belongs to the pre-optimization data, the block to be determined is considered to belong to the space to be optimized; if the case image data belongs to the post-optimization data, the block to be determined is considered not to belong to the space to be optimized.
[0048] VI. Recognition Result Interaction and Feedback; The output target built environment's hierarchical recognition space dataset to be optimized is imported into a holographic sand table display device equipped with a wearable 3D motion capture system and a large display screen with a resolution of 8K or higher. The distribution of commercial streets, residential streets, and industrial streets belonging to the central area, transition area, and edge area of the target built environment to be optimized is displayed, and human-computer virtual interaction is conducted and feedback information is recorded; The information of streets that are not considered to be spaces to be optimized in the feedback is fed back to the corresponding hierarchical case database, and the recognition results of the spaces to be optimized in the corresponding hierarchical level are updated.
[0049] Example
[0050] The technical solution of the present invention will be described in detail below using Shanghai as an example, as shown in Figures 2 and 3:
[0051] (1) Built environment data collection and preprocessing, specifically including:
[0052] (1.1) Taking the Lujiazui Financial District of Shanghai as the core data collection area, a UAV equipped with a lidar system (with an accuracy of 50mm) and a multispectral camera was used to acquire three-dimensional morphological data covering an area of 12.8 square kilometers. Among them, the building morphological data includes building height, density, and land area; the green space morphological data includes vegetation type, quantity, green space area, and green space height; and the road morphological data includes road length, land area, and road grade.
[0053] (1.2) Data cleaning was performed using the CloudCompare platform to remove abnormal point cloud data caused by signal reflection (approximately 230 million points were cleaned per day). The cleaned data was then imported into the CityEngine platform for voxelization processing; the voxel volume was set to 1 cubic decimeter (1000×1000×1000mm). 3 The voxel network was constructed using an octree data structure. Each voxel attribute field contains three-dimensional morphological information such as building height (0.1m accuracy), vegetation canopy thickness (0.05m accuracy), and road level (coded according to the municipal road classification standard). Finally, a voxel network dataset containing 128 million voxels was generated for the Lujiazui Financial District.
[0054] (2) Preliminary classification of the built environment morphology, specifically including:
[0055] (2.1) Obtain the boundary vector data (coordinate system: Shanghai local coordinate system 2000) of 36 blocks in Lujiazui Financial District from the official website of Shanghai Municipal Planning and Natural Resources Bureau. Establish a spatial index in ArcGIS Pro platform and extract the corresponding voxel network data of each block. Define three federated learning subtasks: central area (building density ≥ 60%), transition area (30% ≤ building density < 60%), and edge area (building density < 30%). Configure NVIDIA A100 GPU nodes for local training in each subtask.
[0056] (2.2) In the multi-task Bayesian federated learning framework, the global iteration count was set to 60 times and the local training epoch count was set to 50 times. The Matérn 3 / 2 covariance function was used to construct a multidimensional covariance matrix. The building height variation index was calculated by the standard deviation within the block (σ≥15m is high variation), and the green space plant morphology richness was evaluated by the Shannon index (H'≥2.5 is high richness). The dynamic weight adjustment module set the initial weights as follows: building voxels 0.6, road voxels 0.3, and green space voxels 0.1. During the training process, the weights were automatically optimized according to the gradient descent direction. Finally, the building weight in the central area (such as Yincheng Middle Road block) was increased to 0.82, and the green space weight in the peripheral area (such as Binjiang Forest Park block) was increased to 0.68.
[0057] (3) Revise the morphological classification of the built environment, specifically including:
[0058] (3.1) Spatial topology verification was performed on the preliminary classification results. A Delaunay triangulation was established in the QGIS platform to analyze the spatial adjacency relationships of the blocks. It was found that the area of the three transition zone blocks on the north side of Pudong Avenue was insufficient (0.78 square kilometers), and the correction procedure was initiated. The three-dimensional morphological indicators of the target block sets were extracted: sky openness (SVF = 0.62), building density (48%), and road intersection density (12 / km). 2 The feature matching was performed with the adjacent central area (average SVF = 0.31, building density 68%) and edge area (SVF = 0.75, building density 22%).
[0059] (3.2) The cosine similarity algorithm was used to calculate the similarity of feature vectors. The similarity score between the target street set and the central area was 0.72, and the score between the target street set and the edge area was 0.35. According to the principle of maximum similarity, the street set was corrected to a transition area. After correction, the Lujiazui Financial District finally formed a three-level morphological classification system of 8.2 square kilometers of central area (including the core area of Xiaolujiazui), 3.1 square kilometers of transition area (including Zhuyuan Commercial Area), and 1.5 square kilometers of edge area (including Binjiang Ecological Area).
[0060] (4) Preliminary identification of the space to be optimized, specifically including:
[0061] (4.1) Use a near-Earth satellite with a storage capacity of more than 100T to carry a multispectral camera and a lidar system with an accuracy of more than 10m to collect nighttime light data of the built environment of the target in Shanghai within one natural month;
[0062] (4.2) Based on the land use function attributes in the Shanghai Urban Land Use Planning Map issued by the Shanghai Municipal Planning Department in accordance with the latest national "Classification of Urban Land Use and Standards for Planning and Construction Land Use", commercial blocks, residential blocks and industrial blocks in the target built environment in the central area, transition area and edge area are extracted respectively, and the corresponding average nighttime light data is matched according to the coordinates;
[0063] (4.3) For blocks in the same tier and with the same land use function, each tier and function block is divided into three intervals based on the average nighttime light data. The division of blocks in different tiers and functions is carried out independently. For the average nighttime light data collected from each tier and function block, outliers are removed and the data is normalized. Rules are set for model training, with the average nighttime light data of blocks of the same function type in the central area, transition area, and edge area decreasing in that order. For commercial land in Shanghai, nighttime light data from 6 pm to 10 pm is collected. For residential land in Shanghai, nighttime light data from 6 pm to 12 am is collected. For industrial land in Shanghai, nighttime light data from 6 pm to 6 am is collected.
[0064] (4.4) Train the XGBoost model using the preprocessed dataset to learn how to map blocks to three nighttime activity intervals: high, medium, and low. After the model training is completed, apply it to the Shanghai test dataset to verify the classification accuracy. Extract the average nighttime light data, the corresponding grade, and the land use function attribute of the blocks to be classified in Shanghai, input them into the XGBoost model, and output the nighttime activity interval to which the block belongs.
[0065] (5) Spatial hierarchical recognition to be optimized, specifically including:
[0066] (5.1) For the initial space to be optimized in the central area, use a drone equipped with a multispectral camera to collect image data of the corresponding street blocks, and perform a surround acquisition every 10 meters in height;
[0067] (5.2) Construct a central area case database, which contains image data before and after optimization of no less than 500 cases of street appearance optimization in the central area, and perform feature labeling on the image data before and after optimization;
[0068] (5.3) For the blocks in the preliminary central area to be optimized, image data preprocessing is performed. The image is divided into 10*10 pixel blocks using Vision Transformer. The blocks are linearly embedded in the high-dimensional space to form a set of feature vectors. These feature vectors are input into the Transformer encoder for information interaction and feature fusion to obtain the image feature codes of the blocks. The similarity between the feature codes is calculated to determine the case image data that is most similar to the block to be judged. If the case image data belongs to the data before optimization, the block to be judged is determined to belong to the space to be optimized. If the case image data belongs to the data after optimization, the block to be judged is determined not to belong to the space to be optimized.
[0069] (6) Interaction and feedback of recognition results
[0070] (6.1) The output of the hierarchical identification of the spatial dataset to be optimized of the target built environment in Shanghai is imported into a holographic sand table display device equipped with a wearable 3D motion capture system and a large display screen with a resolution of 8K or higher. The distribution of commercial blocks, residential blocks and industrial blocks in Shanghai belonging to the central area, transition area and edge area is displayed, and human-computer virtual interaction is carried out and feedback information is recorded.
[0071] (6.2) Feed back the information of the blocks that are not considered to be spaces to be optimized in the feedback to the case database of the corresponding level, and update the identification results of the spaces to be optimized of the corresponding level.
[0072] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.
Claims
1. A method for identifying the built environment space to be optimized based on a morphological hierarchical model, characterized in that, Includes the following steps: Step 1: Construction Environment Data Collection and Preprocessing Using drones with storage capacity of 5T or more, equipped with multispectral cameras and LiDAR systems with an accuracy of 100mm or more, three-dimensional morphological data of the built environment of the target are collected and identified, specifically including building morphology data, green space morphology data, and road morphology data. The data is cleaned to remove outliers. The data is standardized into a voxel network, with a voxel volume of 1 cubic decimeter. Each voxel stores the three-dimensional morphological information of the building, green space, or road at the corresponding spatial location, generating a voxel network dataset. Step Two: Preliminary Classification of Built Environment Morphology Collect street boundary vector data of the target built environment and obtain a voxel network dataset within its coordinate range; based on the voxel network dataset, use multi-task Bayesian federated learning to divide the street boundary vector data of the target built environment into a central area, a transition area and an edge area. Specifically, three sub-tasks are first defined: identifying the central area, transition area, and edge area of the target built environment; on local devices, the three sub-tasks are jointly modeled through a multi-output Gaussian process, the training set is used to optimize the model parameters, a prior distribution is introduced, and the three-dimensional combination features of building voxels, green space voxels, and road voxels in the voxel network data are captured. Then, the posterior distributions on different devices are uploaded to the global processor for aggregation to form a global MOGP prior, which is then distributed back to each local device for the next round of training. After no less than 50 iterations, the model converges, and the street boundary vector data of the target built environment is classified into the central area, transition area, and edge area. Step 3: Revise the morphological gradation of the built environment Based on the hierarchical results of the street boundary vector data of the target built environment, connect the streets belonging to the same partition and located in the spatial neighborhood to form several street sets; if the total area of the streets in the street set is not less than 1 square kilometer, the verification is passed. For areas smaller than 1 square kilometer, the classification of the area will be revised. During the correction process, if all adjacent block sets of this block set belong to the same partition, then this block set is also assigned to this partition; if the blocks surrounding this block set belong to different partitions, the target block set is compared to which adjacent block set has the most similar three-dimensional combination features of the voxel network, and this block set is assigned to the partition to which the most similar block set belongs. Step 4: Preliminary identification of areas to be optimized Using a near-Earth satellite with over 100T of storage space, equipped with a multispectral camera and a lidar system with an accuracy of over 10m, nighttime light data of the target built environment was collected over one calendar month. Based on land use attributes, commercial, residential, and industrial blocks within the target built environment of the central, transition, and edge zones were extracted, and their corresponding nighttime light data averages were matched according to coordinates. For blocks at the same tier and with the same land use attribute, the XGBoost model was used to divide each tier and function block into three intervals based on the average nighttime light data. Blocks in the lowest interval were identified as the space to be optimized. Based on land use status information, blocks with a land use status of "under construction" or "completed but not yet operational" were removed, resulting in a preliminary dataset of the commercial, residential, and industrial blocks to be optimized within the central, transition, and edge zones. Step 5: Spatial hierarchy recognition to be optimized For the initial optimization space in the central area, a drone equipped with a multispectral camera was used to collect image data of the corresponding coordinate blocks. A circling acquisition was performed every 10 meters in height. ViT was used, based on the central area case database, to determine whether the streetscape of various blocks in the central area needed optimization. For the initial optimization space in the transition area, a street view acquisition vehicle equipped with a 360° panoramic camera was used to collect streetscape image data of the corresponding coordinate blocks. CNN was used, based on the transition area case database, to determine whether the streetscape of various blocks in the transition area needed optimization. For the initial optimization space in the edge area, satellite image data of the corresponding coordinate blocks was acquired via near-Earth satellites. Stereo image pairs were used to capture images of the target at 60° intervals to obtain overlapping stereo images. Matching and calculation were performed to extract the 3D information of the graphic elements of the target area. CNN was used, based on the edge area case database, to determine whether the appearance of various blocks in the edge area needed optimization. This resulted in a hierarchical recognition dataset of the optimization spaces for commercial, residential, and industrial blocks within the central, transition, and edge areas. Step Six: Interaction and Feedback on Recognition Results The target built environment's hierarchical recognition dataset of spaces to be optimized, output from step five, is imported into a holographic sand table display device equipped with a wearable 3D motion capture system and a large display screen with a resolution of 8K or higher. This device displays the distribution of commercial, residential, and industrial blocks belonging to the central, transition, and edge areas of the target built environment, allowing for virtual human-computer interaction and recording feedback information. Information on blocks deemed not to be spaces to be optimized is fed back to the corresponding hierarchical case database, updating the corresponding hierarchical recognition results of spaces to be optimized.
2. The method for identifying the built environment space to be optimized based on a morphological hierarchical model according to claim 1, characterized in that, Step two involves capturing the three-dimensional combined features of building voxels, green space voxels, and road voxels in the voxel network data. Specifically, this is achieved by employing three sub-tasks: multi-output Gaussian process joint modeling, to capture the three-dimensional combined features. First, based on the voxel network dataset, a prior distribution is defined, embedding the spatial distribution patterns of building, green space, and road voxels into the covariance function. This function quantifies the spatial correlation of different voxel types, covering average building height, building height variation, green space coverage, green space height variation, green space plant morphology richness, road network density, and road network connectivity, as shown in the table below. This allows for the construction of a multidimensional covariance matrix, reflecting the statistical regularities of the three-dimensional combination characteristics; During the local device training phase, when optimizing the MOGP model parameters using the training set, a class weight factor is introduced to assign dynamic weights to building, green space and road voxels respectively. By maximizing the marginal likelihood function, the weights are adaptively adjusted to strengthen the dominant features. The building and road voxel densities are higher in the central area, and their weights account for a larger proportion of the covariance calculation, decreasing sequentially in the transition and edge areas. The green space voxel density is higher in the edge area, and its weights account for a larger proportion of the covariance calculation, decreasing sequentially in the edge and central areas. During training, the posterior distribution of each subtask is updated through variational inference, and the parameters of each device are fused using a weighted average method to generate a global prior distribution. After at least 50 iterations, once the model converges, the global prior is used to characterize the differences in the three-dimensional combined features of different zones.
3. The method for identifying the built environment space to be optimized based on a morphological hierarchical model according to claim 2, characterized in that, In step three, the target block set is compared to determine which adjacent block set has the closest three-dimensional combination features of the voxel network. Specifically, the following indicators of the voxel network in this block set are extracted: three-dimensional fractal dimension, sky openness, plot ratio, building density, green patch density, and road intersection density, as shown in the table below. A feature matching algorithm is used to compare the extracted three-dimensional combined features with the features of adjacent street sets one by one to obtain the similarity score between each adjacent street set and the target street set; the tier type to which the street set with the highest score belongs is selected as the tier type to which the target street set belongs.
4. The method for identifying the built environment space to be optimized based on a morphological hierarchical model according to claim 3, characterized in that, In step four, each functional block of each level is divided into three intervals based on the average nighttime light data. This is done independently by dividing the blocks of different levels and functions. The average nighttime light data of each functional block of each level is normalized after removing outliers. Rules are set for model training, with the average nighttime light data of blocks of the same functional type decreasing sequentially in the central area, transition area, and edge area. Nighttime light data from 6 PM to 10 PM is collected for commercial land, from 6 PM to midnight for residential land, and from 6 PM to 6 AM for industrial land. The XGBoost model is trained using the preprocessed dataset to learn how to map blocks to three nighttime activity intervals: high, medium, and low. Once the model training is complete, it is applied to the test dataset to verify the classification accuracy. The average nighttime light data, the classification level, and the land use function attributes of the blocks to be classified are extracted and input into the XGBoost model to output the nighttime activity range to which the block belongs.
5. The method for identifying the built environment space to be optimized based on a morphological hierarchical model according to claim 4, characterized in that, In step five, it is determined whether the streetscape of various blocks in the central area needs to be optimized. Specifically, this is done by constructing a case database of the central area, which contains image data before and after optimization of no less than 500 cases of streetscape optimization of blocks in the central area; image data preprocessing is performed on the blocks in the central area to be optimized, and the image is divided into 10*10 pixel blocks using Vision Transformer. The blocks are linearly embedded in a high-dimensional space to form a set of feature vectors, which are input into the Transformer encoder for information interaction and feature fusion to obtain the image feature encoding of the blocks. Calculate the similarity between feature codes, determine the case image data that is most similar to the block to be determined. If the case image data belongs to the data before optimization, the block to be determined is determined to belong to the space to be optimized; if the case image data belongs to the data after optimization, the block to be determined is determined not to belong to the space to be optimized.