A method, system and medium for automatically extracting typical patterns of residential space form
By generating a unified outer contour through morphological closing and opening operations, and combining multi-scale site extraction and subspace clustering, the problems of boundary inconsistency and multi-scale site identification in the quantitative analysis of residential spatial morphology are solved, and efficient generation of simulation templates for wind environment and pollution diffusion is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA UNIV OF GEOSCIENCES (WUHAN)
- Filing Date
- 2026-01-27
- Publication Date
- 2026-06-05
AI Technical Summary
Existing technologies struggle to uniformly define boundaries in the quantitative analysis of residential spatial morphology, adapt to the spatial characteristics of multi-scale locations, and lack unified statistics on the influence of multiple wind directions and height layers, resulting in false positives, false negatives, and insufficient interpretability of clustering, making it difficult to output parameterized templates that can be directly used for wind environment or pollution diffusion simulation.
A unified outer contour is generated by morphological closing and opening operations. Combined with multi-scale site extraction and subspace clustering, the windward interface is identified by multi-wind direction rotation and height stratification. A parameterized geometric template is constructed, and a morphological template that can be directly used for simulation is output.
It achieves a unified boundary definition for residential spatial morphology, improves the robustness of site spatial identification and the interpretability of results, reduces the simulation computing resource requirements, and enhances the efficiency of wind environment and pollution diffusion analysis.
Smart Images

Figure CN122156690A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of urban and regional information engineering technology, specifically involving an automatic extraction method, system and medium for typical patterns of residential spatial morphology. It is based on mathematical morphology and subspace clustering for residential spatial morphology modeling, wind environment or pollution diffusion analysis and simulation case library construction, which belongs to the intersection of urban microclimate assessment and data-driven planning auxiliary decision-making. Background Technology
[0002] With the acceleration of urbanization, the ventilation efficiency and air pollutant dispersion capacity of high-density built-up areas have received increasing attention. The spatial morphology of building complexes, including their size, layout, orientation, and height distribution, is a key factor influencing near-surface wind environment and pollution dispersion processes. To quantify the wind-blocking effect of building complexes, existing studies typically use morphological indicators such as the frontal area index and apply them to urban ventilation corridor identification and microclimate optimization. However, most of these methods are based on approximations of fixed wind direction and single height layers, making it difficult to accurately reflect the characteristics of the windward interface corresponding to variable wind directions and different height layers (such as pedestrian level, middle and upper floors of buildings) in real urban environments. Their applicability and accuracy are limited at the scale of residential areas, a typical urban functional unit.
[0003] In the quantitative analysis of residential spatial morphology, the following technical challenges currently exist: First, the lack of a unified and objective standard for defining the external outline of residential areas leads to inconsistent statistical methods for indicators such as interface density and line coverage among different cases, resulting in poor comparability. Second, the morphology and scale of open spaces within residential areas (such as courtyards, plazas, and alleyway nodes) vary significantly, making it easy to cause false positives or false negatives when using single-scale spatial filtering methods. Furthermore, most existing studies employ principal component analysis combined with global dimensionality reduction and clustering strategies such as K-means clustering. While these methods can identify patterns in the overall dimension, they struggle to identify localized and interpretable typical patterns in subspaces composed of certain dimensions. This makes it difficult to convert the results into parametric design templates or simulation examples, limiting their direct reuse in planning practice.
[0004] Therefore, there is an urgent need in this field for a method that can automatically and systematically extract typical patterns of residential spatial morphology. This method should be able to unify the definition of residential boundaries, adapt to the spatial characteristics of multi-scale sites, comprehensively consider the influence of multi-wind direction and building height stratification, and finally output a parametric morphological template with clear physical meaning that can directly support the selection of wind environment or pollution diffusion simulation. Summary of the Invention
[0005] This invention addresses the problems in residential-scale wind environment / pollution diffusion analysis, such as inconsistent boundary calibers, false positives and false negatives due to differences in spatial scale, lack of unified statistics for multi-wind-direction and multi-height layers, and insufficient interpretability of high-dimensional index clustering. It proposes an automatic extraction method, system, and medium for typical patterns of residential spatial morphology, realizing an integrated process of "unified boundary - multi-scale sites - layered windward interfaces - subspace clustering", and outputting parameterized templates that can be directly used for simulation selection.
[0006] An automatic extraction method for typical patterns of residential spatial morphology, comprising: Step S1: Obtain building information data including residential area boundaries, building outlines and number of floors, set the raster resolution Δ, and generate a building binary map I, a residential area space raster U and a floor number raster H; Step S2: Perform morphological closing operation on the binary architectural image I and fill the holes by incrementally increasing the structural element radius r, where r is incremented by steps. Increment, determine the enclosed area The minimum connected radius r* is used to achieve full connectivity for the first time, and will As the unified outer outline of the residential area; in: In the formula, I represents a binary raster image of a building; A circular or approximately circular structuring element with radius r; and These represent morphological dilation and erosion operations, respectively. This is an operator for filling holes in a binary region.
[0007] Step S3: Perform a morphological opening operation on the residential space grid U according to the preset set of radius scales S of structural elements, and extract candidate open space regions with different radii of structural elements. Each After expansion, the intersection with the binary building graph I yields the set of building indices that enclose the site. The consistency score and overall stability score of the enclosed building combination around the same open space candidate area identified under different structural element radii are evaluated. Step S4: Set the direction set Θ, rotate the floor number grid H, and divide it into at least three layers according to height; scan each layer column by column along the windward orthogonal direction, identify the windward leading edge pixels, and calculate the windward interface density. Line bonding rate and gap spacing set ; in: In the formula, This represents the proportion of the number of columns of the leading edge pixels on the windward side of the k-th slice facing the wind direction θ out of the total number of columns, used to characterize the interface coverage; θ is the scanning or incoming flow direction; k is the height slice index; θ represents the grid slice of the k-th layer in the direction θ; W represents the number of vertical columns of slices set along the direction perpendicular to the wind direction. This is an indicator function (it takes a value of 1 when a column has leading edge cells, and 0 otherwise); This represents the proportion by which the uniform outer contour of the residential area is "attached" to the windward side facing the direction of the wind, θ, by the slice of the k-th floor of the building. The total length of the line segment that coincides with the leading edge pixel of the building's k-th floor slice on the windward side facing the wind direction θ on the unified outer contour of the residential area. To ensure a uniform overall perimeter of the outer contour; is the pixel length of the j-th consecutive zero-value gap; M is the number of gap segments; Δ is the raster resolution; This represents the actual width of the j-th gap.
[0008] Step S5: Construct a morphological index system that includes individual buildings, enclosures, locations, and windward interfaces. Divide the indexes into multiple index groups with morphological significance, and discretize and pre-cluster the indexes in each index group to obtain categorical variables. Step S6: Incorporate the categorical variables and continuous variables into a multidimensional attribute space for grid partitioning, identify dense cells based on grid cell density, and merge adjacent dense cells using the minimum description length criterion to output a typical pattern set and its minimum feature description.
[0009] Step S7: For each of the typical patterns, generate a parametric geometric template that includes a uniform outer contour, dominant building type, windward interface parameters, and site scale level.
[0010] Further, in step S2, the minimum connectivity radius The method of determination is as follows: In the formula, Indicates the region The number of connected components; and fill holes with an area smaller than a set threshold.
[0011] Furthermore, in step S2, the area threshold is the actual area corresponding to 50 to 500 pixels.
[0012] Furthermore, in step S2, the step Δr is taken as an integer multiple of the grid resolution, and the optimal solution among the parallel solutions of r* with the smallest outer contour perimeter is taken as the standard.
[0013] Furthermore, in step S3, the scale set S covers from the smallest scale... To the maximum scale The scope; the enclosed building index set By open space candidate areas The result is obtained by intersecting the expanded binary graph with the building's binary graph I.
[0014] Furthermore, in step S4, the angular step of the direction set Θ is 45° or 30°; the height stratification includes a low-rise area, a mid-rise area, and a high-rise area, which are divided based on the number of building floors or based on the standard floor height. The converted absolute height.
[0015] Furthermore, in step S4, a minimum threshold is set for the length of the continuous zero-value gap. To filter out the tiny gaps caused by noise.
[0016] Furthermore, in step S5, the index group includes at least the building type proportion group, the enclosure and depth group, the site space group, the windward interface group, and the orientation combination group; before pre-clustering, each index in the group is discretized by equal frequency or entropy basis, and converted into 3 to 9 levels.
[0017] Furthermore, in step S6, the cost function for merging based on the minimum description length criterion is expressed as: In the formula, R represents the candidate high-dimensional dense region; For a structured representation of region R; The bit length required to encode the structure itself; The bit length required to encode the data residual under given structural conditions; The total description length of region R.
[0018] An automatic extraction system for typical patterns of residential spatial morphology, characterized in that it includes a processor and a memory, wherein the memory stores a computer program, and the computer program, when executed by the processor, implements any of the methods described above.
[0019] A computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements any one of the methods described above.
[0020] An electronic device includes a processor, a memory, and a communication interface. The memory stores a program executable on the processor, which, when executed, causes the electronic device to perform the method described above. The electronic device is further configured to... Its parameter template is provided to the external simulation platform through the communication interface.
[0021] Compared with the prior art, the beneficial effects of the present invention are: 1. By adaptively determining the minimum connectivity radius, a unified and continuous residential area outline is generated, which solves the problem of incomparability of morphological indicators (such as line coverage rate and interface density) caused by different boundary definition standards, and provides a standardized basis for batch analysis and evaluation across plots and data sources.
[0022] 2. By employing multi-scale morphological opening operations, the system extracts open spaces of different scales, from alleyway nodes to central courtyards. Combined with the enclosed building backscan mechanism, it effectively reduces false detections and false negatives caused by single-scale filtering, and significantly improves the robustness of space recognition.
[0023] 3. By introducing multi-wind-direction rotation and three-segment height-layer slicing, the windward leading edge on different wind directions and height layers is automatically identified, and indicators with clear wind environment physical significance, such as interface density, line coverage rate and gap spacing distribution, are output, so that morphological characteristics are directly related to the analysis process of ventilation, pollutant diffusion, etc.
[0024] 4. By adopting a two-level strategy of "index group pre-clustering + CLIQUE subspace clustering", typical patterns can be discovered on local subsets of high-dimensional index space, and the minimum feature description expressed by the "dimension-value interval" rule is output. The results are intuitive and interpretable, and can be easily transformed into parametric geometric templates or design guidelines.
[0025] 5. By generating geometric templates that bind key morphological parameters for each typical pattern and directly connecting them to CFD or pollution diffusion simulation platforms, the number of simulation samples can be significantly reduced, the computational resource requirements can be lowered, and batch and rapid evaluation of residential morphology-wind field analysis can be achieved.
[0026] Significant differences from existing technologies include: Compared to the overall clustering approach based on PCA dimensionality reduction and K-means, this invention performs density discrimination and rule extraction on dimensional subsets, enabling the output of executable interval rules that correspond one-to-one with simulation templates; Compared to the calculation of frontal area index for single wind direction or single layer, this invention provides a systematic and automatic identification of multi-wind direction and layered interfaces, and calculates the line coverage rate and gap statistics with a unified outer contour diameter; Compared to empirical site extraction, this invention establishes a mapping of "site-building combination" through multi-scale opening operations and backscanning, improving the interpretability and reusability of the results. Attached Figure Description
[0027] Figure 1 Treatment of expansion corrosion for different building outlines; Figure 2 Skeletalization processing of building outline data; Figure 3 Classify the spatial form of some building outlines; Figure 4 Examples of multi-scale residential site extraction results; Figure 5 This is a pre-clustering of the building orientation combination type index group. Detailed Implementation
[0028] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. These embodiments are implemented based on the technical solution of the present invention, and provide detailed implementation methods and processes. However, the scope of protection of the present invention is not limited to the following embodiments.
[0029] Example 1 The method of the present invention consists of steps such as data preparation, unified outer contour generation, open space extraction and enclosure back scan, multi-wind direction layered windward interface identification, index group pre-clustering, subspace clustering and minimum feature description derivation, parameterized template generation and simulation docking, forming a complete processing flow from raw residential area data to "interpretable pattern → parameter template".
[0030] Step S1: Data Acquisition and Rasterization This step is the data preparation phase of the entire method, which aims to convert the raw vector geographic data into raster data suitable for morphological analysis.
[0031] First, obtain building information data including residential area boundaries, building outlines, and number of floors. This data primarily consists of building outline vectors within the residential area or plot, and its attribute fields should include the number of floors or height information. If only height 'h' is available, then a given standard floor height should be used. (Usually taking 3 meters) the number of floors Converted to If there is no explicit boundary for the residential area, the initial outline can be the plot boundary, the road-enclosed area, or the land use control line, and then corrected for uniformity by calculating the minimum connectivity radius r*. Based on this, a raster resolution Δ=1m is set, and a binary building map I (building area is 1, non-building area is 0) is generated according to a unified coordinate system based on the building outline; a residential space raster U (accessible space is 1, obstacles are 0) is generated based on the residential area boundary; and a floor count raster H (each cell stores the height of the floors or equivalent floors) is generated based on the number of building floors. Simultaneously, nearest neighbor interpolation is used for buildings with missing floor counts, and abnormal heights are identified and backfilled with the corresponding quantiles using the IQR method to ensure the stability of subsequent statistics. The binary building map I, the residential space raster U, and the floor count raster H are equivalent to three layers, displaying the building outline, the residential area boundary, and the number of building floors, respectively.
[0032] Step S2: Generate a uniform outer contour using adaptive closing operation Perform morphological closing operations on the binary architectural image I, along with a hole filling operation; In the formula, I represents a binary raster image of a building; A circular or approximately circular structuring element with radius r; where radius r is incremented by steps. Increasing, Take an integer multiple of the raster resolution (in this embodiment) To generate enclosed regions at different scales r : and These represent morphological dilation and erosion operations, respectively. An operator for filling holes in a binary region; This represents the enclosed region obtained after closure operations and hole filling at scale r.
[0033] During the hole-filling process (a hole is a background area (0 value) completely surrounded by foreground pixels, i.e., an area completely enclosed by a building; replacing such a hole from 0 to 1 is hole filling), for areas smaller than a set threshold... All holes are filled, and the hole filling area threshold is met. Preferably, it corresponds to the actual area of [50, 500] pixels to effectively filter out interference caused by small gaps (such as tiny holes) in the built-up area.
[0034] Furthermore, according to each scale The minimum connected radius r* is determined by the number of connected components, where a connected component is a set of pixels that can be connected to each other through adjacent pixels according to the 8-neighborhood rule. The number of such a set is denoted as the number of connected components. In the formula, Indicates the region The number of connected components; In order to make The minimum radius for achieving overall connectivity for the first time (i.e., performing "closing operation + hole filling" on the binary image of the building, with the structural element radius r increasing incrementally, and the minimum r when the resulting "enclosed area" first becomes overall connectivity (number of connected components = 1) is r*, which determines the outer contour of the building / residential area. The scale of this outer contour is used for the calculation of the line-to-line ratio, interface density, etc. in the subsequent statistics of the windward interface).
[0035] Step S3: Automatically identify candidate open space areas within the residential area boundary, establish a structured topological relationship between the same candidate open space area and its surrounding enclosed building combination, and evaluate the stability of this relationship through multi-scale analysis; S31. Perform a morphological opening operation on the residential space grid U according to the preset set of radius scales S of the structural elements to extract candidate open space regions corresponding to the radii of different structural elements in the residential space grid U. : Wherein, U represents the residential space grid; For radius Circular structural elements; Let be the set of radius scales of the structuring elements used in the open operation, covering from the smallest scale. To the maximum scale The range, of which , (However, in this embodiment, we take...) The opening operation can filter out narrow channels with a width less than s in the residential space grid, and retain only candidate areas of connected open spaces with a width not less than s. This represents the candidate open space area extracted by open operation under the radius scale s of the structural element, which can effectively highlight plazas, atriums and other nodal spaces with a width greater than s.
[0036] S32. Subsequently, for each Perform morphological reconstruction and execute a dilation operation. The boundary of the candidate open space is extended outward by a structural element radius *s* to cover the buildings immediately adjacent to its edge. The result is then intersected with the binary building graph *I* to obtain the set of enclosed building combination indices that enclose the candidate open space. : In the formula, This indicates finding the intersection, meaning that both the pixels before and after the operator are foreground pixels.
[0037] By reversing the location of buildings surrounding the open space candidate area from the open space candidate area, a topological relationship between the open space candidate area and its enclosing building combination is established.
[0038] S33. Evaluate the consistency score of the surrounding building combination of the same open space candidate area identified under different structural element radii; This includes: further adjustments within the set S of structuring element radius scales. , Quantify the similarity between any two scales. , ( < , Jaccard similarity of the set of building indices under the largest radius in the set of structural element radius scales S, and defined as a stability score: In the formula, , These are the structuring element radius scales. and A set of indexes for enclosed building combinations in candidate areas of enclosed open spaces; This represents the number of buildings in the intersection of two sets of enclosed building combination indexes; This represents the number of all unique buildings in the index set of two enclosed building combinations. The number of elements in the set; This indicates the proportion of enclosed building combinations in the enclosed building combinations at the two structural element radius scales. The larger the proportion, the more consistent the enclosed buildings identified by the two structural element radius scales are.
[0039] S34. Furthermore, for all radii scales S of the structuring elements... < The average of the combinations yields the overall stability score. : In the formula, The number of radii of structuring elements in the set of radii of structuring elements, S; Indicates all Summation of combinations; The overall stability score is given for the enclosed building combination surrounding the candidate area of the same open space at the radial scale of multiple structural elements, and It can serve as auxiliary information for pattern selection and interpretation, and provide a basis for subsequent spatial pattern interpretation and morphological stability analysis.
[0040] Step S4: Identification of the windward interface based on direction deflection and height stratification Set the direction set Θ, with an angular step of 45° (i.e., 8 wind directions). To obtain more refined analysis results, the refinement can be set to 30° (12 wind directions). The angle is formed by rotating clockwise from due north to the direction of the incoming wind.
[0041] Using due north as the wind direction, rotate the floor grid H counterclockwise around the image center by an angle. ,get The floor grid H is divided into three slices based on floor thresholds (i.e., common low-rise / multi-story buildings, multi-story residential buildings, and mid-rise residential buildings): (1-2 floors) (3-6 floors) and (7 floors and above) or at a given floor height When the absolute height range is [0,2], ]、[2 6 ]、[>6 Slice the material.
[0042] After rotation In the middle, each layer is sliced along a direction perpendicular to the direction of the incoming wind. Divide the slice into W columns and scan each column sequentially from the windward side. Record the leading edge pixels of each slice (in a rotating slice, when scanning column by column along the windward direction, the first building pixel encountered by each column from the windward side is recorded as the leading edge pixel), and calculate the windward interface density. Line bonding rate and gap spacing set ,in: In the formula, k is the slice height layer index. Let W be the grid slice of the k-th layer along the direction of the incoming wind θ, and let W be the number of vertical columns of slices set along the direction perpendicular to the direction of the incoming wind. This is an indicator function (it takes a value of 1 when a column has leading edge cells, and 0 otherwise). This represents the proportion of the number of columns of the leading edge pixels on the windward side of the k-th slice facing the wind direction θ out of the total number of columns, used to characterize the degree of interface coverage.
[0043] Line bonding rate Defined as: In the formula, The total length of the line segment that coincides with the leading edge pixel of the building's k-th floor slice on the windward side facing the direction θ, on the unified outer contour of the residential area. The total perimeter of the unified outer outline of the residential area, It represents the proportion by which the uniform outer contour of the residential area is "attached" to the windward side facing the direction θ of the building, where the building's k-th floor slice is sliced.
[0044] Notch spacing set Its actual width is expressed as: In the formula, Let this be the set of pixels of length denoted as notches (notches refer to pixels with a value of 0 (no leading edge pixels in the column / the segment of the interface is not covered by the building) identified on the windward side of the k-th floor slice of the building facing the wind direction θ) in a one-dimensional scan sequence (by column) of the windward interface or its corresponding interface line grid. Let M be the pixel length of the j-th consecutive zero-value gap, M be the number of gap segments, and Δ be the raster resolution. Let be the actual width of the j-th segment gap, where a zero-value gap refers to a continuous interval formed by the aforementioned zero-value gap pixels in the sequence, and its length is denoted as the pixel length of the gap. Multiplying this length by Δ yields the actual width. To suppress false detections, a minimum threshold is set for the length of continuous zero-value gaps. To filter out tiny gaps caused by noise, a 3×1 sliding window majority vote can be used to smooth out leading edge jitter, thereby obtaining a more stable windward interface line.
[0045] Step S5: Construction of Morphological Indicator System and Pre-clustering of Indicator Groups First, based on the building's binary map and skeleton information, basic morphological indicators for individual buildings are extracted, including base area, major and minor axes of the circumscribed rectangle, compactness, eccentricity, aspect ratio, convex hull ratio, and the number of skeleton bifurcations and endpoints. These indicators are calculated using a binary region attribute algorithm. Compactness measures the degree to which the shape approximates a circle, and is generally expressed using... Calculation, where Represents the area of a shape; This represents the perimeter of the shape; for the major and minor axes a and b of the fitted ellipse, the eccentricity is calculated using the formula: Where 'a' represents the length of the major axis of the smallest bounding rectangle (half the length along the principal axis), and 'b' represents the length of the minor axis of the smallest bounding rectangle (half the length perpendicular to the principal axis); the convex hull ratio is the ratio of the area of the actual shape of the object to the area of its smallest convex hull; the skeleton bifurcation number is obtained by skeletonizing the building binary image to S+, and counting the 8-neighborhood connections N of the skeleton pixels: if N≥3, it is a bifurcation point; if N=1, it is an endpoint. The total number of bifurcation points or its density is "skeleton bifurcation number / bifurcation density". The building skeleton binary image S... + Discrete convolution with a 3×3 sliding window kernel K is used to extract bifurcation and endpoint information: In the formula, S + The image represents a binary raster image of the building's framework; K is a 3×3 sliding window kernel used to count 8-neighborhood connections; * indicates discrete convolution; N is the convolution result, in S + A pixel position with a value of 1 is determined to be a bifurcation point if the neighborhood count N≥3, and an endpoint if N=1. The bifurcation density and endpoint density are then used as supplementary features to be included in the index group.
[0046] Next, macroscopic morphological indicators related to the overall layout of the residential area are calculated. These indicators characterize the degree of concentration or dispersion of building clusters within the residential area, as well as the depth of enclosure. This includes the uniform outer contour generated in step S2. Define the aspect ratio as follows: In the formula, To and The number of building base outlines that intersect or are cut by it; This represents the total number of buildings within the residential area. A higher value indicates that the building layout is more concentrated on the outer perimeter of the residential area; a lower value indicates that the buildings tend to be laid out inwards.
[0047] To more precisely characterize the enclosed attributes of a building complex, "enclosure depth" is further defined as a joint feature. This feature consists of two parts: the proportion of non-outline buildings, i.e. Secondly, all these non-outline buildings are unified into a single outer outline. The statistical distribution of the Euclidean shortest distance (such as mean, standard deviation, quantiles, etc.). This joint feature comprehensively reflects the depth information of the "quantity" and "location" of some buildings within the residential area.
[0048] Furthermore, the orientation combination characteristics of residential buildings are quantified. Rose diagrams are used to discretize the primary orientation angles of all buildings in the residential area using equal frequency methods, generating dominant orientation clusters. The cluster with the most samples is extracted as the majority orientation cluster, and the cluster with the second most samples is extracted as the second majority orientation cluster. The codes of these two orientation clusters are then combined to form a composite categorical variable.
[0049] The 26th annual conference on Building Orientations Using Angle Clustering, presented at GIScience Research UK, explains how to use rose diagrams to discretize the primary orientation angles of all buildings in a residential area at equal frequencies, generating dominant orientation clusters and second-most common orientation clusters, and how to jointly encode them to form categorical variables.
[0050] Subsequently, all morphological indicators were divided into multiple indicator groups, including building type proportion group, enclosure and depth group, site space group, windward interface group, and orientation combination group. The enclosure and depth group includes the depth ratio. The enclosing depth joint feature and building density (this item was calculated and measured in the macroscopic morphological index of step S5) are mentioned, wherein the orientation combination group is the majority and second-largest orientation clusters of the joint encoding as the core feature. Each index within the group is discretized using isofrequency or entropy basis methods, converting it into q levels (pre-clustered discrete levels in this embodiment). Preferably, it is 5), and the similarity between categories is calculated based on Jaccard distance or Hamming distance, and then pre-clustering is performed to obtain category variables with high internal consistency, providing structured feature representation for subsequent morphological pattern recognition.
[0051] Step S6: Subspace Clustering and Minimal Feature Description Using each residential area as a point in the multidimensional attribute space, and each categorical variable and each residual continuous variable as a coordinate axis in the multidimensional attribute space, each coordinate axis is... (In this embodiment, the number of subspace clustering grid divisions) The multidimensional attribute space is divided into a large number of hypercube mesh units g by dividing the mesh into several levels.
[0052] The residential density of each grid cell g is: In the formula, X is the set of residential sample areas to be clustered, and g is a hypercube grid cell in the multidimensional attribute space. δ(g) represents the number of elements in the set, and δ(g) characterizes the population density in the hypercube grid cell g.
[0053] Iterate through all hypercube mesh cells g, marking those that satisfy the condition. (Density threshold in this embodiment) We selected hypercube mesh elements with a density of 0.03 as candidate dense elements.
[0054] When labeling dense cells, the idea of subspace clustering is used to traverse from low dimensions. If a cell is already labeled as a candidate dense cell in the attribute space of a certain dimension, the traversal is gradually extended to higher dimensions. Hypercube grid cells that are still labeled as candidate dense cells in higher-dimensional attribute spaces are retained until all dimensions of attribute subspaces are traversed. Hypercube grid cells that are labeled as candidate dense cells in all dimensions are selected as dense cells. The depth-first search (DFS) algorithm is used to connect spatially adjacent dense cells into a connected region, denoted as a candidate high-dimensional dense region R. In this process, each dense grid cell is regarded as a graph node. Adjacent dense cells (sharing faces / edges / corners, it is recommended to specify the adjacency criteria) are connected by edges. DFS starts from any unvisited dense cell and recursively / stack-wise visits all its reachable neighbors to obtain a connected region (connected component).
[0055] The minimum description length (MDL) cost function is used to discriminate and merge adjacent dense regions. The cost function is expressed as follows: In the formula, R represents the candidate high-dimensional dense region; For a structured representation of the candidate high-dimensional dense region R; The bit length required to encode the structure itself; The bit length required to encode the data residual under given structural conditions; The total description length of the candidate high-dimensional dense region R.
[0056] By comparing each candidate high-dimensional dense region The size of the data is used to prioritize retaining regions with shorter total description lengths, thereby determining the final set of typical patterns. For each pattern, a minimal feature description is constructed, comprising "dimension set - value range / category - sample coverage - representative residential area list", from which a set of regions with a relatively short total description length is output. This means that the cluster can describe more samples with shorter rules while avoiding overfitting.
[0057] Furthermore, for each typical model Generate the corresponding parametric geometry template, which includes the following key elements: uniform outer contour. Polysegment approximation, dominant building type sequence generated by S5, dominant orientation cluster distribution, windward interface parameters (including interface density between each direction and layer). Line bonding rate (including the distribution of gap width), the site scale level generated by S3 (each open site candidate area corresponds to the recognition result under one or more structural element radii s, mapped by threshold / quantile, such as alleyway nodes / courtyards / plazas, etc.) and stability score. This allows for pattern interpretation without dimensionality reduction. Templates are exported in JSON or CSV structured format, enabling direct use by external CFD or pollution diffusion simulation platforms, facilitating one-click configuration for simulation selection.
[0058] For each typical pattern template, parametric geometry (boundaries, building height stratification, windward interface parameters, site scale, etc.) is generated and input into the CFD / pollution diffusion model as a representative "prototype residential area," forming a "pattern-simulation result" library. This alleviates the dimensional disaster problem caused by the extremely complex and diverse spatial forms of residential areas. Furthermore, the paper "Research on the Impact of Exogenous Particulate Matter on Outdoor Spaces of Residential Areas and its Improvement" explains how to use this parametric geometric template.
[0059] Compared to simulating each residential area sample individually, this method achieves significant sample compression through pattern templates. The sample compression rate is defined as: In the formula, The number of templates generated, where each region with a shorter total description length is a template. This represents the total number of buildings within the original residential area. This indicates the sample compression ratio achieved through template processing. This mechanism significantly reduces the resource requirements for simulation calculations and improves the analytical efficiency of residential wind environment and pollution diffusion assessments.
[0060] To verify the effectiveness of the method of this invention, a batch processing case containing 200 buildings in a residential area was conducted using the above parameter settings (Δ=1m, Θ is a 45° step, q=5, m=20, τ=0.03), outputting 12 typical pattern templates. Compared with the traditional scheme based on PCA dimensionality reduction combined with K-means clustering (k=12), the method of this invention can reduce the pattern rule length by about 30%–40% while maintaining a similar number of clusters, improve Bootstrap stability (consistency of multiple samplings) by about 10–20 percentage points, and reduce the size of the simulation case library by about 60%–70%. The above values show that this method has significant improvements in pattern interpretability, stability, and simulation efficiency. It should be noted that the specific effects depend on the actual project data.
[0061] Example 2: Refined Analysis Based on 30° Angle Stepping and LiDAR Height Field Based on Example 1, this example further refines the wind direction resolution and altitude data processing methods, with specific improvements as follows: The angular step of the direction set Θ is refined from 45° to 30° to provide more precise wind direction coverage. Simultaneously, an absolute height field is obtained by fusing a digital surface model (DSM) generated from LiDAR data with a digital terrain model (DTM), replacing the building attribute-based floor grid H in the original embodiment. This is done at a given standard floor height. Under the premise of this, the absolute height field is sliced according to the following intervals to form three corresponding layer grids: the lower layer region [0,2] ], Middle layer [2h0,6 ] and high-rise area [>6 The remaining processing flow in this embodiment is the same as in Embodiment 1.
[0062] To reduce interpolation errors caused by raster rotation, a bicubic interpolation algorithm is used for the height field during orientation rotation. Furthermore, the height threshold is adaptively adjusted before binarization or layering to enhance the algorithm's robustness to different data sources.
[0063] In residential areas with a high proportion of high-rise buildings, the advantages of this embodiment are particularly evident: the differences in interface characteristics between different height sections are more significant, with the windward interface density corresponding to the high-rise section being particularly high. and line bonding rate The changes in these features can more effectively characterize the morphological characteristics of wind-damping corridors formed by high-rise buildings, thus providing a more accurate morphological basis for wind environment analysis in high-density residential areas.
[0064] Example 3: Parallel computing and robust control for cross-region batch processing Building upon Example 1, this example proposes a partitioned parallel and block-based computing architecture to address the batch processing needs of cross-regional, large-scale residential area samples. For each independent plot, steps S1 to S6 are executed sequentially, and the following robustness control strategies are integrated at the plot level to improve the system's robustness and output quality in complex scenarios: (1) Boundary anomaly correction: During the generation of a uniform outer contour, if a minimum connected contour is detected... If the perimeter of a polygon suddenly increases and the proportion of its corresponding concave polygon exceeds a preset threshold, it will automatically revert to the previous feasible value of radius r to avoid unreasonable expansion of the boundary caused by abnormally large-scale closing operations.
[0065] (2) Noise reduction: After the building skeletonization process, pruning of small branches is carried out, and branches with a length less than the threshold are removed. The skeleton branches are deleted; at the same time, isolated connected regions with too small an area are removed to reduce the interference of raster noise and data defects on the calculation of morphological indicators.
[0066] (3) Missing value imputation: To address the potential missing indicators in the calculation of land parcels, a strategy of K-nearest neighbor (KNN) imputation combined with median backfilling is adopted to ensure the integrity and consistency of the dataset.
[0067] (4) Dual threshold screening mechanism: In the pattern extraction stage, an upper limit is set on the intra-cluster dispersion of the clusters obtained by pre-clustering. This, together with the density threshold τ in subspace clustering, forms a dual threshold criterion, thereby filtering out candidate patterns that are unstable or have low support within a class, and finally retaining only the set of typical patterns with stable statistical properties and clear structure, ensuring the interpretability and practicality of the results.
[0068] This embodiment significantly improves the automation level and reliability of the output results in batch processing scenarios through the above mechanism.
[0069] Example 4: Boundary Case Handling and Failure Recovery Strategies To ensure the robustness of the method of this invention in complex real-world scenarios, this embodiment formulates corresponding processing and recovery strategies for the following typical boundary conditions: (1) Complex terrain containing water bodies or significant elevation differences: When constructing the residential area spatial grid U, a water body vector mask or a steep slope mask generated based on DEM is introduced to mark such areas as impassable spaces (assigned a value of 0), thereby avoiding them from being misidentified as open spaces by morphological opening operations.
[0070] (2) Noise interference caused by ultra-dense grid points: When using excessively high grid resolution (e.g., Δ<0.5m) causes a large amount of fine noise in the building outline, before entering the main process (step S2), perform a small-scale (e.g., r=1) morphological opening or closing operation on the building binary image I to pre-filter the irrelevant pixel-level noise.
[0071] (3) Outline generation of ultra-low density residential areas: If the radius r increases to the preset upper limit, the enclosed area Number of connected components If the value is still greater than 1, a backoff strategy is triggered: starting from the preset maximum r, a reverse decreasing search is performed, and the first value that satisfies the condition is found. The radius r = 1 is used as the effective minimum connected radius, and the "low-density outer contour" label is added to the output results for the residential area for subsequent analysis reference.
[0072] (4) Suppression of the impact of isolated high-rise buildings: For isolated high-rise buildings within a layer, when calculating the line coverage ratio λ and interface density ρ, a weighting coefficient based on the building base area or the distance from adjacent buildings is introduced for adjustment, so as to avoid the disproportionate impact of individual sparse but large-scale buildings on the overall morphological indicators of the residential area.
[0073] Corresponding to the above method, this invention also provides an automatic extraction system for typical patterns of residential spatial morphology. This system adopts a modular design, with each module decoupled via a message bus, supporting parallel processing and hot parameter updates. The system mainly includes: Rasterization module: Used to execute step S1, responsible for receiving vector data and attributes, performing coordinate transformation, interpolation and cleaning, and outputting building binary map I, residential space raster U and floor number raster H.
[0074] Adaptive Closure Operation Module: Used to execute step S2, responsible for managing the structuring element radius sequence, performing closure operations and hole filling, calculating the minimum connected radius r, and generating a uniform outer contour. .
[0075] The site space extraction module is used to execute step S3. It manages the multi-scale set S, performs opening operations on each scale s, and outputs the set of enclosing building indexes corresponding to each site through morphological reconstruction and intersection operations. .
[0076] Windward Interface Analysis Module: Used to execute step S4, responsible for managing the wind direction set. Based on the layer definition, perform raster rotation, leading edge pixel scanning, and calculate the windward interface density ρ, line coverage λ, and notch spacing set. It also supports configurable angle step and height threshold, and provides automatic correction based on leading edge error backtracking.
[0077] Pre-clustering module: Used to execute step S5, responsible for defining indicator groups, performing feature discretization and preliminary category aggregation, and generating pre-clustered category variables.
[0078] CLIQUE Subspace Clustering and Pattern Derivation Module: This module executes step S6, handling high-dimensional space mesh generation, dense cell identification, DFS-based connectivity analysis, MDL-based cluster merging, and ultimately outputting a set of typical patterns. Its minimum feature description and parameterized template.
[0079] Furthermore, its subspace clustering module includes a density calculation unit, a candidate generation unit, a depth-first merging unit, and an MDL discrimination unit, and supports runtime parameter tuning of m and τ.
[0080] This system operates on a hardware environment equipped with an 8-core or higher CPU and 16GB or more of memory. It utilizes a software stack built upon open-source libraries such as GDAL / GEOS, OpenCV, and NumPy / SciPy, and incorporates an object-oriented data access layer and task orchestration module to achieve efficient batch processing. All modules work collaboratively within this environment, realizing a fully automated, integrated pipeline processing from raw data to parameterized templates. Those skilled in the art can adjust the optimal parameters in each step of this technical system to adapt to application scenarios with different accuracy requirements and data characteristics.
[0081] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this invention. It should be understood that the above descriptions are merely specific embodiments of this invention and are not intended to limit the scope of protection of this invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for automatically extracting typical patterns of residential spatial morphology, characterized in that, Includes the following steps: Step S1: Obtain building information data including residential area boundaries, building outlines and number of floors, set the raster resolution Δ, and generate a building binary map I, a residential area space raster U and a floor number raster H; Step S2: Perform morphological closing operation on the binary architectural image I and fill the holes by incrementally increasing the radius r of the structural element, with r incrementing by steps. Increment, determine the enclosed area The minimum connected radius r* is used to achieve full connectivity for the first time, and As the unified outer outline of the residential area; in: In the formula, I represents a binary raster image of a building; A circular or approximately circular structuring element with radius r; and These represent morphological dilation and erosion operations, respectively. An operator for filling holes in a binary region; Step S3: Perform a morphological opening operation on the residential space grid U according to the preset set of radius scales S of structural elements, and extract candidate open space regions corresponding to the radii of different structural elements. ; each After expansion, the intersection with the binary building graph I yields the set of building indices that enclose the site. The consistency score and overall stability score of the enclosed building combination around the same open space candidate area identified under different structural element radii are evaluated. Step S4: Set the direction set Θ, rotate the floor number grid H, and divide it into at least three layers according to height; scan each layer column by column along the windward orthogonal direction, identify the windward leading edge pixels, and calculate the windward interface density. , line bonding rate and gap spacing set ; in: In the formula, This represents the proportion of the number of columns of the leading edge pixels on the windward side of the k-th slice facing the wind direction θ out of the total number of columns, used to characterize the degree of interface coverage; θ is the scanning or incoming flow direction; k is the height slice index; θ represents the grid slice of the k-th layer in the direction θ; W represents the number of vertical columns of slices set along the direction perpendicular to the wind direction. This is an indicator function that takes the value 1 when a column has leading edge cells, and 0 otherwise; This represents the proportion by which the uniform outer contour of the residential area is "attached" to the windward side facing the direction of the wind, θ, by the slice of the k-th floor of the building. The total length of the line segment that coincides with the leading edge pixel of the building's k-th floor slice on the windward side facing the direction of the wind θ on the unified outer contour of the residential area; To ensure a uniform overall perimeter of the outer contour; is the pixel length of the j-th consecutive zero-value gap; M is the number of gap segments; Δ is the raster resolution; This represents the actual width of the j-th segment of the gap; Step S5: Construct a morphological index system that includes individual buildings, enclosures, locations, and windward interfaces. Divide the indexes into multiple index groups with morphological significance, and discretize and pre-cluster the indexes in each index group to obtain categorical variables. Step S6: Incorporate the categorical variables and continuous variables into a multidimensional attribute space for grid partitioning, identify dense cells based on grid cell density, and merge adjacent dense cells using the minimum description length criterion to output a typical pattern set and its minimum feature description. Step S7: For each of the typical patterns, generate a parametric geometric template that includes a uniform outer contour, dominant building type, windward interface parameters, and site scale level.
2. The method as described in claim 1, characterized in that, In step S2, the minimum connectivity radius The method of determination is as follows: In the formula, Indicates the region The number of connected components; and fill holes with an area smaller than a set threshold.
3. The method as described in claim 2, characterized in that, In step S2, the area threshold is the actual area corresponding to 50 to 500 pixels.
4. The method as described in claim 1, characterized in that, In step S3, the scale set S covers from the smallest scale... To the maximum scale The scope; the enclosed building index set By open space candidate areas The result is obtained by intersecting the expanded binary graph with the building's binary graph I.
5. The method as described in claim 1, characterized in that, In step S4, the angular step of the direction set Θ is 45° or 30°; the height stratification includes a low-rise area, a mid-rise area, and a high-rise area, which are divided based on the number of building floors or based on the standard floor height. The converted absolute height.
6. The method as described in claim 1, characterized in that, In step S4, a minimum threshold is set for the length of the continuous zero-value gap. To filter out the tiny gaps caused by noise.
7. The method as described in claim 1, characterized in that, In step S5, the index group includes at least the building type proportion group, the enclosure and depth group, the site space group, the windward interface group, and the orientation combination group; before pre-clustering, each index in the group is discretized by equal frequency or entropy basis and converted into 3 to 9 levels.
8. The method as described in claim 1, characterized in that, In step S6, the merging based on the minimum description length criterion is expressed by the following cost function: In the formula, R represents the candidate high-dimensional dense region; For a structured representation of region R; The bit length required to encode the structure itself; The bit length required to encode the data residual under given structural conditions; The total description length of region R.
9. An automatic extraction system for typical patterns of residential spatial morphology, characterized in that, It includes a processor and a memory, the memory storing a computer program that, when executed by the processor, implements the method as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 8.