A method for automatic identification of urban and suburban areas based on unsupervised machine learning

Through unsupervised machine learning algorithms and Gaussian mixture models (GMM), urban and suburban areas can be automatically identified, solving the problems of frequent manual intervention and complex processes in existing technologies, achieving efficient and precise urban and suburban identification, and supporting urban planning.

CN118968298BActive Publication Date: 2025-09-23TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411016163.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-27
Publication Date
2025-09-23
Estimated Expiration
2044-07-27

AI Technical Summary

Technical Problem

Existing technologies require a lot of manual intervention when identifying urban and suburban areas, have complex workflows, and lack automated and efficient identification methods.

Method used

An unsupervised machine learning algorithm is used to automatically identify the scope of urban and suburban areas by constructing an architectural morphology index system and Gaussian mixture model (GMM) cluster analysis, reducing manual intervention. The multidimensional matrix of architectural morphology characteristics is used for clustering and Bayesian information criterion optimization to ultimately determine the scope of urban and suburban areas.

Benefits of technology

It realizes automatic identification of urban and suburban areas, simplifies the workflow, improves the efficiency and accuracy of identification, can quickly process large-scale data, and provide base map support for urban planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118968298B_ABST
    Figure CN118968298B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for automatically identifying urban and suburban areas using an unsupervised machine learning algorithm. This method uses building outline data to automatically calculate building morphological features, obtains key indicators that reflect the degree of urbanization through feature extraction, and automatically identifies the spatial extent of urban and suburban areas using an unsupervised machine learning algorithm. This method can reduce manual intervention, simplify the workflow for demarcating urban and suburban boundaries, and can finely identify the spatial extent of areas with different degrees of urbanization. The method is highly repeatable and efficient, can quickly process large-scale data, and achieve efficient identification of urban and suburban areas. It provides a base map reference and technical support for the construction of a "one map for national land space planning" in the context of national land space planning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an automatic identification technology for urban and suburban areas based on unsupervised machine learning. Background Art

[0002] The urbanization gradient theory posits that as urbanization intensity changes, the environment exhibits spatial differences, such as transitioning from natural landscapes to highly urbanized areas. Buildings are the most fundamental and sophisticated elements of urban form. Building morphological attributes, such as density, area, shape, and orientation, can well reflect land use and economic activity. Therefore, by identifying changes in building spatial morphology, urbanization intensity can be identified, thereby deriving the spatial extent and corresponding boundaries of regions with varying degrees of urbanization.

[0003] The "Regulations for Determining Urban Areas" (TD / T1064-2021), issued by the Ministry of Natural Resources in 2021, stipulate technical standards for delineating urban areas based on national land survey data and high-precision remote sensing data, taking into account the coverage of public service facilities and the administrative scope of neighborhood and village committees. This provides a sound methodological and technical foundation for subsequent research. Further research is needed to explore methods for automatically identifying the spatial extents of different urbanization gradients, particularly urban and suburban areas, that can reduce manual intervention and streamline workflows. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this paper proposes a method for automatically identifying urban and suburban boundary lines using an unsupervised machine learning algorithm. This method emphasizes the automated identification of urban and suburban boundaries across large spatial scales by building a framework based on architectural form indicators and machine learning algorithms. This approach minimizes manual intervention, facilitates the rapid and efficient formation of a base map for national spatial planning and related urban research and analysis, and provides technical support for digital spatial governance.

[0005] The technical solutions of the present invention are as follows:

[0006] A method for automatically identifying urban and suburban areas based on unsupervised machine learning, characterized by comprising the following steps:

[0007] Step (1): Obtain building outline data within the study area, totaling J spatial units, and calculate the total building area intersecting within a circle with a radius of k meters and the center of each spatial unit j as the center point. , average building perimeter , building orientation angle entropy , average value of building shape index , average building compactness index and the number of buildings Six architectural form characteristics constitute a multidimensional matrix ;

[0008] Step (2): Multidimensional matrix of architectural form characteristics of all spatial units obtained in step (1) Perform Gaussian mixture model GMM cluster analysis;

[0009] Step (3): The mixing coefficient calculated according to step (2) , mean vector and covariance matrix , calculate the corresponding Bayesian Information Criterion

[0010] Step (4): The number of clusters calculated according to step (3) Corresponding , by comparing the number of clusters corresponding to the minimum BIC value as the optimal number of clusters

[0011] Step (5): The optimal number of clusters obtained according to step (4) , and obtain the probability density function of the Gaussian mixture model

[0012] Step (6): Clustering probability calculated according to step (5) , classify each spatial unit j into the type with the highest probability ;

[0013] Step (7): Clustering results of each spatial unit j obtained in step (6) , summarize the total building area of ​​all space units , and in the type Average value , sort from large to small, and get the order , calculate the adjacent data after sorting and Rate of change ;

[0014] Step (8): The rate of change obtained from step (7) ,right Sort from small to large and get the two largest ones The corresponding cluster types are the scope of urban area and suburban area respectively.

[0015] A method for automatically identifying urban and suburban areas based on unsupervised machine learning, comprising the following steps:

[0016] Step (1): Obtain building outline data within the study area, totaling J spatial units, and calculate the total building area intersecting within a circle with a radius of k meters and the center of each spatial unit j as the center point. , average building perimeter , building orientation angle entropy , average value of building shape index , average building compactness index and the number of buildings Six architectural form characteristics constitute a multidimensional matrix , the specific formula is as follows:

[0017]

[0018] In formula 1 is the total building area of ​​space unit j, For the circle The area of ​​the building.

[0019]

[0020] In formula 2 is the average building perimeter of space unit j, For the circle The perimeter of the building.

[0021]

[0022] In formula 3 is the building orientation angle entropy value of spatial unit j, is the probability distribution of the building orientation angle k.

[0023]

[0024] In formula 4 is the average value of the building shape index of spatial unit j, To include buildings The minimum circumscribed circle area of ​​the outline polygon.

[0025]

[0026] In formula 5 is the average building compactness index of space unit j in the year, For buildings Buildings within the minimum circumcircle of the outline polygon circumference.

[0027]

[0028] In formula 6 A multidimensional matrix containing 6 architectural morphological characteristics for each spatial unit j.

[0029] Step (2): Multidimensional matrix of architectural form characteristics of all spatial units obtained in step (1) Perform Gaussian mixture model GMM cluster analysis, number of clusters The value is 1, 2, ..., m, and the mixing coefficient corresponding to different cluster numbers is obtained. , mean vector and the covariance matrix , the formula is as follows:

[0030]

[0031] In formula 7 For spatial unit j belongs to The posterior probability of the class, Therefore is the mean, is the covariance matrix of the multivariate normal distribution.

[0032]

[0033] In formula 8 It is The mixing coefficient of clusters satisfies and , is an indicator variable, if the space unit j belongs to If the class is valid, it is 1; otherwise, it is 0.

[0034]

[0035] In formula 9 is the mean vector.

[0036]

[0037] In formula 10 is the covariance matrix for each cluster, is a vector The cosquare matrix is ​​obtained by multiplying .

[0038] Step (3): The mixing coefficient calculated according to step (2) , mean vector and the covariance matrix , calculate the corresponding Bayesian Information Criterion , the formula is as follows:

[0039]

[0040] In formula 11 The number of clusters is The corresponding maximum likelihood value.

[0041]

[0042] In formula 12, The number of clusters is The corresponding Bayesian Information Criterion (BIC) is used.

[0043] Step (4): The number of clusters calculated according to step (3) Corresponding , by comparing the number of clusters corresponding to the minimum BIC value as the optimal number of clusters , the formula is as follows:

[0044]

[0045] In formula 13, is the optimal number of clusters, To obtain the minimum algorithm.

[0046] Step (5): The optimal number of clusters obtained according to step (4) , respectively calculated The corresponding clustering probability density function of the Gaussian mixture model is obtained to obtain the probability density function of the Gaussian mixture model , the formula is as follows:

[0047]

[0048] In formula 14 The number of clusters is The probability density function of the Gaussian mixture model.

[0049] Step (6): Clustering probability calculated according to step (5) , classify each spatial unit j into the type with the highest probability The formula is as follows:

[0050]

[0051]

[0052] In formula 15 for The corresponding probability density is, Therefore is the mean, is the covariance matrix of the multivariate normal distribution.

[0053] Step (7): Clustering results of each spatial unit j obtained in step (6) , respectively summarize the different =1,2,3,……, The sum of the building areas of all space units in the type ,get . For 1,2,3,……, Type Average value Sort from large to small to get the order , calculate the adjacent data after sorting and Rate of change The formula is as follows:

[0054]

[0055] Step (8): The rate of change obtained from step (7) ,right Sort from small to large and get the two largest ones The corresponding cluster types are used as the range of urban areas and suburban areas respectively. The formula is as follows:

[0056]

[0057]

[0058] In formula 18 for The i value corresponding to the maximum value is the urban area. for The i value corresponding to the second largest value is the suburban area.

[0059] This invention combines high-precision remote sensing data with unsupervised machine learning technology to automatically calculate the morphological and spatial characteristics of buildings, accurately identifying areas with different degrees of urbanization, and providing a powerful tool for urban planning and management. Specific beneficial effects are:

[0060] (1) By adopting unsupervised machine learning algorithms, human intervention is reduced, workflows are simplified, and automatic identification of urban and suburban areas is achieved;

[0061] (2) Through the cluster analysis algorithm, the spatial scope of different urbanization gradient areas can be finely identified, including but not limited to central urban areas, old urban areas, suburban areas, remote suburbs and rural areas, etc. The technical method is highly scalable.

[0062] (3) The proposed method is highly repeatable and efficient, capable of rapidly processing large-scale data and achieving efficient identification of urban and suburban areas. It provides a base map reference and technical support for the construction of a "national spatial planning map" in the context of national spatial planning. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 : The technical roadmap of the present invention.

[0064] Figure 2 : Characteristics of six building forms in Shanghai in 2016.

[0065] Figure 3 : Seven urban morphological types in Shanghai in 2016.

[0066] Figure 4 : Schematic diagram of Shanghai’s urban and suburban areas in 2016. DETAILED DESCRIPTION

[0067] The technical solution provided by this application will be further described below in conjunction with specific embodiments and accompanying drawings. The advantages and features of this application will become more apparent with reference to the following description.

[0068] The present invention relates to an automatic identification method for urban and suburban areas based on unsupervised machine learning.

[0069] The case study is Shanghai.

[0070] Figure 2 : Characteristics of 6 types of buildings in Shanghai in 2016

[0071] Figure 3 : Seven types of urban forms in Shanghai in 2016

[0072] Figure 1 The steps of the embodiment method are as follows:

[0073] A method for automatically identifying urban and suburban areas based on unsupervised machine learning, comprising the following steps:

[0074] Step (1): Obtain the building outline data for 2016 within the study area, totaling 3,430,994,784 (J) spatial units, and calculate the total building area intersecting within a circle with a radius of 250 meters and the center of each spatial unit j. , average building perimeter , building orientation angle entropy , average value of building shape index , average building compactness index and the number of buildings Six architectural form characteristics constitute a multidimensional matrix , see attached Figure 2 The specific formula is as follows:

[0075]

[0076] In formula 1 is the total building area of ​​space unit j, For the circle The area of ​​the building.

[0077]

[0078] In formula 2 is the average building perimeter of space unit j, For the circle The perimeter of the building.

[0079]

[0080] In formula 3 is the building orientation angle entropy value of spatial unit j, is the probability distribution of the building orientation angle k.

[0081]

[0082] In formula 4 is the average value of the building shape index of spatial unit j, To include buildings The minimum circumscribed circle area of ​​the outline polygon.

[0083]

[0084] In formula 5 is the average building compactness index of space unit j in the year, For buildings Buildings within the minimum circumcircle of the outline polygon circumference.

[0085]

[0086] In formula 6 A multidimensional matrix containing 6 architectural morphological characteristics for each spatial unit j.

[0087] Step (2): Multidimensional matrix of architectural form characteristics of all spatial units obtained in step (1) Perform Gaussian mixture model GMM cluster analysis, number of clusters The value is 1, 2, ..., 20, and the mixing coefficient corresponding to different cluster numbers is obtained. , mean vector and the covariance matrix , the formula is as follows:

[0088]

[0089] In formula 7 For spatial unit j belongs to The posterior probability of the class, Therefore is the mean, is the covariance matrix of the multivariate normal distribution.

[0090]

[0091] In formula 8 It is The mixing coefficient of clusters satisfies and , is an indicator variable, if the space unit j belongs to If the class is valid, it is 1; otherwise, it is 0.

[0092]

[0093] In formula 9 is the mean vector.

[0094]

[0095] In formula 10 is the covariance matrix for each cluster, is a vector The cosquare matrix is ​​obtained by multiplying .

[0096] Step (3): The mixing coefficient calculated according to step (2) , mean vector and the covariance matrix , calculate the corresponding Bayesian Information Criterion , the formula is as follows:

[0097]

[0098] In formula 11 The number of clusters is The maximum likelihood value corresponding to .

[0099]

[0100] In formula 12, The number of clusters is The corresponding Bayesian Information Criterion (BIC) is used.

[0101] Step (4): The number of clusters calculated according to step (3) Corresponding , the number of clusters corresponding to the minimum BIC value is obtained by comparison as the optimal number of clusters 7 ( ), the formula is as follows:

[0102]

[0103] In formula 13, is the optimal number of clusters, To obtain the minimum algorithm.

[0104] Step (5): According to the optimal number of clusters 7 obtained in step (4), calculate The corresponding clustering probability density function of the Gaussian mixture model is obtained to obtain the probability density function of the Gaussian mixture model , the formula is as follows:

[0105]

[0106] In formula 14 The number of clusters is The probability density function of the Gaussian mixture model.

[0107] Step (6): Clustering probability calculated according to step (5) , classify each spatial unit j into the type with the highest probability , see attached Figure 3 The formula is as follows:

[0108]

[0109]

[0110] In formula 15 for The corresponding probability density is, Therefore is the mean, is the covariance matrix of the multivariate normal distribution.

[0111] Step (7): Clustering results of each spatial unit j obtained in step (6) , respectively summarize the different =1,2,3,……, The sum of the building areas of all space units in the type ,get . For 1,2,3,……, Type Average value Sort from large to small to get the order , calculate the adjacent data after sorting and Rate of change The formula is as follows:

[0112]

[0113] Step (8): The rate of change obtained from step (7) ,right Sort from small to large and get the two largest ones The corresponding cluster types are as the scope of urban area and suburban area, see the attached Figure 4 The formula is as follows:

[0114]

[0115]

[0116] In formula 18 for The i value corresponding to the maximum value is the urban area. for The i value corresponding to the second largest value is the suburban area.

[0117] This technical solution uses building outline data to automatically calculate building morphological characteristics. Through feature extraction, it derives key indicators reflecting urbanization levels. It then applies an unsupervised machine learning algorithm to automatically identify the spatial extents of urban and suburban areas. This method reduces manual intervention, simplifies the delineation workflow for urban and suburban boundaries, and enables refined identification of the spatial extents of areas with varying degrees of urbanization. The method is highly reproducible and efficient, capable of rapidly processing large amounts of data and effectively identifying urban and suburban areas. It provides a basemap reference and technical support for the development of a "Single Map for National Land and Space Planning" within the context of national land and space planning.

Claims

1. A method for automatically identifying urban and suburban areas based on unsupervised machine learning, characterized in that: The following steps are involved: Step (1): Obtain building outline data within the study area, totaling J spatial units, and calculate the total building area intersecting within a circle with a radius of k meters and the center of each spatial unit j as the center point. , average building perimeter , building orientation angle entropy , average value of building shape index , average building compactness index and the number of buildings Six architectural form characteristics constitute a multidimensional matrix ; Step (2): Multidimensional matrix of architectural form characteristics of all spatial units obtained in step (1) Perform Gaussian mixture model GMM cluster analysis; Step (3): The mixing coefficient calculated according to step (2) , mean vector and the covariance matrix , calculate the corresponding Bayesian Information Criterion Step (4): The number of clusters calculated according to step (3) Corresponding , by comparing the number of clusters corresponding to the minimum BIC value as the optimal number of clusters Step (5): The optimal number of clusters obtained according to step (4) , and obtain the probability density function of the Gaussian mixture model Step (6): Clustering probability calculated according to step (5) , classify each spatial unit j into the type with the highest probability ; Step (7): Clustering results of each spatial unit j obtained in step (6) , summarize the total building area of ​​all space units , and in the type Average value , sort from large to small, and get the order , calculate the adjacent data after sorting and Rate of change ; Step (8): The rate of change obtained from step (7) ,right Sort from small to large and get the two largest ones The corresponding cluster types are the scope of urban area and suburban area respectively.

2. The method for automatically identifying urban and suburban areas based on unsupervised machine learning according to claim 1, wherein: In step (1), the six architectural features form a multidimensional matrix , specifically: The formulas are as follows: In formula 1 is the total building area of ​​space unit j, For the circle The area of ​​the building; In formula 2 is the average building perimeter of space unit j, For the circle The perimeter of the building; In formula 3 is the building orientation angle entropy value of spatial unit j, is the probability distribution of the building orientation angle k; In formula 4 is the average value of the building shape index of spatial unit j, To include buildings The minimum circumscribed circle area of ​​the outline polygon; In formula 5 is the average building compactness index of space unit j in the year, For buildings Buildings within the minimum circumcircle of the outline polygon circumference; In formula 6 A multidimensional matrix containing 6 architectural morphological characteristics for each spatial unit j.

3. The method for automatically identifying urban and suburban areas based on unsupervised machine learning according to claim 1, wherein: Step (2): Multidimensional matrix of architectural form characteristics of all spatial units obtained in step (1) Perform Gaussian mixture model GMM cluster analysis, number of clusters The value is 1, 2, ..., m, and the mixing coefficient corresponding to different cluster numbers is obtained. , mean vector and the covariance matrix , the formula is as follows: In formula 7 For spatial unit j belongs to The posterior probability of the class, Therefore is the mean, is the covariance matrix of the multivariate normal distribution; In formula 8 It is The mixing coefficient of clusters satisfies and , is an indicator variable, if the space unit j belongs to If the class is 1, otherwise it is 0; In formula 9 is the mean vector; In formula 10 is the covariance matrix for each cluster, is a vector The cosquare matrix is ​​obtained by multiplying .

4. The method for automatically identifying urban and suburban areas based on unsupervised machine learning according to claim 2, wherein: Step (3): The mixing coefficient calculated according to step (2) , mean vector and the covariance matrix , calculate the corresponding Bayesian Information Criterion , the formula is as follows: In formula 11 The number of clusters is The maximum likelihood value corresponding to ; In formula 12, The number of clusters is The corresponding Bayesian Information Criterion (BIC) is used.

5. The method for automatically identifying urban and suburban areas based on unsupervised machine learning according to claim 2, wherein: Step (4): The number of clusters calculated according to step (3) Corresponding , by comparing the number of clusters corresponding to the minimum BIC value as the optimal number of clusters , the formula is as follows: In formula 13, is the optimal number of clusters, To obtain the minimum algorithm.

6. The method for automatically identifying urban and suburban areas based on unsupervised machine learning according to claim 2, wherein: Step (5): The optimal number of clusters obtained according to step (4) , respectively calculated The corresponding clustering probability density function of the Gaussian mixture model is obtained to obtain the probability density function of the Gaussian mixture model , the formula is as follows: In formula 14 The number of clusters is The probability density function of the Gaussian mixture model.

7. The method for automatically identifying urban and suburban areas based on unsupervised machine learning according to claim 2, wherein: Step (6): Clustering probability calculated according to step (5) , classify each spatial unit j into the type with the highest probability ; The formula is as follows: In formula 15 for The corresponding probability density is, Therefore is the mean, is the covariance matrix of the multivariate normal distribution.

8. The method for automatically identifying urban and suburban areas based on unsupervised machine learning according to claim 2, wherein: Step (7): Clustering results of each spatial unit j obtained in step (6) , respectively summarize the different =1,2,3,……, The sum of the building areas of all space units in the type ,get , respectively for 1,2,3,……, Type Average value Sort from large to small to get the order , calculate the adjacent data after sorting and Rate of change ; The formula is as follows: 。 9. The method for automatically identifying urban and suburban areas based on unsupervised machine learning according to claim 2, wherein: Step (8): The rate of change obtained from step (7) ,right Sort from small to large and get the two largest ones The corresponding cluster types are used as the range of urban areas and suburban areas respectively; the formula is as follows: In formula 18 for The i value corresponding to the maximum value is the urban area. In formula 19, for The i value corresponding to the second largest value is the suburban area.

Citation Information

Patent Citations

  • Method and system for automatically partitioning urban spatial form

    CN109492796A

  • High-precision dynamic identification method for urban marginal area spatial range

    CN116645012A