A city classification method for constructing a beautiful city differentiation index system

By performing feature compression and data confidence correction on raster data, and combining spatial coordinates, an unsupervised machine learning method is used for city classification. This solves the problem of inaccurate classification in the processing of multi-source heterogeneous data and achieves more accurate city classification results.

CN122471253APending Publication Date: 2026-07-28INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-25
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

Existing city classification methods fail to effectively handle multi-source heterogeneous data, resulting in the high-dimensional features and low confidence indices of raster data affecting city classification results. Furthermore, they fail to take into account attribute similarity and geographical proximity, leading to inaccurate classification results.

Method used

A raster feature statistical compression method based on discrete probability distribution is used to compress raster data features. The index weights are corrected by combining data confidence and spatial coordinates are introduced. Clustering is performed through unsupervised machine learning to form a set of urban feature vectors for urban classification.

Benefits of technology

It achieves city classification within a unified feature space, reduces the impact of high-dimensional raster features and low confidence indices, takes into account attribute similarity and geographical proximity, and improves the accuracy and stability of classification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122471253A_ABST
    Figure CN122471253A_ABST
Patent Text Reader

Abstract

The application provides a city classification method for constructing a beautiful city differentiation index system, and relates to the technical field of city classification data processing. The application takes a city unit as an object, acquires multi-source heterogeneous data of an ecological environment protection dimension and a city development and construction dimension, performs discrete probability distribution feature compression on grid data, and extracts index features from the remaining data. Based on data confidence correction of a basic weight, a city feature vector set is constructed by introducing a spatial coordinate. Binary classification results of each dimension are obtained through unsupervised clustering, and then a subdivided type and city categories of double-high type, high-low type, low-high type and double-low type are formed. The application can reduce the influence of grid high-dimensional features, low-confidence indexes and missing spatial proximity relationships on city classification results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of urban classification data processing technology, and in particular to an urban classification method for constructing a differentiated index system for beautiful cities. Background Technology

[0002] Urban ecological environment assessment and city type identification typically employ an indicator system construction, indicator data processing, and classification result output as their basic technical approaches. Existing technologies have disclosed schemes for classifying city types based on economic, resource, and environmental factors, and forming ecological city indicator systems under different city types. For example, Chinese invention patent application publication CN101777170A discloses a method for constructing an ecological city indicator system based on city classification. This method classifies city types according to a three-factor matrix composed of economic, resource, and environmental factors, and determines corresponding indicators for different city types. Chinese invention patent application publication CN104537597A discloses a technical method for diagnosing the rationality of urban spatial patterns. This method constructs a multi-level indicator system and classifies the rationality level of cities based on a comprehensive diagnostic index. These methods illustrate that existing urban evaluation technologies typically rely on multiple indicators to construct evaluation models and output city types or levels through indicator calculations or comprehensive indices.

[0003] With the increasing use of remote sensing raster data, environmental monitoring data, annual statistical data, and geospatial data in urban assessment, urban classification methods are no longer suitable for relying solely on single indicator values ​​or linear composite indices. Air environment, water environment, ecological status, and urban development and construction level correspond to different data sources and different data formats. Raster data can reflect continuous spatial distribution, but has a large number of pixels and high raw dimensionality; station monitoring data can reflect local environmental conditions, but spatial representativeness is affected by station location; annual time series data can reflect annual changes within a city, but it is difficult to express differences within the city; binary classification data can express judgments about natural conditions, but the amount of information is relatively limited. Before multi-source heterogeneous data enters the same urban classification model, it is necessary to complete unified feature representation, indicator influence degree correction, and spatial coordinate introduction processing to enable different cities to form comparable air, water, ecological, and urban development and construction classification results within the same characteristic space.

[0004] However, existing city evaluation and classification methods often focus on indicator system design, weight assignment, and comprehensive index calculation, neglecting the technical processing of multi-source heterogeneous data before it enters the clustering model. Directly using raster data in calculations at the raw pixel level can easily lead to excessively high feature dimensions and reduce the classification stability between city units. If annual time-series data, station monitoring data, raster data, and binary classification data are directly fused with equal weights or fixed weights, it is difficult to eliminate the differences in stability, spatial representativeness, and information content among different data types. Existing clustering processes also tend to determine city categories solely based on attribute similarity, failing to consider city longitude and latitude as independent variables, resulting in cities with similar attribute values ​​but geographically distant locations being grouped into the same category. Therefore, a city classification method is needed that can statistically compress and represent raster data, adjust the influence of indicators based on data confidence levels, and simultaneously consider attribute similarity and geographical proximity. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide a city classification method for constructing a differentiated index system for beautiful cities. By performing feature compression, confidence correction, and spatial coordinate introduction on multi-source heterogeneous data, the classification results for dimensions such as air, water, ecology, and urban development and construction can be formed within a unified feature space, reducing the impact of high-dimensional raster features, low-confidence indicators, and missing spatial proximity relationships on the city classification results.

[0006] To achieve the above objectives, the present invention provides the following solution:

[0007] A city classification method for constructing a differentiated index system for beautiful cities, including:

[0008] Using city units as the classification object, multi-source heterogeneous data were acquired to characterize the dimensions of ecological environment protection and urban development and construction. The ecological environment protection dimension includes three sub-dimensions: air, water, and ecology. The multi-source heterogeneous data includes annual time series data, station monitoring data, raster data, and binary classification data.

[0009] The raster data is compressed and characterized using a raster feature statistical compression method based on discrete probability distribution to obtain raster dimensionality reduction features. Corresponding indicator features are extracted from annual time series data, site monitoring data and binary classification data.

[0010] The basic weights of each indicator are set based on expert knowledge, and the basic weights are adjusted by combining the confidence of data of different data types to obtain the hybrid weights based on the confidence of multi-source data.

[0011] In the dimensions of air, water, ecology, and urban development and construction, the corresponding raster dimensionality reduction features and indicator features are fused according to the mixed weights, and spatial coordinates are introduced as independent variables to form the urban feature vector set of the corresponding dimensions.

[0012] The feature vector sets of each city are input into an unsupervised machine learning method for clustering, resulting in binary classification results for each city in the dimensions of air, water, ecology, and urban development and construction.

[0013] Based on the binary classification results, each city is classified into a sub-category consisting of good or poor air quality, good or poor water quality, good or poor ecology, and good or poor development. If a city is judged to be at a high level in two or more of the three sub-dimensions of air, water and ecology, it is judged to be at a high level in ecological and environmental protection; otherwise, it is judged to be at a low level in ecological and environmental protection.

[0014] Based on the binary classification results of high-level or low-level ecological and environmental protection and urban development and construction, each city is divided into high-and-low type, high-low type, low-high type, or low-and-low type.

[0015] The present invention discloses the following beneficial effects:

[0016] This invention acquires multi-source heterogeneous data on both ecological environmental protection and urban development and construction dimensions simultaneously during the same city classification process. Furthermore, it limits the ecological environmental protection dimension to three sub-dimensions: air, water, and ecology. This ensures that annual time-series data, station monitoring data, raster data, and binary classification data are no longer used as isolated indicators for separate judgments, but rather enter subsequent feature fusion and clustering processing after the dimension affiliation is clearly defined. Therefore, the city classification results can simultaneously reflect environmental status, natural conditions, and construction levels, avoiding the problem of existing methods losing differences between different dimensions when relying on only a single comprehensive indicator to form classification results.

[0017] This invention employs a raster feature statistical compression method based on discrete probability distribution to compress and represent raster data. The resulting raster dimensionality-reduced features are then used in conjunction with other indicator features for classification. This processing method preserves the overall morphological information of the pixel value distribution in the raster data, reducing the problem of excessively high feature dimensionality caused by directly involving original pixels in calculations. This allows raster data to enter the city feature vector set in a feature format suitable for comparing city units. For data with continuous spatial distribution characteristics, such as annual average wind speed and vegetation index, the above processing can reduce the interference of differences in original spatial resolution on clustering results.

[0018] This invention sets basic weights for each indicator based on expert knowledge and adjusts these weights by incorporating the confidence levels of different data types, resulting in a hybrid weighting based on the confidence levels of multi-source data. This scheme does not simply assign equal weights to annual time-series data, site monitoring data, raster data, and binary classification data; instead, it allows differences in stability, spatial representativeness, and information content among different data types to be considered in determining the degree of influence of the indicators. Therefore, data with lower confidence levels or weaker stability are less likely to excessively influence the classification results, while data with higher confidence levels can maintain a more reasonable contribution in the corresponding dimensions.

[0019] This invention fuses features across the dimensions of air, water, ecology, and urban development and construction, and introduces spatial coordinates as independent variables into the corresponding urban feature vector sets. Then, unsupervised machine learning methods are used for clustering. This process ensures that city classification is not only based on the similarity of indicator attributes but also takes into account the geographical proximity between cities, reducing the likelihood of cities with similar attribute values ​​but geographically distant locations being mechanically grouped into the same category. For urban objects in ecological environment governance that are significantly affected by topography, climate, water systems, and regional transport, the introduction of spatial coordinates makes the classification results closer to the actual spatial distribution.

[0020] This invention, after obtaining binary classification results through clustering across various dimensions, categorizes each city into subcategories based on whether its air quality is good or poor, its water quality is good or poor, its ecology is good or poor, and its development is good or poor. Furthermore, it determines whether ecological environmental protection is at a high or low level based on the number of high-level results in the three sub-dimensions of air, water, and ecology. This is then combined with the binary classification results from the urban development and construction dimension to obtain dual-high, high-low, low-high, or dual-low types. This classification results in at least four distinct comprehensive types, and each comprehensive type can be traced back to the binary classification results of the air, water, ecology, and urban development and construction dimensions, avoiding the problem of unclear classification criteria caused by simply outputting comprehensive scores. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 A flowchart of the method provided in an embodiment of the present invention;

[0023] Figure 2 A schematic diagram of the multi-year average rainfall in the Beijing-Tianjin-Hebei region provided for embodiments of the present invention;

[0024] Figure 3This is a schematic diagram of the green coverage rate of urban built-up areas in the Beijing-Tianjin-Hebei region, provided as an embodiment of the present invention.

[0025] Figure 4 A schematic diagram illustrating the development and construction level of cities in the Beijing-Tianjin-Hebei region, provided for embodiments of the present invention. Detailed Implementation

[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0027] The purpose of this invention is to provide a city classification method for constructing a differentiated index system for beautiful cities. By introducing raster feature statistical compression, data confidence correction, and spatial coordinates into the clustering process, city indicators from different sources and with different data forms can participate in the classification together, reducing classification bias caused by excessively high original raster dimensions, differences in indicator stability, and simple attribute clustering.

[0028] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0029] In this embodiment, to avoid misunderstandings due to the use of both Chinese names and English abbreviations for the same indicator, the relevant terms are explained as follows: GDP refers to Gross Domestic Product; GDP per capita refers to per capita Gross Domestic Product; Energy consumption per unit of GDP refers to energy consumption per unit of Gross Domestic Product; Water consumption per unit of GDP refers to water consumption per unit of Gross Domestic Product; Carbon dioxide emissions per unit of GDP refers to carbon dioxide emissions per unit of Gross Domestic Product; Building land use area per unit of GDP refers to building land use area per unit of Gross Domestic Product; COD refers to Chemical Oxygen Demand; PM2.5 refers to urban fine particulate matter; NDVI refers to the vegetation index. In this embodiment, the indicator names involving the above-mentioned English abbreviations are interpreted according to their corresponding Chinese names.

[0030] Figure 1 A flowchart of the method provided in the embodiments of the present invention, such as Figure 1 As shown, this invention provides a city classification method for constructing a differentiated index system for beautiful cities, including:

[0031] Step 100: Using city units as the classification object, obtain multi-source heterogeneous data to represent the ecological environment protection dimension and the urban development and construction dimension; the ecological environment protection dimension includes three sub-dimensions: air, water, and ecology, and the multi-source heterogeneous data includes annual time series data, station monitoring data, raster data, and binary classification data;

[0032] Step 200: The raster data is compressed and characterized using a raster feature statistical compression method based on discrete probability distribution to obtain raster dimensionality reduction features, and corresponding indicator features are extracted from annual time series data, site monitoring data and binary classification data.

[0033] Step 300: Set the basic weights of each indicator based on expert knowledge, and adjust the basic weights in combination with the confidence of data of different data types to obtain the mixed weights based on the confidence of multi-source data;

[0034] Step 400: In the dimensions of air, water, ecology and urban development and construction, respectively, the corresponding raster dimensionality reduction features and indicator features are fused according to the mixed weights, and spatial coordinates are introduced as independent variables to form the urban feature vector set of the corresponding dimensions.

[0035] Step 500: Input the feature vector sets of each city into an unsupervised machine learning method for clustering to obtain binary classification results for each city in the dimensions of air, water, ecology and urban development and construction;

[0036] Step 600: Based on the binary classification results, each city is classified into a sub-category consisting of good or bad air, good or bad water, good or poor ecology, and good or poor development. If two or more of the three sub-dimensions of air, water, and ecology are judged as high level, the city is judged as high level of ecological and environmental protection; otherwise, the city is judged as low level of ecological and environmental protection.

[0037] Step 700: Combining the results of the binary classification of high-level or low-level ecological and environmental protection and urban development and construction, each city is divided into high-two-high, high-low, low-high, or low-two-low types.

[0038] Specifically, the goal of this embodiment is to scientifically classify cities from the perspective of ecological environment governance, providing guidance for constructing differentiated indicator systems and development goals for different types of beautiful cities. The classification is based on two main dimensions: the level of urban ecological environment protection and the level of urban development and construction. The ecological environment protection dimension is further subdivided into three sub-dimensions: air, water, and ecology. Each of the three sub-dimensions (air, water, and ecology) selects 2-4 indicators related to natural endowments and ecological environment status (see Table 1), collectively constituting 9 indicators for the urban ecological environment protection dimension; the urban development and construction dimension is defined by 8 indicators (see Table 1).

[0039] Table 1. Selection of Indicators for City Clustering

[0040]

[0041] In one embodiment, this invention uses urban units as the classification object and constructs an urban classification process around the dimensions of ecological environmental protection and urban development and construction. The ecological environmental protection dimension includes three sub-dimensions: air, water, and ecology. The urban classification process includes two stages: dimensional clustering and comprehensive classification. In the dimensional clustering stage, cluster analysis is performed on air, water, ecology, and urban development and construction respectively to obtain the relative performance of each city in each dimension. In the comprehensive classification stage, based on the clustering results of air, water, ecology, and urban development and construction, each city is first classified into one of thirty-two sub-types consisting of good or poor air quality, good or poor water quality, good or poor ecology, and good or poor development. Then, based on the overall performance of the ecological environmental protection dimension and the urban development and construction dimension, each city is divided into two types: high-quality and low-quality, high-low-quality, low-high-quality, or low-quality and low-quality.

[0042] Specifically, unsupervised machine learning methods were used for cluster analysis on the three sub-dimensions of air, water, and ecology, as well as the dimension of urban development and construction. These unsupervised machine learning methods could employ the k-means algorithm or a Gaussian mixture model. During the clustering process, spatial coordinates were introduced as independent variables into the corresponding dimension's city feature vector set, ensuring that the city classification results were influenced by both attribute similarity and geographical proximity. This process makes it easier for cities with similar attributes and geographical proximity to be grouped into the same category in the corresponding dimension, avoiding the problem of determining city categories solely based on the similarity of indicator values. The number of clusters for each dimension was set to two, used to divide cities into high-level and low-level categories in the corresponding dimension. A high-level category indicates that the city's ecological environment protection status in the air, water, or ecology sub-dimensions is relatively good, or its performance in the urban development and construction dimension is relatively high; a low-level category indicates that the city's ecological environment protection status in the air, water, or ecology sub-dimensions is relatively poor, or its performance in the urban development and construction dimension is relatively low.

[0043] Before forming the city feature vector set, feature extraction and fusion are performed on multi-source heterogeneous data representing the dimensions of ecological environment protection and urban development and construction. The multi-source heterogeneous data includes station monitoring data, annual time-series data, raster data, and binary classification data. Station monitoring data is used to represent monitoring indicators such as surface water quality. Annual time-series data is used to represent annual changes in indicators such as air quality, socio-economic indicators, and urban construction indicators. Raster data is used to represent indicators with spatially continuous distribution characteristics, such as vegetation index and average annual wind speed. Binary classification data is used to represent whether the city's geographical and natural conditions are conducive to the diffusion of air pollutants. After feature extraction, different data types are fused according to their dimension affiliation (air, water, ecology, or urban development and construction) to form a multi-dimensional city feature vector set corresponding to each dimension.

[0044] During the feature fusion process across various dimensions, different variables within each dimension are assigned hybrid weights based on the confidence levels of multi-source data. These hybrid weights are based on fundamental weights determined by expert knowledge and adjusted for data confidence levels. Annual time-series data, site monitoring data, raster data, and binary classification data differ in stability, spatial representativeness, and information content; the lower the data confidence level, the lower the impact of the corresponding indicator after adjustment. Through this processing, data from different sources and in different formats can participate in city classification according to their impact level commensurate with data quality, reducing the interference of low-stability or low-representative data on clustering results.

[0045] After completing clustering across all dimensions, a sub-category is generated for each city based on the binary classification results for air, water, ecology, and urban development and construction. The binary classification results for the air sub-dimensional determine whether a city has good or poor air quality; for the water sub-dimensional, they determine whether it has good or poor water quality; for the ecology sub-dimensional, they determine whether it has good or poor ecology; and for the urban development and construction dimension, they determine whether it has good or poor development. Combining these four categories of results results in thirty-two sub-categories: good or poor air quality, good or poor water quality, good or poor ecology, and good or poor development. This sub-category retains the classification source for each city across the four dimensions, avoiding unclear classification criteria due to outputting only a comprehensive rating.

[0046] Based on the 32 subcategories, a comprehensive assessment of the ecological and environmental protection dimension is further conducted. If a city is rated high in two or more of the three sub-dimensions (air, water, and ecology), it is classified as having a high level of ecological and environmental protection. If a city is rated high in only one or none of the sub-dimensions, it is classified as having a low level of ecological and environmental protection. Subsequently, four comprehensive types are generated by combining the binary classification results of the urban development and construction dimension. Cities with both high levels of ecological and environmental protection and high levels in urban development and construction are classified as "Double High"; cities with high levels of ecological and environmental protection and low levels in urban development and construction are classified as "High-Low"; cities with low levels of ecological and environmental protection and high levels in urban development and construction are classified as "Low-High"; and cities with both low levels of ecological and environmental protection and low levels in urban development and construction are classified as "Double Low".

[0047] The above clustering results provide a classification basis for the differentiated evaluation of beautiful city construction. For cities with low ecological and environmental protection pressure, the ecological and environmental evaluation indicators focus on status indicators to reflect the current level of ecological and environmental performance; for cities with high ecological and environmental protection pressure, the ecological and environmental evaluation indicators focus on process indicators to reflect the relative degree of improvement in the ecological and environmental situation. For cities with high levels of urban development and construction, the evaluation indicators related to low-carbon construction and green production focus on status indicators; for cities with relatively weak levels of urban development and construction, the evaluation indicators related to low-carbon construction and green production focus on process indicators to reflect the construction progress and degree of improvement. Through the above classification results, different cities can be assigned different evaluation orientations, avoiding the situation where cities with significant differences in natural endowment, ecological pressure, and construction stage are treated under the same evaluation rules.

[0048] In one implementation, taking advantage of the high spatial resolution, large number of pixels, and high original feature dimensionality of raster data, a raster feature statistical compression method based on discrete probability distribution is used to compress and represent the raster data features. Specifically, the value range of raster data within the overall study area is first divided into multiple intervals. In the embodiment for the Beijing-Tianjin-Hebei region, the number of intervals is set to two hundred, and each interval is used to establish a discrete probability distribution. Subsequently, using the built-up area or administrative boundary of each city as a mask, all effective pixel values ​​within the corresponding city unit are extracted to form a spatial sample set. This spatial sample set is denoted as... , This represents the number of valid pixels within the corresponding city unit.

[0049] After obtaining the spatial sample set, the frequency density of the spatial sample set falling into each interval is calculated, and a discrete probability distribution vector is constructed accordingly. The discrete probability distribution vector is denoted as... , The discrete probability distribution vector, representing the number of intervals, satisfies the condition that the sum of its components is one. This discrete probability distribution vector characterizes the shape features of the pixel value distribution function within a city unit. Based on this discrete probability distribution vector, the mean, median, 10th percentile, 25th percentile, 75th percentile, and 90th percentile are further extracted as supplementary statistical feature values. The discrete probability distribution vector is combined with these supplementary statistical feature values ​​to form the low-dimensional feature vector corresponding to this raster data layer.

[0050] Furthermore, the low-dimensional feature vectors are subjected to dimensionality reduction processing to obtain the raster dimensionality-reduced features of the raster data layer. In a specific embodiment, the dimensionality reduction processing can employ an autoencoder or principal component analysis. An autoencoder is a nonlinear dimensionality reduction method based on neural networks, suitable for preserving the nonlinear variation characteristics in the raster data distribution. Principal component analysis is a linear dimensionality reduction method, suitable for extracting the main variance directions in the raster features. The dimensionality of the reduced features can be compressed to five to eight dimensions. Through the above processing, raster data such as annual average wind speed and vegetation index no longer directly enter the clustering model in their original pixel form, but instead participate in subsequent feature fusion as the raster dimensionality-reduced features corresponding to city units.

[0051] In one implementation, the unsupervised machine learning method may employ a distance-based clustering algorithm. The distance-based clustering algorithm may be the k-means algorithm. The k-means algorithm takes the city feature vector set formed after the aforementioned feature extraction and fusion as input. The city feature vector set is denoted as... , The number of cities participating in the classification, Indicates the first The city's performance in terms of air, water, ecology, or urban development and construction. 3D feature vectors The dimension is the concatenation of all variable feature vectors within the corresponding dimension. The number of clusters is denoted as . In this embodiment, Taking two is used to divide cities into two categories, high-level and low-level, in the corresponding dimension.

[0052] The k-means algorithm partitions data by minimizing the sum of squared Euclidean distances between a sample and its cluster center. The objective function can be expressed as:

[0053]

[0054] in, The objective function value; For the first One cluster; For the first The cluster centers of each cluster; For belonging to the first The k-means algorithm iteratively updates the cluster centers and sample affiliations during training until the objective function converges. After convergence, it outputs the category label for each city in the corresponding dimension. and each cluster center The category attribution labels are used to generate binary classification results for each city in the dimensions of air, water, ecology, or urban development and construction.

[0055] In another implementation, the unsupervised machine learning method can employ a probability distribution-based clustering model. This probability distribution-based clustering model can be a Gaussian mixture model. The Gaussian mixture model takes the same set of city feature vectors as the k-means algorithm as input and assumes that each city feature vector is generated by a mixture of multiple Gaussian components. For the... City feature vectors The probability density function of a Gaussian mixture model can be expressed as:

[0056]

[0057] in, For the set of model parameters; For the first The mixture weights of the Gaussian components are such that the sum of the mixture weights is one. For the first The mean vector of Gaussian components; For the first The covariance matrix of Gaussian components; For and A defined Gaussian distribution density function. Number of clusters. In this implementation, two values ​​are used to output two categories: high level and low level.

[0058] The model parameters of the Gaussian mixture model are estimated using the expectation-maximization (EM) algorithm. The EEM algorithm first calculates the posterior probability of a city's feature vector belonging to each Gaussian component based on the current model parameters. Then, it updates the model parameters based on these posterior probabilities and repeats this process until the log-likelihood function meets the convergence condition. After model convergence, it outputs a probability vector indicating which city belongs to which category. This probability vector is denoted as... Indicates the first The city belongs to the first The posterior probabilities of each category are calculated. The city's category is determined based on the maximum posterior probability, thus obtaining the city category label. Compared to distance-based clustering algorithms, Gaussian mixture models can characterize the distribution of different categories in the feature space through mean vectors and covariance matrices, and represent the city category affiliation in probabilistic form. They are suitable for situations where there is uncertainty when annual time series data, site monitoring data, raster data, and binary classification data are all involved in classification.

[0059] In one implementation, after obtaining the probability results of cities on the corresponding dimensions using a Gaussian mixture model, a secondary correction can be performed on cities with unstable category classifications. Specifically, for any one of the following sub-dimensions—air, water, ecology, or urban development and construction—the first... The posterior probabilities of each city belonging to the high-level and low-level categories are calculated, and the posterior probabilities of the th city are calculated. The city in the Probability differences across dimensions:

[0060]

[0061] in, For the first The city in the The probability difference across each dimension is used to characterize the clustering confidence of the city in the corresponding dimension; For the first The city in the Posterior probability of belonging to the high-level category in each dimension; For the first The city in the The posterior probability of belonging to the low-level category in the dimension; the th The dimensions can be sub-dimensions of air quality, water quality, ecology, or urban development and construction. The smaller the probability difference, the weaker the difference in the city's affiliation between high-level and low-level categories, and the closer the category affiliation is to the cluster boundary; the larger the probability difference, the more stable the city's category affiliation in the corresponding dimension.

[0062] Furthermore, setting up with the first Each dimension corresponds to a pre-set information threshold. When the probability difference Below the preset information threshold At that time, the first The city was identified as the [number]th Border cities in each dimension. For each border city, a set of neighboring cities is determined based on its spatial coordinates. This set of neighboring cities can consist of cities whose spatial distance from the border city is less than a preset distance threshold, or it can consist of a preset number of cities that are spatially closest to the border city. Subsequently, the values ​​of each city in the set of neighboring cities are counted in the [missing dimension]. The classification results across multiple dimensions determine the majority category within the set of neighboring cities. This processing is confined to the same dimension to avoid cross-influence between classification results for air, water, ecology, and urban development and construction. Exemplarily, the pre-set confidence threshold... According to the first The distribution of probability differences for each city in each dimension is determined. In one implementation, the distribution of probability differences for each city in each dimension is determined. The probability differences of all cities within each dimension are sorted from smallest to largest, and the probability differences corresponding to preset quantiles are used as the preset confidence threshold. The preset quantile can be taken from the 10th percentile to the 30th percentile, so that cities with smaller differences in posterior probability are identified as border cities, and cities whose category classification has already been stabilized are avoided from being included in the secondary correction range.

[0063] Furthermore, the border cities in the first The initial category in each dimension is compared with the majority category in the set of neighboring cities. When the initial category matches the majority category, the initial category of the border city is retained; when the initial category does not match the majority category, and the probability difference is... Below the preset information threshold At that time, the border cities are classified according to the majority category in the first... The classification is corrected across all dimensions to obtain a corrected binary classification result. This corrected binary classification result is used to generate subcategories such as good or poor air quality, good or poor water quality, good or poor ecology, and good or poor development. These subcategories are then used to determine whether ecological environmental protection is at a high or low level, or whether it falls into two categories: high-high, high-low, low-high, or low-low. This process reduces category jumps caused by similar posterior probabilities in border cities, ensuring that the classification results, while retaining probabilistic clustering results, further conform to the category consistency of geographically neighboring cities.

[0064] In one implementation, to ensure that indicators corresponding to different types of data participate in feature fusion according to data quality differences, this invention employs a hybrid weighting based on multi-source data confidence to determine the influence of each indicator in city classification. The weight of each indicator is jointly determined by expert knowledge and data confidence. Considering the differences in inherent uncertainty between different types of data such as annual time-series data and raster data, expert knowledge is first used to determine the prior basic weights of each indicator, and then the prior basic weights are corrected based on data confidence. The higher the data uncertainty, the smaller the corrected weight of the corresponding indicator. In one implementation, after determining the basic weights of each indicator, a data quality correction factor is further introduced to correct the basic weights to obtain the final weights of each indicator. The final weights are the hybrid weights in the claims. For the first... The first indicator, the first The basic weight of each indicator is denoted as . The first The data quality correction factor for each indicator is denoted as... The first The final weight of each indicator is denoted as: .

[0065] Specifically, first according to the first The data quality correction factor is determined by the data stability of each indicator. The data quality correction factor Calculate according to the following formula:

[0066]

[0067] in, For the first Data quality correction factor for each indicator; For the first The coefficient of variation of the first indicator, which is used to reflect the first... The data stability of the first indicator. A larger coefficient of variation indicates greater stability of the first indicator. The higher the volatility of an indicator's data, the lower its data stability, and the smaller the corresponding data quality correction factor.

[0068] Furthermore, utilizing the aforementioned data quality correction factor For the basic weights The adjustments are made, and the weights of each indicator after the adjustments are normalized to obtain the result. The final weight of each indicator The final weight Calculate according to the following formula:

[0069]

[0070] in, For the first The final weight of each indicator; For the first The basic weight of each indicator; For the first Data quality correction factor for each indicator; This is the sum of the products of the basic weights of each indicator within the same dimension and the data quality correction factor. Through the above normalization process, the final sum of the weights of each indicator is one.

[0071] Through the above processing, the indicators corresponding to annual time-series data, site monitoring data, raster data, and binary classification data can participate in feature fusion according to differences in data stability. Indicators with lower data stability have reduced weights after data quality correction, while indicators with higher data stability maintain a relatively high level of influence within their corresponding dimensions. This allows the subsequent construction of the city feature vector set to simultaneously reflect both expert prior judgments and differences in confidence levels among multi-source data.

[0072] In one implementation, this embodiment uses 2022 data as the basis for city classification and uses city units as the basic spatial unit for analysis. The city classification process includes three stages. The first stage classifies and identifies the ecological environment protection dimension and the urban development and construction dimension, with the ecological environment protection dimension including three sub-dimensions: air, water, and ecology. The second stage analyzes the combined pattern of cities in the ecological environment protection dimension and the urban development and construction dimension based on the classification and identification results of each dimension. The third stage comprehensively considers the performance of cities in the ecological environment protection dimension and the urban development and construction dimension, classifying cities into high-two-high, high-low, low-high, or low-two-low types.

[0073] Within the air sub-dimension of the aforementioned ecological environmental protection dimension, four indicators are selected: annual average wind speed, natural conditions, urban fine particulate matter concentration, and the proportion of days with good or excellent air quality. Annual average wind speed and natural conditions characterize the diffusion conditions of urban air pollutants, while urban fine particulate matter concentration and the proportion of days with good or excellent air quality characterize the current state of urban air pollution. These four indicators collectively form the classification basis of the air sub-dimension, ensuring that the urban characteristics of the air sub-dimension encompass both differences in natural background conditions and differences in air pollution load.

[0074] In this embodiment, annual average wind speed is used as a meteorological indicator affecting the dispersion capacity of air pollutants in the air sub-dimension classification. When the annual average wind speed is high, the likelihood of continuous accumulation of air pollutants within a local space is lower; when the annual average wind speed is low, the dispersion capacity of air pollutants is relatively limited, and the risk of local accumulation is relatively high. Taking the 2022 annual average wind speed data of the Beijing-Tianjin-Hebei region as an example, the grid value of the annual average wind speed in the study area ranges from 1.32771 m / s to 4.17948 m / s. The annual average wind speed is relatively high in northwestern cities and eastern coastal cities of the Beijing-Tianjin-Hebei region, and relatively low in central cities. This spatial distribution indicates that there are differences in the dispersion capacity of air pollutants among different urban units, and annual average wind speed can serve as an important input indicator characterizing dispersion conditions in the air sub-dimension.

[0075] In the classification process for the air sub-dimension, annual average wind speed is used as raster data for feature extraction. First, using urban built-up areas or administrative boundaries as masks, effective pixel values ​​within the corresponding urban units are extracted. Then, a raster feature statistical compression method based on discrete probability distribution is used to compress and represent the effective pixel values, forming corresponding raster dimensionality-reduced features. These raster dimensionality-reduced features are used to characterize the spatial distribution of annual average wind speed within the corresponding urban units, and to avoid excessively high feature dimensionality caused by directly entering the original raster pixels into the clustering model.

[0076] Furthermore, the grid-reduced features corresponding to the annual average wind speed are fused with the indicator features corresponding to natural conditions, urban fine particulate matter concentration, and the proportion of days with good or excellent air quality. Spatial coordinates are introduced as independent variables to form a city feature vector set for the air sub-dimensional. Subsequently, this city feature vector set for the air sub-dimensional is input into an unsupervised machine learning method for clustering, obtaining a binary classification result for each city in the air sub-dimensional. Based on the binary classification result, each city is determined to have either good or poor air quality in the air sub-dimensional.

[0077] In one embodiment, in the classification process of the air sub-dimension, in addition to the annual average wind speed, natural conditions are also used as indicators characterizing the conditions for the diffusion of air pollutants. These natural conditions correspond to the binary classification data in the claims and are used to characterize whether the urban geographical natural conditions are conducive to the diffusion of air pollutants. These natural conditions can be determined based on the city's geographical pattern, topographic features, and location relative to land and sea. Geographical pattern, topographic features, and location relative to land and sea constitute the inherent natural basis for the accumulation and diffusion of air pollutants, and can reflect the differences in air flow, vertical exchange, and pollutant retention among different urban units.

[0078] Taking the Beijing-Tianjin-Hebei region as an example, the study area is bordered by the Taihang Mountains to the west, the Yanshan Mountains to the north, and the North China Plain to the central and eastern parts, presenting an overall spatial pattern of mountains surrounding the plain and plains opening to the south. The northwest and north of the study area have relatively high elevations and more pronounced topographic relief; the central, eastern, and southern parts are mainly plains with relatively low and flat terrain. This spatial pattern creates a transitional topographic condition from a semi-enclosed basin to a plain, which to some extent restricts near-surface air flow, making pollutants more likely to accumulate and remain in the piedmont plains and low-wind-speed areas. Therefore, the overall topographic conditions of the Beijing-Tianjin-Hebei region are not conducive to the rapid dispersion of pollutants.

[0079] Different cities have different locations, resulting in varying conditions for the dispersion of air pollutants. Qinhuangdao, Chengde, and Zhangjiakou, influenced by their mountainous terrain, higher elevation, or transitional location between mountains and the sea, have relatively favorable natural conditions for air pollutant dispersion. Other cities, to some extent, are affected by mountain barriers, low wind speeds on plains, or insufficient near-surface exchange capacity, making it difficult for air pollutants to dissipate quickly. Based on the above geographical patterns, topographic features, and location relative to land and sea, a binary assessment of the air pollutant dispersion conditions for each city unit was conducted, yielding the results of the natural condition assessment for each city, as shown in Table 2.

[0080] Table 2 Results of Natural Condition Assessment for Cities in the Beijing-Tianjin-Hebei Region

[0081] In Table 2, cities with relatively favorable natural conditions for air pollutant dispersion are denoted as favorable, and cities with relatively unfavorable natural conditions for air pollutant dispersion are denoted as unfavorable. The results of the natural condition determination are used as binary classification data in the air sub-dimension for subsequent feature fusion.

[0082] In this embodiment, urban fine particulate matter concentration is used as an indicator reflecting the level of urban air pollution in the air sub-dimension classification. The urban fine particulate matter concentration in this embodiment is the annual average concentration of PM2.5. PM2.5 is a significant pollutant affecting the health risks of air pollution, and the annual average concentration of PM2.5 can characterize the fine particulate matter pollution load of urban units on an annual scale. Taking the 2022 data of the Beijing-Tianjin-Hebei region as an example, the annual average concentrations of PM2.5 in Zhangjiakou, Chengde, and Qinhuangdao were all within the secondary limit range stipulated in the "Ambient Air Quality Standards," with the transitional secondary limit being 30 micrograms per cubic meter; the annual average concentrations of PM2.5 in Beijing and Langfang were both lower than the average level of 37 micrograms per cubic meter in the Beijing-Tianjin-Hebei region; the annual average concentrations of PM2.5 in other cities were relatively higher. These differences indicate that there are significant differences in the fine particulate matter pollution load among different urban units, and the annual average concentration of PM2.5 can serve as an important input indicator characterizing the current state of air pollution in the air sub-dimension.

[0083] In this embodiment, the proportion of days with good or excellent air quality is used as an annual time-series indicator reflecting the air quality status and is included in the air sub-dimension classification. The proportion of days with good or excellent air quality can characterize the stability of air quality at the annual scale for urban units. Taking the 2022 data of the Beijing-Tianjin-Hebei region as an example, the proportion of days with good or excellent air quality generally shows a spatial pattern of decreasing from north to south. Chengde and Zhangjiakou have the highest proportions, followed by Qinhuangdao, Beijing, and Tangshan, while southern cities such as Shijiazhuang, Xingtai, and Handan have relatively low proportions. These differences indicate that different cities have significant differences in the sustainability of air quality, and the proportion of days with good or excellent air quality can, together with the annual average PM2.5 concentration, characterize the current state of air pollution.

[0084] In the feature construction of the air sub-dimensionality, annual average wind speed is used as raster data for feature extraction, natural conditions are used as binary classification data for feature extraction, and urban fine particulate matter concentration and the proportion of days with good air quality are used as annual time-series data for feature extraction. Specifically, the raster data corresponding to annual average wind speed is first compressed and characterized to obtain raster dimensionality reduction features corresponding to annual average wind speed; then, the index features corresponding to natural conditions, urban fine particulate matter concentration, and the proportion of days with good air quality are extracted. Subsequently, the raster dimensionality reduction features and the index features are fused according to a mixed weight based on the confidence of multi-source data, and spatial coordinates are introduced as independent variables to form the urban feature vector set of the air sub-dimensionality.

[0085] Based on four indicators—annual average wind speed, natural conditions, annual average PM2.5 concentration, and the proportion of days with good or excellent air quality—a set of urban feature vectors for the air quality sub-dimensional is constructed, and cluster analysis is performed on the urban air environment status. In this embodiment, a Gaussian mixture model is used to cluster the urban feature vector set for the air quality sub-dimensional, dividing each city into two categories: good air quality and poor air quality. The clustering results show that Qinhuangdao, Chengde, and Zhangjiakou are identified as cities with good air quality, while the remaining cities are identified as cities with poor air quality. Cities with good air quality typically have strong atmospheric pollutant diffusion capabilities, making it difficult for pollutants to accumulate in local spaces for a long period, resulting in a relatively low overall pollution load; cities with poor air quality typically have relatively limited diffusion conditions or a higher pollution load. The classification results of the air quality sub-dimensional can reflect the comprehensive differences between natural background conditions and the current pollution status. In the water sub-dimensional of the ecological environment protection dimension, three indicators are selected: multi-year average rainfall, the proportion of surface water with good or excellent quality, and surface water quality monitoring indicators. The proportion of surface water with good or excellent quality is the proportion that reaches or is better than Class III water quality. Multi-year average rainfall is used to characterize urban water resource conditions, the proportion of surface water with excellent quality is used to reflect the overall level of urban surface water quality, and surface water quality monitoring indicators are used to characterize the comprehensive characteristics of urban surface water quality. These three indicators together constitute the classification basis of the water sub-dimension, so that the urban characteristics of the water sub-dimension include both differences in natural background conditions and differences in water environment quality.

[0086] In this embodiment, multi-year average rainfall is included as an important natural factor affecting water environmental quality in the water sub-dimension classification. Multi-year average rainfall can reflect, to some extent, the ability of urban water bodies to dilute pollutants and their water renewal capacity. Cities with higher multi-year average rainfall generally have stronger water renewal conditions, which are conducive to pollutant dilution and transport; cities with lower multi-year average rainfall have relatively limited water renewal capacity and relatively weaker pollutant dilution conditions.

[0087] Figure 2 This is a schematic diagram illustrating the multi-year average rainfall in the Beijing-Tianjin-Hebei region, as described in this embodiment of the invention. Figure 2 It can be seen that the average annual rainfall in the Beijing-Tianjin-Hebei region from 2015 to 2024 was generally at a lower-middle level compared to the national average, with some differences among different cities. The average annual rainfall in the Beijing-Tianjin-Hebei region exhibits a spatial pattern of relatively higher rainfall in coastal areas and relatively lower rainfall in inland areas. Coastal cities such as Qinhuangdao and Tangshan have relatively high average annual rainfall, while northwestern cities such as Zhangjiakou and Chengde have relatively low average annual rainfall. These spatial differences indicate that the natural conditions of water resources are not consistent among different urban units, and the average annual rainfall can serve as an important input indicator representing natural endowments in the water sub-dimension.

[0088] In the feature construction of the water sub-dimensional dimension, corresponding indicator features are extracted from the multi-year average rainfall, the proportion of surface water with excellent quality, and surface water quality monitoring indicators. Subsequently, based on a mixed weighting method using multi-source data confidence, the indicator features corresponding to the multi-year average rainfall, the proportion of surface water with excellent quality, and the surface water quality monitoring indicators are fused, and spatial coordinates are introduced as independent variables to form a city feature vector set for the water sub-dimensional dimension. After inputting the city feature vector set of the water sub-dimensional dimension into an unsupervised machine learning method, binary classification results are obtained for each city in the water sub-dimensional dimension, and the city is determined to have good or poor water quality based on this.

[0089] In the water sub-dimension of the ecological and environmental protection dimension, the proportion of surface water with excellent or good quality serves as an important indicator for measuring water environmental quality and participates in the water sub-dimension classification. This proportion refers to the percentage of cross-sections that meet or exceed the Class III water quality standard. Class III water mainly applies to secondary protection zones of centralized drinking water surface water sources, overwintering grounds for fish and shrimp, migration channels, aquaculture areas, and swimming areas, where all water quality indicators meet the Class III water quality standards in the "Surface Water Environmental Quality Standard." A higher proportion of excellent or good surface water quality indicates a better overall surface water quality in the corresponding urban unit. Taking the 2022 data from the Beijing-Tianjin-Hebei region as an example, Qinhuangdao, Hengshui, and Xingtai had a proportion of excellent or good surface water quality exceeding 80%, while Tianjin and Cangzhou had proportions below 60%, with other cities falling between these ranges. These differences indicate that there are variations in the overall surface water quality level among different urban units, and the proportion of excellent or good surface water quality can serve as an important input indicator for characterizing the overall water quality level in the water sub-dimension.

[0090] In this embodiment, surface water quality monitoring indicators are used to characterize the comprehensive characteristics of urban surface water quality. These indicators utilize data from automatic monitoring stations, selecting four metrics: turbidity, dissolved oxygen, ammonia nitrogen, and chemical oxygen demand (COD), to characterize real-time water quality changes. Turbidity reflects the impact of suspended particulate matter on water transparency; lower turbidity values ​​indicate clearer water. Dissolved oxygen reflects the oxygen content in the water; higher dissolved oxygen values ​​are more conducive to maintaining the stability of the aquatic ecosystem. Ammonia nitrogen reflects the level of ammonia nitrogen pollution in the water; higher ammonia nitrogen values ​​indicate a higher degree of pollution. COD reflects the degree of influence of organic pollutants on the water; higher COD values ​​indicate a higher degree of organic pollution.

[0091] Taking the monitoring results of major water quality indicators in the Beijing-Tianjin-Hebei region in 2022 as an example, the water quality in most cities remained at a high level, but some cities or sections still showed relatively high levels of water quality indicators. Turbidity in some sections of Chengde and Cangzhou exceeded 100 NTU; ammonia nitrogen concentration in some sections of Shijiazhuang and Cangzhou exceeded 1 mg / L; and the chemical oxygen demand (COD) in Tianjin was significantly higher than in other cities. These monitoring results indicate that different urban units exhibit differences in turbidity, dissolved oxygen, ammonia nitrogen, and COD, and the aforementioned surface water quality monitoring indicators can serve as important input indicators characterizing the comprehensive water quality features within the water sub-dimension.

[0092] Based on multi-year average rainfall, the proportion of surface water with excellent quality, and surface water quality monitoring indicators, a city feature vector set for the water sub-dimension is constructed, and cluster analysis is performed on the urban water environment status. In this embodiment, cities are divided into two categories: good water quality and poor water quality. The clustering results show that Tianjin, Cangzhou, Langfang, Shijiazhuang, and Xingtai are identified as poor water quality cities, while the other cities are identified as good water quality cities. Cities with good water quality typically have favorable natural conditions and high water quality. They are mostly located in the Yanshan to Taihang Mountains or piedmont areas, with relatively high altitudes, strong natural water regeneration capacity, and relatively low population and industrial density, resulting in less anthropogenic pollutant discharge. Cities with poor water quality typically have high pollution loads or relatively limited self-purification capacity of their water bodies. Therefore, the binary classification results of the water sub-dimension can reflect the comprehensive differences in natural water resource conditions and surface water quality status.

[0093] Within the ecological sub-dimension of the ecological environmental protection dimension, two indicators are selected: vegetation index and green coverage rate of built-up areas. The selection of these two indicators is shown in Table 1. In this embodiment, the vegetation index is the Normalized Difference Vegetation Index (NDVI). The vegetation index is used to characterize the city's natural ecological baseline, while the green coverage rate of built-up areas is used to characterize the level of urban green space construction. These two indicators together constitute the classification basis of the ecological sub-dimension, ensuring that the urban characteristics of the ecological sub-dimension include both differences in natural baseline conditions and differences in the level of urban green space construction.

[0094] In this embodiment, the vegetation index is used as an important indicator to measure the vegetation cover of a region in the ecological sub-dimension classification. The higher the vegetation index value, the higher the vegetation cover and the stronger the ecosystem stability. The vegetation index typically ranges from -1 to 1; when the vegetation index is less than zero, it often corresponds to non-vegetated surfaces such as water bodies, clouds, or snow cover; when the vegetation index is between 0 and 0.2, it often corresponds to bare land or sparse vegetation; when the vegetation index is between 0.2 and 0.5, it usually corresponds to medium-coverage vegetation such as grassland or farmland; when the vegetation index is greater than 0.5, it often corresponds to high-coverage vegetation types such as forests.

[0095] Taking the vegetation index data of the Beijing-Tianjin-Hebei region in 2022 as an example, differences exist among cities. Mountainous areas such as Chengde, northern Beijing, and western Baoding have relatively high vegetation indices; plains areas such as Beijing city proper, Tianjin, Shijiazhuang, and Handan have relatively low vegetation indices; Zhangjiakou's overall vegetation index is at a moderate level. Northern Zhangjiakou, located on the Bashang Plateau, has vegetation mainly consisting of grassland and shrubland, resulting in a vegetation index relatively lower than that of mountainous areas with higher forest cover. These spatial differences indicate that the natural ecological foundations of different urban units are not consistent, and the vegetation index can serve as an important input indicator representing the natural ecological foundation in the ecological sub-dimension.

[0096] In this embodiment, the green coverage rate of built-up areas is used as an important indicator to measure the level of urban greening construction in the ecological sub-dimension classification. The green coverage rate of built-up areas is the proportion of various types of green coverage area within the urban built-up area to the total area of ​​the built-up area. A higher green coverage rate indicates a higher level of greening in the corresponding city and better ecological environmental protection and construction.

[0097] Figure 3 This is a schematic diagram illustrating the green coverage rate of urban built-up areas in the Beijing-Tianjin-Hebei region, as described in an embodiment of the present invention. Figure 3 It can be seen that, except for Tianjin, the green coverage rate of built-up areas in other cities in the Beijing-Tianjin-Hebei region all exceeded 40% in 2022. Beijing, Handan, and Langfang had relatively high green coverage rates, all exceeding 47%. These differences indicate that there are variations in the level of urban greening construction among different city units, and that the green coverage rate of built-up areas can serve as an important input indicator for characterizing the level of urban green space construction in the ecological sub-dimension.

[0098] Based on vegetation index and green coverage rate of built-up areas, a set of urban feature vectors for the ecological sub-dimensional is constructed, and cluster analysis is performed on the urban ecological environment protection status. In this embodiment, cities are divided into two categories: "good ecological environment" and "poor ecological environment." The clustering results show that Beijing, Chengde, Hengshui, and Qinhuangdao are identified as "good ecological environment" cities, while the remaining cities are identified as "poor ecological environment" cities. Cities with good ecological environment typically have a good natural ecological foundation or a high level of urban greening. Cities with poor ecological environment typically have relatively low vegetation coverage or high intensity of urban construction. Therefore, the binary classification results of the ecological sub-dimensional can reflect the comprehensive differences in natural ecological foundation and urban green space construction level.

[0099] In the dimension of urban development and construction, eight indicators were selected: GDP per capita, urbanization level, industrial added value per capita, built-up area per capita, energy consumption per unit of GDP, water consumption per unit of GDP, carbon dioxide emissions per unit of GDP, and building land use area per unit of GDP. The specific indicators are shown in Table 1. These eight indicators characterize the comprehensive performance of urban development and construction from two aspects: the level of urban economic development and the level of intensive development. They reflect not only differences in urban economic development levels but also differences in resource and energy utilization efficiency, land use efficiency, and the degree of intensive construction.

[0100] Figure 4 This is a schematic diagram illustrating the urban development and construction levels of the Beijing-Tianjin-Hebei region in an embodiment of the present invention. Figure 4 It can be seen that different cities differ in terms of GDP per capita, urbanization level, industrial added value per capita, built-up area per capita, energy consumption per unit of GDP, water consumption per unit of GDP, carbon dioxide emissions per unit of GDP, and building land use area per unit of GDP.

[0101] In terms of GDP per capita, Beijing's GDP per capita is significantly higher than other cities, placing it at the forefront of the region in terms of urban development and construction. Tianjin and Tangshan also have relatively high GDP per capita, while other cities have relatively lower GDP per capita. Urbanization level reflects the stage of urbanization; a higher value indicates a higher degree of urbanization and a greater concentration of population and economic activity in cities. Beijing and Tianjin both have urbanization levels exceeding 85%, significantly higher than other cities. Shijiazhuang, Zhangjiakou, Langfang, and Tangshan also have relatively high urbanization levels, all exceeding 65%. Other cities have relatively low urbanization levels.

[0102] In terms of per capita industrial added value, there are significant differences among cities in the Beijing-Tianjin-Hebei region. Tangshan has the highest per capita industrial added value, indicating a strong supporting role of industrial development in economic growth; Tianjin and Beijing also have relatively high per capita industrial added value; while cities like Zhangjiakou and Baoding have relatively lower per capita industrial added value. Per capita built-up area is used to characterize urban spatial expansion. Cities like Xingtai, Cangzhou, and Chengde have relatively high per capita built-up area, indicating relatively dispersed urban spaces; while Beijing and Tianjin have relatively low per capita built-up area, indicating relatively intensive urban land use.

[0103] Regarding energy consumption per unit of GDP, the lower the value, the lower the dependence of economic development on energy. Beijing has the lowest energy consumption per unit of GDP, with significantly better energy utilization efficiency than other cities; Langfang and Baoding also have relatively low energy consumption per unit of GDP; Tangshan, Handan, and Chengde have higher energy consumption per unit of GDP. Water consumption per unit of GDP reflects the water resource utilization efficiency of economic activities. Beijing has the lowest water consumption per unit of GDP and the highest water use efficiency; Tianjin and Langfang also have relatively high water use efficiency; Hengshui and Xingtai have significantly higher water consumption per unit of GDP than other cities, indicating relatively lower water resource utilization efficiency.

[0104] In terms of carbon dioxide emissions per unit of GDP, there are significant differences among cities in the Beijing-Tianjin-Hebei region. Beijing and Zhangjiakou have the lowest carbon dioxide emissions per unit of GDP, indicating relatively low carbon intensity; Tianjin and Baoding are at a medium level; Handan, Tangshan, and Chengde have relatively high carbon dioxide emissions per unit of GDP. Building land use area per unit of GDP reflects the land input intensity corresponding to economic output. Langfang has the highest building land use area per unit of GDP, indicating relatively extensive land use; Xingtai, Chengde, and Hengshui are also at relatively high levels; Beijing and Tianjin have the lowest building land use area per unit of GDP, indicating relatively high land use efficiency.

[0105] Based on the aforementioned eight indicators, a set of urban feature vectors is constructed along the dimensions of urban development and construction, and cluster analysis is performed on the levels of urban development and construction. In this embodiment, cities are divided into two categories: well-developed and poorly developed. The clustering results show that Beijing, Tianjin, Shijiazhuang, and Zhangjiakou are identified as well-developed cities. These cities have a high level of economic development, a high degree of urbanization, a relatively complete industrial development foundation, and perform well in terms of resource and energy utilization efficiency and intensive land use. The other nine cities are identified as poorly developed cities. These cities have a relatively low overall level of economic development, and there is still room for improvement in their urbanization process and industrial support capacity. They also exhibit a certain degree of extensive characteristics in terms of energy, water resources, and land use.

[0106] Furthermore, the cluster analysis results of the three sub-dimensions of air, water, and ecology covered by the ecological environment protection dimension are summarized and combined with the cluster analysis results of the urban development and construction dimension to conduct a comprehensive analysis at the city scale. This aims to identify the combined characteristics of the ecological environment protection level and the urban development and construction level, forming a combined pattern of ecological environment protection and urban development and construction. This combined pattern is used for further city type classification.

[0107] In this embodiment, based on the feature vectors extracted by principal component analysis, a Gaussian mixture model was used for cluster analysis, and the clustering results are shown in Table 4. Chengde and Qinhuangdao were classified as high-level in all three sub-dimensions of air, water, and ecology, indicating that the basic ecological environment conditions of these cities are generally good. Zhangjiakou was classified as high-level in the air and water sub-dimensions, but low-level in the ecology sub-dimension. The natural background conditions in northern Zhangjiakou, dominated by grasslands and shrubs, resulted in a vegetation index that was relatively lower than that of mountainous areas with higher forest coverage. Beijing and Hengshui were classified as high-level in the water and ecology sub-dimensions, but low-level in the air sub-dimension. Baoding, Handan, and Tangshan were classified as high-level only in the water sub-dimension. Tianjin, Cangzhou, Langfang, Shijiazhuang, and Xingtai were classified as low-level in all three sub-dimensions of air, water, and ecology. Table 3 shows the city clustering results in this embodiment of the invention, which lists the binary classification results of each city in the dimensions of air, water, ecology, and urban development and construction.

[0108] Table 3 City Clustering Results

[0109]

[0110] Based on the comprehensive judgment rules for ecological and environmental protection status, the overall ecological and environmental protection status of cities is assessed. A city is classified as having a high level of ecological and environmental protection when it is rated as high in two or more of the three sub-dimensions (air, water, and ecology); a city is classified as having a low level of ecological and environmental protection when it is rated as high in only one or zero sub-dimensions. According to Table 3, Beijing, Zhangjiakou, Chengde, Hengshui, and Qinhuangdao are classified as cities with a high level of ecological and environmental protection, while the remaining cities are classified as cities with a low level. Clustering results for the urban development and construction dimension show that Beijing, Tianjin, Shijiazhuang, and Zhangjiakou are classified as cities with a high level of urban development and construction, while the remaining cities are classified as cities with a low level. Combining the comprehensive judgment results for the ecological and environmental protection dimensions, different cities exhibit various combinations of performance across the two dimensions.

[0111] Furthermore, based on the comprehensive performance of cities in the dimensions of ecological environmental protection and urban development and construction, cities are divided into four combination types. Based on the results of principal component analysis to extract feature vectors and Gaussian mixture model for cluster analysis, cities in the Beijing-Tianjin-Hebei region are classified. The classification results of cities in the Beijing-Tianjin-Hebei region according to their comprehensive performance in the dimensions of ecological environmental protection and urban development and construction are shown in Table 4, including two high-performing cities, two low-performing cities, three high-low cities, and six low-performing cities.

[0112] Table 4 shows the combination pattern of cities in the dimensions of ecological environment protection and urban development and construction in the embodiments of the present invention, which is used to list the judgment results of each city in the dimensions of ecological environment protection, urban development and construction, and comprehensive city category.

[0113] Table 4. Combination Pattern of Cities in Ecological Environmental Protection and Urban Development and Construction Dimensions

[0114]

[0115] According to Table 4, the "Double High" type cities include Beijing and Zhangjiakou. These cities have high levels in both ecological environment protection and urban development and construction. When constructing the indicator system, the evaluation indicators can focus on state indicators. The "Low-High" type cities include Tianjin and Shijiazhuang. These cities have high levels in urban development and construction but low levels in ecological environment protection. Ecological environment-related indicators can focus on process indicators, while low-carbon construction and green production-related indicators can focus on state indicators. The "High-Low" type cities include Chengde, Qinhuangdao, and Hengshui. These cities have high levels in ecological environment protection but low levels in urban development and construction. Urban construction and production-related indicators can focus on process indicators, while low-carbon construction and green production-related indicators can focus on state indicators. The "Double Low" type cities include Baoding, Cangzhou, Handan, Tangshan, Xingtai, and Langfang. These cities have low levels in both ecological environment protection and urban development and construction. When constructing the indicator system, the evaluation indicators can focus on process indicators.

[0116] In this embodiment, the basic weight settings for each indicator are shown in Table 5. Table 5 shows the selection of city classification indicators and basic weight settings for beautiful city construction in this embodiment of the invention, which lists the basic weights corresponding to each indicator in the dimensions of air, water, ecology, and urban development and construction. The basic weights in Table 5 are used for the mixed weight calculation based on the confidence level of multi-source data. Each indicator first participates in the weight initialization according to the basic weight, and then is corrected by combining the data quality correction factor to obtain the mixed weight of the corresponding indicator in feature fusion.

[0117] Table 5. Selection of City Classification Indicators and Setting of Basic Weights for Beautiful City Construction

[0118]

[0119] In each dimension of clustering, city longitude and latitude are introduced as spatial coordinates into the city feature vector set. The basic weights for city longitude and city latitude are both 0.05.

[0120] The beneficial effects of this invention are as follows:

[0121] (1) This invention unifies the feature representation of annual time-series data, site monitoring data, raster data, and binary classification data, and performs feature fusion in the dimensions of air, water, ecology, and urban development and construction, so that urban indicators from different sources, scales, and data forms can enter the same urban feature vector set. At the same time, it reduces the impact of low-stability data on the classification results by using mixed weights based on the confidence of multi-source data, and introduces urban longitude and latitude as spatial coordinates into the clustering process, so that the urban classification results simultaneously reflect attribute similarity and geographical proximity. Thus, this invention can further obtain comprehensive urban categories such as high-air quality, high-low quality, low-high quality, and low-low quality, based on the subdivided results of good or poor air quality, good or poor water quality, good or poor ecology, and good or poor development. The classification basis is clear, the data source is traceable, and it can provide a stable urban type foundation for the construction of a differentiated indicator system.

[0122] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0123] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A city classification method for constructing a differentiated index system for beautiful cities, characterized in that, include: Using city units as the classification object, multi-source heterogeneous data are obtained to characterize the ecological environment protection dimension and the urban development and construction dimension. The ecological environment protection dimension includes three sub-dimensions: air, water, and ecology. The multi-source heterogeneous data includes annual time-series data, station monitoring data, raster data, and binary classification data. The raster data is compressed and characterized using a raster feature statistical compression method based on discrete probability distribution to obtain raster dimensionality reduction features, and corresponding indicator features are extracted from the annual time series data, the site monitoring data, and the binary classification data. The basic weights of each indicator are set based on expert knowledge, and the basic weights are modified by combining the confidence of data of different data types to obtain the hybrid weights based on the confidence of multi-source data. In the dimensions of air, water, ecology, and urban development and construction, the corresponding grid dimensionality reduction features and index features are fused according to the hybrid weights, and spatial coordinates are introduced as independent variables to form a set of urban feature vectors for the corresponding dimensions. The feature vector sets of each city are input into an unsupervised machine learning method for clustering, resulting in binary classification results for each city in the dimensions of air, water, ecology, and urban development and construction. Based on the binary classification results, each city is classified into a sub-category consisting of good or bad air quality, good or bad water quality, good or poor ecology, and good or poor development. If two or more of the three sub-dimensions of air, water, and ecology are judged as high level, the city is judged as having a high level of ecological and environmental protection; otherwise, the city is judged as having a low level of ecological and environmental protection. Based on the high or low level of ecological environmental protection and the binary classification results of urban development and construction, each city is divided into two types: high-high, high-low, low-high, or low-high.

2. The city classification method for constructing the differentiated index system for beautiful cities according to claim 1, characterized in that, Acquire multi-source heterogeneous data to characterize the dimensions of ecological environmental protection and urban development and construction, including: Obtain annual average wind speed, natural conditions, urban fine particulate matter concentration, and the percentage of days with good air quality to characterize air sub-dimensions; The multi-year average rainfall, the proportion of surface water with excellent quality, and surface water quality monitoring indicators used to characterize water sub-dimensions are obtained. These surface water quality monitoring indicators include turbidity, dissolved oxygen, ammonia nitrogen, and chemical oxygen demand. Obtain vegetation indices and green coverage rates of built-up areas to characterize ecological sub-dimensions; Obtain the following metrics to characterize urban development and construction: GDP per capita, urbanization level, industrial added value per capita, built-up area per capita, energy consumption per unit of GDP, water consumption per unit of GDP, carbon dioxide emissions per unit of GDP, and building land use area per unit of GDP.

3. The city classification method for constructing the differentiated index system for beautiful cities according to claim 1, characterized in that, The raster data is compressed and characterized using a raster feature statistical compression method based on discrete probability distribution to obtain raster dimensionality reduction features, including: The value range of the raster data within the overall study area is divided into multiple intervals; Using urban built-up areas or administrative boundaries as masks, effective pixels within the corresponding urban units are extracted to form a spatial sample set. Based on the frequency density of the spatial sample set in each interval, a discrete probability distribution vector is constructed. The discrete probability distribution vector is combined with the statistical feature values ​​to form the low-dimensional feature vector corresponding to the raster data; The low-dimensional feature vector is subjected to dimensionality reduction processing to obtain the raster dimensionality-reduced feature.

4. The city classification method for constructing the differentiated index system for beautiful cities according to claim 3, characterized in that, The discrete probability distribution vector is combined with statistical feature values ​​to form a low-dimensional feature vector corresponding to the raster data, including: The discrete probability distribution vector is used as a distribution feature to characterize the distribution shape of pixel values; Based on the spatial sample set, statistical characteristic values ​​are determined to characterize the central tendency and quantile distribution of pixel values; The distribution characteristics and the statistical characteristic values ​​are combined to form the low-dimensional feature vector.

5. The city classification method for constructing the differentiated index system for beautiful cities according to claim 1, characterized in that, Extracting relevant indicator features from the annual time-series data, the site monitoring data, and the binary classification data, including: Extract the following indicators from the annual time-series data: urban fine particulate matter concentration, proportion of days with good air quality, multi-year average rainfall, proportion of surface water with good quality, green coverage rate of built-up areas, per capita GDP, urbanization level, per capita industrial added value, per capita built-up area, energy consumption per unit of GDP, water consumption per unit of GDP, carbon dioxide emissions per unit of GDP, and building land use area per unit of GDP. Extract the corresponding index features of turbidity, dissolved oxygen, ammonia nitrogen, and chemical oxygen demand from the monitoring data of the aforementioned sites; The index features corresponding to the natural conditions are extracted from the binary classification data, and the natural conditions are used to characterize whether they are conducive to the diffusion of air pollutants.

6. The city classification method for constructing the differentiated index system for beautiful cities according to claim 1, characterized in that, Based on expert knowledge, basic weights for each indicator are set, and these basic weights are then adjusted by incorporating data confidence levels from different data types, resulting in a hybrid weighting based on multi-source data confidence levels, including: The basic weights of each indicator are set based on expert knowledge; The data quality correction factor for each indicator is determined based on the coefficient of variation of each indicator, wherein the coefficient of variation is used to characterize the data stability of the corresponding indicator. The data quality correction factor is used to correct the basic weights of the corresponding indicators to obtain the corrected weights of the corresponding indicators. The modified weights of each indicator are normalized to obtain the mixed weights of each indicator.

7. The city classification method for constructing the differentiated index system for beautiful cities according to claim 2, characterized in that, In the dimensions of air, water, ecology, and urban development and construction, the corresponding raster dimensionality reduction features and indicator features are fused according to the hybrid weights, and spatial coordinates are introduced as independent variables to form a set of urban feature vectors for the corresponding dimensions, including: In the air sub-dimension, the grid dimensionality reduction features corresponding to the annual average wind speed and the index features corresponding to natural conditions, urban fine particulate matter concentration and the proportion of days with good air quality are fused according to the hybrid weight, and longitude and latitude are introduced as the spatial coordinates to form the urban feature vector set of the air sub-dimension. In the water sub-dimension, the multi-year average rainfall, the proportion of surface water with excellent quality and the surface water quality monitoring indicators are fused according to the mixed weights, and the spatial coordinates are introduced to form a city feature vector set in the water sub-dimension. In the ecological sub-dimension, the raster dimensionality reduction features corresponding to the vegetation index and the indicator features corresponding to the green coverage rate of the built-up area are fused according to the hybrid weight, and the spatial coordinates are introduced to form a set of urban feature vectors for the ecological sub-dimension. In the urban development and construction dimension, the indicators corresponding to per capita GDP, urbanization level, per capita industrial added value, per capita built-up area, energy consumption per unit of GDP, water consumption per unit of GDP, carbon dioxide emissions per unit of GDP, and building land use area per unit of GDP are integrated according to the mixed weights, and the spatial coordinates are introduced to form an urban feature vector set for the urban development and construction dimension.

8. The city classification method for constructing the differentiated index system for beautiful cities according to claim 1, characterized in that, The feature vector sets of each city are input into an unsupervised machine learning method for clustering, resulting in binary classification results for each city across the dimensions of air, water, ecology, and urban development and construction, including: The feature vector sets of each city are input into a distance-based clustering algorithm, and the number of clusters is set to two. The cluster centers and sample affiliations are iteratively updated with the goal of minimizing the distance between a sample and its corresponding cluster center. After the clustering process meets the convergence condition, the category affiliation label of each city in the corresponding dimension is output. Based on the category affiliation labels, the binary classification results for each city in the dimensions of air, water, ecology, and urban development and construction are obtained.

9. The city classification method for constructing the differentiated index system for beautiful cities according to claim 1, characterized in that, The feature vector sets of each city are input into an unsupervised machine learning method for clustering, resulting in binary classification results for each city across the dimensions of air, water, ecology, and urban development and construction, including: The feature vector sets of each city are input into a probability distribution-based clustering model, and the number of categories is set to two. Estimate the model parameters of the probability distribution-based clustering model; After the clustering process meets the convergence condition, the probability result of each city belonging to different categories is output. Based on the probability results, the category of each city in the corresponding dimension is determined, and the binary classification results of each city in the dimensions of air, water, ecology and urban development and construction are obtained.

10. The city classification method for constructing the differentiated index system for beautiful cities according to claim 1, characterized in that, Based on the aforementioned high or low level of ecological environmental protection, and the binary classification results of the urban development and construction dimension, each city is divided into two types: high-high and low-high, high-low and low-high, or low-high and low-high, including: When a city is determined to be at a high level in ecological and environmental protection, and the binary classification result of the city's development and construction dimension is also high, the city will be classified as a "dual-high" type. When a city is determined to be at a high level of ecological and environmental protection, and the binary classification result of the city's development and construction dimension is low, the city is classified as either high or low. When a city is determined to have a low level of ecological and environmental protection, and the binary classification result of the city's development and construction dimension is high, the city will be classified as a low-high type. When a city is determined to have a low level of ecological and environmental protection, and the binary classification result of the city's development and construction dimension is also low, the city will be classified as a "double low" type.